kingjones777's picture
add model card
e929d06 verified
|
Raw
History Blame Contribute Delete
9.99 kB
metadata
license: bsd-2-clause
base_model:
  - pipecat-ai/phonellm-alpha-1
  - nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
language:
  - en
tags:
  - gguf
  - rocmfp4
  - rocmfpx
  - strix-halo
  - gfx1151
  - vulkan
  - nemotron
  - mixture-of-experts
  - voice-agent
  - tool-use
  - function-calling
  - pipecat

PhoneLLM Alpha 1 β€” ROCmFP4 for AMD Strix Halo (gfx1151)

The model here is not our work. PhoneLLM Alpha 1 is by Daily / the Pipecat team β€” pipecat-ai/phonellm-alpha-1 β€” a full-parameter fine-tune of NVIDIA Nemotron 3 Nano 30B-A3B. This repository adds only the ROCmFP4/ROCmFPX quantisation ladder for AMD Strix Halo and the measurements below. Go star their repo. Pipecat also ship an official NVFP4 build for NVIDIA Blackwell: pipecat-ai/phonellm-alpha-1-nvfp4.

A hybrid Mamba-Transformer MoE voice-agent model β€” 30B total, 3.5B active β€” quantised to run on a Ryzen AI Max+ 395 (Radeon 8060S, gfx1151). One binary, both backends: HIP (ROCm) and Vulkan are a runtime -dev flag, not a rebuild.

PhoneLLM is built for one job: call the right tool at the right time, with thinking disabled, at phone-call latency. That shapes how we verified it β€” see Verification.

Which file should I use?

File ftype Size Head Notes
Q4_0_ROCMFP4_STRIX_LEAN 106 15.91 GiB q8_0 Flagship β€” start here. Smallest tier that keeps full tool behaviour. Also published on its own: PhoneLLM-Alpha-1-ROCmFP4-STRIX-LEAN-GGUF
Q4_0_ROCMFP4_FAST 103 15.83 GiB q8_0 Smallest file; same tool score as the flagship
Q4_0_ROCMFP4_COHERENT 102 16.91 GiB q8_0 Highest-precision 4-bit tier
Q6_0_ROCMFPX_AGENT 114 27.26 GiB q8_0 6-bit
Q8_0_ROCMFPX 111 30.37 GiB q8_0 Plain Q8 reference
Q8_0_ROCMFPX_AGENT 115 30.84 GiB q8_0 8-bit, agent-tuned tensor map

Take the 15.91 GiB flagship. Across our probe the 30 GiB Q8 tiers score no better than the 16 GiB 4-bit tiers (see Verification). On a 128 GB Strix Halo that leaves real headroom to co-host your ASR and TTS models on the same box β€” which is the point, since PhoneLLM is the LLM stage of a voice pipeline, not a speech model (it is text-in / text-out; you still need STT and TTS).

Quick start

llama-server -m PhoneLLM-Alpha-1-Q4_0_ROCMFP4_STRIX_LEAN.gguf \
  -dev ROCm0 -fa on -ngl 999 -fit off -np 1 \
  -c 32768 -b 4096 -t 8 --jinja \
  --host 0.0.0.0 --port 8080

Run it the way Pipecat recommend the source model: temperature=0 and thinking disabled.

{"chat_template_kwargs": {"enable_thinking": false}}

llama.cpp resolves this model's chat format as peg-native; tool calls come back as proper tool_calls on /v1/chat/completions with --jinja.

β›” You need a ROCmFPX build β€” stock llama.cpp will NOT load these files

ROCmFP4/ROCmFPX use ggml tensor types 100–119; upstream's table stops at 43. Build with both backends:

cmake -S . -B build-hipvk -DCMAKE_BUILD_TYPE=Release \
  -DGGML_HIP=ON -DGGML_VULKAN=ON -DGGML_NATIVE=ON \
  -DVulkan_GLSLC_EXECUTABLE=/usr/bin/glslc -DAMDGPU_TARGETS=gfx1151
cmake --build build-hipvk -j
build-hipvk/bin/llama-server --list-devices   # must list ROCm0 AND Vulkan0

Head protection β€” why every tier uses a q8_0 head

Our usual ladder protects output.weight with q6_K on the 4-bit tiers. That is impossible on this model. hidden_size is 2688, and K-quants use 256-element superblocks:

2688 % 256 = 128        β†’  ggml.c:8024: GGML_ASSERT(start % type_traits[type].blck_size == 0) failed
[1/401] output.weight - [2688, 131072], bf16, converting to q6_K ..   SIGABRT

q8_0 uses 32-element blocks and 2688 % 32 == 0, so every tier here carries a q8_0 head β€” higher precision than our usual q6_K, at a cost of roughly +88 MB on the head tensor.

We caught this as a clean natural experiment in a single run: the three q6_K-head tiers aborted in ~2 s while the q8_0-head tier built normally, same source, same binary, same moment. Note that llama-quantize --dry-run does not catch it β€” the dry run planned all 401 tensors and printed a clean 60247 MiB β†’ 17223 MiB (4.58 BPW) summary. The assert only fires once real data is written.


Verification

Every tier is checked for load, coherence, and β€” because this is the whole point of PhoneLLM β€” tool calling, using the vendor-recommended mode (temperature=0, enable_thinking: false).

The tool probe is deliberately adversarial: it includes a case where the model must not call anything, and a multi-turn case, because the failure Pipecat built this model to avoid is an agent that says "yes, I've booked that" without emitting a call.

All tiers, greedy (temperature=0), enable_thinking: false, --jinja, chat format peg-native.

Tier Size Loads Coherent Tool probe
Q4_0_ROCMFP4_STRIX_LEAN 15.91 GiB βœ… ~10 s βœ… 3/5
Q4_0_ROCMFP4_FAST 15.83 GiB βœ… ~10 s βœ… 3/5
Q4_0_ROCMFP4_COHERENT 16.91 GiB βœ… ~10 s βœ… 2/5
Q6_0_ROCMFPX_AGENT 27.26 GiB βœ… ~20 s βœ… 3/5
Q8_0_ROCMFPX_AGENT 30.84 GiB βœ… ~25 s βœ… 2/5
Q8_0_ROCMFPX 30.37 GiB βœ… ~20 s βœ… 3/5
BF16 source (control) 58.8 GiB βœ… βœ… 1/5

6/6 tiers load and stay coherent. There is no precision-dependent degradation: the 30 GiB Q8 tiers score the same as the 16 GiB 4-bit tiers, and every quantised tier scores at or above the BF16 control. If quantisation were damaging tool calling, the Q8 tiers would lead. They do not β€” so pick on size.

Flagship detail (STRIX_LEAN), 5 adversarial cases:

PASS  booking       book_table {"name":"Chen","party_size":2,"time":"19:00"}   ← normalised "7pm" β†’ 19:00
PASS  escalate      transfer_to_human {"reason":"Customer has called multiple times ..."}
PASS  no-tool       (correctly emitted NO call)
FAIL  availability  (no call β€” the one unambiguous miss)
FAIL  multiturn     check_availability {"date":"Saturday","party_size":4}

Read 3/5 carefully β€” the rubric is strict and opinionated. The multiturn "failure" is the model checking availability before booking, which is defensible agent behaviour; we counted it wrong because our expected answer demanded a booking. The no-tool pass matters most: the model declined to invent a call when none was warranted, which is the failure mode PhoneLLM exists to avoid. Treat these as a smoke test that the tool path survives quantisation, not as a benchmark score β€” for a real score use Pipecat's PhoneBench.

⚠️ An honest limitation: we could not establish a BF16 baseline on this hardware

We ran the BF16 GGUF as a control arm and it misbehaves on gfx1151 when tools are attached β€” the same prompt that a quantised tier answers with a correct book_table call returns, from BF16, either a degenerate repetition loop or an unrelated non-sequitur. Without tools, BF16 is coherent.

So we can report what the quantised tiers do, but we cannot publish a "delta vs BF16" the way Pipecat report NVFP4 (PhoneBench 72.06 β†’ 71.51). Anyone quoting a quality delta for these files against BF16 on this hardware would be quoting a broken control. We have not root-caused it (candidates: gfx1151 bf16 compute, or the peg-native tool-template path); it is flagged here rather than papered over.


Reproduction block

A number without its binary is a rumour.

Host Ryzen AI Max+ 395 (Strix Halo), Radeon 8060S, gfx1151, 128 GB unified
Build ROCmFPX fork @ e7712358806055c70a9753b070202b0cc7c637e3
GGML_HIP=ON GGML_VULKAN=ON GGML_NATIVE=ON, Release, AMDGPU_TARGETS=gfx1151
llama-server sha256 e860b763d5e8f496ddb4b01236d8a8f37ae09dda456b5d076ab3abe2d9deb9fc
Source pipecat-ai/phonellm-alpha-1, 13 safetensors shards, 58.8 GiB
Converted convert_hf_to_gguf.py --outtype bf16 β†’ 401 tensors, 63.18 GB, arch nemotron_h_moe
Quantise llama-quantize --output-tensor-type q8_0 <bf16> <out> <ftype> 12
Serve (verification) -dev ROCm0 -fa on -ngl 999 -fit off -np 1 -c 32768 -b 4096 -t 8 --jinja

Not measured

  • Perplexity (the source is a voice-agent fine-tune; wikitext PPL is a poor proxy and we would rather publish nothing than a misleading number).
  • PhoneBench β€” that is Pipecat's harness; we did not run it.
  • Vision β€” text-only model, no projector.
  • Context beyond 32768 (the source supports 262144).
  • Decode throughput per tier.

License and attribution

Released under BSD 2-Clause, matching the source. The source is itself a derivative of an NVIDIA Nemotron Open Model License work β€” see LICENSE_NVIDIA.txt in the upstream repo.

Acknowledgements

Daily / Pipecat for PhoneLLM and for publishing an honest PhoneBench methodology. NVIDIA for Nemotron 3 Nano and the hybrid Mamba-Transformer architecture. The ROCmFPX project for the FP4/FPX tensor types and the Strix Halo kernels that make these files possible.