Jonandrop commited on
Commit
3533a71
·
verified ·
1 Parent(s): 7ae8757

docs: add Performance section (M5 Pro 64GB, range/honest), vision graft zero-cost finding, M5 Max historical note

Browse files
Files changed (1) hide show
  1. README.md +42 -1
README.md CHANGED
@@ -28,7 +28,7 @@ This repository restores the multimodal vision tower that is typically stripped
28
  ## Architecture
29
 
30
  - **Body**: 4-bit affine, group size 64 (from [Youssofal/Qwen3.6-27B-MTPLX-Optimized-Speed](https://huggingface.co/Youssofal/Qwen3.6-27B-MTPLX-Optimized-Speed))
31
- - **MTP sidecar (mtp/weights.safetensors)**: Prequantized MTP heads.
32
  - **Vision tower (vision_tower.safetensors, 333 tensors, bf16 unquantized)**: Extracted verbatim from the official checkpoint's own `model.visual.*` weights (shards 7 and 8 of [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B)) and renamed to the `vision_tower.*` prefix for MTPLX native VLM compatibility (~879 MB).
33
 
34
  ## Usage (MTPLX)
@@ -40,6 +40,47 @@ mtplx start --model <path-to-this-model-dir> --port 8092 \
40
 
41
  Serves an OpenAI-compatible endpoint supporting both text and image input payloads via `POST /v1/chat/completions`.
42
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
43
  ## Sources And Attribution
44
 
45
  | Component | Source | Revision / License |
 
28
  ## Architecture
29
 
30
  - **Body**: 4-bit affine, group size 64 (from [Youssofal/Qwen3.6-27B-MTPLX-Optimized-Speed](https://huggingface.co/Youssofal/Qwen3.6-27B-MTPLX-Optimized-Speed))
31
+ - **MTP sidecar (mtp/weights.safetensors)**: Prequantized MTP heads with a 3-bit affine group-64 draft-only LM head.
32
  - **Vision tower (vision_tower.safetensors, 333 tensors, bf16 unquantized)**: Extracted verbatim from the official checkpoint's own `model.visual.*` weights (shards 7 and 8 of [Qwen/Qwen3.6-27B](https://huggingface.co/Qwen/Qwen3.6-27B)) and renamed to the `vision_tower.*` prefix for MTPLX native VLM compatibility (~879 MB).
33
 
34
  ## Usage (MTPLX)
 
40
 
41
  Serves an OpenAI-compatible endpoint supporting both text and image input payloads via `POST /v1/chat/completions`.
42
 
43
+ The MTPLX CLI defaults to `mtp_history_policy=committed`, which this model
44
+ requires for healthy deep-position acceptance. The model's recommended
45
+ draft sampler is temperature 0.7 (see `recommended_draft_sampler` in
46
+ `mtplx_runtime.json`); pass `--draft-temperature 0.7` to match.
47
+
48
+ ## Performance (measured)
49
+
50
+ Apple M5 Pro 64 GB, MTPLX 1.0.4, temperature 0.6, `mtp_history_policy=committed`,
51
+ `--draft-temperature 0.7` (model recommended), thinking OFF, warm (4 warmup
52
+ prompts discarded), 8 measured prompts from the calibration_coding suite,
53
+ max_tokens=192, median tok/s.
54
+
55
+ | Depth | tok/s (wall-clock e2e) | tok/s (decode-only) | speedup vs AR (e2e) | acceptance pos1/2/3 |
56
+ |---|---|---|---|---|
57
+ | AR (`--no-mtp`) | 13.7 | 14.5 | 1.00x | - |
58
+ | MTP depth 1 | 19.7 | 20.9 | 1.44x | 0.923 |
59
+ | MTP depth 2 | 19.9 | 20.8 | 1.46x | 0.950 / 0.753 |
60
+ | MTP depth 3 | 18.1 | 19.7 | 1.32x | 0.900 / 0.759 / 0.648 |
61
+
62
+ Two throughput definitions are reported:
63
+
64
+ - **tok/s (wall-clock e2e)** = `generated_tokens / total_elapsed`, includes
65
+ prompt prefill. The real end-to-end speed and the honest denominator for
66
+ the speedup ratio.
67
+ - **tok/s (decode-only)** = `generated_tokens / decode_elapsed`, excludes
68
+ prefill.
69
+
70
+ Acceptance is stable across runs: greedy draft (temp 0.0) gives
71
+ 0.874/0.740/0.636, recommended draft (temp 0.7) gives 0.900/0.759/0.648,
72
+ both within noise. Vision graft has zero acceptance cost relative to the
73
+ text-only source (Youssofal/Qwen3.6-27B-MTPLX-Optimized-Speed measures
74
+ identical acceptance 0.874/0.740/0.636 under the same config).
75
+
76
+ Earlier baked `mtplx_runtime.json` figures (acceptance 1.0/0.98/0.94, ~63
77
+ tok/s) were measured on Apple M5 Max 128 GB and are not reproducible on M5
78
+ Pro 64 GB due to lower memory bandwidth; they are preserved in the
79
+ `historical_m5max` field of `mtplx_runtime.json` for traceability. The
80
+ vision tower is lazy-loaded (only when a request carries an image), so
81
+ text-only throughput is unaffected by it. Numbers are indicative, not a
82
+ benchmark suite.
83
+
84
  ## Sources And Attribution
85
 
86
  | Component | Source | Revision / License |