Ornith 1.5 35B A3B — hybrid MLX + vision + MTPLX

Vision-preserving hybrid MLX quantization of Ornith-1.5-35B-A3B, built from upstream revision fbb995a. It is intended for Apple Silicon and includes the native vision tower and MTP sidecar.

Model format

  • 28.12 GB; Qwen3.5 MoE multimodal (35B total, 3B active)
  • Language body: 4-bit affine, group size 32
  • Attention-sensitive modules: 8-bit affine, group size 64
  • Recurrent linear-attention projections, vision tensors, and MTP tensors: BF16
  • Context metadata: 262,144 tokens; effective tested context: 65,536 tokens
  • Included: tokenizer, chat template, image/video processor metadata, build recipe, and conversion receipt

The evaluated public weights were anonymously verified byte-identical to model revision 08ce2da.

Recommended runtime

MTPLX 2.7.1 selected turbo depth 1 (D1). On the qualification prompt, D1 reached 116.95 decode tok/s versus 63.12 tok/s autoregressive, a 1.85× speedup, with both quality checks passing. Use precise KV cache for the benchmarked configuration.

mtplx serve \
  --model Shiftedx/ornith-1.5-35b-a3b-attention8-bf16recurrence-vision-mtplx \
  --profile turbo \
  --generation-mode mtp \
  --load-mtp \
  --depth 1 \
  --reasoning on \
  --reasoning-effort medium

Shiftedx Harness

A separate three-trial Shiftedx Harness evaluation scored 24/90 (26.7%) direct and 85/90 (94.4%) with the harness. This measures a deployment policy layer, not a change to the model weights.

Vision usage

python -m mlx_vlm.generate \
  --model Shiftedx/ornith-1.5-35b-a3b-attention8-bf16recurrence-vision-mtplx \
  --image image.jpg \
  --prompt "Describe this image." \
  --max-tokens 256

Limitations

  • This is an experimental quantized runtime artifact; behavior may differ from the BF16 parent.
  • Full BF16 parent parity and the 262,144-token metadata limit were not tested.
  • Strict JSON output can include Markdown fences without an additional policy layer.
  • Benchmark results are specific to the linked revisions, runtime settings, and host.
  • Review the upstream model card for intended use, training details, license, and safety considerations.

Shiftedx Bench post-publication qualification

This table was generated from the frozen lightweight quant gate after the model weights were published. Categories remain separate; the benchmark does not produce a composite intelligence score.

Lane Passed Accuracy Mean wall time Mean decode Peak active memory
Quality 8/10 80.0% 10.16 s 108.50 tok/s 42.45 GiB
Long context 13/15 86.7% 38.92 s 94.28 tok/s 45.12 GiB
Tool calling 6/6 100.0% 1.72 s 84.64 tok/s 43.65 GiB
Agentic 1/2 50.0% 4.65 s — tok/s
Vision 1/4 25.0% 1.45 s 108.00 tok/s 42.96 GiB
  • Tested model revision: 08ce2daeea94986b447acd6691b76110856de49e
  • Benchmark: Shiftedx Bench v0.3.0
  • Context lengths represented: 4,096, 16,384, 65,536, 131,072 prompt tokens; effective tested context: 65,536 tokens
  • Runtime contract: MTPLX 2.7.1 turbo D1; thinking on/medium; temperature=1.0, top_p=0.95, top_k=20; native tool prompt; tokenizer chat template; KV cache off; MTP depth 1
  • Host: Apple M4 Max, 64 GiB unified memory
  • Total measured request wall time: 710.85 seconds
  • 260,096-token status: not run; it is outside the lightweight quant gate.

Scores are specific to the linked model revision, benchmark revision, runtime contract, and host. Changing weight precision, KV-cache precision, reasoning mode, template, or speculative depth creates a different benchmark candidate.

Downloads last month
479
Safetensors
Model size
8B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Shiftedx/ornith-1.5-35b-a3b-attention8-bf16recurrence-vision-mtplx

Quantized
(68)
this model

Collection including Shiftedx/ornith-1.5-35b-a3b-attention8-bf16recurrence-vision-mtplx