File size: 4,278 Bytes
63b94e5 3adb20c c3dd288 63b94e5 c3dd288 58dc802 c3dd288 a9dbc81 63b94e5 268f632 63b94e5 c3dd288 63b94e5 c3dd288 63b94e5 c3dd288 63b94e5 c3dd288 63b94e5 c3dd288 63b94e5 c3dd288 63b94e5 c3dd288 63b94e5 c3dd288 3adb20c 8014fa2 63b94e5 c3dd288 63b94e5 8f23fd6 c3dd288 63b94e5 8f23fd6 63b94e5 8f23fd6 3adb20c 8f23fd6 f8a3527 63b94e5 3adb20c c3dd288 63b94e5 c3dd288 63b94e5 c3dd288 63b94e5 3c64222 c3dd288 3adb20c c3dd288 8014fa2 c3dd288 3adb20c c3dd288 63b94e5 cac5b5c c3dd288 63b94e5 c3dd288 3adb20c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 | ---
license: mit
language:
- en
pipeline_tag: text-generation
tags:
- fastflowlm
- q4nx
- npu
- qwen3.6-moe
- 35b-a3b
- ornith
- quantization
base_model:
- ornith-ai/Ornith-1.0-35B
base_model_relation: quantized
quantized_by: Atomic-Germ
---
# *IF YOU USE COMMUNITY QWEN MODELS DO NOT UPGRADE TO FLM v1.0.2+*
# Ornith-1.0-35B - Q4NX for FastFlowLM (AMD Ryzen AI XDNA2)
Ornith-1.0-35B is converted to Q4NX for hardware-accelerated inference with FastFlowLM on AMD Ryzen AI NPUs.
## What is Q4NX?
Q4NX is FastFlowLM's native packed-quantization format - a rearranged Q4_1 layout tuned for the NPU matrix engine's tile sizes and memory access patterns. It is **not** a GGUF file and it does not run on llama.cpp or Ollama; it is meant exclusively for the [FastFlowLM](https://fastflowlm.com) engine on AMD Ryzen AI NPUs.
## Requirements
- FastFlowLM >= 0.9.46 (`flm` CLI)
- AMD Ryzen AI processor with **XDNA2 (NPU2)** - Strix Point / Ryzen AI 300
series or later
- Linux with the XRT NPU stack installed
- ~47 GB of unified system memory (Q4NX weights + activations + KV cache)
## Files
| File | Purpose |
|---|---|
| model.q4nx | Quantized Q4NX text weights |
| config.json | FastFlowLM model configuration |
| tokenizer.json | Tokenizer |
| tokenizer_config.json | Special tokens and chat template |
| chat_template.jinja | Chat template |
| README.md | |
## Install and run
This repository works with `flm-add`, a small installer that copies the model
into the FastFlowLM user directory and registers the tag. It never
modifies the system FastFlowLM install.
`pip install flm-add` or `uv tool install flm-add`
```bash
uv tool install flm-add
unset FLM_CONFIG_PATH FLM_XCLBIN_PATH && flm-add Atomic-Germ/Ornith-1.0-35B-A3B-NPU2 --tag ornith1.0-moe:35b-a3b --family qwen3.6-moe
FLM_CONFIG_PATH="$HOME/.config/flm/model_list.json" FLM_XCLBIN_PATH="$HOME/.config/flm" flm run ornith1.0-moe:35b-a3b
```
| Context Length | TTFT (s) | Prefill Speed (tok/s) | Decoding Speed (tok/s)
|---|---|---|---|
| 1k | 12.355 ± 0.122 | 79.07 ± 0.78 | 10.41 ± 0.01
| 2k | 16.849 ± 0.142 | 115.41 ± 0.97 | 10.36 ± 0.07
| 4k | 24.852 ± 0.035 | 156.09 ± 0.22 | 10.14 ± 0.10
| 8k | 40.241 ± 0.385 | 192.55 ± 1.84 | 9.74 ± 0.07
| 16k | 73.398 ± 1.227 | 211.03 ± 3.53 | 9.02 ± 0.07
| 32k | 148.780 ± 0.630 | 208.11 ± 0.89 | 7.70 ± 0.03
---
## Kernels
FastFlowLM's NPU kernels (xclbins) are closed source and are not shipped in this repository. This model uses the **`qwen3.6-moe`** engine family and is shape-identical to the official **`qwen3.6-moe:35b-a3b`** model (`Qwen3.6-Moe-35BA3B-NPU2`). Point the runtime's xclbin path at the matching `xclbins` directory (or ship your own) before running.
## Model
- Registry tag: `ornith1.0-moe:35b-a3b`
- Engine family: `qwen3.6-moe`
- Kernel source: official fastflowlm `qwen3.6-moe:35b-a3b`
- Context length: 262,144 tokens (from config)
- Hidden size: 2048
- Layers: 40
- Intermediate size: 512
- Vocabulary: 248320
- `model.q4nx` size: 23.2 GB
- Base model: [ornith-ai/Ornith-1.0-35B](https://huggingface.co/ornith-ai/Ornith-1.0-35B)
- License: other
ChatCompletionChunk: {"id":"chatcmpl-81cc726ee0a82b1bf2703a65","object":"chat.completion.chunk","created":1786483208,"model":"ornith-moe:35b-a3b","system_fingerprint":"fp_7076fd14a68716c5","choices":[{"index":0,"delta":{"content":null},"finish_reason":"stop"}],"usage":{"prompt_tokens":11140,"completion_tokens":330,"total_tokens":11470,"active_kv_tokens":11470,"max_kv_token_capacity":32768,"kv_token_occupancy_rate_percentage":35.003662109375,"load_duration":1.082e-06,"prefill_duration_ttft":74.225606656,"decoding_duration":33.763416,"prefill_speed_tps":150.0829767768493,"decoding_speed_tps":9.773892546891583}}
## Original model card
See the upstream model card for training details, benchmarks, and upstream
usage. This repository only contains the Q4NX conversion for FastFlowLM.
- Upstream card: [ornith-ai/Ornith-1.0-35B](https://huggingface.co/ornith-ai/Ornith-1.0-35B) |