File size: 4,278 Bytes
63b94e5
3adb20c
c3dd288
 
63b94e5
c3dd288
 
 
 
 
 
 
58dc802
c3dd288
 
a9dbc81
 
63b94e5
268f632
63b94e5
c3dd288
63b94e5
c3dd288
63b94e5
c3dd288
63b94e5
c3dd288
63b94e5
c3dd288
63b94e5
c3dd288
 
 
 
 
63b94e5
c3dd288
63b94e5
c3dd288
 
 
 
 
 
3adb20c
8014fa2
63b94e5
c3dd288
63b94e5
8f23fd6
 
c3dd288
63b94e5
8f23fd6
 
63b94e5
8f23fd6
3adb20c
8f23fd6
 
f8a3527
63b94e5
3adb20c
 
 
 
 
 
 
 
 
 
 
c3dd288
63b94e5
c3dd288
63b94e5
c3dd288
63b94e5
3c64222
c3dd288
3adb20c
c3dd288
 
 
8014fa2
c3dd288
3adb20c
c3dd288
 
63b94e5
cac5b5c
 
c3dd288
63b94e5
c3dd288
 
3adb20c
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
---
license: mit
language:
- en
pipeline_tag: text-generation
tags:
- fastflowlm
- q4nx
- npu
- qwen3.6-moe
- 35b-a3b
- ornith
- quantization
base_model:
- ornith-ai/Ornith-1.0-35B
base_model_relation: quantized
quantized_by: Atomic-Germ
---
# *IF YOU USE COMMUNITY QWEN MODELS DO NOT UPGRADE TO FLM v1.0.2+*

# Ornith-1.0-35B - Q4NX for FastFlowLM (AMD Ryzen AI XDNA2)

Ornith-1.0-35B is converted to Q4NX for hardware-accelerated inference with FastFlowLM on AMD Ryzen AI NPUs.

## What is Q4NX?

Q4NX is FastFlowLM's native packed-quantization format - a rearranged Q4_1 layout tuned for the NPU matrix engine's tile sizes and memory access patterns. It is **not** a GGUF file and it does not run on llama.cpp or Ollama; it is meant exclusively for the [FastFlowLM](https://fastflowlm.com) engine on AMD Ryzen AI NPUs.

## Requirements

- FastFlowLM >= 0.9.46 (`flm` CLI)
- AMD Ryzen AI processor with **XDNA2 (NPU2)** - Strix Point / Ryzen AI 300
  series or later
- Linux with the XRT NPU stack installed
- ~47 GB of unified system memory (Q4NX weights + activations + KV cache)

## Files

| File | Purpose |
|---|---|
| model.q4nx | Quantized Q4NX text weights |
| config.json | FastFlowLM model configuration |
| tokenizer.json | Tokenizer |
| tokenizer_config.json | Special tokens and chat template |
| chat_template.jinja | Chat template |
| README.md | |

## Install and run

This repository works with `flm-add`, a small installer that copies the model
into the FastFlowLM user directory and registers the tag. It never
modifies the system FastFlowLM install.

`pip install flm-add` or `uv tool install flm-add`

```bash
uv tool install flm-add
unset FLM_CONFIG_PATH FLM_XCLBIN_PATH && flm-add Atomic-Germ/Ornith-1.0-35B-A3B-NPU2 --tag ornith1.0-moe:35b-a3b --family qwen3.6-moe
FLM_CONFIG_PATH="$HOME/.config/flm/model_list.json" FLM_XCLBIN_PATH="$HOME/.config/flm" flm run ornith1.0-moe:35b-a3b

```

| Context Length |              TTFT (s) |      Prefill Speed (tok/s) |     Decoding Speed (tok/s)
|---|---|---|---|
|             1k |      12.355 ±   0.122 |          79.07 ±      0.78 |       10.41 ±      0.01
|             2k |      16.849 ±   0.142 |         115.41 ±      0.97 |       10.36 ±      0.07
|             4k |      24.852 ±   0.035 |         156.09 ±      0.22 |       10.14 ±      0.10
|             8k |      40.241 ±   0.385 |         192.55 ±      1.84 |        9.74 ±      0.07
|            16k |      73.398 ±   1.227 |         211.03 ±      3.53 |        9.02 ±      0.07
|            32k |     148.780 ±   0.630 |         208.11 ±      0.89 |        7.70 ±      0.03

---

## Kernels

FastFlowLM's NPU kernels (xclbins) are closed source and are not shipped in this repository. This model uses the **`qwen3.6-moe`** engine family and is shape-identical to the official **`qwen3.6-moe:35b-a3b`** model (`Qwen3.6-Moe-35BA3B-NPU2`). Point the runtime's xclbin path at the matching `xclbins` directory (or ship your own) before running.

## Model

- Registry tag: `ornith1.0-moe:35b-a3b`
- Engine family: `qwen3.6-moe`
- Kernel source: official fastflowlm `qwen3.6-moe:35b-a3b`
- Context length: 262,144 tokens (from config)
- Hidden size: 2048
- Layers: 40
- Intermediate size: 512
- Vocabulary: 248320
- `model.q4nx` size: 23.2 GB
- Base model: [ornith-ai/Ornith-1.0-35B](https://huggingface.co/ornith-ai/Ornith-1.0-35B)
- License: other

ChatCompletionChunk: {"id":"chatcmpl-81cc726ee0a82b1bf2703a65","object":"chat.completion.chunk","created":1786483208,"model":"ornith-moe:35b-a3b","system_fingerprint":"fp_7076fd14a68716c5","choices":[{"index":0,"delta":{"content":null},"finish_reason":"stop"}],"usage":{"prompt_tokens":11140,"completion_tokens":330,"total_tokens":11470,"active_kv_tokens":11470,"max_kv_token_capacity":32768,"kv_token_occupancy_rate_percentage":35.003662109375,"load_duration":1.082e-06,"prefill_duration_ttft":74.225606656,"decoding_duration":33.763416,"prefill_speed_tps":150.0829767768493,"decoding_speed_tps":9.773892546891583}}

## Original model card

See the upstream model card for training details, benchmarks, and upstream
usage. This repository only contains the Q4NX conversion for FastFlowLM.
- Upstream card: [ornith-ai/Ornith-1.0-35B](https://huggingface.co/ornith-ai/Ornith-1.0-35B)