--- base_model: - inclusionAI/Ling-3.0-flash tags: - mlx - omlx - llm - oq - quantized - mtp - 4-bit library_name: mlx license: mit --- # Ling-3.0-flash-oQ4e-mtp This model was quantized using [oQ](https://github.com/jundot/omlx/releases#release-v0.5.8.dev1) (oMLX v0.5.8.dev1) mixed-precision quantization. MTP included but currently is not supported by oMLX. Original model: [Ling-3.0-flash-fp8](https://huggingface.co/inclusionAI/Ling-3.0-flash-fp8) This model for M3, M4, M5. **For Mac M1, M2 try** [this](https://huggingface.co/d1sl1ke/Ling-3.0-flash-oQ4e-fp16-mtp). ## Benchmarks Mac Studio M4 Max 40c, 128 GB | Parameter | Value | |----------|----------| | **Model** | `Ling-3.0-flash-oQ4e-mtp` (70.0 GB VRAM) | | **Benchmark Context** | Code (Mixed) | | **Generation Length** | 128 tokens (fixed) | ### Metrics (Single Request) | Test | TTFT (ms) | TPOT (ms/tok) | pp TPS | tg TPS | E2E Latency | Throughput | Peak Mem | |------|-----------|---------------|--------|--------|-------------|------------|----------| | **pp8192/tg128** | 8,874.6 | 18.14 | 923.1 tok/s | 55.6 tok/s | 11.182 s | 744.1 tok/s | 70.18 GB | | **pp32768/tg128** | 40,822.9 | 19.77 | 802.7 tok/s | 51.0 tok/s | 43.340 s | 759.0 tok/s | 74.59 GB | | **pp65536/tg128** | 99,803.4 | 22.06 | 656.7 tok/s | 45.7 tok/s | 102.614 s | 639.9 tok/s | 80.54 GB | ## Quantization details - **Model type**: bailing_hybrid - **Bits**: 4 - **Group size**: 64 - **Format**: MLX safetensors