Ling-3.0-tiny-GGUF / README.md
NANI-Nithin's picture
Upload README.md with huggingface_hub
dca7ccb verified
|
Raw
History Blame
11.2 kB
---
title: Ling-3.0-tiny-GGUF
library_name: llama.cpp
model_type: quantized
base_model: inclusionAI/Ling-3.0-tiny
datasets: []
tags: [gguf, quantized, text-generation, moe, bailing-hybrid, reasoning, thinking, agentic, llama.cpp]
---
# Ling-3.0-tiny GGUF (llama.cpp)
[![Model Type](https://img.shields.io/badge/Model%20Type-Quantized%20GGUF-3B82F6?style=flat-square)](https://github.com/ggerganov/llama.cpp)
[![Backend](https://img.shields.io/badge/Backend-llama.cpp-FF7800?style=flat-square&logo=cpp)](https://github.com/ggerganov/llama.cpp)
[![Size](https://img.shields.io/badge/Total%20Size-139.31%20GB-FA8532?style=flat-square)](https://huggingface.co/NANI-Nithin/Ling-3.0-tiny-GGUF)
Quantized GGUF files for **Ling-3.0-tiny**, IBM's lightweight hybrid reasoning MoE model optimized for deployment with [llama.cpp](https://github.com/ggerganov/llama.cpp). This model delivers strong reasoning and agentic capabilities at low inference cost through an efficient hybrid architecture combining KDA and MLA attention with a sparse MoE FFN.
## Model Overview
**Ling-3.0-tiny** is a 7.9B parameter model with only 1.3B activated parameters per token, designed for efficient local and edge deployment. It features:
- **Efficient Hybrid-Linear Architecture:** 3:1 alternating stacking of KDA and MLA (3 KDA layers followed by 1 MLA layer per 4-layer block) with a sparse MoE FFN comprising 128 routed experts
- **Native Hybrid Reasoning and Agentic Capabilities:** Supports both fast responses and multi-step reasoning through configurable thinking mode
- **Local and Edge Deployment:** Validated on NVIDIA DGX Spark, Apple Silicon MacBook, and Mac mini for capable reasoning without datacenter-class GPUs
### Key Capabilities
- **Parameter-Efficient MoE:** Only 1.3B of 7.9B parameters activated per token for balanced performance and efficiency
- **Hybrid Attention:** Combines KDA (Kimi Delta Attention) and MLA (Multi-Head Latent Attention) for efficient long-context processing
- **Fast Inference:** Reaches ~100-105 tokens/s on DGX Spark and 86-90 tokens/s on M4 Pro MacBook with FP8
- **Memory Efficient:** ~8.34 GB peak memory usage at 8K context length
- **Thinking Mode:** Native chain-of-thought reasoning with per-request configurability
## Available Quantizations
| File | Size | Quality | Recommended Use |
|------|------|---------|-----------------|
| **Ling-3.0-tiny-BF16.gguf** | 14.72 GB | Full precision source. Every quant below is cut from this file. | Original model, maximum quality |
| **Ling-3.0-tiny-F16.gguf** | 14.72 GB | Full precision source. | Alternative full precision |
| **Ling-3.0-tiny-Q8_0.gguf** | 7.83 GB | Effectively lossless. Use when disk and RAM are not the constraint. | Highest quality, less compression |
| **Ling-3.0-tiny-Q6_K.gguf** | 6.05 GB | Near-lossless; the last stop before quality becomes measurable. | Balanced quality/size |
| **Ling-3.0-tiny-Q5_K_M.gguf** | 5.25 GB | Very good quality, noticeably smaller than Q6_K. | Good trade-off |
| **Ling-3.0-tiny-Q5_K_S.gguf** | 5.11 GB | Slightly smaller than Q5_K_M for a slight quality cost. | Smaller footprint |
| **Ling-3.0-tiny-Q5_1.gguf** | 5.55 GB | Legacy. Prefer Q5_K_M. | Historical compatibility |
| **Ling-3.0-tiny-Q5_0.gguf** | 5.11 GB | Legacy. Prefer Q5_K_M. | Historical compatibility |
| **Ling-3.0-tiny-Q4_K_M.gguf** | 4.49 GB | The usual default. Best quality-per-byte for most people. | Default choice |
| **Ling-3.0-tiny-Q4_K_S.gguf** | 4.24 GB | A little smaller than Q4_K_M, a little worse. | Smaller footprint |
| **Ling-3.0-tiny-IQ4_NL.gguf** | 4.22 GB | Non-linear 4-bit; good on hardware without fast K-quant kernels. | Specialized hardware |
| **Ling-3.0-tiny-IQ4_XS.gguf** | 3.99 GB | Best sub-4.5bpw option; usually beats Q4_K_S at a smaller size. | Size-critical |
| **Ling-3.0-tiny-Q4_1.gguf** | 4.66 GB | Legacy. Prefer Q4_K_M. | Historical compatibility |
| **Ling-3.0-tiny-Q4_0.gguf** | 4.22 GB | Legacy round-to-nearest. Prefer Q4_K_M unless a runtime needs this. | Historical compatibility |
| **Ling-3.0-tiny-MXFP4_MOE.gguf** | 4.39 GB | MoE-only 4-bit microscaling format for the expert tensors. | Specialized MoE deployment |
| **Ling-3.0-tiny-Q3_K_L.gguf** | 3.86 GB | Small, with real quality loss. Usable when RAM is tight. | Tight RAM constraints |
| **Ling-3.0-tiny-Q3_K_M.gguf** | 3.58 GB | Smaller again; noticeable degradation. | Memory-constrained |
| **Ling-3.0-tiny-IQ3_M.gguf** | 3.31 GB | Strong at ~3.7bpw, clearly better than Q3_K_M. | Quality-conscious sizing |
| **Ling-3.0-tiny-IQ3_S.gguf** | 3.27 GB | Slightly smaller than IQ3_M. | Compact version |
| **Ling-3.0-tiny-Q3_K_S.gguf** | 3.27 GB | Aggressive. Prefer IQ3_M at a similar size. | Maximum compression |
| **Ling-3.0-tiny-IQ3_XS.gguf** | 3.11 GB | Aggressive but coherent. | Extreme compression |
| **Ling-3.0-tiny-IQ3_XXS.gguf** | 2.91 GB | Very aggressive; the last coherent step down. | Extreme compression |
| **Ling-3.0-tiny-Q2_K.gguf** | 2.78 GB | Very small, heavily degraded. For experimentation. | Experimental only |
| **Ling-3.0-tiny-IQ2_M.gguf** | 2.52 GB | The smallest size most people find usable. | Memory-constrained |
| **Ling-3.0-tiny-Q2_K_S.gguf** | 2.59 GB | Smaller than Q2_K, at a further quality cost. | Even smaller |
| **Ling-3.0-tiny-IQ2_S.gguf** | 2.31 GB | Below the usual usability line. | Extreme compression |
| **Ling-3.0-tiny-IQ2_XS.gguf** | 2.27 GB | Experimental. | Experimental only |
| **Ling-3.0-tiny-IQ2_XXS.gguf** | 2.06 GB | Experimental. | Experimental only |
| **Ling-3.0-tiny-Q2_0.gguf** | 2.28 GB | Extreme, group-64. Included for completeness. | Historical compatibility |
| **Ling-3.0-tiny-IQ1_M.gguf** | 1.80 GB | Extreme. Expect substantial degradation. | Maximum compression |
| **Ling-3.0-tiny-IQ1_S.gguf** | 1.64 GB | Extreme. Expect substantial degradation. | Maximum compression |
| **Ling-3.0-tiny-Q1_0.gguf** | 1.21 GB | Extreme. Included for completeness. | Historical compatibility |
*All files are cut from the BF16 source (14.72 GB). Each quant is independently uploaded and deleted immediately after successful upload to minimize peak disk usage.*
## Model Architecture
Ling-3.0-tiny features a unique hybrid architecture combining KDA/MLA attention with a sparse MoE FFN:
### Attention Mechanism
- **KDA (Kimi Delta Attention):** 3 layers per 4-layer block
- **MLA (Multi-Head Latent Attention):** 1 layer per 4-layer block
- **Hybrid Stacking:** 3:1 ratio for efficient long-context processing
### Feed-Forward Network
- **MoE (MultiplE Experts):** 128 total experts
- **Routed Experts:** 8 activated per token
- **Shared Expert:** 1 additional expert
- **Activation Efficiency:** Only 1.3B of 7.9B parameters activated per token
### Core Components
- **Layers:** 24 total
- **Hidden Size:** 1536
- **Vocab Size:** 157184
- **Position Embedding:** Rotary Position Embedding (RoPE)
- **Precision:** bfloat16 (source)
## Inference
### Generation Parameters
> **Important:** Use `temperature=1.0` and `top_p=0.95` across **all tasks and serving backends**, including general chat, reasoning, and tool calling.
| Parameter | Value | Notes |
|:----------|:------|:------|
| `temperature` | `1.0` | Required for all modes |
| `top_p` | `0.95` | Nucleus sampling threshold |
| `top_k` | `20` | Recommended for stable generation |
| `max_new_tokens` | `8192` | Thinking mode (increase for complex reasoning) |
| `max_new_tokens` | `2048` | Non-thinking mode |
| `do_sample` | `True` | Required when temperature > 0 |
### Thinking Modes
| Mode | Template Parameters | Behavior |
|:-----|:-------------------|:---------|
| **Thinking** (default) | `enable_thinking=True` | Full chain-of-thought reasoning inside `<think>...</think>` |
| **Non-thinking** | `enable_thinking=False` | Direct answer with no reasoning overhead |
| **Low-effort** | `enable_thinking=True, low_effort=True` | Brief reasoning for simpler queries |
## Serving with llama.cpp
### Basic Usage
```bash
# Download a quantization (recommended: Q4_K_M)
huggingface-cli download NANI-Nithin/Ling-3.0-tiny-GGUF Ling-3.0-tiny-Q4_K_M.gguf --local-dir .
# Run inference
llama-cli -m Ling-3.0-tiny-Q4_K_M.gguf -p "Hello, how are you?"
```
Or use the Hugging Face Hub integration:
```bash
# Serve with llama-server
llama-server -hf NANI-Nithin/Ling-3.0-tiny-GGUF:Q4_K_M
# Run with model path
llama-cli -hf NANI-Nithin/Ling-3.0-tiny-GGUF:Q4_K_M -p "Hello"
```
### API Usage
```bash
# Start the server
curl -s http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "granite-4.2-3b-Q4_K_M",
"messages": [{"role": "user", "content": "Explain quantum computing in simple terms"}],
"temperature": 1.0,
"top_p": 0.95,
"max_tokens": 8192
}'
```
## Quick Start Example
```bash
# Download the model
huggingface-cli download NANI-Nithin/Ling-3.0-tiny-GGUF Ling-3.0-tiny-Q4_K_M.gguf
# Run inference
./llama-cli -m Ling-3.0-tiny-Q4_K_M.gguf -p "What is the Riemann hypothesis?"
```
## Technical Details
- **Source Model:** inclusionAI/Ling-3.0-tiny
- **Author:** inclusionAI
- **License:** MIT
- **Architecture:** BailingMoeV3ForCausalLM (hybrid KDA/MLA + sparse MoE)
- **Parameters:** 7.9B total, 1.3B activated per token
- **Experts:** 128 total (8 routed + 1 shared per token)
- **Context Length:** 131K tokens (natively supports 128K)
- **Created:** August 10, 2026
- **Languages:** Multiple languages supported
## Usage Notes
1. **Disk Space:** Download one quantization at a time. Each file ranges from 1.21 GB (Q1_0) to 14.72 GB (BF16 source).
2. **Memory Requirements:** Varies by quantization. Q4_K_M requires ~4 GB VRAM, Q8_0 requires ~8 GB VRAM.
3. **MoE Optimization:** The sparse MoE architecture provides efficient inference while maintaining broad capabilities.
4. **Thinking Mode:** Enable `enable_thinking=True` to get chain-of-thought reasoning. Set to `False` for faster, direct answers.
## Performance Characteristics
Ling-3.0-tiny achieves impressive efficiency:
- **FP8 Performance:** ~100-105 tokens/s on DGX Spark, 86-90 tokens/s on M4 Pro MacBook
- **Memory Usage:** ~8.34 GB peak at 8K context length
- **Agentic Performance:** Score of 25 on Artificial Analysis Intelligence Index v4.1.1
- **End-to-End Latency:** ~18 seconds for 500-token response including reasoning
## Model Card Information
This GGUF repo contains quantized versions of [Ling-3.0-tiny](https://huggingface.co/inclusionAI/Ling-3.0-tiny), featuring a unique hybrid architecture combining KDA/MLA attention with sparse MoE for efficient reasoning and agentic capabilities. The quantization was performed using llama.cpp's quantization pipeline, preserving the model's MoE efficiency while reducing size for deployment.
For the full source model documentation, including detailed training methodology, evaluation benchmarks, and advanced deployment recipes (SGLang, vLLM, Ollama), refer to the source repo: [inclusionAI/Ling-3.0-tiny](https://huggingface.co/inclusionAI/Ling-3.0-tiny).
---
**Note:** These files are optimized for llama.cpp and are not compatible with vLLM, SGLang, or the Transformers library in their current format. For those frameworks, use the source model `inclusionAI/Ling-3.0-tiny`.