--- library_name: llama.cpp base_model: LiquidAI/LFM2.5-1.2B-Instruct-DSpark license: other license_name: lfm1.0 license_link: LICENSE pipeline_tag: text-generation tags: - speculative-decoding - dspark - lfm2 - draft-model - gguf ---
Liquid AI
Try LFMDocsLEAPDiscord
# LFM2.5-1.2B-Instruct-DSpark-GGUF GGUF build of [`LiquidAI/LFM2.5-1.2B-Instruct-DSpark`](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-DSpark) for **llama.cpp** (DSpark speculative decoding is in mainline, ggml-org/llama.cpp #25173). This is a standalone **draft sidecar**: it carries only the drafter (5 attention layers, rank-256 Markov head, confidence head, block size 9). Token embeddings and the LM head are shared from the target model at load time, so it must be paired with a [LFM2.5-1.2B-Instruct-GGUF](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF) target file. Find more information about LFM2.5-DSpark in our [blog post](https://www.liquid.ai/blog/lfm2.5-dspark). ## 📦 Files | file | quant | size | notes | |---|---|---:|---| | `LFM2.5-1.2B-Instruct-DSpark-F16.gguf` | F16 | 594 MB | best accept length, recommended when memory allows | | `LFM2.5-1.2B-Instruct-DSpark-Q8_0.gguf` | Q8_0 | 315 MB | accept length −2% vs F16 | | `LFM2.5-1.2B-Instruct-DSpark-Q4_K_M.gguf` | Q4_K_M | 174 MB | accept length −3% vs F16, smallest recommended — sub-4-bit draft quants measurably hurt both accept length and throughput | Draft quantization changes speed only marginally (the drafter is a small share of each cycle); choose by memory budget. The target model quant is the main speed/quality lever and is independent of this file. ## 🏃 How to run (llama.cpp) ```bash llama-server -m LFM2.5-1.2B-Instruct-F16.gguf \ -md LFM2.5-1.2B-Instruct-DSpark-F16.gguf \ --spec-type draft-dspark --spec-draft-n-max 10 --spec-draft-n-min 0 \ -fa on -ngl 99 ``` The block size is read from the sidecar metadata (n-max is clamped to it). Speculative decoding is **exact**: the target verifies every proposed token, so greedy output equals the target alone; per-response `timings` report `draft_n` / `draft_n_accepted`. Other models in the LFM2.5-DSpark GGUF family: | Draft (GGUF) | Target (GGUF) | |---|---| | [LFM2.5-1.2B-Instruct-DSpark-GGUF](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-DSpark-GGUF) | [LFM2.5-1.2B-Instruct-GGUF](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF) | | [LFM2.5-2.6B-DSpark-GGUF](https://huggingface.co/LiquidAI/LFM2.5-2.6B-DSpark-GGUF) | [LFM2.5-2.6B-GGUF](https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF) | | [LFM2.5-8B-A1B-DSpark-GGUF](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-DSpark-GGUF) | [LFM2.5-8B-A1B-GGUF](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-GGUF) | ## 📊 Acceptance and benchmarks See [`LiquidAI/LFM2.5-1.2B-Instruct-DSpark`](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-DSpark) for acceptance-length tables (H100 and Apple silicon) and target benchmarks. ## 📬 Contact - Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai) - If you are interested in custom solutions with edge deployment, please contact [our sales team](https://www.liquid.ai/contact). ## Citation ```bibtex @article{liquidAI202626B, author = {Liquid AI}, title = {LFM2.5-2.6B: Agents Everywhere}, journal = {Liquid AI Blog}, year = {2026}, note = {www.liquid.ai/blog/lfm2-5-2-6b}, } ``` ```bibtex @article{liquidAI2026dspark, author = {Liquid AI}, title = {LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook}, journal = {Liquid AI Blog}, year = {2026}, note = {www.liquid.ai/blog/lfm2.5-dspark}, } ```