tugot17's picture
Rename GGUF files: remove Draft-v1 from filenames (#1)
9235d67
|
Raw
History Blame Contribute Delete
4.35 kB
---
library_name: llama.cpp
base_model: LiquidAI/LFM2.5-1.2B-Instruct-DSpark
license: other
license_name: lfm1.0
license_link: LICENSE
pipeline_tag: text-generation
tags:
- speculative-decoding
- dspark
- lfm2
- draft-model
- gguf
---
<div align="center">
<img
src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png"
alt="Liquid AI"
style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;"
/>
<div style="display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;">
<a href="https://playground.liquid.ai/"><strong>Try LFM</strong></a> β€’
<a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> β€’
<a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> β€’
<a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a>
</div>
</div>
# LFM2.5-1.2B-Instruct-DSpark-GGUF
GGUF build of [`LiquidAI/LFM2.5-1.2B-Instruct-DSpark`](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-DSpark) for **llama.cpp** (DSpark speculative decoding is in mainline, ggml-org/llama.cpp #25173).
This is a standalone **draft sidecar**: it carries only the drafter (5 attention layers, rank-256 Markov head, confidence head, block size 9). Token embeddings and the LM head are shared from the target model at load time, so it must be paired with a [LFM2.5-1.2B-Instruct-GGUF](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF) target file.
Find more information about LFM2.5-DSpark in our [blog post](https://www.liquid.ai/blog/lfm2.5-dspark).
## πŸ“¦ Files
| file | quant | size | notes |
|---|---|---:|---|
| `LFM2.5-1.2B-Instruct-DSpark-F16.gguf` | F16 | 594 MB | best accept length, recommended when memory allows |
| `LFM2.5-1.2B-Instruct-DSpark-Q8_0.gguf` | Q8_0 | 315 MB | accept length βˆ’2% vs F16 |
| `LFM2.5-1.2B-Instruct-DSpark-Q4_K_M.gguf` | Q4_K_M | 174 MB | accept length βˆ’3% vs F16, smallest recommended β€” sub-4-bit draft quants measurably hurt both accept length and throughput |
Draft quantization changes speed only marginally (the drafter is a small share of each cycle); choose by memory budget. The target model quant is the main speed/quality lever and is independent of this file.
## πŸƒ How to run (llama.cpp)
```bash
llama-server -m LFM2.5-1.2B-Instruct-F16.gguf \
-md LFM2.5-1.2B-Instruct-DSpark-F16.gguf \
--spec-type draft-dspark --spec-draft-n-max 10 --spec-draft-n-min 0 \
-fa on -ngl 99
```
The block size is read from the sidecar metadata (n-max is clamped to it). Speculative decoding is **exact**: the target verifies every proposed token, so greedy output equals the target alone; per-response `timings` report `draft_n` / `draft_n_accepted`.
Other models in the LFM2.5-DSpark GGUF family:
| Draft (GGUF) | Target (GGUF) |
|---|---|
| [LFM2.5-1.2B-Instruct-DSpark-GGUF](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-DSpark-GGUF) | [LFM2.5-1.2B-Instruct-GGUF](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF) |
| [LFM2.5-2.6B-DSpark-GGUF](https://huggingface.co/LiquidAI/LFM2.5-2.6B-DSpark-GGUF) | [LFM2.5-2.6B-GGUF](https://huggingface.co/LiquidAI/LFM2.5-2.6B-GGUF) |
| [LFM2.5-8B-A1B-DSpark-GGUF](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-DSpark-GGUF) | [LFM2.5-8B-A1B-GGUF](https://huggingface.co/LiquidAI/LFM2.5-8B-A1B-GGUF) |
## πŸ“Š Acceptance and benchmarks
See [`LiquidAI/LFM2.5-1.2B-Instruct-DSpark`](https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct-DSpark) for acceptance-length tables (H100 and Apple silicon) and target benchmarks.
## πŸ“¬ Contact
- Got questions or want to connect? [Join our Discord community](https://discord.com/invite/liquid-ai)
- If you are interested in custom solutions with edge deployment, please contact [our sales team](https://www.liquid.ai/contact).
## Citation
```bibtex
@article{liquidAI202626B,
author = {Liquid AI},
title = {LFM2.5-2.6B: Agents Everywhere},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2-5-2-6b},
}
```
```bibtex
@article{liquidAI2026dspark,
author = {Liquid AI},
title = {LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook},
journal = {Liquid AI Blog},
year = {2026},
note = {www.liquid.ai/blog/lfm2.5-dspark},
}
```