Spark-X2.5-4B GGUF

Community GGUF quantizations of XHToken/Spark-X2.5-4B.

☕ If you find this GGUF useful, please consider
buying me a coffee.
Your support helps cover the GPU costs of future releases.
Thank you for supporting this work.

About Spark-X2.5-4B

Spark-X2.5-4B is a compact, general-purpose language model developed by SparkLLM. According to the upstream authors, the model is designed for:

  • General-purpose capabilities: conversation, writing, translation, reasoning, and coding
  • Tool use and agentic workflows
  • Native context window of up to 1M tokens
  • More than 200 supported languages
  • Efficient hybrid attention: one full-attention layer combined with three sliding-window attention layers
  • Broad ecosystem support: llama.cpp, vLLM, SGLang, MLX, Ollama, and LM Studio

The upstream model is also designed for broad hardware compatibility and efficient long-context inference.

For the original model architecture, training details, benchmarks, and official usage instructions, see the official model card.

This repository is a quantization-only release for local inference. No model training or fine-tuning was performed.

Files

Quantization File size (GiB) A10M generation token/s Validation Recommendation / Notes
Q8_0 4.07 67.5 Load/generate pass High quality; reference quantization.
Q6_K 3.15 79.3 Load/generate pass High quality; near-reference quality.
Q5_K_M 2.77 87.6 Load/generate pass Daily use; strong quality/size balance.
Q4_K_M 2.42 92.9 Load/generate pass Recommended default for general use.
Q3_K_M 2.02 79.9 Load/generate pass Lower-memory profile; validate your workload.
Q2_K 1.66 95.9 Load/generate pass Aggressive low-memory profile.
IQ2_XS 1.39 106.8 Load/generate pass Experimental; validate carefully.
IQ1_M 1.20 39.9 Load/generate pass Very-low-bit option.
Q1_0 0.76 69.9 Load/generate pass Experimental / legacy minimum-memory option.

Q1/Q2 can lose instruction following, reasoning, and tool-call reliability. Select based on available memory and validate on the workload that matters to you.

The A10M generation figures were measured with single-stream llama-bench on an NVIDIA A10M.

License and attribution

The upstream model is released under Apache License 2.0. Preserve the upstream attribution and license when redistributing these derivative files. This is a community GGUF quantization, not an official XHToken/SparkLLM release or endorsement.

Checksums are available in SHA256SUMS.txt.

Downloads last month
1,633
GGUF
Model size
4B params
Architecture
spark2_5
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ngquocvinh/Spark-X2.5-4B-GGUF

Quantized
(19)
this model