Spark-X2.5-1.7B GGUF

Community GGUF quantizations of XHToken/Spark-X2.5-1.7B. This repository contains nine quantized files for local inference. No training or fine-tuning was performed.

☕ If you find this GGUF useful, please consider
buying me a coffee.
Your support helps cover the GPU costs of future releases.
Thank you for supporting this work.

Files

Quantization File size (GiB) A10M generation token/s Validation Recommendation / Notes
Q8_0 1.70 175.54 Load/generate pass High quality.
Q6_K 1.31 199.07 Load/generate pass High quality.
Q5_K_M 1.17 223.04 Load/generate pass Daily use.
Q4_K_M 1.03 241.87 Load/generate pass Recommended default.
Q3_K_M 0.87 203.60 Load/generate pass Lower-memory profile.
Q2_K 0.74 231.70 Load/generate pass Aggressive low-memory profile.
IQ2_XS 0.61 245.03 Load/generate pass Experimental.
IQ1_M 0.54 252.02 Load/generate pass Experimental.
Q1_0 0.40 336.65 Load/generate pass Experimental / legacy minimum-memory option.

The A10M generation figures were measured with single-stream llama-bench on an NVIDIA A10M.

Q1/Q2 and the IQ variants can lose instruction following, reasoning, and tool-call reliability. Validate the chosen file on the workload that matters to you.

License and attribution

The upstream model is released under Apache License 2.0. Preserve the upstream attribution and license when redistributing these derivative files. This is a community GGUF quantization, not an official XHToken/SparkLLM release or endorsement.

Checksums are available in SHA256SUMS.txt.

Downloads last month
1,169
GGUF
Model size
2B params
Architecture
spark2_5
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ngquocvinh/Spark-X2.5-1.7B-GGUF

Quantized
(9)
this model