WhiskyAKM commited on
Commit
6dca099
·
verified ·
1 Parent(s): 1ca3c18

Upload folder using huggingface_hub

Browse files
.gitattributes CHANGED
@@ -41,3 +41,6 @@ minicpm5-2b-Q5_K_S.gguf filter=lfs diff=lfs merge=lfs -text
41
  minicpm5-2b-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
42
  minicpm5-2b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
43
  minicpm5-2b-bf16.gguf filter=lfs diff=lfs merge=lfs -text
 
 
 
 
41
  minicpm5-2b-Q6_K.gguf filter=lfs diff=lfs merge=lfs -text
42
  minicpm5-2b-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
43
  minicpm5-2b-bf16.gguf filter=lfs diff=lfs merge=lfs -text
44
+ minicpm5-2b-dspark-Q4_K_M.gguf filter=lfs diff=lfs merge=lfs -text
45
+ minicpm5-2b-dspark-Q8_0.gguf filter=lfs diff=lfs merge=lfs -text
46
+ minicpm5-2b-dspark-bf16.gguf filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -49,6 +49,16 @@ MiniCPM5-2B is designed for local assistants, coding agents, tool-use workflows,
49
  | `minicpm5-2b-Q4_K_S.gguf` | Q4_K_S | 1.4 GB | Smaller, slight quality loss |
50
  | `minicpm5-2b-Q4_0.gguf` | Q4_0 | 1.4 GB | Legacy 4-bit, broad compatibility |
51
 
 
 
 
 
 
 
 
 
 
 
52
  ## Usage
53
 
54
  ### llama.cpp CLI
@@ -70,6 +80,17 @@ MiniCPM5-2B is designed for local assistants, coding agents, tool-use workflows,
70
 
71
  The GGUF also works with **Ollama** and **LM Studio**.
72
 
 
 
 
 
 
 
 
 
 
 
 
73
  ### Thinking Mode
74
 
75
  The model supports deep-thinking output. You can control it per request via the chat template, e.g. with an OpenAI-compatible API:
@@ -98,6 +119,7 @@ These GGUF files were created from the BF16 source model using `llama-quantize`
98
  ## Acknowledgements
99
 
100
  - Original model: [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B)
 
101
  - Quantization tool: [llama.cpp](https://github.com/ggml-org/llama.cpp)
102
 
103
  ## License
 
49
  | `minicpm5-2b-Q4_K_S.gguf` | Q4_K_S | 1.4 GB | Smaller, slight quality loss |
50
  | `minicpm5-2b-Q4_0.gguf` | Q4_0 | 1.4 GB | Legacy 4-bit, broad compatibility |
51
 
52
+ ### DSpark Draft Model (Speculative Decoding)
53
+
54
+ | File | Quantization | Size | Use Case |
55
+ | :------------------------------- | :----------- | :---- | :---------------------------------------------- |
56
+ | `minicpm5-2b-dspark-bf16.gguf` | BF16 | 623 MB | Draft model, full precision |
57
+ | `minicpm5-2b-dspark-Q8_0.gguf` | Q8_0 | 334 MB | Draft model, near-lossless |
58
+ | `minicpm5-2b-dspark-Q4_K_M.gguf` | Q4_K_M | 186 MB | Draft model, smallest footprint |
59
+
60
+ These are the [MiniCPM5-2B-DSpark](https://huggingface.co/openbmb/MiniCPM5-2B-DSpark) draft models, trained for speculative decoding with MiniCPM5-2B. They accelerate generation without changing the target model's outputs.
61
+
62
  ## Usage
63
 
64
  ### llama.cpp CLI
 
80
 
81
  The GGUF also works with **Ollama** and **LM Studio**.
82
 
83
+ ### Speculative Decoding (DSpark)
84
+
85
+ Pair the target model with a DSpark draft model to speed up inference. Draft model quality has minimal impact on output, so smaller quants (e.g. Q4_K_M) are usually fine:
86
+
87
+ ```bash
88
+ ./llama-server \
89
+ -m minicpm5-2b-Q4_K_M.gguf \
90
+ -md minicpm5-2b-dspark-Q4_K_M.gguf \
91
+ --host 0.0.0.0 --port 8080
92
+ ```
93
+
94
  ### Thinking Mode
95
 
96
  The model supports deep-thinking output. You can control it per request via the chat template, e.g. with an OpenAI-compatible API:
 
119
  ## Acknowledgements
120
 
121
  - Original model: [openbmb/MiniCPM5-2B](https://huggingface.co/openbmb/MiniCPM5-2B)
122
+ - Draft model: [openbmb/MiniCPM5-2B-DSpark](https://huggingface.co/openbmb/MiniCPM5-2B-DSpark)
123
  - Quantization tool: [llama.cpp](https://github.com/ggml-org/llama.cpp)
124
 
125
  ## License
minicpm5-2b-dspark-Q4_K_M.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5ad0b44d4facfc422ececd107992562e09ded4912e5c4432256d872c02b40e48
3
+ size 194103104
minicpm5-2b-dspark-Q8_0.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:4ed1f40c1b346a94a504451acbfd7a01a6ec9aa9a87b4a5704c052beee5bd732
3
+ size 349218624
minicpm5-2b-dspark-bf16.gguf ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8cf18d0d5838ec2c92dc330f7ecc8c16dcc508ec1eb5e7a515f3b816a9c0c139
3
+ size 652732224