add Ktransformers cookbook to model card

#35
Files changed (1) hide show
  1. README.md +14 -0
README.md CHANGED
@@ -126,6 +126,20 @@ sglang serve \
126
  --swa-full-tokens-ratio 0.1 \
127
  ```
128
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
129
  ## How to Run Locally
130
 
131
  Please refer to the [inference](inference/README.md) folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.
 
126
  --swa-full-tokens-ratio 0.1 \
127
  ```
128
 
129
+ ## How to Run with KTransformers
130
+
131
+ KTransformers supports CPU–GPU heterogeneous inference by running native MXFP4 experts with KT-Kernel and using SGLang as the serving engine. After installing KTransformers and downloading the checkpoint, this example launches it on a single RTX 4090:
132
+
133
+ ```bash
134
+ kt run /path/to/DeepSeek-V4-Flash-0731 \
135
+ --gpu-experts 10 --cpu-threads 60 --numa-nodes 2 \
136
+ --tp 1 --kt-method MXFP4 \
137
+ --max-total-tokens 16384 --max-running-requests 2 \
138
+ --mem-fraction-static 0.85
139
+ ```
140
+
141
+ See the [KTransformers recipe](https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/DeepSeek-V4-Flash.md) for installation, validated hardware, and advanced launch options.
142
+
143
  ## How to Run Locally
144
 
145
  Please refer to the [inference](inference/README.md) folder for detailed instructions on running DeepSeek-V4 locally, including model weight conversion and interactive chat demos.