Instructions to use unsloth/DeepSeek-V4-Flash-0731-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: llama cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./llama-cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL # Run inference directly in the terminal: ./build/bin/llama-cli -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Use Docker
docker model run hf.co/unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
- LM Studio
- Jan
- Ollama
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Ollama:
ollama run hf.co/unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
- Unsloth Desktop
- Pi
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Docker Model Runner:
docker model run hf.co/unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
- Lemonade
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Run and chat with the model
lemonade run user.DeepSeek-V4-Flash-0731-GGUF-UD-Q4_K_XL
List all available models
lemonade list
- Hermes Agent
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use unsloth/DeepSeek-V4-Flash-0731-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "unsloth/DeepSeek-V4-Flash-0731-GGUF:UD-Q4_K_XL" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
DAE know how to corretly run DSpark with part of the model offloaded to RAM using llama.cpp?
My build is RTX 4090 and RTX 6000 Pro, 120GB VRAM + 32GB RAM. The model I am using is Q4 UD-IQ4-NL 137GB quant. I am currently using mtp draft model and I get around 30-40 tokens per second which is surprising to me. I thought that if I can swap mtp with DSpark then I might get another 10 tokens boosts, but the result was far different. Using the following setting the model loads, but pp and token generation is extremely slow. Think 1 - 2 tokens per second output.
llama-server
--no-warmup
--model /home/pk7677/.cache/huggingface/hub/models--unsloth--DeepSeek-V4-Flash-0731-GGUF/snapshots/fbbb5b93fb787c21338159b0af3318bb3f4d9768/UD-IQ4_NL/DeepSeek-V4-Flash-0731-UD-IQ4_NL-00001-of-00004.gguf
--model-draft /home/pk7677/.cache/huggingface/hub/models--unsloth--DeepSeek-V4-Flash-0731-GGUF/snapshots/fbbb5b93fb787c21338159b0af3318bb3f4d9768/dspark-DeepSeek-V4-Flash-0731-Q8_0.gguf
--spec-type draft-dspark
--spec-draft-n-max 3
--device-draft CUDA1
--host 0.0.0.0
--port 8000
--fit on --tensor-split 2,3
--n-cpu-moe 13
--main-gpu 1
--split-mode layer
--ctx-size 65536
--flash-attn on
--threads 16
--cont-batching
--temp 1.0 --top-p 0.95 --top-k 0 --min-p 0
--jinja
--batch-size 2048
--ubatch-size 2048
--alias 'DeepSeek-V4-Flash-0731'
Does anyone know how to set this up correctly?
adding dspark is putting another 11gigs of load to the system. So if you put that in vram as you should, 10g more is offloaded to system ram. system ram being so much slower is causing more slow down than dspark can help with. 1-2 tg speed means something is terribly wrong though, like cpu only speed.
FWIW I would drop the n cpu moe 13 and use the -fitc and -fitt params. You put the dspark on CUDA1, so list your devices in order and then set the -fitt large enough to leave space for the dspark model,
--devices CUDA0,CUDA1
-fitt 400,11000
-fitc 65536
adjust as needed if you get oom messages, I am not sure if rocm and cuda use exact same or not. For me I set fitt to 10800 to get the full precision dspark on a specific gpu.
I also did not specify split mode.
adding dspark is putting another 11gigs of load to the system. So if you put that in vram as you should, 10g more is offloaded to system ram. system ram being so much slower is causing more slow down than dspark can help with. 1-2 tg speed means something is terribly wrong though, like cpu only speed.
FWIW I would drop the n cpu moe 13 and use the -fitc and -fitt params. You put the dspark on CUDA1, so list your devices in order and then set the -fitt large enough to leave space for the dspark model,
--devices CUDA0,CUDA1
-fitt 400,11000
-fitc 65536adjust as needed if you get oom messages, I am not sure if rocm and cuda use exact same or not. For me I set fitt to 10800 to get the full precision dspark on a specific gpu.
I also did not specify split mode.
Thank you. Didn't know about these flags, they help bring my pp up, and no more annoying manual tensor split.
There are smaller dflash quants that may be worth trying