Please wait for official announcement before using!

#5
by danielhanchen - opened
Unsloth AI org

Hey guys, as said previously, the GGUFs are not finalized until we announce them, there is more issues with the the inference implementation than we originally anticipated. We are trying to see if can push PRs to llama.cpp to fix the issues.

It would be great if in the future you could create your own repository with an updating branch for testing.

Unsloth AI org

It would be great if in the future you could create your own repository with an updating branch for testing.

We can't as Hugging Face as upload limits for private storage. Next time we might use another account. Nevertheless the issue was not the GGUFs, but rather the backend implementation itself.

Thanks for the answer, by repository I specified about GitHub with actual branch. Now we need to compile several alternative patches for running with 8-bit KV settings, MOE offloading on multiple GPUs, etc. for each DS4 tests.
Regards.

Unsloth AI org
edited Jul 8

We found DeepSeek-V4 issues in llama.cpp that caused gibberish after the 2nd turn. The cause was broken prompt caching. To run correctly, please use the latest llama.cpp version.
We also improved the DeepSeek-V4 chat jinja template, and tested over 4000 conversations to be equivalent with the official baseline.

Guide: https://unsloth.ai/docs/models/deepseek-v4
GGUF: https://huggingface.co/unsloth/DeepSeek-V4-Flash-GGUF

You can run DeepSeek-V4-Flash with all our fixes and Thinking toggles via Unsloth Studio:
deepseek-v4-flash in unsloth studio

llama.cpp added DeepSeek V4 support in #24162 - we noticed that when using any GGUF from any provider, multi turn conversations would not function well when compared to DS4's Hugging Face baseline. llama.cpp uses --ctx-checkpoints N which allowed it to do prefix caching to save inference costs. Instead of re-processing every prompt again on the 2nd, nth ask, we can use KV caching. However we found DS4 needed --ctx-checkpoints 0 or else you will get gibberish. Please use the latest version of llama.cpp to get fixes.

Engine Score Calculation Tool selection Parallel Tools Multi Turn tools Nested tools
Official code 15/15 3 3 3 3 3
Any provider 4/15 1 2 0 0 1
After our fix 15/15 3 2 3 3 3

Thanks guys and feel free to support our Tweet, Linkedin post or Reddit post

CC: @LaikaFramework @bacchio @klwjack @s3nh @ParadigmComplex @nv-rush @AlanSilvaTech @CHHORVORN

shimmyshimmer changed discussion status to closed

Hey hey,

I am a bit curious and need advice. (My setup is 8 RTX3090s)
Before official llama.cpp received the fixes I ran your Q8 GGUF with following llama.cpp fork https://github.com/fairydreaming/llama.cpp.git

  • With that fork, I am able to let it run Q8 GGUF at 350k context with -b 2048 -ub256

  • Now I tried to run the same Q8 GGUF with the latest llama.cpp official release I get OOM with the same startup command. I have to lower the context to 100k to get it running.

This is my startup command I use with llama fork is as follows.

./build/bin/llama-server
-m /mnt/extra/models/deepseek4flash/DeepSeek-V4-Flash-UD-Q8_K_XL-00001-of-00005.gguf
--host 0.0.0.0
--port 8788
--alias MainLLM
-ngl 99
-fa on
--no-mmap
--jinja
--ctx-checkpoints 0
-c 250000
-b 2048
-ub 256
--cache-type-k q8_0
--cache-type-v q8_0
--parallel 1

maglat

Try just applying the latest commit "pull/25402/head:deepseek-v4-checkpointing-fix" from llama.cpp as a patch to https://github.com/fairydreaming/llama.cpp/tree/dsv4 to work with checkpoints.
There is no actual branch yet exist.

Sign up or log in to comment