DAE know how to corretly run DSpark with part of the model offloaded to RAM using llama.cpp?

#40
by LukeC110 - opened

My build is RTX 4090 and RTX 6000 Pro, 120GB VRAM + 32GB RAM. The model I am using is Q4 UD-IQ4-NL 137GB quant. I am currently using mtp draft model and I get around 30-40 tokens per second which is surprising to me. I thought that if I can swap mtp with DSpark then I might get another 10 tokens boosts, but the result was far different. Using the following setting the model loads, but pp and token generation is extremely slow. Think 1 - 2 tokens per second output.

llama-server
--no-warmup
--model /home/pk7677/.cache/huggingface/hub/models--unsloth--DeepSeek-V4-Flash-0731-GGUF/snapshots/fbbb5b93fb787c21338159b0af3318bb3f4d9768/UD-IQ4_NL/DeepSeek-V4-Flash-0731-UD-IQ4_NL-00001-of-00004.gguf
--model-draft /home/pk7677/.cache/huggingface/hub/models--unsloth--DeepSeek-V4-Flash-0731-GGUF/snapshots/fbbb5b93fb787c21338159b0af3318bb3f4d9768/dspark-DeepSeek-V4-Flash-0731-Q8_0.gguf
--spec-type draft-dspark
--spec-draft-n-max 3
--device-draft CUDA1
--host 0.0.0.0
--port 8000
--fit on --tensor-split 2,3
--n-cpu-moe 13
--main-gpu 1
--split-mode layer
--ctx-size 65536
--flash-attn on
--threads 16
--cont-batching
--temp 1.0 --top-p 0.95 --top-k 0 --min-p 0
--jinja
--batch-size 2048
--ubatch-size 2048
--alias 'DeepSeek-V4-Flash-0731'

Does anyone know how to set this up correctly?

adding dspark is putting another 11gigs of load to the system. So if you put that in vram as you should, 10g more is offloaded to system ram. system ram being so much slower is causing more slow down than dspark can help with. 1-2 tg speed means something is terribly wrong though, like cpu only speed.

FWIW I would drop the n cpu moe 13 and use the -fitc and -fitt params. You put the dspark on CUDA1, so list your devices in order and then set the -fitt large enough to leave space for the dspark model,
--devices CUDA0,CUDA1
-fitt 400,11000
-fitc 65536

adjust as needed if you get oom messages, I am not sure if rocm and cuda use exact same or not. For me I set fitt to 10800 to get the full precision dspark on a specific gpu.
I also did not specify split mode.

adding dspark is putting another 11gigs of load to the system. So if you put that in vram as you should, 10g more is offloaded to system ram. system ram being so much slower is causing more slow down than dspark can help with. 1-2 tg speed means something is terribly wrong though, like cpu only speed.

FWIW I would drop the n cpu moe 13 and use the -fitc and -fitt params. You put the dspark on CUDA1, so list your devices in order and then set the -fitt large enough to leave space for the dspark model,
--devices CUDA0,CUDA1
-fitt 400,11000
-fitc 65536

adjust as needed if you get oom messages, I am not sure if rocm and cuda use exact same or not. For me I set fitt to 10800 to get the full precision dspark on a specific gpu.
I also did not specify split mode.

Thank you. Didn't know about these flags, they help bring my pp up, and no more annoying manual tensor split.

Sign up or log in to comment