DeepSeek-V4-Flash-0731 GGUF โ€” Quantized by BatiAI

BatiFlow bati.cpp deepseek MIT

The 2026-07-31 refresh of DeepSeek-V4-Flash โ€” quantized from the official weights, verified in Korean. 284B-A13B MoE with CSA+HCA hybrid attention. Runs on a single high-RAM Mac.

๐Ÿ“ฆ Quantizations

Quant Size Shards Best for
Q3_K_M 135 GB 4 M4 Max 128GB (tight) โ†’ 192GB Mac Studio
Q4_K_M 172 GB 4 M2 Ultra 192GB / server

Built directly from official DeepSeek weights via a Q8_0 intermediate, BatiAIโ€‘signed (general.author: BatiAI). Q5+ intentionally omitted โ€” the size/benefit tradeoff doesn't land on any real Mac tier.

๐Ÿ†• What changed in 0731. Same architecture as the original V4โ€‘Flash (43 layers, 256 experts, vocab 129280) with refreshed weights โ€” and a new storage format: experts are stored INT8 with UE8M0 scales (scale_fmt: ue8m0) instead of plain FP8. That is why these files are slightly larger than our earlier V4โ€‘Flash GGUFs, and why conversion needs torch>=2.7 (E8M0 dtype).

โœ… Verified (this build, Q3_K_M)

Captured greedy runs โ€” not copied from the base model card:

Test Output
Math 127+58= โ†’ 185
Korean "์„œ์šธ์€ ๋Œ€ํ•œ๋ฏผ๊ตญ์˜ ์ˆ˜๋„๋กœ, ์—ญ์‚ฌ์™€ ํ˜„๋Œ€๊ฐ€ ๊ณต์กดํ•˜๋Š” ์—ญ๋™์ ์ธ ๋„์‹œ์ž…๋‹ˆ๋‹ค. / ๋ถˆ๊ณ ๊ธฐ๋Š” ๋‹ฌ์ฝคํ•œ ๊ฐ„์žฅ ์–‘๋…์— ์žฌ์šด ์†Œ๊ณ ๊ธฐ๋ฅผ ๋ถˆ์— ๊ตฌ์›Œ๋‚ด๋Š”โ€ฆ / ๋น„๋น”๋ฐฅ์€ ๋ฐฅ ์œ„์— ๋‹ค์–‘ํ•œ ๋‚˜๋ฌผ๊ณผ ๊ณ ๊ธฐ, ๊ณ ์ถ”์žฅ์„ ์–น์–ดโ€ฆ / ๊น€์น˜๋Š” ๋ฐฐ์ถ”๋ฅผ ์†Œ๊ธˆ์— ์ ˆ์—ฌ ๊ณ ์ถง๊ฐ€๋ฃจ์™€ ์ “๊ฐˆ ๋“ฑ์œผ๋กœ ์–‘๋…ํ•ด ๋ฐœํšจ์‹œํ‚จโ€ฆ" โ€” fluent, zero token mixing or loops
Tool call {"tool":"get_weather","args":{"city":"๋ถ€์‚ฐ"}} โ€” exact JSON

๐Ÿš€ Usage

โš™๏ธ Runtime: batiai/bati.cpp โ€” DeepSeekโ€‘V4's CSA+HCA hybrid attention (deepseek4) is not in mainline llama.cpp. Build our fork.

hf download batiai/DeepSeek-V4-Flash-0731-GGUF "DeepSeek-V4-Flash-0731-Q3_K_M-*.gguf" --local-dir ./v4f

./llama-cli -m ./v4f/DeepSeek-V4-Flash-0731-Q3_K_M-00001-of-00004.gguf -ngl 99 -c 8192

โš ๏ธ Two things the base repo doesn't tell you (we hit both)

  1. No chat template ships with this model โ€” neither the original nor the GGUF has one. Use DeepSeek's format explicitly:
    <๏ฝœUser๏ฝœ>your question here<๏ฝœAssistant๏ฝœ>
    
  2. CUDA offload used to abort in ggml_cuda_op_concat (GGML_ASSERT(src0->type == GGML_TYPE_F32)) because the backend advertised concat support for types its kernel can't handle. Fixed in bati.cpp โ€” the CUDA backend now reports concat support only for F32, so those nodes fall back to CPU instead of crashing. Use bati.cpp at or after that fix, or run CPU/Metal. (Apple Metal users were never affected.)

โœจ What BatiAI did

  • Quantized from the official DeepSeek weights โ€” never a reโ€‘quant of someone else's GGUF
  • Handled the new UE8M0 / INT8โ€‘expert storage format end to end
  • Verified in Korean โ€” most GGUF publishers never check this
  • Fixed the runtime bug we found along the way, in the open, in bati.cpp
  • BatiAI metadataโ€‘signed

๐Ÿ“œ License โ€” MIT

Base model ยฉ DeepSeek, MIT. Quantized weights redistributed under the same terms.


Who we are. BatiAI builds onโ€‘device Korean AI โ€” BatiFlow runs LLMs, speechโ€‘toโ€‘text (batisay), document OCR (batisee) and speaker diarization locally on a Mac. No audio, no documents, no prompts leave the device. Browse the whole line at huggingface.co/batiai.

Downloads last month
251
GGUF
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for batiai/DeepSeek-V4-Flash-0731-GGUF

Quantized
(193)
this model

Collection including batiai/DeepSeek-V4-Flash-0731-GGUF