Instructions to use mlx-community/DeepSeek-V4-Flash-0731-2.4bit-mixed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/DeepSeek-V4-Flash-0731-2.4bit-mixed with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("mlx-community/DeepSeek-V4-Flash-0731-2.4bit-mixed") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use mlx-community/DeepSeek-V4-Flash-0731-2.4bit-mixed with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "mlx-community/DeepSeek-V4-Flash-0731-2.4bit-mixed" --prompt "Once upon a time"
- Atomic Chat
Optiq serve encountered an error during startup
#2
by sqh11 - opened
Python 3.14 installs pip install mlx optiq==0.4.19, and after starting, issues a request error.
Start command:
optiq serve \
--model "${MODEL_NAME}" \
--stream-experts \
--host "${HOST}" \
--port "${PORT}"
Error log:
127.0.0.1 - - [14/Aug/2026 02:07:29] "POST /v1/chat/completions HTTP/1.1" 200 -
INFO:root:Prompt Cache: 0 sequences, 0.00 GB
INFO:root:- assistant: 0 sequences, 0.00 GB
INFO:root:- user: 0 sequences, 0.00 GB
INFO:root:- system: 0 sequences, 0.00 GB
INFO:root:Prompt processing progress: 0/34
libc++abi: terminating due to uncaught exception of type std::runtime_error: There is no Stream(gpu, 1) in current thread.
Have you encountered the same problem?