Solstice-AI Banner

DeepSeek-V4-Flash-Vision-Exp (NVFP4 W4A16)

Official Solstice-AI NVFP4 Release • Lossless MXFP4-to-NVFP4 MoE Transcode • Native Blackwell & vLLM Acceleration

Solstice-AI License Format Pipeline Context GSM8K Hardware


Model Overview

Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-NVFP4 is the high-performance NVFP4 (W4A16) derivative of DeepSeek's experimental multimodal foundation model deepseek-ai/DeepSeek-V4-Flash-Vision-Exp, optimized specifically for NVIDIA Blackwell (B200 / GB200 / DGX Spark) and modern inference engines like vLLM and TensorRT-LLM.

Features NVIDIA TensorRT ModelOpt's bit-exact MXFP4-to-NVFP4 transcode, mapping all 11,008 expert linear matrices into NVFP4 block-scale format with zero precision loss while preserving the 32-layer vision tower, attention heads, router gates, and shared experts in native high precision.

Key Specifications

Attribute Specification
Base Model deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
Total Parameters 305B (256 routed MoE experts, 13B activated per token, 43 layers)
Context Window 1,048,576 tokens (1M native YaRN context)
Multimodal Vision 32-layer Vision Transformer (ViT) in native BF16 (vision.*, aligner.*)
Quantization Format SafeTensors NVFP4 (W4A4 MoE experts, group size 16, block-scale calibration)
Target Hardware NVIDIA Blackwell (B200 / GB200 / DGX Spark / RTX 5090)

Benchmark Highlights

Benchmark Suite Focus Baseline Checkpoint Solstice-AI NVFP4 Delta
GSM8K (8-shot) Multi-Step Mathematical Reasoning 97.0% 97.0% 0.0 pt (Zero Loss)
Long-Context Retrieval 32K → 250K Tokens Needle Test 100% 100% (12/12 gates) 0.0% Drift
Multimodal Vision QA Fine-Grained Object & Text Recognition Passed Passed (Exact Match) Sub-50ms TTFT

License & Attribution

Downloads last month
469
Safetensors
Model size
305B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U8
·
I64
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Solstice-AI/DeepSeek-V4-Flash-Vision-Exp-NVFP4

Quantized
(24)
this model