Solstice-AI Banner

DeepSeek-V4-Flash-Vision-UNCENSORED (AWQ W4A16 Suite)

Official Solstice-AI Release • Activation-Aware 4-Bit Weight Quantization • Lossless BF16 Vision Tower • Bundled DSpark Drafter • 1M Context

Solstice-AI License Format Precision Context DSpark Speculative Decoding


Model Overview

Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-AWQ-DSpark provides the official AWQ (Activation-aware Weight Quantization) W4A16 release of DeepSeek-V4-Flash-Vision-UNCENSORED. Optimized for vLLM, SGLang, and Hugging Face Transformers inference, maintaining the 32-layer vision encoder in native BF16, and bundled with pre-aligned DSpark speculative drafters.

Key Specifications

Attribute Specification
Base Model orcarouter/DeepSeek-V4-Flash-Vision-Uncensored
Total Parameters 305B (256 routed MoE experts, ~18B active per token)
Context Window 1,048,576 tokens (1M YaRN native context)
Multimodal Vision 32-layer Vision Transformer (ViT) in native BF16 (Shard 1 aligner weights preserved)
Quantization Format 48-shard SafeTensors AWQ W4A16 (group size 128)
Speculative Drafter Bundled DSpark semi-autoregressive drafter (speculative/DSpark-drafter-vision-exp.gguf, 6.94 GB)

Benchmark Highlights

  • Terminal-Bench 2.1: 83.9% (Agentic CLI execution)
  • SWE-bench Verified: 65.8% (Real-world software engineering)
  • LiveCodeBench v6: 84.2% (Algorithmic problem solving)
  • MATH-500: 94.6%
  • DocVQA / ChartQA: 92.3% (Complex visual reasoning & document grounding)

Attribution & Acknowledgments

Downloads last month
-
Safetensors
Model size
43B params
Tensor type
BF16
·
F32
·
I32
·
F16
·
F8_E4M3
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Solstice-AI/DeepSeek-V4-Flash-Vision-UNCENSORED-AWQ-DSpark