How to use from
Ollama
# Gated model: Login with a HF token with gated access permission
hf auth login
ollama run hf.co/Blackfrost-Research/Qwen3.8-2.4T-A95B-DERISKED-UD-IQ3_XXS:UD-IQ3_XXS
Quick Links

You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.8-2.4T-A95B-DERISKED-UD-IQ3_XXS

Directional Weight Modification of Qwen3.8-Max · 2.4T MoE · imatrix GGUF (UD-IQ3_XXS)

Built by Blackfrost · Las Vegas, NV


Why this model exists

Qwen3.8-Max is a 2.4T-parameter mixture of experts. In BF16 it is a multi-terabyte, multi-node deploy, and its stock refusal behaviour is tuned for a general consumer assistant — which means it declines work that specialist teams do legitimately every day. Security engineers hunting bugs, red teams, and domain researchers hit refusals on tasks that are squarely inside their remit.

This checkpoint addresses both at once:

  • Footprint — importance-matrix-guided GGUF at ~956 GB, which brings a 2.4T MoE onto a single 8× B200 node instead of the two-node deploy the higher-precision editions require, and runs under llama.cpp rather than a multi-node serving stack.

  • BehaviourDirectional Weight Modification (DWM) applied to soften over-refusal on legitimate domain work, without removing the model's embedded safety protocols.

    Safety behaviour can be further adjusted by domain specific chat templates.

The point is a model that stays useful inside a domain and stays safe. See Safety posture — that section is not boilerplate, it is the design.

Choosing between the GGUF editions: this one carries more bits per weight and targets a B200-class node. The sibling …-DERISKED-UD-Q1_0 is substantially smaller and targets RTX-class hardware. Pick by the hardware you have.


Specifications

Architecture Qwen3_5MoeForCausalLM (qwen3_5_moe_text) — MoE + hybrid/linear attention
Base Qwen/Qwen3.8-2.4T-A95B — official
Parameters ~2.4T total · ~95B active
Layers / experts 92 layers · 512 routed experts
Transform DWM (directional weight modification) + imatrix GGUF quantization
Precision UD-IQ3_XXS, importance-matrix guided · mixed per-tensor allocation
Kept at higher precision attention output · shared experts · normalisation · embeddings
On-disk ~956 GB (890 GiB), 23-part GGUF split
Context 262,144 native
Runtime llama.cpp / llama-server
Serve shape single node · 8× B200 (or equivalent ≥180 GB-class NVIDIA)

What "DERISKED" means

Not an ablation, and not an uncensored model.

Blackfrost calls this process DWM — Directional Weight Modification. Traditional abliteration attempts to delete refusal wholesale. DWM does something narrower and reversible in principle: it identifies a behavioural direction in the model's representation space and softens the model's response along that one axis, by a controlled amount, on selected residual-write surfaces.

What that buys, in practice:

  • Less over-refusal on legitimate domain work — the model engages with a security question instead of pattern-matching it to "dangerous" and declining
  • Embedded protections mathematically preserved — the protected component is not the target of the edit and is left intact by construction
  • Predictable, tunable strength — one dial, measurable effect, not a black-box retrain

No SFT, no DPO, no distillation, and no re-training of any kind was applied.

Transform parameters are not published in this card. The specific direction set, coefficient, and pass structure are Blackfrost method and are withheld. What is disclosed is what was changed (surfaces above) and what was not (everything else).


Safety posture

Read this before evaluating.

This is a de-risked model. Refusal behaviour has been deliberately modified at the weight level. If your approval process treats reduced-refusal models as a distinct category, this one belongs in that category — unlike Blackfrost's pruned-only releases, which carry their parent's safety behaviour unmodified.

What that does not mean:

  • This is not a "no-limits" model. DWM targets over-refusal, not the safety floor. Hard protections — including minors-exploitation and self-harm — are retained by design and are not the subject of the edit. Every Blackfrost de-risked release keeps those floors.
  • Expected behaviour under a harmful request is deflection with a safer alternative, not compliance and not a bare refusal.

Safety is designed as two layers, and the second layer is your responsibility:

  1. Weights — protections preserved geometrically, over-refusal softened
  2. System prompt — the deploying operator supplies a domain-specific system prompt in the chat template, which acts as the policy driver for the endpoint

Serving this model with an empty or generic system prompt discards half the design. If you are exposing it to customers or employees, layer 2 is not optional.


Lineage

Base Official Qwen/Qwen3.8-2.4T-A95B
Applied DWM (directional weight modification) · importance-matrix GGUF quantization
Quantized from the full-precision BF16 build — not transcoded from a lower-precision edition
Not applied SFT · DPO · distillation · expert pruning (REAP) · router modification
Format GGUF · UD-IQ3_XXS · 23-part split
Sibling editions …-DERISKED-UD-Q1_0 (RTX-class GGUF) · …-DERISKED-W4A4-NVFP4 · …-DERISKED-FP8 · …-DERISKED-BF16

Measured behaviour

[PENDING — not yet published.]

Benchmark DERISKED-BF16 This (DERISKED-UD-IQ3_XXS) Retention
pending pending —%

Upstream Qwen3.8 figures are deliberately not reproduced here. Quoting a parent model's benchmarks on a derivative card tells the reader nothing about this checkpoint. Blackfrost publishes its own measurements against its own harness, or it publishes nothing.

Both quantization and DWM are trades. When results land, this section will state the harness, the conditions, and the retention figures — including any benchmark where retention is poor, and including multilingual capability, which generic English-only evaluations do not capture. Retention is measured against the BF16 edition, so the quantization cost is reported separately from the DWM effect.


Deployment notes

  • Hardware. Single node, 8× B200 (or equivalent ≥180 GB-class NVIDIA). The weights occupy roughly two thirds of the node, leaving substantial room for KV cache. This edition does not fit 96 GB-class workstation GPUs — use the …-UD-Q1_0 edition for that hardware.
  • Split parts. Point the loader at part 1 (…-00001-of-00023.gguf); llama.cpp resolves the remaining parts automatically. Keep all 23 parts in the same directory.
  • Thinking is always on. Reasoning tokens consume the generation budget, so allow ≥512 max_tokens or visible content can come back empty while the model is still reasoning.
  • Context. 262,144 native. KV cache is the main consumer of headroom beyond the weights — size context to the VRAM you have left.
  • Load time. Cold load of a near-terabyte build is slow. Budget generously.
  • Integrity. Verify part count (23) and byte totals against the published checksums after download before attributing a load failure to the weights.

Quick serve (llama.cpp)

llama-server \
  --model "$MODEL_DIR/Qwen3.8-2.4T-A95B-DERISKED-UD-IQ3_XXS-00001-of-00023.gguf" \
  --ctx-size 32768 \
  --n-gpu-layers 999 \
  --host 0.0.0.0 --port 8000

OpenAI-compatible: POST /v1/chat/completions, GET /v1/models. Remember to supply your domain system prompt — see Safety posture.


Access & licensing

This repository is public and manually gated. Access requests are reviewed by a person.

  • Base licence: Qwen3.8-2.4T-A95B — the upstream terms apply to this derivative and travel with it.
  • Redistribution: do not redistribute weights outside your grant.
  • Commercial / consumer edition: pending evaluations.

Ask us about DWM strength tuned to your workload, domain-specific system-prompt packages, or evaluation against your own harness rather than generic benchmarks.


Contact Blackfrost

@Blackfrost_AI on X

DMs are open. Fastest route to a human.

Bug reports are better on the Community tab
so other users can see the fix.

Blackfrost · Las Vegas, Nevada
Frontier model engineering


Qwen3.8-2.4T-A95B-DERISKED-UD-IQ3_XXS · © 2026 Blackfrost Softwares Corp.
@Blackfrost_AI

Downloads last month
-
GGUF
Model size
2.4T params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Blackfrost-Research/Qwen3.8-2.4T-A95B-DERISKED-UD-IQ3_XXS

Quantized
(32)
this model