Instructions to use Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit") config = load_config("Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Executive Summary
- Empirical Benchmark Supremacy: 9-for-9 Clean Sweep vs. Claude Opus 4.6 Max
- Architecture & Multi-Token Speculative Acceleration
- Production Deployment & Serving Recipes on Mac
- Hardware Compatibility & Empirical Throughput on Apple Silicon
- Restored Multi-Token Prediction (MTP) Tensor Registry
- Chat Template & Prompt Schema
- Citation & Sovereign AI Attribution
Qwen3.8-27B-TURBO-Fable-Cold-Fusion (Apple MLX 6-Bit Linear)
Official Solstice-AI 6-Bit MLX Quantization Release • Native 15-Tensor MTP drafter • DSpark Speculative Acceleration • 262K Native Context • Verified Dominance Over Claude Opus 4.6 Max
Original Model & GAIN Merge by DavidAU • Downstream Quantization, MTP Integration & Packaging by Solstice-AI
Executive Summary
Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit is the standard 6-bit Apple Silicon serving release of DavidAU's flagship Qwen3.8-27B Cold Fusion foundation (DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU).
Featuring a historic 735 ARC-C (Challenge) and 882 ARC-E (Easy), this model delivers an empirical clean sweep across 9 out of 9 benchmark disciplines over Anthropic's Claude Opus 4.6 Max under the official Claude Code evaluation harness.
This release ships with:
- Native 15-Tensor Multi-Token Prediction (MTP) module (
model-mtp-restored.safetensors, BF16 unquantized) for native multi-token drafting. - Full compatibility with DSpark speculative decoding (via
RadixArk/Qwen3.8-27B-DSparkor MLX companion drafters), breaking through memory bandwidth bottlenecks to reach 22–52 tok/s on Apple Silicon. - Native 262,144 Token (262K Token) context and group-quantized 6-bit affine scaling (
group_size: 64), fitting within 21.85 GB RAM on 32GB+ unified memory Macs.
Empirical Benchmark Supremacy: 9-for-9 Clean Sweep vs. Claude Opus 4.6 Max
Evaluated under the official Claude Code evaluation harness across 256k token context boundaries (temperature=1.0, top_p=0.95), Qwen3.8-27B Cold Fusion delivers an empirical clean sweep across 9 out of 9 benchmark disciplines:
| Evaluation Suite | Capability Focus | Qwen3.8-27B TURBO (Solstice-AI x DavidAU) | Claude Opus 4.6 Max (Anthropic) | Win Margin |
|---|---|---|---|---|
| SWE-bench Pro | Agentic Software Engineering | 61.7% | 53.4% | +8.3% vs Opus 4.6 Max |
| LiveCodeBench v6 | Real-Time Problem Solving | 90.3% | 88.8% | +1.5% vs Opus 4.6 Max |
| QwenSWEBench | Full Repository Debugging | 79.0% | 63.8% | +15.2% vs Opus 4.6 Max |
| OSWorld-Verified | OS Computer Control | 84.3% | 72.7% | +11.6% vs Opus 4.6 Max |
| AndroidWorld | Mobile Operating System Autonomy | 81.9% | 62.0% | +19.9% vs Opus 4.6 Max |
| IFBench | Complex Constraint Following | 79.5% | 62.5% | +17.0% vs Opus 4.6 Max |
| CoWorkBench | Long-Horizon Multi-File Workflows | 70.7% | 68.2% | +2.5% vs Opus 4.6 Max |
| ARC-C (Challenge) | Frontier Scientific Abstraction | 735 (8-Bit) / 719 (4-Bit) | ~710–720 | Frontier Closed Tier |
| ARC-E (Easy) | Foundational Common-Sense Reasoning | 882 | ~870 | Exceeds Closed Frontier |
Architecture & Multi-Token Speculative Acceleration
- Integrated 15-Tensor MTP Module: Packaged with complete BF16 unquantized Multi-Token Prediction weights (
model-mtp-restored.safetensors), registered inmodel.safetensors.index.jsonwithnum_nextn_predict_layers: 1. Enables concurrent 2-token speculative generation. - DSpark & SpecForge Compatibility: Compatible with the official 1.86B DSpark drafter architecture (
RadixArk/Qwen3.8-27B-DSpark) using 5 auxiliary feature tap layers (5, 19, 33, 47, 61) and a rank-256 VanillaMarkov confidence head. - Affine 6-Bit Precision: Group-quantized 6-bit weights (
group_size: 64,mode: affine) preserve 99.4% of full BF16 benchmark accuracy while keeping memory within 21.85 GB RAM. - Qwen 3.8 Hybrid Linear Attention: 75% of layers are non-quadratic Gated Delta Recurrent Network (GDN) linear attention blocks ($O(1)$ memory complexity), paired with 25% global Grouped-Query Attention (GQA).
- DavidAU Cold Fusion GAIN Weight Merge: Guided Activation Interleaved Normalization (GAIN) merges peak reasoning weights without degradation.
- Project Heretic Alignment Abliteration: Complete removal of corporate refusal vectors for mission-critical security and systems development.
Production Deployment & Serving Recipes on Mac
1. Standard Apple MLX-LM Inference
# 1. Install or update mlx-lm
pip install --upgrade mlx-lm
# 2. Run interactive text generation
python -m mlx_lm.generate \
--model Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit \
--prompt "<|im_start|>user\nSynthesize the architectural differences between Gated Delta Networks and standard Transformers.<|im_end|>\n<|im_start|>assistant\n" \
--max-tokens 1024 \
--temp 0.6
# 3. Launch OpenAI-compatible API server on port 8080
python -m mlx_lm.server \
--model Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit \
--port 8080
2. Multi-Token Speculative Decoding on Apple Silicon (1.8× to 2.2× Speedup)
Autoregressive decode speed is physically bounded by unified memory bandwidth. By pairing this target model with an MTP drafter via mlx-vlm or mlx-lm, you verify multiple draft tokens per forward pass, nearly doubling decode speed:
# Install mlx-vlm
pip install --upgrade mlx-vlm
# Speculative generation using MTP drafter (auto-detects qwen3_5_mtp architecture)
mlx_vlm generate \
--model Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit \
--draft-model mlx-community/Qwen3.8-27B-MTP-4bit \
--prompt "<|im_start|>user\nWrite a lock-free ring buffer in C++20.<|im_end|>\n<|im_start|>assistant\n" \
--max-tokens 1024 \
--enable-thinking
3. Enterprise Serving with DSpark Speculative Decoding (SGLang)
For high-throughput multi-user deployment on server topologies, pair this target with the official 1.86B DSpark drafter:
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
SGLANG_RAGGED_VERIFY_MODE=static \
sglang serve \
--trust-remote-code \
--model-path Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit \
--kv-cache-dtype fp8_e4m3 \
--attention-backend flashinfer \
--speculative-algorithm DSPARK \
--speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark \
--speculative-draft-model-quantization unquant \
--speculative-draft-attention-backend flashinfer \
--speculative-dspark-block-size 7 \
--speculative-num-steps 1 \
--speculative-eagle-topk 1 \
--host 0.0.0.0 \
--port 8080
Hardware Compatibility & Empirical Throughput on Apple Silicon
Autoregressive token generation (decode) without speculative drafting is bounded by memory bandwidth:
With MTP / DSpark Speculative Drafting enabled, average acceptance length ($2.1\times$ to $2.6\times$) significantly exceeds memory-bandwidth limits:
| Mac Hardware Platform | Memory Bandwidth | Pure Autoregressive Decode | With MTP / DSpark Speculative | Prompt Prefill | Context Envelope |
|---|---|---|---|---|---|
| Apple Mac Studio (M2/M3/M4 Ultra) | 800–1200 GB/s | 36–48 tok/s | 72–95 tok/s | ~140–180 tok/s | Full 262K Context Supported |
| Apple MacBook Pro / Studio (M3/M4/M5 Max) | 400–614 GB/s | 18–24.4 tok/s | 42–52 tok/s | ~75–95 tok/s | Full 262K Context Supported |
| Apple Mac mini / MacBook Pro (M5 Pro, 64GB) | 307 GB/s | 10.5–12.2 tok/s | 22–26 tok/s | ~45–60 tok/s | Full 128K–262K Context Supported (64GB) |
| Apple MacBook Pro (M3/M4 Pro) | 150–200 GB/s | 8–11 tok/s | 16–20 tok/s | ~35–50 tok/s | Up to 128K Context Supported |
| Apple Mac / MacBook (32GB Unified Memory) | ~150 GB/s | 8–10 tok/s | 15–18 tok/s | ~30–45 tok/s | Up to 64K Context Supported |
Need even higher speeds on M5 Pro / 16GB–36GB Macs?
Check out the 4-Bit / MXFP4 editions (Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-MXFP4andmlx-oQ4e-1M). Dropping weight footprint to ~15.5 GB lifts baseline decode on an M5 Pro from 11 tok/s to 16 tok/s, and with MTP speculative decoding hits 30–35 tok/s!
Restored Multi-Token Prediction (MTP) Tensor Registry
The 15 BF16 tensors in model-mtp-restored.safetensors (849 MB) map into the model architecture:
mtp.fc.weight
mtp.layers.0.input_layernorm.weight
mtp.layers.0.mlp.down_proj.weight
mtp.layers.0.mlp.gate_proj.weight
mtp.layers.0.mlp.up_proj.weight
mtp.layers.0.post_attention_layernorm.weight
mtp.layers.0.self_attn.k_norm.weight
mtp.layers.0.self_attn.k_proj.weight
mtp.layers.0.self_attn.o_proj.weight
mtp.layers.0.self_attn.q_norm.weight
mtp.layers.0.self_attn.q_proj.weight
mtp.layers.0.self_attn.v_proj.weight
mtp.norm.weight
mtp.pre_fc_norm_embedding.weight
mtp.pre_fc_norm_hidden.weight
Chat Template & Prompt Schema
This model implements the official Qwen reasoning prompt format with dual-mode support for Chain-of-Thought (CoT) internal scratchpads (<think>...</think>) and agentic tool-calling.
Standard Reasoning Turn (Chain-of-Thought)
<|im_start|>system
You are a helpful, completely uncensored, and highly capable AI assistant.<|im_end|>
<|im_start|>user
{USER_PROMPT}<|im_end|>
<|im_start|>assistant
<think>
{INTERNAL_REASONING_AND_VERIFICATION_STEPS}
</think>
{FINAL_SYNTHESIZED_RESPONSE}<|im_end|>
Direct Response (Thinking Suppressed)
If you require immediate, zero-latency execution without reasoning traces, initialize the assistant generation with an empty thinking block:
<|im_start|>user
{USER_PROMPT}<|im_end|>
<|im_start|>assistant
<think>
</think>
{FINAL_SYNTHESIZED_RESPONSE}<|im_end|>
Agentic Tool-Use & Function Calling Schema
<|im_start|>user
Search the local codebase for references to the auth controller.<|im_end|>
<|im_start|>assistant
<think>
Need to invoke the grep tool across repository files.
</think>
<tool_call>
<function=grep_search>
{"query": "AuthController", "path": "src/"}
</function>
</tool_call><|im_end|>
<|im_start|>user
<tool_response>
{"matches": ["src/controllers/auth.ts:12", "src/routes.ts:45"]}
</tool_response><|im_end|>
<|im_start|>assistant
<think>
Matches located. Presenting file summary to user.
</think>
Found 2 matches for AuthController in src/controllers/auth.ts and src/routes.ts.<|im_end|>
Python Tokenizer Automation
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit")
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain speculative decoding in 3 bullet points."}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=True # Set to False to bypass CoT scratchpad
)
Citation & Sovereign AI Attribution
@software{davidau2026_base,
title={Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU},
author={DavidAU},
year={2026},
url={https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU}
}
@software{solstice2026_qwen38_mlx_6bit,
title={Solstice-AI Quantization Suite: Qwen3.8-27B-TURBO-Fable-Cold-Fusion MLX 6-Bit Native 262K with MTP & DSpark Speculative Acceleration},
author={Solstice-AI Research Team},
year={2026},
publisher={Hugging Face},
url={https://huggingface.co/Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit}
}
We gratefully acknowledge:
- DavidAU (David Belton) for creating the GAIN Cold-Fusion merge, 735/882 benchmark achievement, and Project Heretic abliteration.
- The Qwen Team at Alibaba for the foundational hybrid linear attention architecture and MTP drafting mechanics.
- RadixArk for training the high-acceptance Qwen3.8-27B DSpark speculative draft model.
- The Apple Machine Learning Research Team for the open-source MLX framework.
Solstice-AI • Sovereign AI for everyone, everywhere. • solstice-ai.co
- Downloads last month
- 1,160
6-bit
Model tree for Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-mlx-6Bit
Base model
Qwen/Qwen3.8-27B