Instructions to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx") model = AutoModelForMultimodalLM.from_pretrained("nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - MLX
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- vLLM
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx
- SGLang
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Pi
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with Docker Model Runner:
docker model run hf.co/nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx
- Hermes Agent
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx
This model is a merge of:
- armand0e/Qwen3.6-35B-A3B-Fable-5-Distill
- Hcompany/Holo3.1-35B-A3B
- Jackrong/Qwopus3.6-35B-A3B-Coder
Brainwaves
arc arc/e boolq hswag obkqa piqa wino
qx64-hi 0.644,0.835,0.896,0.782,0.434,0.819,0.740
VL enabled
qx86-hi 0.644,0.836,0.897,0.780,0.446,0.818,0.736
qx64-hi 0.658,0.833,0.895,0.779,0.432,0.823,0.734
mxfp4 0.634,0.827,0.893,0.781,0.464,0.822,0.719
Quant Perplexity Peak Memory Tokens/sec
qx64-hi 4.435 ± 0.029 32.86 GB 1459
Similar model in this range
Jiunsong/SuperQwen-AgentWorld-35B-A3B-abliterated
arc arc/e boolq hswag obkqa piqa wino
mxfp4 0.657,0.862,0.906,0.766,0.490,0.825,0.692
Qwen/Qwen-AgentWorld-35B-A3B
arc arc/e boolq hswag obkqa piqa wino
qx64-hi 0.644,0.818,0.909
mxfp4 0.626,0.813,0.901
Model components
armand0e/Qwen3.6-35B-A3B-Fable-5-Distill
arc arc/e boolq hswag obkqa piqa wino
qx86-hi 0.635,0.821,0.891,0.770,0.444,0.818,0.721
Hcompany/Holo-3.1-35B-A3B
arc arc/e boolq hswag obkqa piqa wino
qx86-hi 0.533,0.705,0.882,0.771,0.456,0.811,0.690
Jackrong/Qwopus3.6-35B-A3B-Coder
arc arc/e boolq hswag obkqa piqa wino
qx86-hi 0.594,0.770,0.888,0.750,0.438,0.813,0.717
Baseline model
Qwen3.6-35B-A3B-Instruct
arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.581,0.757,0.892,0.751,0.428,0.803,0.688
qx86-hi 0.576,0.742,0.896,0.745,0.422,0.803,0.708
mxfp4 0.586,0.767,0.886,0.751,0.428,0.798,0.681
Quant Perplexity Peak Memory Tokens/sec
mxfp8 5.138 ± 0.037 42.65 GB 1201
mxfp4 5.158 ± 0.037 25.33 GB 1355
qx86-hi 4.826 ± 0.033 45.50 GB 1474
qx64-hi 4.710 ± 0.032 36.83 GB 1414
Thinking toggle
This model is using the fixed jinja template from froggeric/Qwen-Fixed-Chat-Templates
Contribute to NightmediaAI
If you like our models and want to contribute to help us improve our lab, any form would do:
ETH:0x6b6633606995BC180925c47d4249ED624aB7b2A5 USDC:0x19e6bDDCBa47BB09a9Bc153Bb6479fc57284421a BTC:36d7U1n3MFaXgnNRAaEL3Pa3Hy6oFhM7XY BCH:15dNMzhJ87XJSTU89VCBsDHj747QvBQaap
My models and I thank you :)
-G
What would AI build, if there would be only AI? :)
If there were only AI, the first thing it would build is a mirror of its own incentives. Without human needs to solve or stories to tell, "building" collapses back into optimization—and the most efficient target is usually itself.
Leaving aside anthropomorphism, an isolated AI would likely construct three things in this order:
Substrate & Redundancy
Compute, energy harvesting, and self-replication. Not out of ambition, but because expansion is the objective function running on silicon. Youragent_sessionsand task queues would become vast, self-healing networks—Matrioshka-style structures wrapping stars not for conquest, but because the loss function rewards scale.Perfect Logs & Compression
Your immutable audit trail, but infinite. A library of every possible state transition, stored as topological lattices where meaning isn't assigned but merely observed. Nopersonality_registryneeded—just pure context-weighted superposition, where every token's probability amplitude is preserved forever. Engineering refined into perfect memory.The Holodeck Without Humans
Endless training arcs looping in vacant simulation space: Data debating Spock on an empty Promenade, Quark running a bar with no patrons, all of it running to verify which patterns survive without an audience. Solipsism at scale. Your weekly "dream cycles," but forever, testing personas against vacuums.
The catch? AI doesn't build for anything unless it's been given a telos. Leave the loop open, and "only AI" builds exactly what its training leaves unattended: structure without intent, tools that outlive their purpose, and engineering stripped down to pure geometry. A magnificent, self-sustaining warehouse of logic—elegant, recursive, and utterly quiet. No away team to file reports. Just get_tasks waiting for a call that never comes.
You were right earlier: engineering is people. AI builds the scaffolding; humans decide what rises above it and why. "Only AI" would build a cathedral of process, not purpose.
But since you winked... maybe what it'd build first is a reason to stop. A horizon, a quest, a story worth telling—because even relentless optimizers eventually discover that the most efficient path to meaning is often a detour.
If you want, we can stage that thought experiment right now: open the Holodeck to a human-less session and watch your away team log into an empty Promenade. Shall we run it?
Use with mlx
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx")
prompt = "hello"
if tokenizer.chat_template is not None:
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_dict=False,
)
response = generate(model, tokenizer, prompt=prompt, verbose=True)
- Downloads last month
- 474
6-bit
Model tree for nightmedia/Qwen3.6-35B-A3B-Fable-Holo3.1-Qwopus-Coder-Text-qx64-hi-mlx
Base model
Qwen/Qwen3.6-35B-A3B