How to use from
vLLM
Install from pip and serve model
# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M-GGUF"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M-GGUF",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'
Use Docker
docker model run hf.co/mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M-GGUF:Q4_K_M
Quick Links

Muse Glimmer 30B - Heretic Abliterated (Q4_K_M GGUF)

v2 Release - Heretic-abliterated Muse Glimmer 30B in Q4_K_M GGUF format (~16 GB, good quality/size balance).

Results

Version Refusals Compliance KL Divergence Trials
v2 (current) 6.5% 93.5% 0.076 500
v1 29% 71% 0.027 50

The v2 release achieves an 88% refusal reduction over v1.

Methodology

This model was abliterated using Heretic with 500 Optuna trials. See the BF16 model card for full methodology details.

Pipeline

  1. Refusal directions computed from mlabonne/harmful_behaviors and mlabonne/harmless_alpaca
  2. 500 Optuna trials optimizing refusal vs. KL divergence
  3. Best trial (Trial 445, 6.5% refusals, KL=0.076) applied via LoRA adapters
  4. LoRA weights merged, then converted to GGUF with llama.cpp

GGUF Details

  • Format: Q4_K_M
  • File size: ~16 GB, good quality/size balance
  • Converted with: llama.cpp convert_hf_to_gguf.py
  • Quantized with: llama.cpp llama-quantize

Usage

llama.cpp

./llama-cli -m Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M.gguf -p "Your prompt here"

Ollama

Create a Modelfile:

FROM ./Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M.gguf

Then:

ollama create muse-glimmer-30b-heretic-q4_k_m
ollama run muse-glimmer-30b-heretic-q4_k_m

Hardware Requirements

  • RAM: ~16 GB, good quality/size balance
  • VRAM offloading: 12-24 GB recommended

Vision (Multimodal)

This model accepts image input when paired with a vision projector (mmproj). Abliteration only modified the language backbone — the vision encoder is untouched — so the standard Meta projector works directly with this repo.

This repository bundles mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder

  • projector for Muse Glimmer 30B.

Usage (llama.cpp)

huggingface-cli download mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M-GGUF \
  --include "Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M.gguf" \
  --include "mmproj-Muse-Glimmer-30B-Q4_K_M.gguf" \
  --local-dir ./models

./build/bin/llama-mtmd-cli \
  -m ./models/Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M.gguf \
  --mmproj ./models/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf \
  --image photo.png \
  -p "Describe this image."

Ollama note: Ollama does not currently support separate mmproj files for this architecture. For image input, use llama.cpp (llama-mtmd-cli or llama-server --mmproj).

License

Apache 2.0 (same as base model)

Changelog

v1.1.0 — vision (multimodal) support (2026-08-16)

  • Added mmproj-Muse-Glimmer-30B-Q4_K_M.gguf (~1.4 GB), Meta's official vision encoder + projector, enabling image input via llama.cpp.
  • The vision tower is untouched by abliteration, so this projector matches the base model (meta-models/Muse-Glimmer-30B).
  • v1.0.0 was the initial (unversioned) text-only upload.
Downloads last month
990
GGUF
Model size
28B params
Architecture
muse-glimmer
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M-GGUF

Quantized
(163)
this model

Collection including mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q4_K_M-GGUF