--- license: apache-2.0 pipeline_tag: text-generation language: - en tags: - heretic - abliterated - uncensored - Muse-Glimmer - 30B - Heretic - GGUF - q6_k base_model: meta-models/Muse-Glimmer-30B quantized_by: mlasli --- # Muse Glimmer 30B - Heretic Abliterated (Q6_K GGUF) **v2 Release** - Heretic-abliterated Muse Glimmer 30B in Q6_K GGUF format (~22 GB, very good quality). ## Results | Version | Refusals | Compliance | KL Divergence | Trials | |---------|----------|------------|---------------|--------| | **v2 (current)** | **6.5%** | **93.5%** | **0.076** | 500 | | v1 | 29% | 71% | 0.027 | 50 | The v2 release achieves an **88% refusal reduction** over v1. ## Methodology This model was abliterated using **[Heretic](https://github.com/d3nd3/heretic)** with 500 Optuna trials. See the [BF16 model card](https://huggingface.co/mlasli/Muse-Glimmer-30B-Heretic-Abliterated-BF16) for full methodology details. ### Pipeline 1. Refusal directions computed from `mlabonne/harmful_behaviors` and `mlabonne/harmless_alpaca` 2. 500 Optuna trials optimizing refusal vs. KL divergence 3. Best trial (Trial 445, 6.5% refusals, KL=0.076) applied via LoRA adapters 4. LoRA weights merged, then converted to GGUF with llama.cpp ## GGUF Details - **Format**: Q6_K - **File size**: ~22 GB, very good quality - **Converted with**: llama.cpp `convert_hf_to_gguf.py` - **Quantized with**: llama.cpp `llama-quantize` ## Usage ### llama.cpp ```bash ./llama-cli -m Muse-Glimmer-30B-Heretic-Abliterated-Q6_K.gguf -p "Your prompt here" ``` ### Ollama Create a Modelfile: ```dockerfile FROM ./Muse-Glimmer-30B-Heretic-Abliterated-Q6_K.gguf ``` Then: ```bash ollama create muse-glimmer-30b-heretic-q6_k ollama run muse-glimmer-30b-heretic-q6_k ``` ## Hardware Requirements - **RAM**: ~22 GB, very good quality - **VRAM offloading**: 12-24 GB recommended ## Vision (Multimodal) This model accepts image input when paired with a vision projector (`mmproj`). Abliteration only modified the language backbone — the vision encoder is untouched — so the standard Meta projector works directly with this repo. This repository bundles `mmproj-Muse-Glimmer-30B-Q4_K_M.gguf` (~1.4 GB), Meta's official vision encoder + projector for Muse Glimmer 30B. ### Usage (llama.cpp) ```bash huggingface-cli download mlasli/Muse-Glimmer-30B-Heretic-Abliterated-Q6_K-GGUF \ --include "Muse-Glimmer-30B-Heretic-Abliterated-Q6_K.gguf" \ --include "mmproj-Muse-Glimmer-30B-Q4_K_M.gguf" \ --local-dir ./models ./build/bin/llama-mtmd-cli \ -m ./models/Muse-Glimmer-30B-Heretic-Abliterated-Q6_K.gguf \ --mmproj ./models/mmproj-Muse-Glimmer-30B-Q4_K_M.gguf \ --image photo.png \ -p "Describe this image." ``` > **Ollama note**: Ollama does not currently support separate `mmproj` files > for this architecture. For image input, use llama.cpp (`llama-mtmd-cli` or > `llama-server --mmproj`). ## License Apache 2.0 (same as base model) ## Changelog ### v1.1.0 — vision (multimodal) support (2026-08-16) - Added `mmproj-Muse-Glimmer-30B-Q4_K_M.gguf` (~1.4 GB), Meta's official vision encoder + projector, enabling image input via llama.cpp. - The vision tower is untouched by abliteration, so this projector matches the base model (`meta-models/Muse-Glimmer-30B`). - v1.0.0 was the initial (unversioned) text-only upload.