Mellum2-12B-A2.5B-Instruct-GGUF

This repository contains GGUF quantizations of JetBrains/Mellum2-12B-A2.5B-Instruct for efficient local inference with llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes.

Model details

  • Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
  • Format: GGUF
  • Architecture: Mellum2
  • Task type: Instruction-following assistant
  • Context length: 131,072 tokens
  • License: Apache 2.0

Available quantizations

File Quantization Size Notes
Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf Q4_K_M 8.07 GB Lower-memory deployment with good quality-to-size tradeoff
Mellum2-12B-A2.5B-Instruct-Q5_K_M.gguf Q5_K_M 9.21 GB Higher-quality 5-bit deployment with a modest size increase

Intended use

This model is intended for:

  • local chat and assistant workflows
  • coding assistance
  • tool-use experiments
  • CPU-friendly or memory-constrained inference
  • users who want a balance between quality and memory usage

Notes

These quantizations were produced with llama.cpp. During conversion, some tensors may require fallback quantization depending on the model architecture, which is expected and does not prevent successful inference.

Usage

llama.cpp

./llama-server -m Mellum2-12B-A2.5B-Instruct-Q5_K_M.gguf --jinja --port 8000

Ollama

Create a Modelfile pointing to the GGUF file:

FROM ./Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf

Then run:

ollama create mellum2-q4 -f Modelfile
ollama run mellum2-q4

License

Released under the same license as the base model: Apache 2.0.

Acknowledgments

  • Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
  • Quantization performed with llama.cpp
Downloads last month
4,861
GGUF
Model size
12B params
Architecture
mellum
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF

Quantized
(26)
this model