How to use from
Pi
Start the llama.cpp server
# Install llama.cpp:
brew install llama.cpp
# Start a local OpenAI-compatible server:
llama serve -hf NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF:
Configure the model in Pi
# Install Pi:
npm install -g @earendil-works/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "llama-cpp": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF:"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

Mellum2-12B-A2.5B-Instruct-GGUF

This repository contains GGUF quantizations of JetBrains/Mellum2-12B-A2.5B-Instruct for efficient local inference with llama.cpp, Ollama, LM Studio, and other GGUF-compatible runtimes.

Model details

  • Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
  • Format: GGUF
  • Architecture: Mellum2
  • Task type: Instruction-following assistant
  • Context length: 131,072 tokens
  • License: Apache 2.0

Available quantizations

File Quantization Size Notes
Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf Q4_K_M 8.07 GB Lower-memory deployment with good quality-to-size tradeoff
Mellum2-12B-A2.5B-Instruct-Q5_K_M.gguf Q5_K_M 9.21 GB Higher-quality 5-bit deployment with a modest size increase

Intended use

This model is intended for:

  • local chat and assistant workflows
  • coding assistance
  • tool-use experiments
  • CPU-friendly or memory-constrained inference
  • users who want a balance between quality and memory usage

Notes

These quantizations were produced with llama.cpp. During conversion, some tensors may require fallback quantization depending on the model architecture, which is expected and does not prevent successful inference.

Usage

llama.cpp

./llama-server -m Mellum2-12B-A2.5B-Instruct-Q5_K_M.gguf --jinja --port 8000

Ollama

Create a Modelfile pointing to the GGUF file:

FROM ./Mellum2-12B-A2.5B-Instruct-Q4_K_M.gguf

Then run:

ollama create mellum2-q4 -f Modelfile
ollama run mellum2-q4

License

Released under the same license as the base model: Apache 2.0.

Acknowledgments

  • Base model: JetBrains/Mellum2-12B-A2.5B-Instruct
  • Quantization performed with llama.cpp
Downloads last month
4,861
GGUF
Model size
12B params
Architecture
mellum
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for NANI-Nithin/Mellum2-12B-A2.5B-Instruct-GGUF

Quantized
(26)
this model