Text Generation
Transformers
GGUF
English
llama-cpp
quantized
atlas-coder
atlas-coder-2
qwen2.5-coder
sub-1b
coding
code-generation
ollama
lm-studio
conversational
Instructions to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
- SGLang
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with Ollama:
ollama run hf.co/Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with Docker Model Runner:
docker model run hf.co/Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
- Lemonade
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Atlas-Coder-2-0.5B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Pluto-AI-Labs/Atlas-Coder-2-0.5B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| license: apache-2.0 | |
| language: | |
| - en | |
| tags: | |
| - gguf | |
| - llama-cpp | |
| - quantized | |
| - atlas-coder | |
| - atlas-coder-2 | |
| - qwen2.5-coder | |
| - sub-1b | |
| - coding | |
| - code-generation | |
| - ollama | |
| - lm-studio | |
| base_model: Siddh07ETH/Atlas-Coder-2-0.5B | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| # β‘ Atlas-Coder-2-0.5B GGUF | |
| > Quantized GGUF releases of **Atlas-Coder-2-0.5B**, optimized for **llama.cpp**, **Ollama**, and **LM Studio**. | |
| --- | |
| ## π EvalPlus Strict Benchmarks | |
| Atlas-Coder-2 is a **0.5B parameter coding model** fine-tuned on **50,000 execution-verified Python programming samples**. | |
| Despite its compact size, it achieves **state-of-the-art performance among strictly sub-1B coding models** on the EvalPlus benchmark suite. | |
| | Benchmark | Atlas-Coder-2 | Qwen2.5-Coder-0.5B | DeepSeek-Coder-1.3B | Llama-3.2-1B | | |
| |------------|--------------:|-------------------:|--------------------:|-------------:| | |
| | **HumanEval+** | **36.6% π₯** | 34.1% | 35.4% | 12.2% | | |
| | **MBPP+** | **43.9% π₯** | 42.1% | 39.8% | 25.6% | | |
| > **Evaluation:** EvalPlus (Pass@1, Greedy Decoding) | |
| --- | |
| ## π¦ Available Quantizations | |
| | File | Quantization | Size | Recommended For | | |
| |------|--------------|------|-----------------| | |
| | Atlas-Coder-2-0.5B-F16.gguf | F16 | 948 MB | Maximum accuracy & benchmarking | | |
| | Atlas-Coder-2-0.5B-Q8_0.gguf | Q8_0 | 506 MB | Near-lossless inference | | |
| | Atlas-Coder-2-0.5B-Q6_K.gguf | Q6_K | 482 MB | Best quality / size balance | | |
| | Atlas-Coder-2-0.5B-Q5_K_M.gguf | Q5_K_M | 401 MB | Recommended for most users β | | |
| | Atlas-Coder-2-0.5B-Q4_K_M.gguf | Q4_K_M | 379 MB | Low-memory devices | | |
| --- | |
| # π Quick Start | |
| ## Ollama | |
| Create a file named **Modelfile** | |
| ```text | |
| FROM ./Atlas-Coder-2-0.5B-Q5_K_M.gguf | |
| SYSTEM "You are Atlas-Coder, an elite AI coding assistant. You write clean, efficient, and well-documented Python code." | |
| PARAMETER temperature 0.2 | |
| PARAMETER top_p 0.95 | |
| PARAMETER repeat_penalty 1.1 | |
| PARAMETER stop "<|im_start|>" | |
| PARAMETER stop "<|im_end|>" | |
| ``` | |
| Create the model: | |
| ```bash | |
| ollama create atlas-coder-2 -f Modelfile | |
| ``` | |
| Run: | |
| ```bash | |
| ollama run atlas-coder-2 "Write a Python function to check whether a string is a palindrome." | |
| ``` | |
| --- | |
| ## LM Studio | |
| 1. Open **LM Studio** | |
| 2. Search for **Siddh07ETH/Atlas-Coder-2-0.5B-GGUF** | |
| 3. Download **Q5_K_M** or **Q6_K** | |
| 4. Load the model and start coding. | |
| --- | |
| ## llama.cpp | |
| ```bash | |
| ./llama-cli \ | |
| -m Atlas-Coder-2-0.5B-Q5_K_M.gguf \ | |
| -p "<|im_start|>system\nYou are a helpful coding assistant.<|im_end|>\n<|im_start|>user\nWrite a Python binary search implementation.<|im_end|>\n<|im_start|>assistant\n" \ | |
| -n 512 \ | |
| --temp 0.2 \ | |
| --top-p 0.95 \ | |
| --repeat-penalty 1.1 | |
| ``` | |
| --- | |
| # π¬ Chat Template | |
| Atlas-Coder-2 uses the standard **ChatML** prompt format. | |
| ```text | |
| <|im_start|>system | |
| {system_message} | |
| <|im_end|> | |
| <|im_start|>user | |
| {user_message} | |
| <|im_end|> | |
| <|im_start|>assistant | |
| ``` | |
| --- | |
| # π Open-Source Ecosystem | |
| **Base Model** | |
| - Siddh07ETH/Atlas-Coder-2-0.5B | |
| **Training Dataset** | |
| - Siddh07ETH/Atlas-Coder-50K-ChatML | |
| - 50,000 execution-verified Python instruction samples in ChatML format. | |
| --- | |
| # π¨βπ» Author | |
| **Siddharth N.R.** | |
| **Pluto AI Research** | |
| --- | |
| # π License | |
| This repository is released under the **Apache 2.0 License**. | |
| The base model (**Qwen2.5-Coder-0.5B-Instruct**) is also licensed under **Apache 2.0**. |