Instructions to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M # Run inference directly in the terminal: llama cli -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M # Run inference directly in the terminal: llama cli -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Use Docker
docker model run hf.co/simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
- SGLang
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with Ollama:
ollama run hf.co/simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
- Unsloth Desktop
- Pi
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with Docker Model Runner:
docker model run hf.co/simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
- Lemonade
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Run and chat with the model
lemonade run user.poc-simorg-coder-30b-a3b-qwen3-lora-v0.1-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Hugging Face publish manifest
Complete list of files in this directory for upload to Hugging Face.
Regenerate staging (hardlinks + copies):
uv run python scripts/prepare_huggingface_publish.py
Upload entire folder (~183 GB) β use upload-large-folder:
$env:HF_XET_HIGH_PERFORMANCE = "1"
uv run hf upload-large-folder simorg-platform/poc-simorg-coder-30b-a3b-qwen3-lora-v0.1 .\huggingface\ `
--repo-type model --num-workers 8
Do not use plain hf upload for the full folder β HF will warn it may fail. See docs/publish.md.
Root files (required)
| File | Purpose |
|---|---|
README.md |
Model card (YAML front matter + usage) |
LICENSE |
Apache License 2.0 (from Qwen base model) |
NOTICE |
Derivative-work attribution |
.gitattributes |
Git LFS rules for *.safetensors and *.gguf |
MANIFEST.md |
This file list |
safetensors/ β merged full model (~60 GB)
All files from merged LoRA + base weights. No scripts, no tokens.
| File |
|---|
config.json |
generation_config.json |
model.safetensors.index.json |
model-00001-of-00013.safetensors |
model-00002-of-00013.safetensors |
model-00003-of-00013.safetensors |
model-00004-of-00013.safetensors |
model-00005-of-00013.safetensors |
model-00006-of-00013.safetensors |
model-00007-of-00013.safetensors |
model-00008-of-00013.safetensors |
model-00009-of-00013.safetensors |
model-00010-of-00013.safetensors |
model-00011-of-00013.safetensors |
model-00012-of-00013.safetensors |
model-00013-of-00013.safetensors |
tokenizer.json |
tokenizer_config.json |
vocab.json |
merges.txt |
added_tokens.json |
special_tokens_map.json |
chat_template.jinja |
Excluded: training scripts, workspace/ cache, HF tokens.
gguf/ β quantized models (~121 GB total)
| File | Approx. size | Notes |
|---|---|---|
qwen3-coder-simorg-Q3_K_M.gguf |
~14 GB | Smallest |
qwen3-coder-simorg-Q4_K_S.gguf |
~16 GB | |
qwen3-coder-simorg-Q4_K_M.gguf |
~17 GB | Recommended |
qwen3-coder-simorg-Q5_K_M.gguf |
~20 GB | |
qwen3-coder-simorg-Q6_K.gguf |
~23 GB | |
qwen3-coder-simorg-Q8_0.gguf |
~30 GB |
Excluded: qwen3-coder-simorg-f16.gguf (local intermediate only; use safetensors/ for full weights). Legacy qwen-simorg-f16.gguf (older Qwen2.5 build).
lora/ β PEFT adapter (~1β2 GB)
| File | Purpose |
|---|---|
adapter_config.json |
LoRA configuration |
adapter_model.safetensors |
LoRA weights |
tokenizer.json |
Tokenizer |
tokenizer_config.json |
|
vocab.json |
|
merges.txt |
|
added_tokens.json |
|
special_tokens_map.json |
|
chat_template.jinja |
Inference chat template |
Excluded: model.safetensors (Ollama-only duplicate), auto-generated PEFT README.md, scripts.
training-data/ β SFT dataset
| File | Entries | Format |
|---|---|---|
train-global.json |
131 | [{"question": "...", "answer": "..."}, ...] |
Add future datasets here before re-running prepare_huggingface_publish.py.
ollama/ β Modelfiles (6 files)
| File | GGUF reference |
|---|---|
Modelfile-qwen-3-Q3-K-M |
../gguf/qwen3-coder-simorg-Q3_K_M.gguf |
Modelfile-qwen-3-Q4-K-S |
../gguf/qwen3-coder-simorg-Q4_K_S.gguf |
Modelfile-qwen-3-Q4-K-M |
../gguf/qwen3-coder-simorg-Q4_K_M.gguf |
Modelfile-qwen-3-Q5-K-M |
../gguf/qwen3-coder-simorg-Q5_K_M.gguf |
Modelfile-qwen-3-Q6-K |
../gguf/qwen3-coder-simorg-Q6_K.gguf |
Modelfile-qwen-3-Q8-0 |
../gguf/qwen3-coder-simorg-Q8_0.gguf |
Paths are relative β run ollama create from the ollama/ directory after clone.
Security checklist
Before upload, confirm:
- No
HF_TOKENorhf_...secrets in any staged file - No
.envor credential files - No Python/shell scripts in this folder
-
adapter_config.jsononly references public base model ID
Total upload size (approximate)
| Section | Size |
|---|---|
safetensors/ |
~60 GB |
gguf/ |
~121 GB |
lora/ |
~1β2 GB |
training-data/ |
<1 MB |
ollama/ |
<1 MB |
| Total | ~183 GB |
Use uv run hf upload or Git LFS; stable connection recommended.