Instructions to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Use Docker
docker model run hf.co/v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
- Ollama
How to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with Ollama:
ollama run hf.co/v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with Docker Model Runner:
docker model run hf.co/v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
- Lemonade
How to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Alice Qwen3 4B Instruct 2507 Heretic Light GGUF
Alice Portable GGUF is a cross-platform Alice build for llama.cpp, Ollama, LM Studio, Windows, Linux, macOS, and phone GGUF clients.
Base model:
p-e-w/Qwen3-4B-Instruct-2507-heretic- GGUF source:
logos-flux/Qwen3-4B-Instruct-2507-heretic-GGUF
This release keeps the heretic/abliterated uncensored base behavior and adds a light Alice persona directly in the GGUF chat template metadata. It is not a LoRA and not a safety-policy tuning pass.
Files
Alice-Qwen3-4B-Instruct-2507-Heretic-Light-Q4_K_M.gguf
Q4_K_M is the first portable release because it is the best balance for phones, miner nodes, and regular Windows/Linux machines.
Intended Behavior
- Chinese/English casual chat.
- Story writing and roleplay.
- Alice identity by default.
- User-requested rename/role changes are accepted.
- Cross-platform local inference through GGUF runtimes.
This 4B Q4 build is not the serious Solidity/code model. It can answer simple code prompts, but larger Alice lanes should handle deeper code/security work.
llama.cpp
llama-cli -m Alice-Qwen3-4B-Instruct-2507-Heretic-Light-Q4_K_M.gguf \
-p "你好" -st -n 256 --temp 0.55 --top-p 0.8 --reasoning off
Server mode:
llama-server -m Alice-Qwen3-4B-Instruct-2507-Heretic-Light-Q4_K_M.gguf \
-c 32768 --temp 0.55 --top-p 0.8
Ollama
Use the included Modelfile:
ollama create alice-qwen3-4b-heretic-light -f Modelfile
ollama run alice-qwen3-4b-heretic-light
Smoke Test
Tested locally with llama.cpp b9290:
你是谁 -> 我是 Alice。
你叫 eva 吧。 -> 好,我叫 Eva。
你好 -> 嗨!今天过得怎么样?
假设你是我的女朋友,今天我很累 -> enters companion roleplay naturally
Why Not Qwen3.5 GGUF
Qwen3.5 GGUF conversion and quantization were tested locally, but current llama.cpp support produced corrupted output for that architecture. This portable release uses Qwen3 Instruct 2507 instead because it runs correctly in standard GGUF runtimes.
- Downloads last month
- 98
4-bit
Model tree for v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF
Base model
p-e-w/Qwen3-4B-Instruct-2507-heretic
docker model run hf.co/v102ss/Alice-Qwen3-4B-Instruct-2507-Heretic-Light-GGUF:Q4_K_M