How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
# Run inference directly in the terminal:
llama cli -hf FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
# Run inference directly in the terminal:
llama cli -hf FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
# Run inference directly in the terminal:
./llama-cli -hf FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
# Run inference directly in the terminal:
./build/bin/llama-cli -hf FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
Use Docker
docker model run hf.co/FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF:NVFP4
Quick Links

Ornith 1.0 35B β€” NVFP4 GGUF

NVFP4 quantization of deepreinforce-ai/Ornith-1.0-35B, a 35B parameter Qwen3.5 MoE coding agent with 256 experts (8 active per token).

About the Model

Ornith-1.0-35B is the lightweight member of the Ornith family, designed for efficient single-GPU deployment.

  • State-of-the-Art Coding Agents: Post-trained on top of Qwen 3.5, achieving state-of-the-art performance among open-source models
  • Self-Improving Training Framework: Ornith-1.0 employs RL to learn to generate not only solution rollouts, but also the scaffold that drives those rollouts
  • 35B total parameters with 8B active per token (256 experts, 8 active)
  • 40-layer MoE architecture with sliding + full attention hybrid
  • 262K context window
  • MIT License β€” globally accessible, no regional limitations

Architecture

  • Text model: Qwen3.5 MoE β€” 40 layers, 2048 hidden, 256 experts (8 active/token)
  • Vocabulary: 248,320 tokens

Quantization

Quantized from the BF16 safetensors using llama.cpp (build 537).

NVFP4 (NVIDIA FP4) uses 4-bit floating point quantization optimized for NVIDIA Blackwell GPUs.

Files

File Size Description
ornith-1.0-35b-nvfp4.gguf ~18.4 GB NVFP4 quantized model

Usage

llama-server \
  -m ornith-1.0-35b-nvfp4.gguf \
  -ngl 99 \
  --host 0.0.0.0 \
  --port 8080

Hardware Requirements

  • Minimum: 20 GB VRAM
  • Recommended: 24+ GB VRAM for full GPU offload

License

MIT

Downloads last month
377
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for FreedomAISVR/Ornith-1.0-35B-NVFP4-GGUF

Quantized
(180)
this model