LFM2.5-2.6B Q4_K_M Fast — Windows CUDA

This repository contains a directly runnable Q4_K_M GGUF of LiquidAI/LFM2.5-2.6B-GGUF and a validated Fast single-request Windows/CUDA profile for llama.cpp.

The model weights are not modified. The same verified GGUF is published in both paired repositories; only the tested runtime profile differs.

Fast profile: inference is accelerated relative to the same Q4_K_M baseline without the Fast runtime settings. The frozen validation gate detected no quality regression and no new failures. This is a measured result for the documented hardware, workloads, and single-request setup—not a universal guarantee for every prompt or runtime.

Measured Fast result

Primary metric: wall-clock decoded tokens per second for one request, without batching. The profile validation used three workloads with five repetitions each (15 runs total, 256 generated tokens per run).

Workload Baseline, tok/s Fast, tok/s Fast vs baseline
Code copy 114.74 226.27 1.972× (+97.2%)
Editorial rewrite 112.09 141.26 1.260× (+26.0%)
Technical summary 113.30 133.67 1.180× (+18.0%)
All 15 runs, mean ± SD 113.38 ± 1.56 167.07 ± 43.51 1.474× (+47.4%)
Independent quality gate Baseline Fast Regression
Passed tasks 10/12 10/12 None measured

The larger Fast standard deviation reflects the deliberately mixed workload set: repetitive code benefits more than free-form editing and summarization. These are profile-validation measurements, not the pending frozen cross-machine benchmark.

Choose the matching profile

Platform Hardware/backend Repository
Windows 11 NVIDIA RTX / CUDA This repository
Ubuntu AMD Strix Halo / Vulkan LFM2.5-2.6B-Ubuntu-Strix-Halo-Vulkan-GGUF

Included weight

File Quantization Size SHA-256
LFM2.5-2.6B-Q4_K_M.gguf Q4_K_M 1,674,454,848 bytes (1.56 GiB) 79fdf00351b46cf26f020aead28d01889886be87c55fa0eb907e6f9b00bfee14

Source revision: b22e29ebf6249a8c9fcdda36914743e9980595c4.

Tested setup

  • Windows 11 laptop
  • NVIDIA GeForce RTX 4060 Laptop GPU, 8 GiB VRAM
  • Intel Core i7-13650HX, 64 GiB RAM
  • NVIDIA driver 591.74
  • official llama.cpp CUDA server container, build 10066 (86a9c79f8)
  • context 8,192, one parallel slot, continuous batching disabled

Download and verify

Install the Hugging Face CLI once, or download the GGUF with the file link above.

py -m pip install -U huggingface_hub

$ModelDir = "C:\Models\LFM2.5-2.6B"
New-Item -ItemType Directory -Force -Path $ModelDir | Out-Null

hf download petr567/LFM2.5-2.6B-Windows-RTX-CUDA-GGUF `
  LFM2.5-2.6B-Q4_K_M.gguf `
  --local-dir $ModelDir

$Expected = "79fdf00351b46cf26f020aead28d01889886be87c55fa0eb907e6f9b00bfee14"
$Actual = (Get-FileHash "$ModelDir\LFM2.5-2.6B-Q4_K_M.gguf" -Algorithm SHA256).Hash.ToLower()
if ($Actual -ne $Expected) { throw "GGUF SHA-256 mismatch" }

Run the validated Fast Windows/CUDA profile

Docker Desktop must be configured for NVIDIA GPU access.

$ModelDir = "C:\Models\LFM2.5-2.6B"

docker run --rm --gpus all `
  -p 127.0.0.1:8080:8080 `
  -v "${ModelDir}:/models:ro" `
  --entrypoint /app/llama-server `
  ghcr.io/ggml-org/llama.cpp@sha256:1b3d1458ccda7287feab41b8001311acc03e24cde99ec0a2908fe83830562f38 `
  -m /models/LFM2.5-2.6B-Q4_K_M.gguf `
  --alias lfm2.5-2.6b-q4_k_m `
  --host 0.0.0.0 --port 8080 `
  -c 8192 -np 1 -ngl 99 `
  -t 14 -tb 14 -b 2048 -ub 512 `
  -fa auto -ctk q8_0 -ctv q8_0 `
  --no-cont-batching --no-cache-prompt --cache-ram 0 `
  --slot-prompt-similarity 0 --jinja --no-webui `
  --spec-type ngram-simple `
  --spec-ngram-simple-size-n 8 `
  --spec-ngram-simple-size-m 64 `
  --spec-ngram-simple-min-hits 1 `
  --spec-draft-n-max 64

The OpenAI-compatible endpoint is available at http://127.0.0.1:8080/v1.

$Body = @{
  model = "lfm2.5-2.6b-q4_k_m"
  messages = @(@{ role = "user"; content = "Write a short hello-world function in Python." })
  max_tokens = 128
  temperature = 0.2
} | ConvertTo-Json -Depth 5

Invoke-RestMethod -Method Post `
  -Uri "http://127.0.0.1:8080/v1/chat/completions" `
  -ContentType "application/json" `
  -Body $Body

LM Studio

The GGUF itself can also be opened in LM Studio. Use an 8,192-token context, maximum GPU offload, and one parallel request. The exact ngram-simple profile above requires a compatible llama.cpp server build; do not assume an arbitrary GUI runtime exposes the same acceleration controls.

Release scope

This release contains the runnable weight and the final launch recipe. The frozen cross-machine benchmark package and its results will be attached in a later revision after verification.

Attribution and license

The license includes a commercial-use revenue threshold. Review the included license before use or redistribution.

Downloads last month
96
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for petr567/LFM2.5-2.6B-Windows-RTX-CUDA-GGUF

Quantized
(3)
this model