FoolDev Claude Fable 5 commited on
Commit
834f5fe
·
1 Parent(s): 98074bd

docs(readme): warn that /v1 can't override num_ctx (OOMs on small hosts)

Browse files

The OpenAI-compatible /v1/chat/completions examples load at the baked
num_ctx 1010000 (v1 has no override knob), which OOMs on <=~48 GB hosts —
the KV cache alone wants ~62 GB. Added a callout beside the examples: use
/api/chat with options.num_ctx, or /save a small-context local tag for
OpenAI clients.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

Files changed (2) hide show
  1. CHANGELOG.md +7 -0
  2. README.md +7 -0
CHANGELOG.md CHANGED
@@ -65,6 +65,13 @@ track the **tooling and documentation**, not the underlying base model.
65
  theoretical footprint at the default is recomputed to ~62 GB KV / ~100 GB total.
66
 
67
  ### Fixed
 
 
 
 
 
 
 
68
  - **`llama-cpp-python` example matches its CPU-only install.** The
69
  `examples/README.md` llama-cpp-python quickstart installed the CPU-only wheel
70
  but then ran with `--gpu-layers 99` — a silent no-op on that build, and a
 
65
  theoretical footprint at the default is recomputed to ~62 GB KV / ~100 GB total.
66
 
67
  ### Fixed
68
+ - **OpenAI `/v1` examples now warn about the `num_ctx` OOM on small hosts.**
69
+ `/v1/chat/completions` has no `num_ctx` knob, so it loads at the baked
70
+ `num_ctx 1010000` and OOMs on ≤~48 GB hosts (the KV cache alone wants ~62 GB) —
71
+ but the examples showed `/v1` with no caveat. Added a callout beside the
72
+ OpenAI-compatible examples: use `/api/chat` with `"options": {"num_ctx": 4096}`,
73
+ or `/save` a small-context local tag for OpenAI clients (`/v1` can't pass the
74
+ override inline).
75
  - **`llama-cpp-python` example matches its CPU-only install.** The
76
  `examples/README.md` llama-cpp-python quickstart installed the CPU-only wheel
77
  but then ran with `--gpu-layers 99` — a silent no-op on that build, and a
README.md CHANGED
@@ -159,6 +159,13 @@ Once the model is loaded (via `ollama run janus`, `lms server`, or `llama-server
159
 
160
  > The examples use `model: "janus"`, the tag from the local build (path B). If you pulled via the TL;DR one-liner instead, use the full tag `hf.co/FoolDev/Janus-35B-HERETIC`, or run `ollama cp hf.co/FoolDev/Janus-35B-HERETIC janus` once to create the short tag.
161
 
 
 
 
 
 
 
 
162
  #### curl
163
 
164
  ```bash
 
159
 
160
  > The examples use `model: "janus"`, the tag from the local build (path B). If you pulled via the TL;DR one-liner instead, use the full tag `hf.co/FoolDev/Janus-35B-HERETIC`, or run `ollama cp hf.co/FoolDev/Janus-35B-HERETIC janus` once to create the short tag.
161
 
162
+ > **On ≤~48 GB hosts, cap `num_ctx` first.** `/v1/chat/completions` (OpenAI-compat) has
163
+ > no `num_ctx` knob, so it loads at the baked **1,010,000** default and OOMs — the KV
164
+ > cache alone wants ~62 GB (see [Hardware requirements](#hardware-requirements)). Either
165
+ > call `/api/chat` with `"options": {"num_ctx": 4096}`, or bake a small-context tag for
166
+ > OpenAI clients: `ollama run hf.co/FoolDev/Janus-35B-HERETIC:Q4_K_M`, then
167
+ > `/set parameter num_ctx 4096` and `/save janus`, and point clients at `janus`.
168
+
169
  #### curl
170
 
171
  ```bash