DJLougen commited on
Commit
cff411b
·
verified ·
1 Parent(s): ec648cf

Professional card: Fireworks training, Ornstein thinking, RL and energy-based FT roadmap

Browse files
Files changed (1) hide show
  1. README.md +29 -24
README.md CHANGED
@@ -11,38 +11,42 @@ pipeline_tag: image-text-to-text
11
  license: apache-2.0
12
  ---
13
 
 
 
14
  # Ornstein3.8-27B
15
 
16
- BF16 safetensors of **Ornstein3.8-27B** a rank-32 LoRA trained on [Fireworks AI](https://fireworks.ai) and merged into [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). Native vision-language model (`Qwen3_5ForConditionalGeneration`) with hybrid linear + full attention (Gated DeltaNet) and a 262K context window. Fine-tuning on Fireworks was straightforward.
17
 
18
- This checkpoint is meant to inject **Ornstein thinking** into Qwen3.8-27B. It is an early merge, not a finished quality release. Formal evals will be appended to this card as they land; further methods will drive the real quality gains.
19
 
20
- GGUF quants (Q8_0 / Q6_K / Q4_K_M) plus an mmproj live in **[GestaltLabs/Ornstein3.8-27B-GGUF](https://huggingface.co/GestaltLabs/Ornstein3.8-27B-GGUF)**.
21
 
22
- ## Support This Work
23
 
24
- I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my uploads have been useful to you, consider buying a PhD student a coffee. It goes a long way toward keeping these experiments running.
25
 
26
- **[Support on Ko-fi](https://ko-fi.com/djlougen)**
27
 
28
- ---
29
 
30
- ## Model info
31
 
32
- - **Architecture:** `Qwen3_5ForConditionalGeneration` (linear + full attention interleaved, Gated DeltaNet; native image/video)
33
- - **Parameters:** ~27 B dense
34
- - **Context:** 262,144 tokens (native)
35
- - **Hidden size / layers:** 5120 / 64
36
- - **Attention:** 24 heads, 4 KV heads, head_dim 256
37
- - **MLP intermediate:** 17,408
38
- - **Vocab:** 248,320
39
- - **Precision:** bfloat16 safetensors (11 shards)
40
- - **Vision:** SigLIP-style tower, `out_hidden_size` 5120, patch 16
41
- - **Post-training:** PEFT LoRA rank 32 / α 32 trained on [Fireworks AI](https://fireworks.ai) (`ft-hk031mimxqpvk`), then merged into language-model linears only (vision + MTP left at base)
 
 
42
 
43
  ## Usage
44
 
45
- Requires a recent Transformers with Qwen3.8 / `qwen3_5` support.
46
 
47
  ```python
48
  from transformers import AutoProcessor, AutoModelForImageTextToText
@@ -64,15 +68,15 @@ messages = [
64
  ]
65
  inputs = processor.apply_chat_template(
66
  messages, add_generation_prompt=True, tokenize=True,
67
- return_dict=True, return_tensors="pt"
68
  ).to(model.device)
69
  out = model.generate(**inputs, max_new_tokens=256)
70
  print(processor.decode(out[0], skip_special_tokens=True))
71
  ```
72
 
73
- Text-only chat works the same way with `{"type": "text", ...}` and no image.
74
 
75
- vLLM / SGLang: load this repo as a Qwen3.8 27B VLM (`qwen3_5`). Pin a build that already supports that architecture.
76
 
77
  ## Files
78
 
@@ -82,8 +86,9 @@ vLLM / SGLang: load this repo as a Qwen3.8 27B VLM (`qwen3_5`). Pin a build that
82
  | `model.safetensors.index.json` | weight map, `total_size` 55562855904 |
83
  | `config.json` | `Qwen3_5ForConditionalGeneration` |
84
  | `tokenizer.json` / `tokenizer_config.json` / `vocab.json` / `merges.txt` | tokenizer |
85
- | `chat_template.jinja` | chat + vision + tool-call template |
86
  | `preprocessor_config.json` / `video_preprocessor_config.json` | image/video processor |
 
87
 
88
  ## Related
89
 
@@ -93,4 +98,4 @@ vLLM / SGLang: load this repo as a Qwen3.8 27B VLM (`qwen3_5`). Pin a build that
93
 
94
  ## License
95
 
96
- Apache 2.0 inherited from the Qwen 3.8 base release.
 
11
  license: apache-2.0
12
  ---
13
 
14
+ ![Ornstein3.8-27B](ornstein3.8-27b.jpg)
15
+
16
  # Ornstein3.8-27B
17
 
18
+ BF16 safetensors for **Ornstein3.8-27B**, a vision-language fine-tune of [Qwen/Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B). Architecture is `Qwen3_5ForConditionalGeneration`: interleaved linear and full attention (Gated DeltaNet), native image/video, 262K context.
19
 
20
+ The LoRA was trained on [Fireworks AI](https://fireworks.ai) and merged into the Qwen3.8-27B language stack. Quantized GGUFs (Q8_0, Q6_K, Q4_K_M) and an mmproj are in [GestaltLabs/Ornstein3.8-27B-GGUF](https://huggingface.co/GestaltLabs/Ornstein3.8-27B-GGUF).
21
 
22
+ ## Status
23
 
24
+ This checkpoint injects **Ornstein thinking** into Qwen3.8-27B. It is an early merge, not a finished quality release.
25
 
26
+ This card will be updated with formal evaluation as results land. Planned quality work uses **RL environments** and **energy-based fine-tuning**.
27
 
28
+ ## Support this work
29
 
30
+ I'm a PhD student in visual neuroscience at the University of Toronto. Training and release compute is self-funded (rented H100s and a local DGX Spark). If these artifacts are useful, [Ko-fi](https://ko-fi.com/djlougen) helps keep the experiments running.
31
 
32
+ ## Model details
33
 
34
+ | | |
35
+ |---|---|
36
+ | Architecture | `Qwen3_5ForConditionalGeneration` |
37
+ | Parameters | ~27B dense |
38
+ | Context | 262,144 tokens |
39
+ | Hidden size / layers | 5120 / 64 |
40
+ | Attention | 24 heads, 4 KV heads, head_dim 256 |
41
+ | MLP intermediate | 17,408 |
42
+ | Vocab | 248,320 |
43
+ | Precision | bfloat16, 11 shards |
44
+ | Vision | SigLIP-style tower, `out_hidden_size` 5120, patch 16 |
45
+ | Post-training | PEFT LoRA rank 32, α 32, trained on [Fireworks AI](https://fireworks.ai); merged into language-model linears only (vision and MTP unchanged) |
46
 
47
  ## Usage
48
 
49
+ Requires a Transformers build with Qwen3.8 / `qwen3_5` support.
50
 
51
  ```python
52
  from transformers import AutoProcessor, AutoModelForImageTextToText
 
68
  ]
69
  inputs = processor.apply_chat_template(
70
  messages, add_generation_prompt=True, tokenize=True,
71
+ return_dict=True, return_tensors="pt",
72
  ).to(model.device)
73
  out = model.generate(**inputs, max_new_tokens=256)
74
  print(processor.decode(out[0], skip_special_tokens=True))
75
  ```
76
 
77
+ Text-only chat uses the same template with `{"type": "text", ...}` and no image.
78
 
79
+ vLLM and SGLang: load this repo as a Qwen3.8 27B VLM (`qwen3_5`). Use a build that already supports that architecture.
80
 
81
  ## Files
82
 
 
86
  | `model.safetensors.index.json` | weight map, `total_size` 55562855904 |
87
  | `config.json` | `Qwen3_5ForConditionalGeneration` |
88
  | `tokenizer.json` / `tokenizer_config.json` / `vocab.json` / `merges.txt` | tokenizer |
89
+ | `chat_template.jinja` | chat, vision, and tool-call template |
90
  | `preprocessor_config.json` / `video_preprocessor_config.json` | image/video processor |
91
+ | `ornstein3.8-27b.jpg` | card banner |
92
 
93
  ## Related
94
 
 
98
 
99
  ## License
100
 
101
+ Apache 2.0, inherited from the Qwen 3.8 base release.