I appreciate your merges, @nightmedia - even if we have to fix the MTP ๐
Nick M
veldierin
AI & ML interests
None yet
Recent Activity
new activity 2 days ago
nightmedia/Qwen3.8-27B-Brainwaves-Heretic-mxfp4-mlx:Gemini trace analysisOrganizations
None yet
New activity in DavidAU/Qwen3.8-27B-Cold-Fable-Fusion-GAIN-V1.1-732-Heretic-Uncensored-stage1 about 22 hours ago
Was this finetune done to retain the level of thinking of the base so we can have proper xhigh reasoning?
1
#6 opened about 22 hours ago
by
veldierin
Gemini trace analysis
12
#1 opened 2 days ago
by
nightmedia
MTP enabled: unbounded VRAM growth during inference until CUDA abort (V100)
4
#4 opened 2 days ago
by
fozosan
Model stops answering prematurely
โ 1
2
#1 opened 3 days ago
by
theblackcat
Great model so far
โค๏ธ๐ฅ 3
6
#17 opened 5 days ago
by
mythrime
request for re-quantization w/fixed MTP heads
๐ 1
4
#3 opened 3 days ago
by
veldierin
How it compares to Ornith 1.5 35B?
2
#2 opened 3 days ago
by
MartinPatterson
Mtp quality? ๐ค
5
#7 opened 11 days ago
by
LinkuStarto
APEX quant
10
#1 opened 8 days ago
by
benoe
Waiting for Ornith-1.5-35B-Heretic-MTP-APEX-GGUF
3
#2 opened 6 days ago
by
atfa
VERSION NVFP4 : Would a nvfp4 version possible ?
โ 1
10
#14 opened 9 days ago
by
crazyhenres
Testing it on greenboost-cli with good results
โค๏ธ๐ 2
8
#15 opened 7 days ago
by
hyphaed
replied to nightmedia's post 7 days ago
reacted to nightmedia's post with ๐ 8 days ago
Post
4094
Qwen3.8-27B metrics
It's hard to track all model cards where I post these, so I figured people would get more value out of seeing these in the open.
The performance is as measured on a M4 MBP 128GB, speed may vary depending on your platform.
These are all instruct metrics, generated by including this line in the jinja template:
Then run the test suite to generate the metrics:
This will generate the file:
This is a JSON containing all gathered metrics; for example the q4-hi:
I use the value of acc_norm for metrics, rounded to 3 decimals.
As I get more quants tested, I will add them here.
A complete test run for a single quant takes 7-9 hours depending on quant size, 10-12 hours for BF16 depending on perplexity: this is why you see on my model cards that I usually post the first three, that only take 2-3 hours :)
-G
It's hard to track all model cards where I post these, so I figured people would get more value out of seeing these in the open.
quant arc arc/e boolq hswag obkqa piqa wino
mxfp8 0.591,0.782,0.896,0.746,0.448,0.801,0.711
q8-hi 0.602,0.779,0.896,0.747,0.446,0.793,0.703
q6-hi 0.602,0.775,0.895,0.748,0.448,0.795,0.710
q4-hi 0.604,0.780,0.898,0.744,0.454,0.795,0.708
mxfp4 0.581,0.771,0.889,0.738,0.442,0.798,0.713
1M
mxfp8 0.590,0.787,0.897,0.744,0.446,0.801,0.709
Quant Perplexity Peak Memory Tokens/sec
mxfp8 6.090 ยฑ 0.054 34.74 GB 138
mxfp4 5.952 ยฑ 0.051 21.30 GB 148The performance is as measured on a M4 MBP 128GB, speed may vary depending on your platform.
These are all instruct metrics, generated by including this line in the jinja template:
{%- set enable_thinking = false %}Then run the test suite to generate the metrics:
mlx_lm.evaluate --model MODEL --tasks winogrande boolq arc_challenge arc_easy hellaswag openbookqa piqaThis will generate the file:
eval_MODEL_0.4.9_winogrande_boolq_arc_challenge_arc_easy_hellaswag_openbookqa_piqaThis is a JSON containing all gathered metrics; for example the q4-hi:
"arc_challenge": {
"alias": "arc_challenge",
"acc,none": 0.5819112627986348,
"acc_stderr,none": 0.014413988396996116,
"acc_norm,none": 0.6040955631399317,
"acc_norm_stderr,none": 0.01429122839353657
},I use the value of acc_norm for metrics, rounded to 3 decimals.
As I get more quants tested, I will add them here.
A complete test run for a single quant takes 7-9 hours depending on quant size, 10-12 hours for BF16 depending on perplexity: this is why you see on my model cards that I usually post the first three, that only take 2-3 hours :)
-G
Introducing Unsloth Dynamic v3 Qwen3.8
๐๐ฅ 51
33
#74 opened 8 days ago
by
danielhanchen
Brainwaves MTP Head: Recipe-Reconstruction Approach
โค๏ธ๐ง 3
3
#2 opened 8 days ago
by
veldierin
Status on repo - ready to test now
๐โค๏ธ 1
10
#1 opened 8 days ago
by
veldierin
Plans for bf16 safetensors ?
โค๏ธ 1
6
#2 opened 9 days ago
by
veldierin
Could you swap the base qwen3.8-27b with a properly heretic'ized one for better uncensoring?
13
#3 opened 10 days ago
by
veldierin