Are these quants using Unsloth UD 3.0 or 2.0?

#20
by InfernalDread - opened

Just wanted to clarify it it was a mistake, since it was added to the UD v2.0 collection rather than the UD v3.0

Unsloth AI org

It's half way. V2.5 I'd say. V3 needs slightly more investigation which is why we still haven't uploaded Ornith yet.

It's half way. V2.5 I'd say. V3 needs slightly more investigation which is why we still haven't uploaded Ornith yet.

Something is weird with the new Qwen3.8 quants. Even on simple tasks they never stop thinking. Currently I’m using an exl3 quant which stops thinking almost instantly on simple tasks and thinks for a couple min on harder ones. The Qwen3.8 V3 was thinking for 10x longer and eventually if it ever did stop it would just get the wrong answer anyway. I observed this from Q4 up to Q8.

@Gaboo It really looks like a reasoning effort difference. Didn't you test different templates during that time? I'm saying this because for example the froggeric (and some others based on froggeric one) change the thinking behavior from stock. I didn't look if they change all the reasoning effort prompts, but I remember medium has a prompt for example, while it doesn't at all in the stock template (or Unsloth's one).
Or didn't you test ggufs in non agentic context vs exl3 in agentic? Because they are shorter in agentic on average (at least with preserve thinking, I actually didnt try without). I can really be like just few lines even in xhigh while almost always 10x longer than that in simple chat. Pretty sure you noticed it but just trying to help with what I know

@Gaboo It really looks like a reasoning effort difference. Didn't you test different templates during that time? I'm saying this because for example the froggeric (and some others based on froggeric one) change the thinking behavior from stock. I didn't look if they change all the reasoning effort prompts, but I remember medium has a prompt for example, while it doesn't at all in the stock template (or Unsloth's one).
Or didn't you test ggufs in non agentic context vs exl3 in agentic? Because they are shorter in agentic on average (at least with preserve thinking, I actually didnt try without). I can really be like just few lines even in xhigh while almost always 10x longer than that in simple chat. Pretty sure you noticed it but just trying to help with what I know

I use exl3 and I also noticed it thinks 10x less than Unsloth. I am starting to highly suspect the Unsloth template as to what is causing the loops, not just the quantization. I tried the froggeric's template on unsloth and it seems a lot more well behaved.

Sign up or log in to comment