file size (quantization)

#16
by jacek2024 - opened

Q1S is 10+50+22 about 82GB

Q4_K_XL is 10+50+50+12 about 122GB

is this like GPT-OSS, some parts are quantized differently?

yes, he down projection is MOD128 which is incompatible with IQ1_S

Unsloth AI org

It's Qwen's architecture, you can read here: https://unsloth.ai/docs/models/qwen3.8-next#quantization-analysis

shimmyshimmer changed discussion status to closed

Sign up or log in to comment