Please provide the nvfp4 version.

#3
by 88hoon - opened

Please provide the nvfp4 version.

Hey,

thanks for the interest in the model, i am not really into making quantizations. With lmcompressor you can do that easily on the hardware you are planning to you, it supports a "layer by layer" compression, so you would not need to have the whole model in VRAM the whole time to do the quantization. i recommend at least 512 calibration samples.

I published the BF16 so ppl can feel free to quant as they like.
Secondly, my HF space is limited ... so i am selective about generating quants

Thank you for your reply.

I understand your situation regarding quantization and HF storage limits. I really appreciate you taking the time to explain it.

Thanks again for releasing the BF16 model.

Sign up or log in to comment