aleada/Pixtral-12B-W4A16
Image-Text-to-Text • 13B • Updated • 91 • 1
Weight-only INT4 for RTX 3090/4090-class cards, where FP8 and NVFP4 are emulated or unusable. Full recipe on every card.
Note W4A16 for Ampere Marlin kernels, with the MTP head grafted back and exempted from the quantization config. Ships a measured draft-acceptance rate (86.5%, 1.73x decode) rather than an assumption - other 4-bit builds of this model retain the head but none publish what it actually accepts.