4-bit quant using Intel AutoRound

This is the default W4A16 scheme.

Quantization Details

at time of quantization, the default implied values not listed in the below json are as follows:

{"batch_size": 8, "iters": 200, "seqlen": 2048, "nsamples": 128, "lr": None}

quantization_config.json

{
  "bits": 4,
  "group_size": 128,
  "sym": true,
  "data_type": "int",
  "low_gpu_mem_usage": true,
  "autoround_version": "0.9.2",
  "quant_method": "auto-round",
  "packing_format": "auto_round:auto_gptq"
}
Downloads last month
11
Safetensors
Model size
5B params
Tensor type
I32
BF16
F16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for hoborific/WeirdCompound-v1.6-24b-W4A16-AutoRound

Quantized
(9)
this model