File size: 11,575 Bytes
9182da3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
---
license: other
license_name: minimax-h3-community-license-agreement
license_link: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE
pipeline_tag: image-text-to-video
base_model:
  - MiniMaxAI/MiniMax-H3
  - Comfy-Org/MiniMax-H3
tags:
  - minimax-h3
  - ref2va
  - comfyui
  - diffusion-single-file
  - int8-convrot
  - synchronized-audio-video
  - experimental
language:
  - en
  - zh
inference: false
---

# MiniMax H3 Ref2VA Patch-In HF 1.02

Experimental ComfyUI single-file checkpoint derived from
[MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3) Ref2VA and the
[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3) INT8 ConvRot repack.

> [!CAUTION]
> This is an experimental weight modification, not an official MiniMax release and not a proven
> “de-oil”, “de-wax”, restoration, or quality-fix model. The original checkpoint remains the
> recommended default.

> [!IMPORTANT]
> MiniMax H3 is governed by the
> [MiniMax H3 Community License Agreement](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE),
> which defines excluded territories and mandatory redistribution conditions. Before publishing or
> redistributing this derivative, include the official `LICENSE`, keep the modification notice and
> `NOTICE`, and confirm that the intended distribution method and audience are authorized. A public
> Hugging Face repository may be reachable from excluded territories; a repository gate alone is not
> necessarily geographic access control. This model card is not legal advice.

## 中文说明

这是 MiniMax H3 Ref2VA 的实验性 ComfyUI 单文件衍生模型。它没有训练、微调或蒸馏,只对
视频 patch 输入投影中的 2×2 空间高频分量增加 2%。两组固定条件测试都出现了很弱的皮肤
高频代理正增益,但肉眼仍未确认能够消除 Ref2VA 的油感或蜡感,所以只能作为 EXP 对照模型。

发布前必须阅读上游许可证。该许可证对适用地域、公开分发、商业使用、安全措施、修改声明、
`LICENSE``NOTICE` 都有要求。

## Model description

| Property | Value |
|---|---|
| Task family | MiniMax H3 Base Ref2VA / reference-to-audio-video |
| Format | ComfyUI single-file diffusion model |
| Precision | Native INT8 ConvRot container inherited from the source repack |
| Training | None |
| Modified tensor | `video_patch_proj.weight` only |
| Input patch HF gain | `1.02` |
| Output-head gain | `1.00` — unchanged |
| File size | `34,038,894,550` bytes, approximately 31.70 GiB |
| Status | Experimental A/B candidate |

### What was changed

MiniMax H3 video latents are patchified with a `1 × 2 × 2` patch. For every latent channel, the four
spatial input columns of `video_patch_proj.weight` were transformed in an orthonormal 2×2 Haar basis:

- the DC/common component is preserved at gain `1.00`;
- the three non-DC H/V/D components are multiplied by `1.02`;
- the result is transformed back and written into the input projection;
- no output head, audio tensor, shared Transformer block, VAE, text encoder, or ComfyUI node was changed.

The equivalent 4×4 transform has diagonal `1.015` and off-diagonal `-0.005`. Its all-ones/DC
eigenvector has gain `1.00`; the three orthogonal spatial-detail eigenvectors have gain `1.02`.

This is not an output sharpening filter. It changes the video latent input projection used during
each joint audio-video denoising forward pass. It cannot reconstruct real texture that the model does
not generate.

## Checkpoint integrity

### Source

```text
File: minimax_h3_ref2va_int8_convrot.safetensors
Repository: Comfy-Org/MiniMax-H3
Size: 34,038,894,550 bytes
SHA-256: 9eef934046a0671bc8a5daf87100705e1478419c574cfde70c50fbe6885f76a9
```

The source size and SHA-256 match the LFS object published by
[Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3/blob/main/diffusion_models/minimax_h3_ref2va_int8_convrot.safetensors).

### Modified checkpoint

```text
File: minimax_h3_ref2va_patchin_hf102.safetensors
Size: 34,038,894,550 bytes
SHA-256: 5552621452ff73c9627ece6bd08f08e8c08b6504ee8a753b1aa1e1abcb70d9e4
```

Structural verification:

- safetensors header identical to the source: yes;
- tensor keys and shapes identical to the source: yes;
- changed tensor: `video_patch_proj.weight` only;
- changed byte interval: `[17,879,061,191, 17,881,125,575)`;
- target writeback maximum absolute error: `0`;
- DC-sum maximum absolute error: `9.5367431640625e-07`;
- SHA-256 of all unchanged file regions:
  `82d4866debf06c9b9b7a67d645d1dcac57bdd91eba3b341a55a4e1fe9f687852`.

The machine-readable surgery record is available at [`evidence/surgery.json`](evidence/surgery.json).

## ComfyUI usage

### Required model components

This repository contains only the modified Ref2VA diffusion model. Obtain the remaining H3 components
from the official or ComfyUI repack repositories under their applicable licenses:

- MiniMax H3-compatible Qwen3-VL text encoder;
- `minimax_h3_video_vae_fp16.safetensors`;
- `minimax_h3_audio_vae_fp32.safetensors`;
- a current ComfyUI build with MiniMax H3 and native INT8 ConvRot support.

Suggested layout:

```text
ComfyUI/
└── models/
    ├── diffusion_models/
    │   └── minimax_h3_ref2va_patchin_hf102.safetensors
    ├── text_encoders/
    │   └── qwen3vl_32b_minimax_h3_*.safetensors
    └── vae/
        ├── minimax_h3_video_vae_fp16.safetensors
        └── minimax_h3_audio_vae_fp32.safetensors
```

Load the checkpoint with ComfyUI's **Load Diffusion Model** node and use a Ref2VA/R2V workflow. The
official ComfyUI template is available at
[video_minimax_h3_r2v.json](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json).

### Recommended comparison protocol

For a meaningful A/B, change only the diffusion model:

1. generate once with `minimax_h3_ref2va_int8_convrot.safetensors`;
2. generate again with `minimax_h3_ref2va_patchin_hf102.safetensors`;
3. keep seed, reference media, prompt, dimensions, frame count, sampler, scheduler, shifts, steps, VAEs,
   text encoder and conditioning identical;
4. do not add a Turbo/four-step LoRA when reproducing the validation below;
5. inspect skin, hair, fabric, temporal consistency, identity, composition and audio together.

The local validation used the isolated
[`minimax-h3-audio-T8`](https://github.com/T8mars/comfyui-minimax-h3-audio-T8) ComfyUI nodes with
`dual_clock_euler + native_flow`, video shift `12`, audio shift `3`, and 20 joint denoising steps. That
custom node package is the validation environment, not an architectural requirement of the checkpoint.

## Validation

### Controlled setup

```text
ComfyUI: 0.31.0, commit cbbc9dab1f03d0d9a6caa8a8be7d77a7e37e1e44
Checkpoint family: full Ref2VA INT8 ConvRot
Turbo LoRA: none
Resolution: 512 × 512
Frames: 22 at 24 FPS
Steps: 20
Sampler/scheduler: dual_clock_euler / native_flow
Seeds: 2608103502 and 2608103503
```

The 22-frame test is a short diagnostic, below the official 4–15 second usage range. It does not prove
quality at the recommended 124+ frame range. One centered portrait reference and one prompt were used;
the reference image is not redistributed.

The original checkpoint was generated twice under seed `2608103502`. Decoded video was pixel-identical
and decoded PCM audio was sample-identical, confirming deterministic repeatability for this setup.

### Candidate-to-source ratios

| Seed | Skin gradient | Skin Laplacian | Skin high-pass | Positive skin ROIs | Temporal change | Audio correlation | Audio RMS |
|---:|---:|---:|---:|---:|---:|---:|---:|
| `2608103502` | `1.0325` | `1.0760` | `1.0189` | `2/3` | `1.0321` | `0.9930` | `0.9189` |
| `2608103503` | `1.0266` | `1.0418` | `1.0131` | `3/3` | `1.0134` | `0.9687` | `1.0154` |

The skin high-pass proxy increased by approximately 1.89% and 1.31% across the two seeds, without a
fixed 2×2 parity artifact. This is a weak proxy-level signal only. Manual inspection did not establish
a convincing reduction of oily or waxy skin appearance.

The first seed also showed an approximately 8.1% audio RMS reduction. H3 predicts audio and video in a
joint denoising trajectory, so changing a video-only input projection can still indirectly change the
generated audio.

Full machine-readable results:

- [`evidence/metrics_seed_2608103502.json`](evidence/metrics_seed_2608103502.json)
- [`evidence/metrics_seed_2608103503.json`](evidence/metrics_seed_2608103503.json)

### Visual comparisons

Each comparison is **source / input HF 1.01 / input HF 1.02** from left to right. The 1.01 candidate is
shown only as an additional control and is not included in this repository.

Seed `2608103502`:

![Seed 2608103502 comparison](assets/comparison_seed1_mid.png)

Seed `2608103503`:

![Seed 2608103503 comparison](assets/comparison_seed2_mid.png)

## Limitations

- This is not an official MiniMax model.
- It is not a trained fine-tune, LoRA, distillation, restoration model or calibrated de-wax model.
- Only one tensor was heuristically transformed; the direction was not learned from paired data.
- Evaluation covers two seeds, one portrait reference, one prompt and a short 22-frame diagnostic.
- No official-range 124–362 frame matrix, 768p/2K test, multi-subject test, blind preference study or
  broad prompt evaluation has been completed.
- Small proxy-metric gains may not correspond to perceptual improvement.
- Identity, framing, motion and generated audio may change even when the modified tensor is video-specific.
- Compatibility with Turbo LoRAs, third-party samplers, alternate quantization loaders and older ComfyUI
  versions is not established.
- The source checkpoint remains the recommended default for production or reproducible comparisons.

## Intended use

Intended only for controlled research and A/B evaluation of a small, auditable MiniMax H3 video-input
projection change, subject to the upstream license and acceptable-use policy.

Not intended for claims of guaranteed quality improvement, identity preservation, medical/forensic use,
deceptive impersonation, bypassing safety systems, or any use prohibited by the MiniMax H3 license.

## License and attribution

MiniMax H3 and this Model Derivative are governed by the
[MiniMax H3 Community License Agreement](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE).
The upstream license, including territorial restrictions, redistribution requirements, commercial terms
and acceptable-use policy, takes precedence over this model card.

Before distributing the checkpoint:

1. include an unmodified copy of the official MiniMax H3 `LICENSE` in the repository;
2. include the provided [`NOTICE`](NOTICE);
3. retain a prominent statement that `minimax_h3_ref2va_patchin_hf102.safetensors` is modified;
4. comply with the applicable-territory and downstream-user requirements;
5. verify any commercial-use authorization requirements directly with MiniMax.

Upstream projects:

- [MiniMaxAI/MiniMax-H3](https://huggingface.co/MiniMaxAI/MiniMax-H3)
- [Comfy-Org/MiniMax-H3](https://huggingface.co/Comfy-Org/MiniMax-H3)
- [Official MiniMax H3 GitHub repository](https://github.com/MiniMax-AI/MiniMax-H3)

## Modification notice

`minimax_h3_ref2va_patchin_hf102.safetensors` is a modified file derived from the MiniMax H3 Ref2VA
checkpoint. It was created without training by applying a 1.02 non-DC 2×2 Haar gain to
`video_patch_proj.weight`. This modification is independent and is not endorsed by MiniMax or Comfy-Org.