Download
1 variant available
int4 SafeTensor
MiniMax-H3-Ref2VA-DF-Turbo-W4A8.safetensors
4-bit integer, smallest β’ 12.62 GB
Verified: 15 hours ago

4250 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 90 1 2 3 4 5 6 7 8 9
1.3K0 1 2 3 4 5 6 7 8 9.0 1 2 3 4 5 6 7 8 9K

License:
# β‘ MiniMax-H3 Ref2VA (Safetensors Quantized Checkpoints)
This repository provides optimized Safetensors quantizations of MiniMax-H3-Ref2VA (Reference-to-Video & Audio), engineered specifically for consumer GPUs (16 GB VRAM, RTX 4080 / RTX 4090 / RTX 4090 Mobile / RTX 3090).
Choose between two purpose-built variants depending on your workflow and VRAM budget:
---
## π Variant 1: Digital Forge Turbo W4A8 (All-In-One 4-Step) β This is technically coming soon, this is currently an offload compression version, v2 will be the full model.
The flagship mixed-precision release designed by Digital Forge (DF). It eliminates common INT4 artifacts while mathematically fusing the official LightX2V 4-Step Turbo adapter directly into the model weights.
### π‘οΈ Why Choose DF-Turbo-W4A8?
π *Pre-Baked 4-Step Turbo**: The 4-step Turbo LoRA is fused before quantization. No external LoRA loading requiredβsaves 1.28 GB runtime VRAM and eliminates 624 GEMMs per step!
π‘οΈ *Zero "Zombie Sway" (Pristine BF16 AdaLN)**: All 335 time-modulation layers, LayerNorms, and 1,025 calibration grid points remain in unquantized BF16, ensuring continuous temporal velocity ($\vec{v}$) and smooth, natural character motion.
π― *Zero Reference Feature Drift (INT8 Cross-Attention)**: Uses INT8 per-channel quantization across all cross-attention to_k and to_v layers to preserve fine facial features, eye contact, and clothing textures without outlier clipping.
πΎ *16 GB VRAM Resident (12.62 GiB)**: Fits 100% resident inside 16 GB GPUs with zero PCIe swapping.
### βοΈ Recommended Settings (DF-Turbo-W4A8)
* Sampling Steps: 4
* CFG Guidance: 1.0 (Distilled trajectory)
* Video Shift: 12.0
* Audio Shift: 3.0 (32 kHz synchronized audio)
* Frame Count: 25 to 124 frames
---
## π§ͺ Variant 2: Pure INT4 (Ultra-Low VRAM Experimental)
A pure 4-bit uniform quantization designed for minimal memory consumption.
### π¬ Model Details & Caveats
* VRAM Footprint: *9.80 GiB** (Ultra-compact, maximum memory headroom).
* Quantization Scheme: Symmetric 4-Bit Linear FastInt4Linear) with Group-128 scaling.
* Adapter Support: Requires loading the external LightX2V 4-step Turbo LoRA minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors) if 4-step generation is desired.
β οΈ *Experimental Note**: In extended runs (e.g. 10-second / 124+ frame sequences), pure INT4 quantization on time-modulation layers may exhibit slight identity drift or rhythmic "zombie sway" artifacts. Use DF-Turbo-W4A8 if artifact-free motion is required.
---
## β‘ Performance & Engine Compatibility (Both Models)
* Zero-Copy Loading: Loads via OS virtual memory mapping mmap) in ~0.03 to 0.05 seconds.
* TeaCache Compatible: Full native support for Block-Level TeaCache (~50% compute step bypass).
* Attention Acceleration: Supports SageAttention 2.0 Patched (Fast FP8 with FP32 accumulator), FlashAttention-2, and native PyTorch SDPA.
* VAE Tiling: Recommended VAE tile size of 512 for efficient 3D temporal decoding.
---
## π License & Attribution
Base model licensed under the *MiniMax H3 Community License Agreement**, Copyright Β© 2026 MiniMax.
Powered by *MiniMax H3** & Digital Forge.
