Sign In

Krea 2 Turbo — QuantFunc A4W4 INT4

Updated: Oct 1, 2026

Download

1 variant available

int4 SafeTensor

krea2-turbo-quantfunc-int4-r128.safetensors

4-bit integer, smallest • 8.22 GB

Verified:

Type
Checkpoint Trained
Stats

91

Reviews
Published

Oct 1, 2026

Base Model

Krea 2

Hash
AutoV2
CF465BA398
default creator card background decoration
Followers - 7

7

Likes - 23

23

License:

krea2_show2

Krea 2 Turbo — QuantFunc A4W4 INT4

True 4-bit inference — A4W4 (4-bit activations × 4-bit weights).

QuantFunc's core quantized matrix multiplications use INT4 activations and INT4 weights on its INT4 inference backend. A4W4 describes the compute precision used during inference, alongside the reduced weight size and memory bandwidth demand.

QuantFunc INT4 compresses Krea-2-Turbo to about a quarter of its 16-bit size. In our visual comparisons, composition, detail, color and style stay close to the 16-bit baseline, roughly on par with FP8 and INT8 ConvRot.

Showcase

All images below were generated by Krea-2-QuantFunc-4bit.

Portrait photography Sci-fi scene Portrait Sci-fi Impasto art Watercolor illustration Impasto Watercolor

Fast denoising, fast end-to-end too

High-VRAM

On an RTX 4090, same workflow and generation settings:

Stage QuantFunc INT4 FP8 Speedup Denoising 1.6s 5.0s 3.13x End-to-end 2.5s 6.5s 2.6x

End-to-end includes text encoder, denoising and VAE. Text encoder and VAE are not part of the QuantFunc plugin's acceleration path today. Actual speed varies with resolution, steps, driver, software version and hardware.

Low-VRAM

On 8 GB / 12 GB and similar VRAM-constrained setups, 4-bit weights cut weight-bandwidth demand significantly, adding extra speedup — up to roughly 11x. Actual gains depend on VRAM capacity/bandwidth, offload behavior and generation settings.

Swap one loader, keep your workflow

  1. Install or update ComfyUI-QuantFunc.
  2. Download the r128 or r32 weights.
  3. Swap your model loader for the QuantFunc loader and pick the matching weight file.

Every other node, connection and generation parameter stays as-is.

Krea-2-Turbo is a distilled turbo model — use fewer sampling steps and turn CFG off (guidance 1.0).

RTX 20-series through GB300, one build covers it all

Runs on every NVIDIA SM75+ GPU: RTX 20/30/40/50-series, A100, H100, H200, B100, B200, GB300.

Choose a model

Variant File Size Use case r128 krea2-turbo-quantfunc-int4-r128.safetensors 8.83 GB (8.22 GiB) Quality-first, recommended r32 krea2-turbo-quantfunc-int4-r32.safetensors 8.35 GB (7.77 GiB) Size/VRAM-first

Loading

Weights use QuantFunc's own safetensors format — load with ComfyUI-QuantFunc or the QuantFunc inference engine, not as a drop-in Diffusers checkpoint.

Source & license

This repository distributes derived quantized weights produced from:

Weights follow the Krea 2 Community License — please read the original model's license terms before use.

Official QuantFunc links

Source and weight integrity

Original QuantFunc release: QuantFunc/Krea-2-QuantFunc-4bit.

  • krea2-turbo-quantfunc-int4-r128.safetensors — SHA256: cf465ba39846a5664f0d92ad3c41a8303705d91cb89cbec856371cf5635066bb
  • krea2-turbo-quantfunc-int4-r32.safetensors — SHA256: 89cddbfaeced4fafb80b0cf284bdd29aec85ca204b5a7f5f6e245331c14cc875

Showcase media is reproduced from the original QuantFunc model card. Exact seeds, prompts and rank variants are not supplied for every example. Performance figures are QuantFunc's reported measurements under the stated conditions; they are not guarantees for other workflows.

License terms

These quantized weights are a modified derivative of Krea 2 Turbo; they are not an official or endorsed Krea product. The Krea 2 Community License permits commercial use only below its USD 1 million company-wide annual revenue threshold; otherwise a separate Enterprise License is required. Redistribution must retain the agreement and NOTICE, and deployments must follow its content-filtering and acceptable-use requirements. See the full license agreement.