🛰️ Daily AI Frontier
‹ back to 2026-07-25

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

Research Efficiency & Systems

Ranking

Overall 72
Content 85
Popularity 42

Observed public metrics from 1 member.

Merged summary

TL;DR - KroQuant enables higher-quality W4A4 post-training quantization of diffusion transformers using efficient Kronecker-structured activation transforms. It reduces inference overhead while keeping generated outputs closer to full-precision references.

  • Applies learned invertible transforms to 32-element activation blocks using small tensor-core GEMMs.
  • Stores less than half the parameters of per-channel scaling and runs up to 14% faster than SmoothQuant’s kernel on an MI350 GPU.
  • Uses offline LoRaQ calibration to absorb residual per-weight quantization errors.
  • Outperforms SVDQuant and LoRaQ in similarity to FP outputs across PixArt-ÎŁ, SANA, and FLUX.1-schnell benchmarks while preserving or improving image quality.

Sources (1)

KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

arXiv cs.LG Yann Bouquet, Alireza Khodamoradi, Kristof Denolf, Mathieu Salzmann 2026-07-23 arXiv:2607.21446
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-21 14:36:28.130705 UTC

TL;DR - KroQuant enables higher-quality W4A4 post-training quantization of diffusion transformers using efficient Kronecker-structured activation transforms. It reduces inference overhead while keeping generated outputs closer to full-precision references.

  • Applies learned invertible transforms to 32-element activation blocks using small tensor-core GEMMs.
  • Stores less than half the parameters of per-channel scaling and runs up to 14% faster than SmoothQuant’s kernel on an MI350 GPU.
  • Uses offline LoRaQ calibration to absorb residual per-weight quantization errors.
  • Outperforms SVDQuant and LoRaQ in similarity to FP outputs across PixArt-ÎŁ, SANA, and FLUX.1-schnell benchmarks while preserving or improving image quality.
item →