Skip to content
StableReported 2026-10-02 12:00

ITC-MoE: Importance-guided Token-aware Compression for MoE Diffusion Language Models

A new method called ITC-MoE is introduced for compressing Mixture-of-Experts Diffusion Language Models, addressing the issue of cross-mode non-uniform redundancy.

01

Evidence

  • AarXiv cs.AIPrimary source2026-10-02 12:00
    ITC-MoE consists of two complementary components. First, Importance-guided Adaptive Tucker Compression (IATC)
    View source
  • AarXiv cs.AIPrimary source2026-10-02 12:00
    Specifically, we identify two properties: cross-mode non-uniform redundancy, where parameter redundancy and sensitivity to rank truncation vary across the input, output, and expert modes, and token-wise utilization variation, where hot and cold tokens exhibit distinct spectral characteristics and expert activation patterns. To address these challenges, we propose ITC-MoE, an Importance-guided Tok…
    View source