E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
Paper • 2609.37533 • Published • 67
E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models. LM1B checkpoints: E-MoE and the MDLM / MDLM-MoE baselines.
Note E-MoE: MoE routes as a discrete shared latent (536M total / 139M active).
Note Capacity-matched baseline: same MoE backbone, deterministic routing.
Note Factorized baseline: dense DiT (139M).