Skip to results
MLSift

Titles, abstracts, or an arXiv ID

← Back to results
routineML Systems & EfficiencyMixture-of-Experts2609.04575

Training-Free Halving of Activated Experts in Fine-Grained Mixture-of-Experts Models

Xing Chen, Hengshuai Yao

cs.LG cs.AI

Abstract

Modern fine-grained Mixture-of-Experts (MoE) models route each token to a small number of experts and renormalize their router probabilities. We show that this renormalization implicitly calibrates expert output gain to the training top-k: reducing k at inference changes not only which experts are used but also the strength of the expert branch. We separate these effects by activating the top k1 experts while normalizing by the probability mass of the top k2 experts, introducing one integer with no parameters, training, or measurable compute overhead. On Qwen3.6-35B-A3B, reducing from 8 to 4 experts causes a 4.65-point MMLU drop under standard renormalization but only 0.35 points with k2=16, while halving routed-expert compute. The result replicates on the $11\times$ larger Qwen3.5-397B-A17B, where reducing from 10 to 5 experts loses only 0.55 points with an appropriate reference set. Removing renormalization entirely is catastrophic, showing that preserving a suitable reference mass is crucial. We further find that perplexity and downstream accuracy favor different k2, cautioning against selecting MoE compression settings using unlabeled text alone. Analyses also show that expert identity matters substantially more than expert weighting, while balanced and domain-specialized routing leaves limited room for expert pruning.

Topics

Classified with taxonomy v2 on Mon, 7 Sept 2026.

Report a classification error

Loading the PDF downloads the document. Open it in your browser's viewer, or load it here.

Open PDF