All tags
#Mixture-of-Experts
Content focusing on Mixture-of-Experts (MoE) architectures, sparse activation, gating networks, and scaling laws for efficient LLM training.
1
Models
0
Providers
0
Articles
Top models
View all 1Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

