All tags

#Mixture-of-Experts

Content focusing on Mixture-of-Experts (MoE) architectures, sparse activation, gating networks, and scaling laws for efficient LLM training.

1
Models
0
Providers
0
Articles

Top models

View all 1

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

262,144 ctx