Scalable Mixture-of-Experts Training Infrastructure
The organizations collaborated on releasing Olmo-core 3, a redesigned open training stack. This infrastructure is optimized for scaling mixture-of-experts models into the trillion-parameter range.
Open infrastructure for giant models democratizes supercomputing scale training capabilities beyond closed-source providers.
Source posts · 2
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
Hugging Face introduced Olmo-core 3, an open and scalable training infrastructure designed for large Mixture of Experts models.
Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Olmo-core 3 introduces a redesigned, fully open training stack for efficiently scaling mixture-of-experts models into the trillion-parameter range.
The Allen Institute for AI has introduced Olmo-core 3, an open training infrastructure for scaling mixture-of-experts models.