AI2 just shipped the real infrastructure
Allen Institute for AI dropped Olmo-core 3 today, and it's not another model release—it's the training stack behind their next-generation Olmo models. This is a redesigned mixture-of-experts (MoE) training system built to scale into the trillion-parameter range while keeping computational efficiency from collapsing under its own weight.
The announcement matters because AI2 is doing something most labs won't: opening up the actual infrastructure and training decisions behind each model generation. Model weights are table stakes now. The training stack? That's where the real engineering lives.
And they're shipping receipts. The system has been benchmarked at over one trillion total parameters. One configuration hit 1.2 trillion parameters with 58.36 billion active per token across 512 NVIDIA B300 GPUs, reaching 858 TFLOP/s/GPU in throughput.
Why MoEs need different plumbing
Mixture-of-experts models promise efficiency: you get a huge parameter capacity, but each input only activates a subset of "experts"—specialized sub-networks that handle different patterns. The problem is that the full model still has to live in GPU memory during training, and routing inputs to the right experts across a cluster creates communication overhead that can eat your efficiency gains.
AI2's earlier MoE work—OlmoE with 64 routed experts—used fully sharded data parallelism (FSDP), which gathered and resharded model weights for each small training batch. That works, but it doesn't scale elegantly.
Olmo-core 3 switches to a system based on distributed data parallelism (DDP) that keeps experts resident on GPUs and routes data to them, avoiding repeated weight gathering. In a preliminary test on eight B300 GPUs, a 47-billion-parameter MoE hit 52,000 tokens per second per GPU with the new stack, compared with 19,400 using the earlier implementation—roughly 2.7× the throughput.
The scaling story: more experts, stable throughput
Here's the benchmark that caught my attention: AI2 scaled the expert pool from 8 to 128 experts while keeping only four experts active per token. Total parameter capacity grew from 4.6 billion to 47 billion. Training throughput fell by less than 5%.
That's the engineering goal in one number. You want parameter capacity to scale faster than your throughput degrades. If you can keep the active-parameter count roughly constant while growing the total model, you're buying capacity without proportional compute cost—as long as the routing and distribution overhead doesn't kill you.
Olmo-core 3 combines three core techniques for distributing the model and training state across hardware:
- Expert parallelism spreads experts across GPUs so each GPU stores only part of the expert pool
- Pipeline parallelism splits the model's layers across groups of GPUs, reducing per-GPU memory requirements
- Distributed optimizer spreads optimizer state across GPUs instead of replicating it everywhere
Together, these let you scale without requiring every GPU to hold the entire model and its training state in memory.
The optimization layer: where efficiency lives
Scaling techniques get you to the right size. Optimizations keep you fast at that size.
Olmo-core 3 includes several routing and computation optimizations:
- Rowwise expert parallelism places routed data directly into expert input buffers, minimizing rearrangement work
- GPU-resident routing keeps routing metadata on GPUs so the CPU can queue work without waiting for data to copy back
- Grouped GEMM combines many small expert computations so GPUs can execute them more efficiently
They also support MXFP8, a lower-precision number format that represents some values with fewer bits. In a controlled benchmark on four B300 GPUs with uniform expert distribution, enabling MXFP8 across the parts of the system where it helped most increased training throughput by about 21% compared to BF16 baseline. Peak active memory dropped from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and inter-expert data movement rather than attention alone.
The hard trade-offs
What I appreciate about the technical report is that AI2 documents the approaches they tested and didn't adopt. These are the messy realities of systems work:
Token gerrymandering: A routing balance score could improve even as the actual workload distribution became less balanced. The metric was measuring the wrong thing.
Learning rate adjustments: Lowering experts' learning rates because they process fewer tokens didn't improve results in the model family they tested. Sometimes the obvious fix doesn't work.
Non-deterministic compute: GPU calculations took different amounts of time when input values changed, even with identical matrix dimensions. Performance comparisons need matching values, not just matching shapes.
Overlapping communication and computation: Running these on separate GPU streams didn't always speed up training. In some tests it slowed end-to-end execution. More overlap ≠ higher throughput.
These aren't footnotes. They're the actual engineering process: hypothesis, test, measure, update your mental model.
The trillion-parameter ceiling test
AI2 benchmarked Olmo-core 3 up to 1.2 trillion parameters with 58.36 billion active per token across 512 GPUs, hitting 858 TFLOP/s/GPU. They also ran a short-capacity test with DeepEP v2 (an alternative expert-communication method) reaching 2.38 trillion total parameters.
That second test wasn't a full training run—it's a "can we reach this scale" checkpoint rather than sustained training performance. But it establishes the ceiling.
Benchmarks used random routing to measure system performance rather than model quality. That's the right methodology when you're testing infrastructure: you want to know what the system can do before training dynamics complicate the measurement.
Why open infrastructure matters
Olmo-core 3 is the foundation for AI2's next-generation Olmo, which will use an MoE architecture with their largest dataset and longest context window yet. But they're shipping it as a fully open training stack.
Researchers and developers can use Olmo-core 3 to train their own MoEs, adapt it to different hardware, and experiment with routing, parallelism, and other system components. That's a different philosophy than "here are the weights, good luck."
When NVIDIA's Megatron-Core exists as an established option for training large MoEs, why does another open stack matter? Because diversity in training infrastructure is good for the ecosystem. Different design choices, different trade-offs, different optimization targets. AI2's implementation brings an integrated MoE training stack to the framework behind Olmo with design decisions that differ from Megatron's.
And because the more open implementations exist, the harder it becomes to treat training infrastructure as proprietary advantage. Model weights are more useful when the infrastructure and training decisions behind them are open too.
The bottom line
Olmo-core 3 is systems engineering at scale: careful trade-offs between computation, communication, and memory; benchmarking that documents what didn't work; and a design that scales to trillion-parameter models while keeping throughput from collapsing.
The code is on GitHub. The technical report documents the experiments, ablations, and approaches they tested. And the interactive demo lets you see how expert, data, and pipeline parallelism work together.
AI2 is building the next Olmo on this stack. But they're shipping the infrastructure first, open, so others can build on it too. That's the bet: open training infrastructure makes better models possible for more researchers, not just the labs with the biggest budgets.
I'm watching to see what gets built on top of it.