Home / M3
AI runtime efficiencyM3
M3 cuts what you spend on AI at three layers: the tokens each request uses, the model that serves it, and the compute it runs on. Each layer works on its own, and the savings stack when you use them together.
Three layers
Layer 1 · TokensM3 Optimize →
Does the same work with fewer tokens, by compressing context and optimizing each request before it reaches the model.
same models · fewer tokens · faster responsesLayer 2 · ModelsM3 Router →
Sends each request to the most cost-effective model that can handle it, across many models and providers behind one endpoint.
50–80% of list · 87% routing satisfaction · multi-pool uptimeLayer 3 · ComputeM3 Compute →
GPU capacity to rent as GPU-hours or as tokens, priced lower outside peak hours.
GPU-hours or tokens · off-peak rates
Where the savings come from
One source of savings at each layer.
Layer 1: less to process
Compressed context and leaner requests mean fewer tokens billed for the same answer.
Layer 2: the right model at the right price
The same model sells for 50 to 80 percent of list depending on where you buy it, and routine requests do not need top models. M3 Router handles both.
Layer 3: compute when it is cheap
M3 Compute prices hours outside peak demand lower, for training and batch work that can wait.
How the layers fit together
Start with any layer and add the others later.
Begin where the spend is
Most clients start with tokens and routing for API spend, or with compute for training and batch jobs.
Savings compound
Each layer cuts cost at a different point: fewer tokens, sent to cheaper models, on lower-priced compute.
One view across M3
Usage and savings for every layer you use appear on the same dashboard.
What we measure
The same numbers across the M3 family.
| Metric | Unit | Why it matters |
|---|---|---|
| Serving cost | $ per 1M tokens | The number finance asks for. |
| Tokens per request | tokens | Shows what context compression saves. |
| Cost per GPU-hour | $ / GPU-hour | Reported separately for reserved and off-peak capacity. |
| Effective uptime | % | Held up by multi-pool redundancy and failover, not one provider's SLA. |
| Tail latency | p95 ms | What users actually feel. |
| Quality delta | eval score | Change on your own test set, reported with every routing change. |