Home / Services / B300 GPU Capacity
Compute and modelsB300 GPU Capacity
Most GPU capacity sits idle part of every day, and buyers pay for it either way. Our clusters are sold across time zones so they stay busy around the clock, which is what lets us price off-peak hours lower. Take that capacity as raw GPU-hours for training, or as inference tokens served from the same hardware.
Why it's different
The specifics behind the claim.
Current Blackwell generation
NVIDIA HGX B300 nodes, eight GPUs each, for training and inference at current-generation performance.
Sold across time zones
Night hours here are working hours elsewhere. Keeping the cluster busy around the clock is why off-peak and overnight batch capacity costs less.
GPU-hours or tokens
Lease bare GPU-hours for training and fine-tuning, or buy inference as API tokens served from the same clusters.
What is included
Scope is set per site. These are the building blocks we combine.
Reserved capacity
Committed GPU-hours on a schedule you set, for training runs and steady inference.
Off-peak and overnight batch
Lower-priced hours for batch inference, evaluation and other jobs that can run outside peak time.
Inference as tokens
Skip cluster management entirely and call models through an API, billed per token.
How an engagement runs
Four stages. Each ends with something you can review before the next begins.
Scope
Tell us the GPU volume, timing and whether you want GPU-hours or tokens.
Confirm
We confirm availability, peak and off-peak terms.
Connect
Workloads connect to the allocated capacity or API.
Scale
Add or release capacity as your needs change.