Home / Services / B300 GPU Capacity

Compute and models

B300 GPU Capacity

Most GPU capacity sits idle part of every day, and buyers pay for it either way. Our clusters are sold across time zones so they stay busy around the clock, which is what lets us price off-peak hours lower. Take that capacity as raw GPU-hours for training, or as inference tokens served from the same hardware.

Why it's different

The specifics behind the claim.

B300

Current Blackwell generation

NVIDIA HGX B300 nodes, eight GPUs each, for training and inference at current-generation performance.

24/7

Sold across time zones

Night hours here are working hours elsewhere. Keeping the cluster busy around the clock is why off-peak and overnight batch capacity costs less.

2 ways

GPU-hours or tokens

Lease bare GPU-hours for training and fine-tuning, or buy inference as API tokens served from the same clusters.

What is included

Scope is set per site. These are the building blocks we combine.

Reserved capacity

Committed GPU-hours on a schedule you set, for training runs and steady inference.

Off-peak and overnight batch

Lower-priced hours for batch inference, evaluation and other jobs that can run outside peak time.

Inference as tokens

Skip cluster management entirely and call models through an API, billed per token.

How an engagement runs

Four stages. Each ends with something you can review before the next begins.

  1. Scope

    Tell us the GPU volume, timing and whether you want GPU-hours or tokens.

  2. Confirm

    We confirm availability, peak and off-peak terms.

  3. Connect

    Workloads connect to the allocated capacity or API.

  4. Scale

    Add or release capacity as your needs change.

Tell us about your site, your racks and your timeline.

Start a conversation