Measured, not estimated

What the fleet actually does.

Every number on this page is a real device benchmark, run in-cluster from the control plane. No projections, no marketing math. Where something isn't built yet, we say so.

Benchmarked 2026-09-24 · FP16 8192³ matmul · serving on Qwen2.5-7B (Q4), GPU-resident

Compute
469 TFLOP/s
dense FP16, 11 GPUs
GPU memory
185 GB
across the fleet
GPUs online
11
independent endpoints
Serving
~1,195 tok/s
9-GPU serving pool

The fleet, card by card

Three GPU architectures under one control plane. The seven Blackwell cards land within 1.5% of each other — the fabric, drivers, and cooling all behave identically.

TierCardsArchFP16VRAM
NVIDIA — workhorse7× RTX 5060 TiBlackwell (sm_120)~50 TFLOP/s ea16 GB ea
NVIDIA — Ada1× RTX 4060 TiAda (sm_89)45.8 TFLOP/s16 GB
AMD — modern1× RX 9060 XTRDNA4 (gfx1200)40.4 TFLOP/s16 GB
AMD — HBM22× Radeon Pro VIIgfx90615.5 + 16.6 TFLOP/s16 GB ea
Verified total · 11 GPUs469.9 TFLOP/s185 GB

Serving throughput

Real token generation on a 7B (Q4) model, GPU-resident.

  • Single stream~85 tok/s
  • 4-way concurrent (per GPU)~130 tok/s
  • Prompt prefill>1,200 tok/s
  • Serving pool (9 GPUs)~1,195 tok/s

Serving pool = 7 Blackwell + 1 Ada + 1 RDNA4. The two HBM2 cards are held for compute/embeddings, not this serving pool.

What fits on one 16 GB card

  • FP16 (full precision)~7–8B
  • INT8 / FP8~13B
  • 4-bit (GPTQ/AWQ)~30B
  • 70B-classon the roadmap

Every worker clears the ~30B (4-bit) bar; the NVIDIA tier does it eight times over, in parallel. 70B requires cross-GPU pooling that is designed but not yet shipped — we don't sell it as available.

Designed as many workers, not one giant model

The nodes are separate machines joined by a 2.5 GbE fabric. That's ideal for running many models side by side and serving many users in parallel — the fleet's real shape is 11 independent GPU endpoints, a throughput farm. Splitting one very large model across nodes is bottlenecked by the network, not the silicon; that's why 70B-class serving is a roadmap item with a specific engineering path, not a checkbox we've already ticked.

Method & caveats

Start free trial See pricing