Measured, not estimated
Every number on this page is a real device benchmark, run in-cluster from the control plane. No projections, no marketing math. Where something isn't built yet, we say so.
Benchmarked 2026-09-24 · FP16 8192³ matmul · serving on Qwen2.5-7B (Q4), GPU-resident
Three GPU architectures under one control plane. The seven Blackwell cards land within 1.5% of each other — the fabric, drivers, and cooling all behave identically.
| Tier | Cards | Arch | FP16 | VRAM |
|---|---|---|---|---|
| NVIDIA — workhorse | 7× RTX 5060 Ti | Blackwell (sm_120) | ~50 TFLOP/s ea | 16 GB ea |
| NVIDIA — Ada | 1× RTX 4060 Ti | Ada (sm_89) | 45.8 TFLOP/s | 16 GB |
| AMD — modern | 1× RX 9060 XT | RDNA4 (gfx1200) | 40.4 TFLOP/s | 16 GB |
| AMD — HBM2 | 2× Radeon Pro VII | gfx906 | 15.5 + 16.6 TFLOP/s | 16 GB ea |
| Verified total · 11 GPUs | 469.9 TFLOP/s | 185 GB | ||
Real token generation on a 7B (Q4) model, GPU-resident.
Serving pool = 7 Blackwell + 1 Ada + 1 RDNA4. The two HBM2 cards are held for compute/embeddings, not this serving pool.
Every worker clears the ~30B (4-bit) bar; the NVIDIA tier does it eight times over, in parallel. 70B requires cross-GPU pooling that is designed but not yet shipped — we don't sell it as available.
The nodes are separate machines joined by a 2.5 GbE fabric. That's ideal for running many models side by side and serving many users in parallel — the fleet's real shape is 11 independent GPU endpoints, a throughput farm. Splitting one very large model across nodes is bottlenecked by the network, not the silicon; that's why 70B-class serving is a roadmap item with a specific engineering path, not a checkbox we've already ticked.