BACK TO LAB

GPU hall control room.

A working miniature of a bare-metal GPU cloud: four pods of 8-GPU servers, 1,024 GPUs in all. Racked hardware becomes customer capacity through a real provisioning pipeline: discovery over the BMC, firmware, network install and burn-in. Outside air follows real Lidköping weather and energy follows the real SE3 day-ahead price. Place customer orders, roll firmware behind a canary, then run the drills: trip the cooling, spike the price, storm the fleet with faults. Every button is an API call logged with its HTTP status, and an advisor explains what to do next.

Sim clock
Day 1 · 00:00
Site power
664 kW uncapped
PUE
1.08 IT 614 kW
Utilization
58% 448/768 GPUs
Delivered
448 GPUs full speed
Energy / GPU-h
€0.017 €7/h site
Largest block
256 GPUs 0% fragmented
Customer impact
0 node-minutes lost
HALL FLOOR LIVE

Rows face each other across cold aisles and exhaust into hot aisles, where the haze follows each row's cooling load. Click a node to inspect it, or switch to 2D for the full rack grid and bulk provisioning.

AllocatedAvailableProvisioningOut of service Draining Racked Fault
CAPACITY · GPUs by state 1024 installed
  • Allocated 448
  • Available 320
  • Provisioning 0
  • Out of service 0

Collecting samples. Let the clock run.

FLEET · nodes per state
Nodes per lifecycle state
racked32
discovering0
firmware0
imaging0
burnin0
available40
allocated56
draining0
maintenance0
rma0
API CONSOLE
  1. Nothing yet. Click a node, place an order or let the clock run.
NODE
A01-1 ALLOCATED

Running a customer workload.

Location
Pod A · rack 1 · slot 1
Firmware
1.4.0
Host OS
Installed
Customer
Northwind Labs · ord-1
Power
9.4 kW
GPU temp
67 °C
OPS ADVISOR rule-based
  • RACKED_IDLEinfo 256 GPUs are racked but not provisioned. ▸ Provision them to turn hardware into sellable capacity.
DRILLS
ORDERS · POST /v1/orders
GPUs
Placement
Free320 GPUs
Largest one-pod block256 GPUs
Fragmentation0%
  • Northwind Labs ord-1 · 256 GPUs · pod A
  • Fjord AI ord-2 · 128 GPUs · pod B
  • Aurora Bio ord-3 · 64 GPUs · pod B
CONTROLS
CLOCK
Fault rate×1
×1 is a realistic large-fleet rate: one node fault every few sim days in this hall.
OUTSIDE Lidköping (SE3) air live 2026-10-11 · SE3 price live 2026-10-10 · hourly, UTC SITE (API)
Cooling units online
A 4/4
B 4/4
C 4/4
D 4/4
HOW THIS MAPS TO PRODUCTION

Provisioning

HereA typed state machine with timed steps that can fail into maintenance.

In productionDiscovery and power control over Redfish on each BMC, network boot (PXE/iPXE) into an image service, driven by a bare-metal manager such as MAAS, Ironic or Tinkerbell. Each step persisted and retryable in a workflow engine, so a crashed controller resumes instead of orphaning a node.

Burn-in

HereA four-hour stress step with a fixed failure rate.

In productionGPU diagnostics (DCGM level 3), NCCL all-reduce across the pod to prove the fabric, memory and InfiniBand link checks, with results stored per node so repeat offenders go to RMA.

Telemetry and faults

HerePoisson faults scaled by GPU temperature, auto-drain and backfill.

In productionDCGM exporter and node exporter into Prometheus, XID and ECC events from kernel logs, and alert rules that call the same drain endpoint the operator button does.

Placement

HerePod-contiguous best fit with cooling and power admission.

In productionTopology read from the fabric manager, so jobs land on one non-blocking leaf group, with the same admission checks fed by live facility data instead of a model.

API

HereA pure reducer returning REST-shaped responses and audit events.

In productionThe same contract served by a Go or TypeScript service with an OpenAPI spec, idempotency keys on every mutation, per-tenant auth and an append-only audit table.

Firmware

HereCanary, waves, and an automatic halt on burn-in failures.

In productionUpdates through Redfish UpdateService, gated by the same burn-in signal, with the halt wired to paging so a bad build stops at the canary without anyone watching.

Power and cooling

HerePUE from outside air, lumped row cooling, a site power cap.

In productionLive data from the building and DCIM systems (cooling units and PDUs over Modbus or BACnet), and power capping enforced per node through the BMC or GPU power limits.

Advisor

HereRule-based, same interface as the other control rooms.

In productionThe rules stay as guardrails, with an open-weights model served on the hall’s own GPUs drafting explanations and runbook steps from the same structured state.

Deliberately simplified

  • One node type and one lumped temperature per row; no airflow model.
  • Burn-in compressed to four hours so a sim day stays watchable.
  • No network, storage or BMC failures, and no customer workload behaviour.
  • Fault rates are order-of-magnitude, taken from published large-cluster reports.