A working miniature of a bare-metal GPU cloud: four pods of 8-GPU servers, 1,024 GPUs in all.
Racked hardware becomes customer capacity through a real provisioning pipeline: discovery over
the BMC, firmware, network install and burn-in. Outside air follows real Lidköping weather and
energy follows the real SE3 day-ahead price. Place customer orders, roll firmware behind a
canary, then run the drills: trip the cooling, spike the price, storm the fleet with faults.
Every button is an API call logged with its HTTP status, and an advisor explains what to do
next.
Sim clock
Day 1 · 00:00
Site power
664 kWuncapped
PUE
1.08IT 614 kW
Utilization
58%448/768 GPUs
Delivered
448 GPUsfull speed
Energy / GPU-h
€0.017€7/h site
Largest block
256 GPUs0% fragmented
Customer impact
0node-minutes lost
HALL FLOORLIVE
Rows face each other across cold aisles and exhaust into hot aisles, where the haze
follows each row's cooling load. Click a node to inspect it, or switch to 2D for the
full rack grid and bulk provisioning.
AllocatedAvailableProvisioningOut of serviceDrainingRackedFault
CAPACITY · GPUs by state1024 installed
Allocated 448
Available 320
Provisioning 0
Out of service 0
Collecting samples. Let the clock run.
FLEET · nodes per state
Nodes per lifecycle state
racked
32
discovering
0
firmware
0
imaging
0
burnin
0
available
40
allocated
56
draining
0
maintenance
0
rma
0
API CONSOLE
Nothing yet. Click a node, place an order or let the clock run.
NODE
A01-1ALLOCATED
Running a customer workload.
Location
Pod A · rack 1 · slot 1
Firmware
1.4.0
Host OS
Installed
Customer
Northwind Labs · ord-1
Power
9.4 kW
GPU temp
67 °C
OPS ADVISORrule-based
RACKED_IDLEinfo256 GPUs are racked but not provisioned.▸ Provision them to turn hardware into sellable capacity.
DRILLS
ORDERS · POST /v1/orders
Free320 GPUs
Largest one-pod block256 GPUs
Fragmentation0%
Northwind Labsord-1 · 256 GPUs · pod A
Fjord AIord-2 · 128 GPUs · pod B
Aurora Bioord-3 · 64 GPUs · pod B
CONTROLS
CLOCK
Fault rate×1×1 is a realistic large-fleet rate: one node fault every few sim days in this hall.
OUTSIDELidköping (SE3) air live 2026-10-11 · SE3 price live 2026-10-10 · hourly, UTCSITE (API)
Cooling units online
A4/4
B4/4
C4/4
D4/4
HOW THIS MAPS TO PRODUCTION
Provisioning
HereA typed state machine with timed steps that can fail into maintenance.
In productionDiscovery and power control over Redfish on each BMC, network boot (PXE/iPXE) into an image service, driven by a bare-metal manager such as MAAS, Ironic or Tinkerbell. Each step persisted and retryable in a workflow engine, so a crashed controller resumes instead of orphaning a node.
Burn-in
HereA four-hour stress step with a fixed failure rate.
In productionGPU diagnostics (DCGM level 3), NCCL all-reduce across the pod to prove the fabric, memory and InfiniBand link checks, with results stored per node so repeat offenders go to RMA.
Telemetry and faults
HerePoisson faults scaled by GPU temperature, auto-drain and backfill.
In productionDCGM exporter and node exporter into Prometheus, XID and ECC events from kernel logs, and alert rules that call the same drain endpoint the operator button does.
Placement
HerePod-contiguous best fit with cooling and power admission.
In productionTopology read from the fabric manager, so jobs land on one non-blocking leaf group, with the same admission checks fed by live facility data instead of a model.
API
HereA pure reducer returning REST-shaped responses and audit events.
In productionThe same contract served by a Go or TypeScript service with an OpenAPI spec, idempotency keys on every mutation, per-tenant auth and an append-only audit table.
Firmware
HereCanary, waves, and an automatic halt on burn-in failures.
In productionUpdates through Redfish UpdateService, gated by the same burn-in signal, with the halt wired to paging so a bad build stops at the canary without anyone watching.
Power and cooling
HerePUE from outside air, lumped row cooling, a site power cap.
In productionLive data from the building and DCIM systems (cooling units and PDUs over Modbus or BACnet), and power capping enforced per node through the BMC or GPU power limits.
Advisor
HereRule-based, same interface as the other control rooms.
In productionThe rules stay as guardrails, with an open-weights model served on the hall’s own GPUs drafting explanations and runbook steps from the same structured state.
Deliberately simplified
One node type and one lumped temperature per row; no airflow model.
Burn-in compressed to four hours so a sim day stays watchable.
No network, storage or BMC failures, and no customer workload behaviour.
Fault rates are order-of-magnitude, taken from published large-cluster reports.