The GPU is full
it is holding one gigabyte of anything at all
What people expectThat serving capacity is a compute number. The audience reads 'the GPU is full' as 'the GPU is working', and here it is neither — the arithmetic unit is idle enough to serve thirty-five times as many people, and the only thing that is full is a reservation nobody checked.
The numbers
- layers
- 80
- kv heads
- 8
- head dim
- 128
- dtype bytes
- 2 bytes
- cache budget gi b
- 40 GiB
- arrival rate per sec
- 30 req/s
- avg prompt tokens
- 700 tokens
- avg output tokens
- 200 tokens
- base step ms
- 18 ms
- attn ms per million tokens
- 40 ms
- block tokens
- 16 tokens
- queue limit
- 64 requests
- paged allocation
- 0
- seed
- 20260731
The mechanism is documented; these particular values are authored, chosen to put the interesting threshold somewhere you can see it. The simulation is deterministic and seeded — the same spec renders the same frames on any machine.