GeeCool
GeeCoolWhat remains is concurrency as a floor; what drops is the peak illusion. Efficiency cores sip bandwidth one by one, and twelve channels to 1536GB read as warehouse-scale at this count — as long as tasks stay fine-grained with decent data locality, 576MB of L3 catches most traffic on-package and the waiting is less dire than feared. Feed it a whole large model in batch, though, and bandwidth bottoms out while cores idle for nothing.
Teams with workloads that split into vast numbers of small units and a per-watt throughput target still find this the end of the line; anyone whose metric is per-request speed should never have looked here. Measure the memory-occupancy curve with real jobs before sign-off — let the idle rate of 288 cores, not the brochure, write the acceptance report.