GeeCoolGuest
GeeCoolEnough for a working set of small, hot models and indexes: embedding, scoring and rerank requests touch only a few MB each yet arrive by the hundreds, so after 60 cores and 120 threads flatten the concurrency, 300MB of L3 keeps the frequently hit weights on-die while eight channels to 4096GB cover everything cold. When the hot working set stops leaving cache, throughput follows almost by itself.
For the recall and precision-scoring legs of a search pipeline it is a sound CPU base; long-context generation, which lives on bandwidth and capacity, will spill past what 300MB can hold and belongs on a higher-bandwidth sibling. During load tests watch the L3 hit rate and tune batch size — once hits climb, that cache is earning its place.