// Performance engineering

Optimizing CPU Cache Locality for High-Throughput Financial Systems

At the throughput levels financial systems demand, the processor is rarely computing. It is waiting for memory. Cache locality work is the discipline of making that wait shorter and rarer.

Layout is the algorithm

An array of structures forces the processor to pull entire records into cache when a scan touches one field. Restructuring into a structure of arrays lets a scan stream only the field it needs, often multiplying effective bandwidth several times over on order book and risk traversal workloads.

Hot and cold field separation is the cheaper cousin of the same idea. Split rarely accessed metadata out of the record so the hot portion of each entry fits more densely into a cache line, then measure the change in last-level cache misses rather than assuming the win.

False sharing and alignment

Two counters updated by two threads that share a cache line will serialize invisibly, and the profile shows nothing but slow atomic operations. Padding and aligning per-thread state to cache line boundaries removes the contention outright.

Be deliberate about which structures deserve padding. Aligning everything wastes cache capacity and can hurt more than the false sharing it prevents; align the small number of frequently written per-thread structures and leave read-mostly data dense.

Prefetching and access order

Hardware prefetchers detect linear and constant-stride access patterns well and pointer chasing not at all. Where a structure must be a tree or a hash map, consider a flattened representation with index offsets instead of pointers so traversal becomes predictable.

Software prefetch hints help when the address is known several iterations ahead, such as batch processing a queue of messages. Issue the hint far enough in advance to cover memory latency and verify with counters — misapplied prefetch instructions pollute cache and cost more than they save.

NUMA placement

On multi-socket machines, remote memory access can cost roughly twice a local one. Pin threads to cores, allocate memory on the local node, and keep the data a thread owns on the same node as that thread. Verify with hardware counters for remote memory access rather than trusting the configuration.

// key takeaways

  • Convert scan-heavy arrays of structures into structures of arrays.
  • Separate hot fields from cold metadata to densify cache lines.
  • Pad only frequently written per-thread state; leave read-mostly data dense.
  • Flatten pointer-chasing structures into index-based layouts for predictable strides.
  • Pin threads and allocate node-locally, then verify with remote-access counters.

Working on something like this?

InnerLoop Resources takes on performance, distributed systems, and CI architecture engagements.

Brief our team →