// Performance engineering

Memory Profiling and Leak Detection in Complex C++ Microservices

Most C++ memory incidents in production are not classic leaks. They are fragmentation, unbounded caches, and allocator arena behavior that looks like a leak on a dashboard and behaves like one in an outage.

Separate the four failure modes

Before choosing a tool, classify the symptom. A true leak grows without bound and is never freed. A cache without eviction grows to a bound nobody chose. Fragmentation shows resident memory rising while live bytes stay flat. Allocator arena growth appears under thread churn as per-thread caches multiply.

The distinguishing signal is the relationship between live heap bytes and resident set size. If live bytes are stable and resident memory climbs, stop looking for a leak and start looking at the allocator.

Tooling by environment

In development and continuous integration, address and leak sanitizers catch the majority of genuine defects and should run on every build of the test suite, accepting the slowdown. Thread and undefined behavior sanitizers belong in nightly runs.

In production, sampling heap profilers integrated with the allocator provide call-stack attributed live-bytes data at a few percent overhead. Continuous profiling is preferable to reactive profiling: when memory climbs at three in the morning, you want the profile that already exists.

Reading a heap profile

Compare two profiles taken an hour apart rather than reading one in isolation. The differential shows growth by call stack, which is the actionable view. Focus on stacks whose live bytes grow monotonically across three or more samples.

Then read the code with ownership in mind. Most growth traces to a container that is inserted into and never erased, a shared pointer cycle, or an object registered with a long-lived observer list that has no deregistration path on the error branch.

Fixing fragmentation

When the problem is fragmentation, allocator choice and configuration matter more than application code. Size-class-aware allocators handle mixed workloads well; tuning decay and purge behavior returns pages to the operating system more aggressively. Where object sizes are known and lifetimes are bounded, arena or slab allocation removes the problem structurally instead of managing it.

// key takeaways

  • Classify the symptom first: leak, unbounded cache, fragmentation, or arena growth.
  • Compare live heap bytes against resident set size to tell leaks from fragmentation.
  • Run address and leak sanitizers on every CI build of the test suite.
  • Use differential heap profiles, not single snapshots, to find growing call stacks.
  • Solve fragmentation with allocator tuning or arena allocation, not application patches.

Working on something like this?

InnerLoop Resources takes on performance, distributed systems, and CI architecture engagements.

Brief our team →