// Performance engineering

Eliminating Garbage Collection Pauses in Ultra-Low-Latency Java Applications

In a trading or bidding system, a twelve millisecond collection pause is not a hiccup — it is the entire latency budget spent at once. Eliminating pauses is less about tuning flags and more about not allocating.

Measure pauses before tuning anything

Start with garbage collection logs and a pause histogram rather than an average. Averages hide the distribution that actually hurts; the p99.9 pause is the number that decides whether a request misses its deadline.

Correlate pause timestamps with application-level latency traces. A surprising share of suspected collection problems turn out to be safepoint issues — biased locking revocation, thread dumps, or long-running counted loops that delay the point at which all threads can stop. Time to safepoint is a separate metric and deserves its own dashboard.

Allocation elimination beats collector tuning

The fastest collection is the one with nothing to collect. Profile allocation by call site and attack the top offenders: boxing in hot paths, string formatting in logging that is later filtered, iterator and lambda capture allocation in inner loops, and defensive copies of buffers.

Replace hot-path collections with primitive-specialized structures. Reuse byte buffers through a bounded pool. Where message decoding dominates, a flyweight pattern over a direct buffer avoids materializing objects entirely, which is the standard approach in the lowest-latency Java systems.

Choosing and configuring a collector

Modern low-pause collectors are concurrent and region-based, trading throughput and footprint for sub-millisecond pauses. They are excellent, but they are not a substitute for allocation discipline: a service allocating gigabytes per second will still generate concurrent-cycle pressure and eventual degradation.

Size the heap generously enough that cycles are infrequent, pre-touch pages at startup so first-touch page faults do not appear as latency spikes, and pin the workload's memory behavior with large pages where the platform supports it.

Validating the result

Gate the improvement in continuous integration. A benchmark that replays a recorded production tape and asserts a p99.9 ceiling turns an optimization into a permanent property of the system rather than a one-time win that decays over the next two quarters.

// key takeaways

  • Track p99.9 pause duration and time to safepoint separately from averages.
  • Eliminate hot-path allocation before touching collector flags.
  • Use flyweights over direct buffers to avoid materializing decoded messages.
  • Size heaps generously, pre-touch pages, and use large pages where available.
  • Assert latency ceilings in CI against a replayed production tape.

Working on something like this?

InnerLoop Resources takes on performance, distributed systems, and CI architecture engagements.

Brief our team →