// Performance engineering
Profiling Linux Kernel Bottlenecks in Virtualized Cloud Environments
Kernel profiling on bare metal is a solved problem. On a virtual machine, half the usual instruments are degraded or absent, and the noisiest variable — the neighbor you cannot see — never appears in your own profile.
Establish what the hypervisor is taking
Begin with steal time and involuntary context switches. Sustained steal indicates the hypervisor is scheduling your virtual CPU away, and no amount of application tuning recovers it. On burstable instance families, exhausted credits produce the same symptom on a longer cycle and are frequently misdiagnosed as a memory or garbage collection problem.
Record baseline steal, run queue length, and interrupt rates over a full week before drawing conclusions. Cloud performance is a distribution, not a value.
Tracing with eBPF
eBPF tracing works well in virtualized environments because it depends on kernel instrumentation rather than hardware performance counters. Syscall latency histograms, block I/O latency distributions, run queue delay, and off-CPU analysis together explain most kernel-side bottlenecks.
Off-CPU profiling is the highest-value technique and the most underused. On-CPU flame graphs show where time is spent computing; off-CPU flame graphs show where time is spent waiting, which in cloud workloads is usually the dominant term.
Where virtualization degrades your instruments
Hardware performance counters may be unavailable, partially virtualized, or inaccurate depending on instance type and hypervisor configuration. Cache miss and instructions-per-cycle numbers should be treated as suspect unless the platform documents counter passthrough.
Timekeeping deserves scrutiny too. Verify the clock source and prefer a paravirtualized clock where offered; a clock source that traps to the hypervisor turns frequent timestamp calls into a measurable tax, which shows up as unexplained overhead in any code that times itself.
Network and storage paths
Virtualized network stacks add per-packet overhead that dominates small-message workloads. Interrupt coalescing settings, receive-side scaling queue counts, and driver offload flags are all worth measuring. For storage, distinguish device latency from queueing delay — block layer histograms separate the two, and the fix differs completely depending on the answer.
// key takeaways
- Check steal time and burst credits before profiling application code.
- Prefer eBPF tracing, which does not depend on virtualized hardware counters.
- Use off-CPU flame graphs — waiting, not computing, usually dominates cloud workloads.
- Treat cache and IPC counters as suspect unless counter passthrough is documented.
- Verify the clock source; a trapping clock taxes every timestamp call.
Working on something like this?
InnerLoop Resources takes on performance, distributed systems, and CI architecture engagements.
Brief our team →