Node.js, part 14: the profiler you already have — --cpu-prof, heap snapshots, and the flame graph
Part 14from the Node.js series · 24 parts in all
"It's slow" is not a diagnosis, and adding console.time everywhere is not a
measurement. Node ships a CPU profiler, a heap snapshot mechanism, and a diagnostic report
generator. Part 14 is using them on a running service — the difference between a guess and a
flame graph.
CPU: sample, do not instrument
# Profile from process start, write a .cpuprofile on exit.
node --cpu-prof --cpu-prof-dir=./prof server.js
# Or attach to a running process and start/stop sampling at will.
node --inspect server.js
# then in chrome://inspect -> Profiler -> Start/Stop
The .cpuprofile opens directly in Chrome DevTools, which turns it into a flame
graph. Read it from the top down: the widest frames are where time goes, and the useful
surprise is usually finding a frame you did not write — a JSON parse, a regex, a serialiser in
a driver. That is the whole method: measure, find the widest frame, fix it, measure again.
Heap: find the leak, not the usage
const v8 = require('v8');
const { writeHeapSnapshot } = v8;
// Two snapshots, taken minutes apart, with traffic in between.
writeHeapSnapshot('before.heapsnapshot');
// ... let the service run ...
writeHeapSnapshot('after.heapsnapshot');
Loading both into DevTools' Memory panel and choosing Comparison is what finds a leak: it is not "what is big", it is "what only grows". The three leaks worth checking first are a growing array or Map used as a cache with no eviction, a listener added per request (part 6), and a closure a timer still holds. A rising heap that falls back down under load is just garbage collection working.
The diagnostic report: when the process will not tell you
# On a hang, without a profiler attached:
kill -USR2 <pid> # writes report---*.json with stacks, memory, libuv handles
# Or force it from code when something is wrong.
process.report.writeReport('./report');
This is the tool for the one incident profiling cannot reach: the event loop is blocked, so nothing responds and the profiler endpoint is unreachable. The report includes the JavaScript stack of the blocked thread and the active libuv handles — which is usually enough to name the culprit outright. Next: wiring your own diagnostics into that same story, and then leaving the profiler behind for the tools built in.