Trust is earned, not given

A different perspective

2021-06-15 · Projects

Node.js, part 14: the profiler you already have — --cpu-prof, heap snapshots, and the flame graph

Part 14from the Node.js series · 24 parts in all

"It's slow" is not a diagnosis, and adding console.time everywhere is not a measurement. Node ships a CPU profiler, a heap snapshot mechanism, and a diagnostic report generator. Part 14 is using them on a running service — the difference between a guess and a flame graph.

CPU: sample, do not instrument

# Profile from process start, write a .cpuprofile on exit.
node --cpu-prof --cpu-prof-dir=./prof server.js

# Or attach to a running process and start/stop sampling at will.
node --inspect server.js
#   then in chrome://inspect -> Profiler -> Start/Stop

The .cpuprofile opens directly in Chrome DevTools, which turns it into a flame graph. Read it from the top down: the widest frames are where time goes, and the useful surprise is usually finding a frame you did not write — a JSON parse, a regex, a serialiser in a driver. That is the whole method: measure, find the widest frame, fix it, measure again.

Heap: find the leak, not the usage

const v8 = require('v8');
const { writeHeapSnapshot } = v8;

// Two snapshots, taken minutes apart, with traffic in between.
writeHeapSnapshot('before.heapsnapshot');
// ... let the service run ...
writeHeapSnapshot('after.heapsnapshot');

Loading both into DevTools' Memory panel and choosing Comparison is what finds a leak: it is not "what is big", it is "what only grows". The three leaks worth checking first are a growing array or Map used as a cache with no eviction, a listener added per request (part 6), and a closure a timer still holds. A rising heap that falls back down under load is just garbage collection working.

The diagnostic report: when the process will not tell you

# On a hang, without a profiler attached:
kill -USR2 <pid>          # writes report---*.json with stacks, memory, libuv handles

# Or force it from code when something is wrong.
process.report.writeReport('./report');

This is the tool for the one incident profiling cannot reach: the event loop is blocked, so nothing responds and the profiler endpoint is unreachable. The report includes the JavaScript stack of the blocked thread and the active libuv handles — which is usually enough to name the culprit outright. Next: wiring your own diagnostics into that same story, and then leaving the profiler behind for the tools built in.