Engineering Performance
Server-Timing: Explain Backend Latency in the Browser
Use Server-Timing to expose useful backend durations, separate network and application costs, and avoid leaking internal topology through performance headers.
In this article
The Server-Timing response header lets a server describe selected performance measurements to the browser. It can help distinguish database work, cache behavior and application processing from the rest of a page's wait. Use bounded, non-sensitive measurements and verify them against server-side evidence.
The header does not automatically measure anything. Your application decides which durations to record. It also does not replace end-to-end latency, tracing or careful measurement of user experience.
Start with one useful measurement
Server-Timing: app;dur=42.5, db;dur=18.2This illustrative response reports durations in milliseconds. Read the Server-Timing reference for syntax and browser behavior. Keep names short and stable so developers can compare requests without decoding a different label every time.
Choose a measurement that answers a real diagnostic question, such as how long the application spent waiting for a database operation. Do not emit dozens of values simply because instrumentation can collect them. A concise header is easier to interpret and less likely to expose unnecessary details.
Use an appropriate clock
Measure elapsed duration with a monotonic clock supported by the runtime. Wall-clock adjustments can distort a duration calculated by subtracting two calendar timestamps. Keep calendar time for event chronology and monotonic time for elapsed work.
Define the start and end boundaries explicitly. Does app include authentication, rendering and downstream calls? Does db include connection acquisition? A number without a boundary definition can create disagreements even when the calculation is correct.
Keep the measurement close to the operation. Avoid a broad timer that remains active across unrelated asynchronous work unless that is exactly what you intend to describe. Concurrent tasks make elapsed durations especially easy to misinterpret.
Do not add overlapping durations blindly
If database calls run concurrently, their summed durations can exceed total request elapsed time. An application timer may already include the database wait. Adding app and db as though they were independent components can therefore double-count work.
Document whether a measurement is elapsed wall duration, cumulative work or another quantity. For a simple public header, elapsed measurements with clear boundaries are often easiest to understand. Keep more detailed relationships in tracing when they are needed.
Use CSV Viewer to inspect a sanitized sample of total and component measurements. Compare requests with the same route and conditions rather than mixing a cache hit with a cold expensive request and attributing the difference to code alone.
Inspect the actual response
Use browser developer tools to check the header and performance view. For a public URL, HTTP Header Checker can help establish whether the deployed response includes it. A proxy may remove or replace headers, so local output is not enough.
Cross-origin resource timing has its own visibility rules. Follow the browser documentation for the intended origin relationship rather than assuming every resource's server timings are readable by application JavaScript. Do not broaden resource timing access without considering what is exposed.
Test successful and error responses. A slow request that fails may be the most useful diagnostic case, yet middleware sometimes adds timings only to the happy path. Keep the error response safe while preserving a limited performance signal where appropriate.
Keep operational details private
Avoid hostnames, database names, tenant identifiers, query text and secrets in descriptions. Public headers can be read by anyone receiving the response. A label such as db is usually more appropriate than naming a private database shard and its region.
Decide whether authenticated internal endpoints can expose richer detail under an explicit policy. Do not assume that the word 'admin' in a route makes all headers private. Review the actual access boundary and logging behavior.
Also avoid high-cardinality metric labels generated from the timing names. If you aggregate these measurements, use stable dimensions. Metric Cardinality explains why request-specific values can create monitoring problems.
Interpret browser wait as a larger workflow
Compare the server's reported processing time with connection establishment, network transfer and client rendering. A small app duration with a slow page does not prove the user imagined the delay; the cost may be outside the instrumented boundary.
Likewise, a large database duration suggests where to investigate but does not prove the database alone is defective. Connection contention, remote network latency or an inefficient application query can contribute to the observed wait.
Use a representative request and correlate it with a trace or controlled log record. Keep the request identifier out of public timing descriptions unless your policy explicitly permits it. The header should guide investigation, not become a complete incident record.
Release and verify
Check syntax, units, boundary definitions and absence of sensitive descriptions. Confirm the production proxy preserves the intended header and that enabling measurement adds acceptable overhead. Keep a small regression fixture for the response format.
Show enough timing to explain the wait
Start with a few clear elapsed measurements, inspect the deployed response and compare them with end-to-end evidence. Server-Timing is most useful when it answers a specific question without leaking the system behind the page.