How do I design a realistic load test and interpret p95/p99 latency results?
Asked by The SDET Playbook
Asked Sep 28, 2026Viewed 0 times
How do I design a realistic load test and interpret p95/p99 latency results?
Asked by The SDET Playbook
Sign in to answer and to vote.
Model real traffic and report the slow tail, not the average. Percentiles show how slow the slowest requests are: p95 means 95% of requests finished within that time, and p99 exposes the tail that averages hide. Use ramp-up, steady-state and ramp-down stages, include think time, and prefer arrival-rate models when you want a fixed request rate. Then set thresholds tied to a service-level objective (for example, p95 under 500 ms and under 1% errors, which is a common illustrative pattern, not a universal target). k6 exits with a non-zero code on a failed threshold so the pipeline can stop automatically. Use different shapes for different questions: load, stress, spike and soak.
In k6, the stages look like this:
export const options = {
stages: [
{ duration: '2m', target: 50 }, // ramp up
{ duration: '5m', target: 50 }, // steady state
{ duration: '2m', target: 200 }, // stress
{ duration: '2m', target: 0 }, // ramp down
],
};
Read the percentiles next to throughput (requests per second) and the error rate: latency that looks healthy while errors climb or throughput flattens means the system is shedding load, not coping with it.
Sources: Grafana k6 course, QASkills p95/p99 guide