Guide

LLM Load Testing with k6: Find Your App's Capacity

Measure your AI app against a controllable model mock, with constant arrival rates, p95 latency, error budgets and dropped iterations.

2026-09-28

Cyan light lines converge through a narrow slot in a dark metal test bench, with a small amber light beneath the opening.
A controlled bottleneck makes the limit visible; the measured limit still belongs to the tested app and mock.

1. Define the boundary before measuring capacity

You want to know how much traffic your application can handle while waiting for a model. Start with a controlled substitute: a separate HTTP service that waits 500 milliseconds and returns a fixed answer. Send traffic through your application, including the middleware, queues and database work you want to measure. Calling the substitute directly measures the substitute.

This guide provides a local rehearsal with a tiny forwarding app. Replace that app with your own staging service before drawing conclusions about your product. The mock has no inference engine, token stream or provider quota. Your result describes capacity under the stated simulated delay and response size. It cannot establish a provider's capacity or real streaming latency.

Use one request per iteration and a fixed arrival rate. In a closed execution model, slower responses postpone subsequent iterations, reducing the traffic offered to the service. An open model lets you ask a more useful question here: what happens when new work keeps arriving?

2. Start the model mock and the local app

Prerequisites are a Node.js runtime with built-in fetch and AbortSignal.timeout, and k6 on your PATH. Save the following as mock.mjs. It accepts one endpoint and returns complete JSON responses. The optional failure setting makes every tenth request fail when set to 10; leave it at zero for the baseline.

import http from 'node:http';
const delay = Number(process.env.DELAY_MS ?? 500);
const failEvery = Number(process.env.FAIL_EVERY ?? 0);
let requests = 0;
http.createServer((req, res) => {
  if (req.method !== 'POST' || req.url !== '/generate') {
    res.writeHead(404).end(); return;
  }
  req.resume();
  req.on('end', () => {
    const fail = failEvery > 0 && ++requests % failEvery === 0;
    setTimeout(() => {
      res.writeHead(fail ? 503 : 200,
        {'content-type': 'application/json'});
      res.end(JSON.stringify(fail ? {error: 'injected'} : {answer: 'mock-ok'}));
    }, delay);
  });
}).listen(9090, '127.0.0.1');

Save this second file as app.mjs. It forwards the request body and preserves the mock's status code. This deliberately small adapter gives you a working route before you connect your real application. It has a five-second upstream timeout and no automatic retries.

import http from 'node:http';
const model = process.env.MODEL_URL ?? 'http://127.0.0.1:9090/generate';
http.createServer(async (req, res) => {
  if (req.method !== 'POST' || req.url !== '/chat') {
    res.writeHead(404).end(); return;
  }
  try {
    const chunks = [];
    for await (const chunk of req) chunks.push(chunk);
    const upstream = await fetch(model, {
      method: 'POST', headers: {'content-type': 'application/json'},
      body: Buffer.concat(chunks), signal: AbortSignal.timeout(5000)
    });
    const body = await upstream.text();
    res.writeHead(upstream.status, {'content-type': 'application/json'});
    res.end(body);
  } catch {
    res.writeHead(502, {'content-type': 'application/json'});
    res.end(JSON.stringify({error: 'upstream'}));
  }
}).listen(8080, '127.0.0.1');

Run the commands below in two separate PowerShell terminals. Keep both processes running while k6 executes in a third terminal. Stop them with Ctrl+C when you finish. These servers bind to the local loopback interface.

# Terminal 1
$env:DELAY_MS='500'
$env:FAIL_EVERY='0'
node mock.mjs

# Terminal 2
node app.mjs

For your own app, route its model adapter to the mock's generate endpoint. Adapt the mock response schema if your adapter expects a different shape, and update the assertion accordingly. Confirm the app logs show a mock call. Disable response caching for this experiment or use distinct inputs; otherwise a repeated prompt can measure cache hits instead of the path you intended. Record retries, worker counts and queue limits with the run.

3. Run a small probe, then fixed load plateaus

Save this as load.js. The constant-arrival-rate executor schedules the requested number of iterations per second while virtual users are available. Do not add a sleep at the end: this executor already controls the pacing. Each iteration below makes exactly one HTTP request, without redirects.

import http from 'k6/http';
import { check } from 'k6';
import { Rate } from 'k6/metrics';
const appErrors = new Rate('app_errors');
http.setResponseCallback(http.expectedStatuses(200));
export const options = {
  scenarios: { app: {
    executor: 'constant-arrival-rate',
    rate: Number(__ENV.RATE || 2), timeUnit: '1s',
    duration: __ENV.DURATION || '60s',
    preAllocatedVUs: Number(__ENV.VUS || 200),
    maxVUs: Number(__ENV.VUS || 200), gracefulStop: '10s'
  } },
  thresholds: {
    http_req_duration: ['p(95)<1500'],
    http_req_failed: ['rate<0.01'],
    app_errors: ['rate<0.01'],
    dropped_iterations: ['count==0']
  }
};
export default function () {
  const res = http.post(__ENV.TARGET || 'http://127.0.0.1:8080/chat',
    JSON.stringify({prompt: 'Return the test sentinel.'}),
    {headers: {'Content-Type': 'application/json'}, timeout: '6s', redirects: 0});
  let answer;
  try { answer = res.json('answer'); } catch { answer = null; }
  const ok = check(res, {
    'valid mock response': r => r.status === 200 && answer === 'mock-ok'
  });
  appErrors.add(!ok);
}

The custom error rate counts invalid JSON, an incorrect answer and non-200 responses. That prevents a fast error page with status 200 from passing. The HTTP error rate is also shown, with only status 200 configured as expected. For a real app, replace the sentinel test with an equally specific response contract.

# Third terminal: a 10-second smoke test, then independent 60-second runs
k6 run -e RATE=1 -e DURATION=10s load.js
k6 run -e RATE=2 load.js
k6 run -e RATE=5 load.js
k6 run -e RATE=10 load.js
k6 run -e RATE=20 load.js

Run each command separately and save its output. The smoke test should show successful checks, no application errors and no dropped iterations. If it fails, fix the wiring before adding traffic. Wait for your app's queue to drain between plateaus. Repeat a short warm-up before recorded runs if startup work changes the results, and use the same procedure for every rate.

4. Read four limits together

The proposed acceptance policy is p95 below 1,500 milliseconds, HTTP and application errors each below 1%, and zero dropped iterations per run. These are exercise settings; substitute your service objectives before using the result for a release decision. k6 thresholds turn these conditions into a nonzero exit status when a limit fails.

According to the built-in metric definitions, HTTP request duration covers sending, waiting and receiving, excluding initial connection setup. Its p95 here measures a complete buffered HTTP response. It is not time to first token. Also inspect successful responses separately if failures become fast enough to make the combined percentile look better.

Dropped iterations are work k6 could not start; they are not failed HTTP responses. Too few virtual users or longer iteration durations can cause drops. At 20 starts per second and a six-second request timeout, about 120 concurrent requests could occupy users before allowing for overhead. The proposed 200-user allocation is a starting budget, not proof that the generator has enough resources.

5. Fill in the capacity worksheet

The table below is a worksheet, not a set of measured results. Use one row per 60-second run with 500 milliseconds of simulated delay and zero injected failures. Record completed iterations as well as drops; configured traffic is not evidence that the app received all of it.

Offered iterations/sCompleted iterationsp95, msHTTP / app errors, %Dropped iterationsPass?
2RecordRecordRecord / RecordRecordRecord
5RecordRecordRecord / RecordRecordRecord
10RecordRecordRecord / RecordRecordRecord
20RecordRecordRecord / RecordRecordRecord

Your answer is a bracket: the highest tested passing rate and the first higher failing rate, under these conditions. If every row passes, you have a lower bound, not the maximum capacity. Repeat around the boundary and retain all runs rather than selecting the best one. A 60-second plateau is a short capacity probe; run a longer soak separately to look for growing queues or memory.

6. Prove that failure detection works

Restart the mock with FAIL_EVERY=10 and repeat the low-load run. With the sample app and no other traffic, every tenth mock request returns 503. Over 120 completed requests, the expected injected count is 12 and both error rates should be 10%. This is an answer key for the proposed control, not a reported test result. The error thresholds must fail even if latency stays low.

Then restore zero failures and increase the delay to 2,000 milliseconds. Successful full responses should now breach the 1,500-millisecond latency budget. A passing run suggests caching, incorrect routing or an assertion that is not measuring the intended path. Keep delay and failure changes in separate experiments so you can explain which condition triggered a failure.

If the baseline fails, compare app CPU, queue depth, connection usage and mock health. Check the load generator's CPU and memory too. A shared laptop can constrain all three processes. Increasing virtual users may resolve a generator configuration problem; it does not repair an overloaded app. For the surrounding failure policy, see AI reliability engineering. Use testing with AI to keep the response contract covered as the application changes.

Attach the app revision, machine resources, runtime and k6 versions, mock delay, failure rule, response size, warm-up procedure and all run outputs to your conclusion. The code is supplied as a reproducible exercise; no k6 execution or capacity measurement is claimed for this guide.

Sources checked 2026-09-28