
Covered code can still have weak assertions
A test can execute the exact line that contains a bug and still pass. Consider a requirement for an age gate: within the accepted domain of integer ages, eligibility begins at 18, inclusive. Calling the function at that boundary gives execution coverage. Checking only that the result is a boolean leaves the important question unanswered: did the function accept the person?
This matters when reviewing AI-generated tests because plausible test names and green output can distract you from the assertion. The experiment below uses a deliberately constructed weak test, rather than an undocumented model run. It demonstrates a mechanism you can check in your own suite; it does not measure how frequently AI produces this mistake. For the wider workflow, start with the Testing with AI guide.
A bounded StrykerJS experiment
The production run on 2026-10-11 used Windows and Node v25.5.0, with StrykerJS core and its Mocha runner pinned to 10.0.0 and Mocha pinned to 11.7.5. Only one function was mutated. Both rounds used the same implementation, configuration and dependency lockfile. The first round checked the result type; the second replaced that test with assertions derived from the inclusive boundary requirement.
Save these files in an empty experiment directory. The package manifest pins the direct dependencies; retain the generated lockfile and use npm ci for subsequent installations so transitive resolutions also stay fixed. These versions identify this experiment, rather than a recommendation to update an existing project. Stryker's getting-started documentation describes installation and running the tool.
package.json
{"private":true,"type":"module","scripts":{"test":"mocha test.mjs","mutate":"stryker run"},"devDependencies":{"@stryker-mutator/core":"10.0.0","@stryker-mutator/mocha-runner":"10.0.0","mocha":"11.7.5"}}
eligibility.mjs
export function eligible(age) {
return age >= 18;
}
test.mjs β before
import assert from 'node:assert/strict';
import { eligible } from './eligibility.mjs';
describe('eligibility', () => {
it('returns a boolean at the adult boundary', () => {
assert.equal(typeof eligible(18), 'boolean');
});
});
stryker.config.mjs
export default {
ignorePatterns: ['npm-cache/**', 'reports/**', 'write-article.mjs', 'test-weak.mjs'],
mutate: ['eligibility.mjs'],
testRunner: 'mocha',
mochaOptions: { spec: ['test.mjs'] },
coverageAnalysis: 'perTest',
reporters: ['clear-text', 'json'],
jsonReporter: { fileName: 'reports/mutation.json' },
concurrency: 1,
timeoutMS: 5000,
timeoutFactor: 1.5
};
The configuration reference documents file selection, coverage analysis and timeout settings. Here, per-test coverage lets the report distinguish covered survivors from mutations with no covering test. One worker keeps the experiment small. The timeout configuration uses a 5,000 ms allowance plus 1.5 times baseline net execution time plus measured overhead; it is not a fixed 5,000 ms cap.
npm install --ignore-scripts --no-audit --no-fund
npm test
npm run mutate
On Windows PowerShell, use npm.cmd if your execution policy blocks npm.ps1. Save the first reports/mutation.json as weak.json before changing the test. After the second run, preserve that report as strong.json. Keeping the reports matters: a rounded percentage alone loses the mutant identities and the denominator.
test.mjs β after
import assert from 'node:assert/strict';
import { eligible } from './eligibility.mjs';
describe('eligibility', () => {
it('returns a boolean at the adult boundary', () => {
assert.equal(eligible(17), false);
assert.equal(eligible(18), true);
assert.equal(eligible(19), true);
});
});
Run the same test and mutation commands again. The strengthened test specifies rejection just below the boundary, acceptance at the boundary and acceptance just above it. These are requirement assertions, not assertions copied from the implementation's current return values. The function and its name remain unchanged between rounds.
Observed results
| Round | Killed | Survived | No coverage | Timeout | Mutation score |
|---|---|---|---|---|---|
| Type-only assertion | 1 | 4 | 0 | 0 | 1/5 = 20.00% |
| Requirement assertions | 5 | 0 | 0 | 0 | 5/5 = 100.00% |
In the first report, replacing the inclusive comparison with a strict greater-than comparison survived. The weak test executed the expression at 18, but false still has type boolean. In the second report, the same mutation was killed because the boundary assertion expects true. Neither report contained compile errors, runtime errors or ignored mutants. These counts describe the exact pinned experiment, not a benchmark of generated test suites.
Reading the outcomes and the denominator
Use Stryker's mutant states and metrics to interpret the report. Killed means a test failed with that mutation active. Survived means the tests passed. No coverage means no test exercised that mutant. Timeout means the tests did not finish within the configured allowance; it counts as detected, although it does not establish that an assertion found the fault.
detected = killed + timeout
valid = detected + survived + noCoverage
mutationScore = detected / valid * 100
Compile errors, runtime errors, ignored mutants and pending mutants do not enter that valid denominator. The covered-code score instead divides detected by detected plus survived. Mixing those scores can hide untested paths. In this experiment, no-coverage and timeout counts were zero in both rounds, so the improvement came from the assertions. Preserve those categories when comparing larger runs.
Equivalent mutants and legitimate exclusions
A survivor is an investigation lead. It can reveal a missing input, a weak assertion or a behavior-preserving transformation. Stryker's equivalent-mutant explanation shows why apparently changed code need not change observable behavior. You need to reason about the accepted input domain and externally observable outputs before labelling a survivor equivalent.
For example, under an explicit integer-only domain, comparisons against 18 using greater-than-or-equal and against 17 using greater-than accept the same values. That reasoning would fail if fractional ages were accepted: 17.5 distinguishes them. This is a separate conceptual example, not a mutant observed in our report. The observed change from inclusive to strict comparison at 18 is not equivalent, because 18 itself distinguishes the outcomes.
If a mutation is excluded, record its location, the transformation, the domain argument and who reviewed it. Revisit that decision when the domain changes. Excluding a hard-to-kill mutant solely to raise the score changes what the number measures. An exclusion can be justified for a documented reason; the report should make the narrowed scope visible.
Improve assertions without teaching to the score
When a mutation survives, write down the behavior it would change in user terms before modifying the test. For this function, the question is whether an exactly 18-year-old is accepted. A useful AI review prompt is: βThe requirement accepts integer ages at least 18. Review these tests for missing observable outcomes and propose boundary assertions. Explain each expected value from that requirement.β
The matching evaluation criterion is whether the tests reject an underage input and accept the inclusive boundary and an older input. Asking only for a higher mutation score invites changes aimed at the report instead of the behavior. Independently review the requirement: if it is wrong, a precise test can faithfully preserve the wrong policy.
A high score does not show that every requirement has a test, that invalid inputs are handled, or that the gate is wired correctly into the application. Our function has no input validation or integration coverage. Add those checks when the actual contract requires them, and use an AI regression test plan to identify affected callers. Mutation testing adds evidence about selected transformations; it cannot supply the missing product specification.
When selective mutation testing is worth its runtime
Start with a changed function whose mistake would matter: an authorization boundary, a quota rule or a parser decision. Name the files being mutated, retain the report, and inspect survivors before expanding the run. This concentrates review effort where a concrete fault can be explained. It also leaves explicit gaps: a selected-file run provides no mutation evidence for excluded modules.
Measure wall-clock time in your own CI before making the run a required check. Our tiny experiment does not predict repository-scale runtime, and worker settings or test isolation can change operational behavior. Compare useful findings with the additional feedback delay. If a run mostly produces equivalent transformations, document them and reassess scope rather than repeatedly chasing a perfect percentage.
For a pull request, attach the mutated scope, valid denominator, category counts and the requirement behind each new assertion. A reviewer should be able to explain why the original test missed the boundary and why the new one detects it. That explanation is the durable improvement; the mutation score is a compact record of this particular run.
Sources and experiment record
- StrykerJS getting started: installation and running the tool.
- StrykerJS configuration: file selection, coverage analysis and timeout settings.
- Mutant states and metrics: report categories and the mutation-score denominator.
- Equivalent mutants: transformations that preserve observable behavior.
Checked: 2026-10-11. The observed results above come from our constructed experiment. Recheck the commands and tool definitions when changing the pinned versions.