ai-assisted-todolist

Mutation testing#

Coverage says a line ran. Mutation testing asks a harder question: if that line were changed, would any test notice? PIT compiles small changes into the bytecode — negate a condition, return null, swap a boolean — and reports which ones no test caught.

cd backend
./mvnw verify                                     # runs it as part of the normal build
./mvnw org.pitest:pitest-maven:mutationCoverage   # just this, about 2 seconds
open target/pit-reports/index.html                # which mutants survived, line by line

Currently 21 mutants, all killed. Every kill names the test that caught it, so the report doubles as evidence that the plain unit tests are doing real work rather than merely executing lines.

This one reports; it does not block. Unlike the coverage gates there is no threshold, because three production classes of ten are in scope: a gate there would police a corner of the codebase while saying nothing about the other seven, which is a worse signal than an honest report. backend-ci.yml uploads pit-report as an artifact on every backend run, so the result is readable without running anything locally. Revisit the decision if the scope widens.

PIT cannot run a @QuarkusTest#

This is the constraint that shapes everything else. PIT runs each mutant in a fresh minion JVM, so a @QuarkusTest in scope means a full application boot — with its own Dev Services PostgreSQL and Keycloak containers — per mutant. Measured on this codebase:

Scope Result
AuthProviderResource, covered by a plain unit test 10 mutants, 2 seconds, all killed
UserService, covered by a @QuarkusTest 14 mutants, 163 seconds, most minions timed out
One lightweight @QuarkusTest allowed into scope coverage calculation goes from under 1s to 24s
The whole suite in scope aborts before mutating anything — KeycloakLoginFlowTest cannot start its containers inside the minion, and PIT requires a green suite

Worth knowing: PIT counts a timed-out or errored minion as a killed mutant, so the @QuarkusTest run reported "100% killed" while actually proving nothing. A mutant that merely stops the application booting scores the same as one a test genuinely caught. This is why the scope is narrow rather than merely slow.

The allowlist, and the pressure it applies#

targetTests names the non-Quarkus test classes explicitly instead of matching a pattern, because that way the failure mode is benign: a new plain unit test is simply not mutated until it is listed, whereas a pattern that swept in a new @QuarkusTest would make every run minutes slower without saying so.

targetClasses stays deliberately wide, and the production classes with no fast unit test are excluded one line at a time in excludedClasses. So a new production class is in scope by default and shows up in the report as uncovered mutants instead of being silently absent — deleting the AuthResource exclusion, for instance, adds eight NO_COVERAGE mutants and takes the reported score from 100% to 56%.

Being wide informs rather than enforces, since nothing fails. It only works if somebody reads the report, which is why it is uploaded as an artifact rather than left in target/.

The exclusion list is therefore a to-do list, and it has been collected on: service.* narrowed to service.UserService* when GravatarUrl and GravatarService earned plain unit tests, taking the run from 10 mutants to 21 and all of them killed. Every remaining line is a class whose behaviour is only pinned down by tests too heavy to mutate, and deleting a line is the reward for writing a fast one — remembering to add the new test to targetTests, or it is written but never used.