src/error.c now scores 81.2%: 238 of 293 mutants killed, 204 by a failing test, 24 by failing to compile, and 10 by hanging the suite -- deleting akerr_mutex_init() or the akerr_initializing re-entry guard deadlocks the first test, which is the right answer for a broken lock. Lock deletions are the one survivor category where surviving does not mean harmless, so measure it rather than assume: rebuilt, the surviving "delete the pool lock" mutant fails tests/err_threads_pool.c in 4 of 10 runs and fails under scripts/thread_test.sh in 5 of 5, with no false positive on the unmutated library. The property assertions alone are a coin flip on a missing lock; the sanitizer run is what holds that line. The harness builds mutants with default CMake options and so never sees it -- TODO item 8. Also warn that a sanitized test binary run by hand does not inherit the halt_on_error CTest gives it, and will print a race and still exit 0. That is how the 5-of-5 above first read as 2 of 5. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
7.7 KiB
Mutation testing
The unit tests tell us the library works. Mutation testing tells us the tests work — that they would actually fail if the library were broken.
scripts/mutation_test.py deliberately breaks the library in small ways
("mutants"), one at a time, and runs the whole CTest suite against each broken
copy:
- if the tests fail, the mutant is killed — good, the suite caught it;
- if the tests still pass, the mutant survived — a bug of that shape would slip through, so it points at a missing test.
The mutation score is killed / (killed + survived). A surviving mutant is
a to-do item: write a test that distinguishes the mutant from the original.
Running
No third-party tools are required — just Python 3 and the normal cmake/ctest toolchain. The harness never touches your working tree; it copies the repo to a scratch directory and mutates the copy.
# Default: mutate src/error.c and include/akerror.tmpl.h
scripts/mutation_test.py
# Faster: just the C source
scripts/mutation_test.py --target src/error.c
# See what would run without building anything
scripts/mutation_test.py --target src/error.c --list
# Gate CI: exit non-zero if the score drops below 90%
scripts/mutation_test.py --threshold 90
Via CMake (configures a build first if needed):
cmake --build build --target mutation
Useful flags: --timeout SECONDS (per-suite build+test cap; a mutant that
hangs is counted as killed), --keep (retain the scratch copy for debugging),
--work DIR (use a specific scratch directory), --junit FILE (write a JUnit
XML report — surviving mutants appear as failing test cases).
CI reporting
Both the unit tests and the mutation run emit JUnit XML that CI consumes:
ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml"— note the absolute path;--output-junitotherwise resolves relative to the test dir.scripts/mutation_test.py --junit mutation-junit.xml
.gitea/workflows/ci.yaml runs both and feeds the XML to
mikepenz/action-junit-report (with if: always(), so results publish even
when a gate fails). The reporter runs with annotate_only: true: Gitea does not
implement the Checks API the action uses to create a check run, so creating one
404s (mikepenz/action-junit-report#23). annotate_only skips that call and the
results surface via the job summary (detailed_summary: true) instead. The
generated *-junit.xml files are git-ignored.
Mutation operators
Each mutant changes exactly one location by one of:
| Tag | Operator | Example |
|---|---|---|
| ROR | relational operator | == → !=, < → <=, >= → > |
| LCR | logical connector | && → || |
| BCR | boolean constant | true → false |
| AOR | arithmetic / compound assign | + → -, += → -= |
| ICR | integer literal | 0 → 1, 1 → 0 |
| SDL | statement deletion | err->refcount += 1; → (removed) |
Preprocessor control lines, comments, and the block of error-code / buffer-size
#defines are skipped: mutating those produces equivalent or uninteresting
mutants that only add noise.
Interpreting survivors
Not every survivor is a test gap — some mutants are equivalent (they don't change observable behaviour, e.g. resizing an internal scratch buffer). For each survivor, decide:
- Real gap → add or strengthen a test in
tests/so the mutant is killed, then re-run. - Equivalent mutant → no test can catch it; leave a note. If a specific
line is a persistent source of equivalents, narrow the target with
--targetor extend the skip rules inscripts/mutation_test.py.
Re-run after adding tests and confirm the score went up.
Current status
src/error.c scores 81.2% — 238 of 293 mutants killed (204 by a failing test,
24 by failing to compile, 10 by hanging the suite), 55 surviving. The CI gate is
set to 65% for headroom.
The ten timeout kills are all in the locking: deleting akerr_mutex_init() or
the akerr_initializing re-entry guard deadlocks the very first test, which is
the correct behaviour for a broken lock and is why the harness counts a hang as
a kill.
The remaining survivors are dominated by:
- Equivalent mutants in
akerr_init: deleting thememset/NULLsetup of file-scope statics (AKERR_ARRAY_ERROR,__akerr_last_ditch,__akerr_last_ignored) changes nothing, because C already zero-initializes objects with static storage duration.int oldid = 0;→1is likewise dead: it is overwritten before use, and so is clearingakerr_initializingat the end of initialization — nothing reads that flag once the once-routine has returned. - Lock acquisition (
akerr_mutex_lock/unlockdeletions, and theakerr_init()call at the head of an entry point). These are the one category where a survivor does not mean the mutant is harmless. Removing a lock leaves a real race, and the assertions intests/err_threads_pool.conly fire when the race actually loses: rebuilding the surviving mutant and running that test ten times caught it four times. The same mutant underscripts/thread_test.shfailed five of five, with no false positive on the unmutated library — but the mutation harness builds without sanitizers, so it never sees that. Deleting anakerr_init()call survives for a duller reason: something else has always initialized the library by the time that line runs. - Default logger / handler internals (
vfprintf,va_end, theerrctx == NULLbranch,exit(1)): killing these needs a subprocess-based test that captures a child's stderr and exit code, rather than the in-process capturing logger the other tests use. - Static assertions (
akerr_assert_name_slots_pow2and the occupancy cap it guards): a mutated compile-time assertion that still compiles has no runtime behavior to observe. Unkillable by construction — the assertion is itself the test, andtests/err_maxval.ccovers the runtime consequence. - Hash and probe details in
akerr_status_slot: dropping one of the multiply steps inakerr_status_hashleaves a worse but still correct hash, and probing backwards (slot - 1u) is an equally valid sequence over a power-of-two table. Both are behaviorally equivalent. - The
capacity <= 0guard inakerr_copy_string, which is defensive: both call sites pass a positive constant.
Findings surfaced by mutation testing:
-
Open: the harness builds every mutant with the default CMake options, so a mutant that only breaks under concurrency is judged by a suite running without ThreadSanitizer. Mutating under
-DAKERR_SANITIZE=threadwould close that, and needs a way to pass CMake options through to the mutant build. See "Mutation testing judges concurrency mutants without a sanitizer" inTODO.md. -
Superseded: status names now use a private sparse registry, so the old public
AKERR_MAX_ERR_VALUEceiling and its consumer ABI mismatch no longer exist.tests/err_maxval.ccovers arbitraryintvalues and registry exhaustion. -
Fixed: the open-addressing probe mask (
& (AKERR_STATUS_NAME_SLOTS - 1)) could be mutated to- 0or+ 1— both of which index past the end of the table — without any test noticing.tests/err_maxval.conly asserted that some names registered before the table filled, which a collapsed probe sequence still satisfies. It now requires a substantial number of entries and reads every one of them back by its own distinct name, so a probe that revisits slots fails on both counts.