src/error.c now scores 81.4%: 245 of 301 mutants killed, 211 by a failing test, 24 by failing to compile, and 10 by hanging the suite. The population grew by 8 with akerr_exit(), and all 8 die, as do the 4 in akerr_default_handler_unhandled_error() -- no survivor anywhere in either function. Two of those are worth naming. Mutating the guard to status < 1 (the sentinel-for-zero behaviour this library deliberately does not have) and mutating the NULL-context exit(1) to exit(0) both die, so the tests pin the two ways an exit could start claiming success rather than just the one that was broken. That moves the default handler out of the "needs a subprocess test" survivor category noted here: it has one now. The default *logger* is still in it -- every other test swaps in the capturing logger, so nothing watches what the real one writes to stderr. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
8.1 KiB
Mutation testing
The unit tests tell us the library works. Mutation testing tells us the tests work — that they would actually fail if the library were broken.
scripts/mutation_test.py deliberately breaks the library in small ways
("mutants"), one at a time, and runs the whole CTest suite against each broken
copy:
- if the tests fail, the mutant is killed — good, the suite caught it;
- if the tests still pass, the mutant survived — a bug of that shape would slip through, so it points at a missing test.
The mutation score is killed / (killed + survived). A surviving mutant is
a to-do item: write a test that distinguishes the mutant from the original.
Running
No third-party tools are required — just Python 3 and the normal cmake/ctest toolchain. The harness never touches your working tree; it copies the repo to a scratch directory and mutates the copy.
# Default: mutate src/error.c and include/akerror.tmpl.h
scripts/mutation_test.py
# Faster: just the C source
scripts/mutation_test.py --target src/error.c
# See what would run without building anything
scripts/mutation_test.py --target src/error.c --list
# Gate CI: exit non-zero if the score drops below 90%
scripts/mutation_test.py --threshold 90
Via CMake (configures a build first if needed):
cmake --build build --target mutation
Useful flags: --timeout SECONDS (per-suite build+test cap; a mutant that
hangs is counted as killed), --keep (retain the scratch copy for debugging),
--work DIR (use a specific scratch directory), --junit FILE (write a JUnit
XML report — surviving mutants appear as failing test cases).
CI reporting
Both the unit tests and the mutation run emit JUnit XML that CI consumes:
ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml"— note the absolute path;--output-junitotherwise resolves relative to the test dir.scripts/mutation_test.py --junit mutation-junit.xml
.gitea/workflows/ci.yaml runs both and feeds the XML to
mikepenz/action-junit-report (with if: always(), so results publish even
when a gate fails). The reporter runs with annotate_only: true: Gitea does not
implement the Checks API the action uses to create a check run, so creating one
404s (mikepenz/action-junit-report#23). annotate_only skips that call and the
results surface via the job summary (detailed_summary: true) instead. The
generated *-junit.xml files are git-ignored.
Mutation operators
Each mutant changes exactly one location by one of:
| Tag | Operator | Example |
|---|---|---|
| ROR | relational operator | == → !=, < → <=, >= → > |
| LCR | logical connector | && → || |
| BCR | boolean constant | true → false |
| AOR | arithmetic / compound assign | + → -, += → -= |
| ICR | integer literal | 0 → 1, 1 → 0 |
| SDL | statement deletion | err->refcount += 1; → (removed) |
Preprocessor control lines, comments, and the block of error-code / buffer-size
#defines are skipped: mutating those produces equivalent or uninteresting
mutants that only add noise.
Interpreting survivors
Not every survivor is a test gap — some mutants are equivalent (they don't change observable behaviour, e.g. resizing an internal scratch buffer). For each survivor, decide:
- Real gap → add or strengthen a test in
tests/so the mutant is killed, then re-run. - Equivalent mutant → no test can catch it; leave a note. If a specific
line is a persistent source of equivalents, narrow the target with
--targetor extend the skip rules inscripts/mutation_test.py.
Re-run after adding tests and confirm the score went up.
Current status
src/error.c scores 81.4% — 245 of 301 mutants killed (211 by a failing test,
24 by failing to compile, 10 by hanging the suite), 56 surviving. The CI gate is
set to 65% for headroom.
The ten timeout kills are all in the locking: deleting akerr_mutex_init() or
the akerr_initializing re-entry guard deadlocks the very first test, which is
the correct behaviour for a broken lock and is why the harness counts a hang as
a kill.
The remaining survivors are dominated by:
- Equivalent mutants in
akerr_init: deleting thememset/NULLsetup of file-scope statics (AKERR_ARRAY_ERROR,__akerr_last_ditch,__akerr_last_ignored) changes nothing, because C already zero-initializes objects with static storage duration.int oldid = 0;→1is likewise dead: it is overwritten before use, and so is clearingakerr_initializingat the end of initialization — nothing reads that flag once the once-routine has returned. - Lock acquisition (
akerr_mutex_lock/unlockdeletions, and theakerr_init()call at the head of an entry point). These are the one category where a survivor does not mean the mutant is harmless. Removing a lock leaves a real race, and the assertions intests/err_threads_pool.conly fire when the race actually loses: rebuilding the surviving mutant and running that test ten times caught it four times. The same mutant underscripts/thread_test.shfailed five of five, with no false positive on the unmutated library — but the mutation harness builds without sanitizers, so it never sees that. Deleting anakerr_init()call survives for a duller reason: something else has always initialized the library by the time that line runs. - The default logger (
vfprintf,va_end, and thereturnin the no-stdlib branch): the other tests replaceakerr_log_methodwith the in-process capturing logger, so nothing observes what the default one writes to a real stderr. Killing these needs a test that captures a child's stderr. The handler internals next to them are no longer in this category:tests/err_unhandled_null.candtests/err_exit_status.cread a forked child's exit code, which kills every mutant inakerr_exit()and inakerr_default_handler_unhandled_error()— all twelve of them, including thestatus < 0→status < 1variant that only a test assertingakerr_exit(0)exits 0 can distinguish. - Static assertions (
akerr_assert_name_slots_pow2and the occupancy cap it guards): a mutated compile-time assertion that still compiles has no runtime behavior to observe. Unkillable by construction — the assertion is itself the test, andtests/err_maxval.ccovers the runtime consequence. - Hash and probe details in
akerr_status_slot: dropping one of the multiply steps inakerr_status_hashleaves a worse but still correct hash, and probing backwards (slot - 1u) is an equally valid sequence over a power-of-two table. Both are behaviorally equivalent. - The
capacity <= 0guard inakerr_copy_string, which is defensive: both call sites pass a positive constant.
Findings surfaced by mutation testing:
-
Open: the harness builds every mutant with the default CMake options, so a mutant that only breaks under concurrency is judged by a suite running without ThreadSanitizer. Mutating under
-DAKERR_SANITIZE=threadwould close that, and needs a way to pass CMake options through to the mutant build. See "Mutation testing judges concurrency mutants without a sanitizer" inTODO.md. -
Superseded: status names now use a private sparse registry, so the old public
AKERR_MAX_ERR_VALUEceiling and its consumer ABI mismatch no longer exist.tests/err_maxval.ccovers arbitraryintvalues and registry exhaustion. -
Fixed: the open-addressing probe mask (
& (AKERR_STATUS_NAME_SLOTS - 1)) could be mutated to- 0or+ 1— both of which index past the end of the table — without any test noticing.tests/err_maxval.conly asserted that some names registered before the table filled, which a collapsed probe sequence still satisfies. It now requires a substantial number of entries and reads every one of them back by its own distinct name, so a probe that revisits slots fails on both counts.