Files
libakerror/tests/MUTATION.md
Andrew Kesterson 076b80c846
Some checks failed
libakerror CI Build / cmake_build (push) Successful in 2m48s
libakerror CI Build / coverage (push) Successful in 2m47s
libakerror CI Build / thread_sanitizer (push) Failing after 2m50s
libakerror CI Build / mutation_test (push) Successful in 39m13s
Record the mutation score after the exit-status fix
src/error.c now scores 81.4%: 245 of 301 mutants killed, 211 by a
failing test, 24 by failing to compile, and 10 by hanging the suite.
The population grew by 8 with akerr_exit(), and all 8 die, as do the 4
in akerr_default_handler_unhandled_error() -- no survivor anywhere in
either function.

Two of those are worth naming. Mutating the guard to status < 1 (the
sentinel-for-zero behaviour this library deliberately does not have) and
mutating the NULL-context exit(1) to exit(0) both die, so the tests pin
the two ways an exit could start claiming success rather than just the
one that was broken.

That moves the default handler out of the "needs a subprocess test"
survivor category noted here: it has one now. The default *logger* is
still in it -- every other test swaps in the capturing logger, so
nothing watches what the real one writes to stderr.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 07:57:33 -04:00

8.1 KiB

Mutation testing

The unit tests tell us the library works. Mutation testing tells us the tests work — that they would actually fail if the library were broken.

scripts/mutation_test.py deliberately breaks the library in small ways ("mutants"), one at a time, and runs the whole CTest suite against each broken copy:

  • if the tests fail, the mutant is killed — good, the suite caught it;
  • if the tests still pass, the mutant survived — a bug of that shape would slip through, so it points at a missing test.

The mutation score is killed / (killed + survived). A surviving mutant is a to-do item: write a test that distinguishes the mutant from the original.

Running

No third-party tools are required — just Python 3 and the normal cmake/ctest toolchain. The harness never touches your working tree; it copies the repo to a scratch directory and mutates the copy.

# Default: mutate src/error.c and include/akerror.tmpl.h
scripts/mutation_test.py

# Faster: just the C source
scripts/mutation_test.py --target src/error.c

# See what would run without building anything
scripts/mutation_test.py --target src/error.c --list

# Gate CI: exit non-zero if the score drops below 90%
scripts/mutation_test.py --threshold 90

Via CMake (configures a build first if needed):

cmake --build build --target mutation

Useful flags: --timeout SECONDS (per-suite build+test cap; a mutant that hangs is counted as killed), --keep (retain the scratch copy for debugging), --work DIR (use a specific scratch directory), --junit FILE (write a JUnit XML report — surviving mutants appear as failing test cases).

CI reporting

Both the unit tests and the mutation run emit JUnit XML that CI consumes:

  • ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml" — note the absolute path; --output-junit otherwise resolves relative to the test dir.
  • scripts/mutation_test.py --junit mutation-junit.xml

.gitea/workflows/ci.yaml runs both and feeds the XML to mikepenz/action-junit-report (with if: always(), so results publish even when a gate fails). The reporter runs with annotate_only: true: Gitea does not implement the Checks API the action uses to create a check run, so creating one 404s (mikepenz/action-junit-report#23). annotate_only skips that call and the results surface via the job summary (detailed_summary: true) instead. The generated *-junit.xml files are git-ignored.

Mutation operators

Each mutant changes exactly one location by one of:

Tag Operator Example
ROR relational operator ==!=, <<=, >=>
LCR logical connector &&||
BCR boolean constant truefalse
AOR arithmetic / compound assign +-, +=-=
ICR integer literal 01, 10
SDL statement deletion err->refcount += 1;(removed)

Preprocessor control lines, comments, and the block of error-code / buffer-size #defines are skipped: mutating those produces equivalent or uninteresting mutants that only add noise.

Interpreting survivors

Not every survivor is a test gap — some mutants are equivalent (they don't change observable behaviour, e.g. resizing an internal scratch buffer). For each survivor, decide:

  1. Real gap → add or strengthen a test in tests/ so the mutant is killed, then re-run.
  2. Equivalent mutant → no test can catch it; leave a note. If a specific line is a persistent source of equivalents, narrow the target with --target or extend the skip rules in scripts/mutation_test.py.

Re-run after adding tests and confirm the score went up.

Current status

src/error.c scores 81.4% — 245 of 301 mutants killed (211 by a failing test, 24 by failing to compile, 10 by hanging the suite), 56 surviving. The CI gate is set to 65% for headroom.

The ten timeout kills are all in the locking: deleting akerr_mutex_init() or the akerr_initializing re-entry guard deadlocks the very first test, which is the correct behaviour for a broken lock and is why the harness counts a hang as a kill.

The remaining survivors are dominated by:

  • Equivalent mutants in akerr_init: deleting the memset/NULL setup of file-scope statics (AKERR_ARRAY_ERROR, __akerr_last_ditch, __akerr_last_ignored) changes nothing, because C already zero-initializes objects with static storage duration. int oldid = 0;1 is likewise dead: it is overwritten before use, and so is clearing akerr_initializing at the end of initialization — nothing reads that flag once the once-routine has returned.
  • Lock acquisition (akerr_mutex_lock/unlock deletions, and the akerr_init() call at the head of an entry point). These are the one category where a survivor does not mean the mutant is harmless. Removing a lock leaves a real race, and the assertions in tests/err_threads_pool.c only fire when the race actually loses: rebuilding the surviving mutant and running that test ten times caught it four times. The same mutant under scripts/thread_test.sh failed five of five, with no false positive on the unmutated library — but the mutation harness builds without sanitizers, so it never sees that. Deleting an akerr_init() call survives for a duller reason: something else has always initialized the library by the time that line runs.
  • The default logger (vfprintf, va_end, and the return in the no-stdlib branch): the other tests replace akerr_log_method with the in-process capturing logger, so nothing observes what the default one writes to a real stderr. Killing these needs a test that captures a child's stderr. The handler internals next to them are no longer in this category: tests/err_unhandled_null.c and tests/err_exit_status.c read a forked child's exit code, which kills every mutant in akerr_exit() and in akerr_default_handler_unhandled_error() — all twelve of them, including the status < 0status < 1 variant that only a test asserting akerr_exit(0) exits 0 can distinguish.
  • Static assertions (akerr_assert_name_slots_pow2 and the occupancy cap it guards): a mutated compile-time assertion that still compiles has no runtime behavior to observe. Unkillable by construction — the assertion is itself the test, and tests/err_maxval.c covers the runtime consequence.
  • Hash and probe details in akerr_status_slot: dropping one of the multiply steps in akerr_status_hash leaves a worse but still correct hash, and probing backwards (slot - 1u) is an equally valid sequence over a power-of-two table. Both are behaviorally equivalent.
  • The capacity <= 0 guard in akerr_copy_string, which is defensive: both call sites pass a positive constant.

Findings surfaced by mutation testing:

  • Open: the harness builds every mutant with the default CMake options, so a mutant that only breaks under concurrency is judged by a suite running without ThreadSanitizer. Mutating under -DAKERR_SANITIZE=thread would close that, and needs a way to pass CMake options through to the mutant build. See "Mutation testing judges concurrency mutants without a sanitizer" in TODO.md.

  • Superseded: status names now use a private sparse registry, so the old public AKERR_MAX_ERR_VALUE ceiling and its consumer ABI mismatch no longer exist. tests/err_maxval.c covers arbitrary int values and registry exhaustion.

  • Fixed: the open-addressing probe mask (& (AKERR_STATUS_NAME_SLOTS - 1)) could be mutated to - 0 or + 1 — both of which index past the end of the table — without any test noticing. tests/err_maxval.c only asserted that some names registered before the table filled, which a collapsed probe sequence still satisfies. It now requires a substantial number of entries and reads every one of them back by its own distinct name, so a probe that revisits slots fails on both counts.