Files
libakerror/tests/MUTATION.md
Tachikoma 24469e83d4 Move outstanding work from TODO.md into the issue tracker
Fourteen issues on source.starfort.tech/andrew/libakerror, labelled by kind and
blast radius and milestoned by what they can land in: 2.0.x for anything that
breaks no ABI, 2.1.0 for additive surface, 3.0.0 for the handler-ladder rewrite.
Everything filed carries status::grooming.

Three of the fourteen were not in this file at all. They are defects in this
library that were written down in a consumer's TODO instead: IGNORE() leaking a
context (#14), the coverage target not being namespaced when embedded (#15),
and no akerrorConfigVersion.cmake being installed (#16). libakstdlib carries a
workaround for each. That is the finding worth keeping -- a library whose
consumers record its defects in their own files cannot see them, and the
workaround gets written once per consumer.

TODO.md keeps the reasoning: why the handler ladder is a major-version change,
why a copied akerr_ErrorContext is a trap in two specific fields, why
validating more inputs lowers branch coverage by construction, why the mutation
score is a floor, and why a registered name is returned by pointer.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Andrew Kesterson <andrew@aklabs.net>
2026-08-02 19:00:22 -04:00

8.1 KiB

Mutation testing

The unit tests tell us the library works. Mutation testing tells us the tests work — that they would actually fail if the library were broken.

scripts/mutation_test.py deliberately breaks the library in small ways ("mutants"), one at a time, and runs the whole CTest suite against each broken copy:

  • if the tests fail, the mutant is killed — good, the suite caught it;
  • if the tests still pass, the mutant survived — a bug of that shape would slip through, so it points at a missing test.

The mutation score is killed / (killed + survived). A surviving mutant is a to-do item: write a test that distinguishes the mutant from the original.

Running

No third-party tools are required — just Python 3 and the normal cmake/ctest toolchain. The harness never touches your working tree; it copies the repo to a scratch directory and mutates the copy.

# Default: mutate src/error.c and include/akerror.tmpl.h
scripts/mutation_test.py

# Faster: just the C source
scripts/mutation_test.py --target src/error.c

# See what would run without building anything
scripts/mutation_test.py --target src/error.c --list

# Gate CI: exit non-zero if the score drops below 90%
scripts/mutation_test.py --threshold 90

Via CMake (configures a build first if needed):

cmake --build build --target mutation

Useful flags: --timeout SECONDS (per-suite build+test cap; a mutant that hangs is counted as killed), --keep (retain the scratch copy for debugging), --work DIR (use a specific scratch directory), --junit FILE (write a JUnit XML report — surviving mutants appear as failing test cases).

CI reporting

Both the unit tests and the mutation run emit JUnit XML that CI consumes:

  • ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml" — note the absolute path; --output-junit otherwise resolves relative to the test dir.
  • scripts/mutation_test.py --junit mutation-junit.xml

.gitea/workflows/ci.yaml runs both and feeds the XML to mikepenz/action-junit-report (with if: always(), so results publish even when a gate fails). The reporter runs with annotate_only: true: Gitea does not implement the Checks API the action uses to create a check run, so creating one 404s (mikepenz/action-junit-report#23). annotate_only skips that call and the results surface via the job summary (detailed_summary: true) instead. The generated *-junit.xml files are git-ignored.

Mutation operators

Each mutant changes exactly one location by one of:

Tag Operator Example
ROR relational operator ==!=, <<=, >=>
LCR logical connector &&||
BCR boolean constant truefalse
AOR arithmetic / compound assign +-, +=-=
ICR integer literal 01, 10
SDL statement deletion err->refcount += 1;(removed)

Preprocessor control lines, comments, and the block of error-code / buffer-size #defines are skipped: mutating those produces equivalent or uninteresting mutants that only add noise.

Interpreting survivors

Not every survivor is a test gap — some mutants are equivalent (they don't change observable behaviour, e.g. resizing an internal scratch buffer). For each survivor, decide:

  1. Real gap → add or strengthen a test in tests/ so the mutant is killed, then re-run.
  2. Equivalent mutant → no test can catch it; leave a note. If a specific line is a persistent source of equivalents, narrow the target with --target or extend the skip rules in scripts/mutation_test.py.

Re-run after adding tests and confirm the score went up.

Current status

src/error.c scores 81.4% — 245 of 301 mutants killed (211 by a failing test, 24 by failing to compile, 10 by hanging the suite), 56 surviving. The CI gate is set to 65% for headroom.

The ten timeout kills are all in the locking: deleting akerr_mutex_init() or the akerr_initializing re-entry guard deadlocks the very first test, which is the correct behaviour for a broken lock and is why the harness counts a hang as a kill.

The remaining survivors are dominated by:

  • Equivalent mutants in akerr_init: deleting the memset/NULL setup of file-scope statics (AKERR_ARRAY_ERROR, __akerr_last_ditch, __akerr_last_ignored) changes nothing, because C already zero-initializes objects with static storage duration. int oldid = 0;1 is likewise dead: it is overwritten before use, and so is clearing akerr_initializing at the end of initialization — nothing reads that flag once the once-routine has returned.
  • Lock acquisition (akerr_mutex_lock/unlock deletions, and the akerr_init() call at the head of an entry point). These are the one category where a survivor does not mean the mutant is harmless. Removing a lock leaves a real race, and the assertions in tests/err_threads_pool.c only fire when the race actually loses: rebuilding the surviving mutant and running that test ten times caught it four times. The same mutant under scripts/thread_test.sh failed five of five, with no false positive on the unmutated library — but the mutation harness builds without sanitizers, so it never sees that. Deleting an akerr_init() call survives for a duller reason: something else has always initialized the library by the time that line runs.
  • The default logger (vfprintf, va_end, and the return in the no-stdlib branch): the other tests replace akerr_log_method with the in-process capturing logger, so nothing observes what the default one writes to a real stderr. Killing these needs a test that captures a child's stderr. The handler internals next to them are no longer in this category: tests/err_unhandled_null.c and tests/err_exit_status.c read a forked child's exit code, which kills every mutant in akerr_exit() and in akerr_default_handler_unhandled_error() — all twelve of them, including the status < 0status < 1 variant that only a test asserting akerr_exit(0) exits 0 can distinguish.
  • Static assertions (akerr_assert_name_slots_pow2 and the occupancy cap it guards): a mutated compile-time assertion that still compiles has no runtime behavior to observe. Unkillable by construction — the assertion is itself the test, and tests/err_maxval.c covers the runtime consequence.
  • Hash and probe details in akerr_status_slot: dropping one of the multiply steps in akerr_status_hash leaves a worse but still correct hash, and probing backwards (slot - 1u) is an equally valid sequence over a power-of-two table. Both are behaviorally equivalent.
  • The capacity <= 0 guard in akerr_copy_string, which is defensive: both call sites pass a positive constant.

Findings surfaced by mutation testing:

  • Open: the harness builds every mutant with the default CMake options, so a mutant that only breaks under concurrency is judged by a suite running without ThreadSanitizer. Mutating under -DAKERR_SANITIZE=thread would close that, and needs a way to pass CMake options through to the mutant build. That is issue #10; TODO.md records why the score is a floor rather than a verdict.

  • Superseded: status names now use a private sparse registry, so the old public AKERR_MAX_ERR_VALUE ceiling and its consumer ABI mismatch no longer exist. tests/err_maxval.c covers arbitrary int values and registry exhaustion.

  • Fixed: the open-addressing probe mask (& (AKERR_STATUS_NAME_SLOTS - 1)) could be mutated to - 0 or + 1 — both of which index past the end of the table — without any test noticing. tests/err_maxval.c only asserted that some names registered before the table filled, which a collapsed probe sequence still satisfies. It now requires a substantial number of entries and reads every one of them back by its own distinct name, so a probe that revisits slots fails on both counts.