Files
libakerror/tests/MUTATION.md
Andrew Kesterson 756933c600
Some checks failed
libakerror CI Build / cmake_build (push) Successful in 2m48s
libakerror CI Build / coverage (push) Successful in 2m47s
libakerror CI Build / thread_sanitizer (push) Failing after 2m49s
libakerror CI Build / mutation_test (push) Successful in 38m22s
Record the mutation score and the concurrency mutants it misses
src/error.c now scores 81.2%: 238 of 293 mutants killed, 204 by a
failing test, 24 by failing to compile, and 10 by hanging the suite --
deleting akerr_mutex_init() or the akerr_initializing re-entry guard
deadlocks the first test, which is the right answer for a broken lock.

Lock deletions are the one survivor category where surviving does not
mean harmless, so measure it rather than assume: rebuilt, the surviving
"delete the pool lock" mutant fails tests/err_threads_pool.c in 4 of 10
runs and fails under scripts/thread_test.sh in 5 of 5, with no false
positive on the unmutated library. The property assertions alone are a
coin flip on a missing lock; the sanitizer run is what holds that line.
The harness builds mutants with default CMake options and so never sees
it -- TODO item 8.

Also warn that a sanitized test binary run by hand does not inherit the
halt_on_error CTest gives it, and will print a race and still exit 0.
That is how the 5-of-5 above first read as 2 of 5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:45:20 -04:00

161 lines
7.7 KiB
Markdown

# Mutation testing
The unit tests tell us the library works. **Mutation testing tells us the tests
work** — that they would actually fail if the library were broken.
`scripts/mutation_test.py` deliberately breaks the library in small ways
("mutants"), one at a time, and runs the whole CTest suite against each broken
copy:
* if the tests **fail**, the mutant is **killed** — good, the suite caught it;
* if the tests still **pass**, the mutant **survived** — a bug of that shape
would slip through, so it points at a missing test.
The **mutation score** is `killed / (killed + survived)`. A surviving mutant is
a to-do item: write a test that distinguishes the mutant from the original.
## Running
No third-party tools are required — just Python 3 and the normal
cmake/ctest toolchain. The harness never touches your working tree; it copies
the repo to a scratch directory and mutates the copy.
```sh
# Default: mutate src/error.c and include/akerror.tmpl.h
scripts/mutation_test.py
# Faster: just the C source
scripts/mutation_test.py --target src/error.c
# See what would run without building anything
scripts/mutation_test.py --target src/error.c --list
# Gate CI: exit non-zero if the score drops below 90%
scripts/mutation_test.py --threshold 90
```
Via CMake (configures a build first if needed):
```sh
cmake --build build --target mutation
```
Useful flags: `--timeout SECONDS` (per-suite build+test cap; a mutant that
hangs is counted as killed), `--keep` (retain the scratch copy for debugging),
`--work DIR` (use a specific scratch directory), `--junit FILE` (write a JUnit
XML report — surviving mutants appear as failing test cases).
## CI reporting
Both the unit tests and the mutation run emit JUnit XML that CI consumes:
* `ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml"` — note the
absolute path; `--output-junit` otherwise resolves relative to the test dir.
* `scripts/mutation_test.py --junit mutation-junit.xml`
`.gitea/workflows/ci.yaml` runs both and feeds the XML to
`mikepenz/action-junit-report` (with `if: always()`, so results publish even
when a gate fails). The reporter runs with `annotate_only: true`: Gitea does not
implement the Checks API the action uses to create a check run, so creating one
404s (mikepenz/action-junit-report#23). `annotate_only` skips that call and the
results surface via the job summary (`detailed_summary: true`) instead. The
generated `*-junit.xml` files are git-ignored.
## Mutation operators
Each mutant changes exactly one location by one of:
| Tag | Operator | Example |
|-----|--------------------------------|----------------------------------|
| ROR | relational operator | `==``!=`, `<``<=`, `>=``>` |
| LCR | logical connector | `&&``\|\|` |
| BCR | boolean constant | `true``false` |
| AOR | arithmetic / compound assign | `+``-`, `+=``-=` |
| ICR | integer literal | `0``1`, `1``0` |
| SDL | statement deletion | `err->refcount += 1;`*(removed)* |
Preprocessor control lines, comments, and the block of error-code / buffer-size
`#define`s are skipped: mutating those produces equivalent or uninteresting
mutants that only add noise.
## Interpreting survivors
Not every survivor is a test gap — some mutants are **equivalent** (they don't
change observable behaviour, e.g. resizing an internal scratch buffer). For each
survivor, decide:
1. **Real gap** → add or strengthen a test in `tests/` so the mutant is killed,
then re-run.
2. **Equivalent mutant** → no test can catch it; leave a note. If a specific
line is a persistent source of equivalents, narrow the target with
`--target` or extend the skip rules in `scripts/mutation_test.py`.
Re-run after adding tests and confirm the score went up.
## Current status
`src/error.c` scores 81.2% — 238 of 293 mutants killed (204 by a failing test,
24 by failing to compile, 10 by hanging the suite), 55 surviving. The CI gate is
set to 65% for headroom.
The ten timeout kills are all in the locking: deleting `akerr_mutex_init()` or
the `akerr_initializing` re-entry guard deadlocks the very first test, which is
the correct behaviour for a broken lock and is why the harness counts a hang as
a kill.
The remaining survivors are dominated by:
* **Equivalent mutants** in `akerr_init`: deleting the `memset`/`NULL` setup of
file-scope statics (`AKERR_ARRAY_ERROR`, `__akerr_last_ditch`,
`__akerr_last_ignored`) changes nothing, because C already zero-initializes
objects with static storage duration. `int oldid = 0;``1` is likewise
dead: it is overwritten before use, and so is clearing `akerr_initializing`
at the end of initialization — nothing reads that flag once the once-routine
has returned.
* **Lock acquisition** (`akerr_mutex_lock`/`unlock` deletions, and the
`akerr_init()` call at the head of an entry point). These are the one category
where a survivor does *not* mean the mutant is harmless. Removing a lock
leaves a real race, and the assertions in `tests/err_threads_pool.c` only fire
when the race actually loses: rebuilding the surviving mutant and running that
test ten times caught it **four** times. The same mutant under
`scripts/thread_test.sh` failed **five of five**, with no false positive on
the unmutated library — but the mutation harness builds without sanitizers, so
it never sees that. Deleting an `akerr_init()` call survives for a duller
reason: something else has always initialized the library by the time that
line runs.
* **Default logger / handler internals** (`vfprintf`, `va_end`, the
`errctx == NULL` branch, `exit(1)`): killing these needs a subprocess-based
test that captures a child's stderr and exit code, rather than the in-process
capturing logger the other tests use.
* **Static assertions** (`akerr_assert_name_slots_pow2` and the occupancy cap
it guards): a mutated compile-time assertion that still compiles has no
runtime behavior to observe. Unkillable by construction — the assertion is
itself the test, and `tests/err_maxval.c` covers the runtime consequence.
* **Hash and probe details** in `akerr_status_slot`: dropping one of the
multiply steps in `akerr_status_hash` leaves a worse but still correct hash,
and probing backwards (`slot - 1u`) is an equally valid sequence over a
power-of-two table. Both are behaviorally equivalent.
* **The `capacity <= 0` guard** in `akerr_copy_string`, which is defensive: both
call sites pass a positive constant.
Findings surfaced by mutation testing:
* **Open:** the harness builds every mutant with the default CMake options, so a
mutant that only breaks under concurrency is judged by a suite running without
ThreadSanitizer. Mutating under `-DAKERR_SANITIZE=thread` would close that,
and needs a way to pass CMake options through to the mutant build. See
"Mutation testing judges concurrency mutants without a sanitizer" in
`TODO.md`.
* **Superseded:** status names now use a private sparse registry, so the old
public `AKERR_MAX_ERR_VALUE` ceiling and its consumer ABI mismatch no longer
exist. `tests/err_maxval.c` covers arbitrary `int` values and registry
exhaustion.
* **Fixed:** the open-addressing probe mask (`& (AKERR_STATUS_NAME_SLOTS - 1)`)
could be mutated to `- 0` or `+ 1` — both of which index past the end of the
table — without any test noticing. `tests/err_maxval.c` only asserted that
*some* names registered before the table filled, which a collapsed probe
sequence still satisfies. It now requires a substantial number of entries and
reads every one of them back by its own distinct name, so a probe that
revisits slots fails on both counts.