Files
libakerror/tests/MUTATION.md
Andrew Kesterson 076b80c846
Some checks failed
libakerror CI Build / cmake_build (push) Successful in 2m48s
libakerror CI Build / coverage (push) Successful in 2m47s
libakerror CI Build / thread_sanitizer (push) Failing after 2m50s
libakerror CI Build / mutation_test (push) Successful in 39m13s
Record the mutation score after the exit-status fix
src/error.c now scores 81.4%: 245 of 301 mutants killed, 211 by a
failing test, 24 by failing to compile, and 10 by hanging the suite.
The population grew by 8 with akerr_exit(), and all 8 die, as do the 4
in akerr_default_handler_unhandled_error() -- no survivor anywhere in
either function.

Two of those are worth naming. Mutating the guard to status < 1 (the
sentinel-for-zero behaviour this library deliberately does not have) and
mutating the NULL-context exit(1) to exit(0) both die, so the tests pin
the two ways an exit could start claiming success rather than just the
one that was broken.

That moves the default handler out of the "needs a subprocess test"
survivor category noted here: it has one now. The default *logger* is
still in it -- every other test swaps in the capturing logger, so
nothing watches what the real one writes to stderr.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 07:57:33 -04:00

167 lines
8.1 KiB
Markdown

# Mutation testing
The unit tests tell us the library works. **Mutation testing tells us the tests
work** — that they would actually fail if the library were broken.
`scripts/mutation_test.py` deliberately breaks the library in small ways
("mutants"), one at a time, and runs the whole CTest suite against each broken
copy:
* if the tests **fail**, the mutant is **killed** — good, the suite caught it;
* if the tests still **pass**, the mutant **survived** — a bug of that shape
would slip through, so it points at a missing test.
The **mutation score** is `killed / (killed + survived)`. A surviving mutant is
a to-do item: write a test that distinguishes the mutant from the original.
## Running
No third-party tools are required — just Python 3 and the normal
cmake/ctest toolchain. The harness never touches your working tree; it copies
the repo to a scratch directory and mutates the copy.
```sh
# Default: mutate src/error.c and include/akerror.tmpl.h
scripts/mutation_test.py
# Faster: just the C source
scripts/mutation_test.py --target src/error.c
# See what would run without building anything
scripts/mutation_test.py --target src/error.c --list
# Gate CI: exit non-zero if the score drops below 90%
scripts/mutation_test.py --threshold 90
```
Via CMake (configures a build first if needed):
```sh
cmake --build build --target mutation
```
Useful flags: `--timeout SECONDS` (per-suite build+test cap; a mutant that
hangs is counted as killed), `--keep` (retain the scratch copy for debugging),
`--work DIR` (use a specific scratch directory), `--junit FILE` (write a JUnit
XML report — surviving mutants appear as failing test cases).
## CI reporting
Both the unit tests and the mutation run emit JUnit XML that CI consumes:
* `ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml"` — note the
absolute path; `--output-junit` otherwise resolves relative to the test dir.
* `scripts/mutation_test.py --junit mutation-junit.xml`
`.gitea/workflows/ci.yaml` runs both and feeds the XML to
`mikepenz/action-junit-report` (with `if: always()`, so results publish even
when a gate fails). The reporter runs with `annotate_only: true`: Gitea does not
implement the Checks API the action uses to create a check run, so creating one
404s (mikepenz/action-junit-report#23). `annotate_only` skips that call and the
results surface via the job summary (`detailed_summary: true`) instead. The
generated `*-junit.xml` files are git-ignored.
## Mutation operators
Each mutant changes exactly one location by one of:
| Tag | Operator | Example |
|-----|--------------------------------|----------------------------------|
| ROR | relational operator | `==``!=`, `<``<=`, `>=``>` |
| LCR | logical connector | `&&``\|\|` |
| BCR | boolean constant | `true``false` |
| AOR | arithmetic / compound assign | `+``-`, `+=``-=` |
| ICR | integer literal | `0``1`, `1``0` |
| SDL | statement deletion | `err->refcount += 1;`*(removed)* |
Preprocessor control lines, comments, and the block of error-code / buffer-size
`#define`s are skipped: mutating those produces equivalent or uninteresting
mutants that only add noise.
## Interpreting survivors
Not every survivor is a test gap — some mutants are **equivalent** (they don't
change observable behaviour, e.g. resizing an internal scratch buffer). For each
survivor, decide:
1. **Real gap** → add or strengthen a test in `tests/` so the mutant is killed,
then re-run.
2. **Equivalent mutant** → no test can catch it; leave a note. If a specific
line is a persistent source of equivalents, narrow the target with
`--target` or extend the skip rules in `scripts/mutation_test.py`.
Re-run after adding tests and confirm the score went up.
## Current status
`src/error.c` scores 81.4% — 245 of 301 mutants killed (211 by a failing test,
24 by failing to compile, 10 by hanging the suite), 56 surviving. The CI gate is
set to 65% for headroom.
The ten timeout kills are all in the locking: deleting `akerr_mutex_init()` or
the `akerr_initializing` re-entry guard deadlocks the very first test, which is
the correct behaviour for a broken lock and is why the harness counts a hang as
a kill.
The remaining survivors are dominated by:
* **Equivalent mutants** in `akerr_init`: deleting the `memset`/`NULL` setup of
file-scope statics (`AKERR_ARRAY_ERROR`, `__akerr_last_ditch`,
`__akerr_last_ignored`) changes nothing, because C already zero-initializes
objects with static storage duration. `int oldid = 0;``1` is likewise
dead: it is overwritten before use, and so is clearing `akerr_initializing`
at the end of initialization — nothing reads that flag once the once-routine
has returned.
* **Lock acquisition** (`akerr_mutex_lock`/`unlock` deletions, and the
`akerr_init()` call at the head of an entry point). These are the one category
where a survivor does *not* mean the mutant is harmless. Removing a lock
leaves a real race, and the assertions in `tests/err_threads_pool.c` only fire
when the race actually loses: rebuilding the surviving mutant and running that
test ten times caught it **four** times. The same mutant under
`scripts/thread_test.sh` failed **five of five**, with no false positive on
the unmutated library — but the mutation harness builds without sanitizers, so
it never sees that. Deleting an `akerr_init()` call survives for a duller
reason: something else has always initialized the library by the time that
line runs.
* **The default logger** (`vfprintf`, `va_end`, and the `return` in the
no-stdlib branch): the other tests replace `akerr_log_method` with the
in-process capturing logger, so nothing observes what the default one writes
to a real stderr. Killing these needs a test that captures a child's stderr.
The *handler* internals next to them are no longer in this category:
`tests/err_unhandled_null.c` and `tests/err_exit_status.c` read a forked
child's exit code, which kills every mutant in `akerr_exit()` and in
`akerr_default_handler_unhandled_error()` — all twelve of them, including the
`status < 0``status < 1` variant that only a test asserting
`akerr_exit(0)` exits 0 can distinguish.
* **Static assertions** (`akerr_assert_name_slots_pow2` and the occupancy cap
it guards): a mutated compile-time assertion that still compiles has no
runtime behavior to observe. Unkillable by construction — the assertion is
itself the test, and `tests/err_maxval.c` covers the runtime consequence.
* **Hash and probe details** in `akerr_status_slot`: dropping one of the
multiply steps in `akerr_status_hash` leaves a worse but still correct hash,
and probing backwards (`slot - 1u`) is an equally valid sequence over a
power-of-two table. Both are behaviorally equivalent.
* **The `capacity <= 0` guard** in `akerr_copy_string`, which is defensive: both
call sites pass a positive constant.
Findings surfaced by mutation testing:
* **Open:** the harness builds every mutant with the default CMake options, so a
mutant that only breaks under concurrency is judged by a suite running without
ThreadSanitizer. Mutating under `-DAKERR_SANITIZE=thread` would close that,
and needs a way to pass CMake options through to the mutant build. See
"Mutation testing judges concurrency mutants without a sanitizer" in
`TODO.md`.
* **Superseded:** status names now use a private sparse registry, so the old
public `AKERR_MAX_ERR_VALUE` ceiling and its consumer ABI mismatch no longer
exist. `tests/err_maxval.c` covers arbitrary `int` values and registry
exhaustion.
* **Fixed:** the open-addressing probe mask (`& (AKERR_STATUS_NAME_SLOTS - 1)`)
could be mutated to `- 0` or `+ 1` — both of which index past the end of the
table — without any test noticing. `tests/err_maxval.c` only asserted that
*some* names registered before the table filled, which a collapsed probe
sequence still satisfies. It now requires a substantial number of entries and
reads every one of them back by its own distinct name, so a probe that
revisits slots fails on both counts.