Fourteen issues on source.starfort.tech/andrew/libakerror, labelled by kind and blast radius and milestoned by what they can land in: 2.0.x for anything that breaks no ABI, 2.1.0 for additive surface, 3.0.0 for the handler-ladder rewrite. Everything filed carries status::grooming. Three of the fourteen were not in this file at all. They are defects in this library that were written down in a consumer's TODO instead: IGNORE() leaking a context (#14), the coverage target not being namespaced when embedded (#15), and no akerrorConfigVersion.cmake being installed (#16). libakstdlib carries a workaround for each. That is the finding worth keeping -- a library whose consumers record its defects in their own files cannot see them, and the workaround gets written once per consumer. TODO.md keeps the reasoning: why the handler ladder is a major-version change, why a copied akerr_ErrorContext is a trap in two specific fields, why validating more inputs lowers branch coverage by construction, why the mutation score is a floor, and why a registered name is returned by pointer. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Co-Authored-By: Andrew Kesterson <andrew@aklabs.net>
166 lines
8.1 KiB
Markdown
166 lines
8.1 KiB
Markdown
# Mutation testing
|
|
|
|
The unit tests tell us the library works. **Mutation testing tells us the tests
|
|
work** — that they would actually fail if the library were broken.
|
|
|
|
`scripts/mutation_test.py` deliberately breaks the library in small ways
|
|
("mutants"), one at a time, and runs the whole CTest suite against each broken
|
|
copy:
|
|
|
|
* if the tests **fail**, the mutant is **killed** — good, the suite caught it;
|
|
* if the tests still **pass**, the mutant **survived** — a bug of that shape
|
|
would slip through, so it points at a missing test.
|
|
|
|
The **mutation score** is `killed / (killed + survived)`. A surviving mutant is
|
|
a to-do item: write a test that distinguishes the mutant from the original.
|
|
|
|
## Running
|
|
|
|
No third-party tools are required — just Python 3 and the normal
|
|
cmake/ctest toolchain. The harness never touches your working tree; it copies
|
|
the repo to a scratch directory and mutates the copy.
|
|
|
|
```sh
|
|
# Default: mutate src/error.c and include/akerror.tmpl.h
|
|
scripts/mutation_test.py
|
|
|
|
# Faster: just the C source
|
|
scripts/mutation_test.py --target src/error.c
|
|
|
|
# See what would run without building anything
|
|
scripts/mutation_test.py --target src/error.c --list
|
|
|
|
# Gate CI: exit non-zero if the score drops below 90%
|
|
scripts/mutation_test.py --threshold 90
|
|
```
|
|
|
|
Via CMake (configures a build first if needed):
|
|
|
|
```sh
|
|
cmake --build build --target mutation
|
|
```
|
|
|
|
Useful flags: `--timeout SECONDS` (per-suite build+test cap; a mutant that
|
|
hangs is counted as killed), `--keep` (retain the scratch copy for debugging),
|
|
`--work DIR` (use a specific scratch directory), `--junit FILE` (write a JUnit
|
|
XML report — surviving mutants appear as failing test cases).
|
|
|
|
## CI reporting
|
|
|
|
Both the unit tests and the mutation run emit JUnit XML that CI consumes:
|
|
|
|
* `ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml"` — note the
|
|
absolute path; `--output-junit` otherwise resolves relative to the test dir.
|
|
* `scripts/mutation_test.py --junit mutation-junit.xml`
|
|
|
|
`.gitea/workflows/ci.yaml` runs both and feeds the XML to
|
|
`mikepenz/action-junit-report` (with `if: always()`, so results publish even
|
|
when a gate fails). The reporter runs with `annotate_only: true`: Gitea does not
|
|
implement the Checks API the action uses to create a check run, so creating one
|
|
404s (mikepenz/action-junit-report#23). `annotate_only` skips that call and the
|
|
results surface via the job summary (`detailed_summary: true`) instead. The
|
|
generated `*-junit.xml` files are git-ignored.
|
|
|
|
## Mutation operators
|
|
|
|
Each mutant changes exactly one location by one of:
|
|
|
|
| Tag | Operator | Example |
|
|
|-----|--------------------------------|----------------------------------|
|
|
| ROR | relational operator | `==` → `!=`, `<` → `<=`, `>=` → `>` |
|
|
| LCR | logical connector | `&&` → `\|\|` |
|
|
| BCR | boolean constant | `true` → `false` |
|
|
| AOR | arithmetic / compound assign | `+` → `-`, `+=` → `-=` |
|
|
| ICR | integer literal | `0` → `1`, `1` → `0` |
|
|
| SDL | statement deletion | `err->refcount += 1;` → *(removed)* |
|
|
|
|
Preprocessor control lines, comments, and the block of error-code / buffer-size
|
|
`#define`s are skipped: mutating those produces equivalent or uninteresting
|
|
mutants that only add noise.
|
|
|
|
## Interpreting survivors
|
|
|
|
Not every survivor is a test gap — some mutants are **equivalent** (they don't
|
|
change observable behaviour, e.g. resizing an internal scratch buffer). For each
|
|
survivor, decide:
|
|
|
|
1. **Real gap** → add or strengthen a test in `tests/` so the mutant is killed,
|
|
then re-run.
|
|
2. **Equivalent mutant** → no test can catch it; leave a note. If a specific
|
|
line is a persistent source of equivalents, narrow the target with
|
|
`--target` or extend the skip rules in `scripts/mutation_test.py`.
|
|
|
|
Re-run after adding tests and confirm the score went up.
|
|
|
|
## Current status
|
|
|
|
`src/error.c` scores 81.4% — 245 of 301 mutants killed (211 by a failing test,
|
|
24 by failing to compile, 10 by hanging the suite), 56 surviving. The CI gate is
|
|
set to 65% for headroom.
|
|
|
|
The ten timeout kills are all in the locking: deleting `akerr_mutex_init()` or
|
|
the `akerr_initializing` re-entry guard deadlocks the very first test, which is
|
|
the correct behaviour for a broken lock and is why the harness counts a hang as
|
|
a kill.
|
|
|
|
The remaining survivors are dominated by:
|
|
|
|
* **Equivalent mutants** in `akerr_init`: deleting the `memset`/`NULL` setup of
|
|
file-scope statics (`AKERR_ARRAY_ERROR`, `__akerr_last_ditch`,
|
|
`__akerr_last_ignored`) changes nothing, because C already zero-initializes
|
|
objects with static storage duration. `int oldid = 0;` → `1` is likewise
|
|
dead: it is overwritten before use, and so is clearing `akerr_initializing`
|
|
at the end of initialization — nothing reads that flag once the once-routine
|
|
has returned.
|
|
* **Lock acquisition** (`akerr_mutex_lock`/`unlock` deletions, and the
|
|
`akerr_init()` call at the head of an entry point). These are the one category
|
|
where a survivor does *not* mean the mutant is harmless. Removing a lock
|
|
leaves a real race, and the assertions in `tests/err_threads_pool.c` only fire
|
|
when the race actually loses: rebuilding the surviving mutant and running that
|
|
test ten times caught it **four** times. The same mutant under
|
|
`scripts/thread_test.sh` failed **five of five**, with no false positive on
|
|
the unmutated library — but the mutation harness builds without sanitizers, so
|
|
it never sees that. Deleting an `akerr_init()` call survives for a duller
|
|
reason: something else has always initialized the library by the time that
|
|
line runs.
|
|
* **The default logger** (`vfprintf`, `va_end`, and the `return` in the
|
|
no-stdlib branch): the other tests replace `akerr_log_method` with the
|
|
in-process capturing logger, so nothing observes what the default one writes
|
|
to a real stderr. Killing these needs a test that captures a child's stderr.
|
|
The *handler* internals next to them are no longer in this category:
|
|
`tests/err_unhandled_null.c` and `tests/err_exit_status.c` read a forked
|
|
child's exit code, which kills every mutant in `akerr_exit()` and in
|
|
`akerr_default_handler_unhandled_error()` — all twelve of them, including the
|
|
`status < 0` → `status < 1` variant that only a test asserting
|
|
`akerr_exit(0)` exits 0 can distinguish.
|
|
* **Static assertions** (`akerr_assert_name_slots_pow2` and the occupancy cap
|
|
it guards): a mutated compile-time assertion that still compiles has no
|
|
runtime behavior to observe. Unkillable by construction — the assertion is
|
|
itself the test, and `tests/err_maxval.c` covers the runtime consequence.
|
|
* **Hash and probe details** in `akerr_status_slot`: dropping one of the
|
|
multiply steps in `akerr_status_hash` leaves a worse but still correct hash,
|
|
and probing backwards (`slot - 1u`) is an equally valid sequence over a
|
|
power-of-two table. Both are behaviorally equivalent.
|
|
* **The `capacity <= 0` guard** in `akerr_copy_string`, which is defensive: both
|
|
call sites pass a positive constant.
|
|
|
|
Findings surfaced by mutation testing:
|
|
|
|
* **Open:** the harness builds every mutant with the default CMake options, so a
|
|
mutant that only breaks under concurrency is judged by a suite running without
|
|
ThreadSanitizer. Mutating under `-DAKERR_SANITIZE=thread` would close that,
|
|
and needs a way to pass CMake options through to the mutant build. That is
|
|
issue #10; `TODO.md` records why the score is a floor rather than a verdict.
|
|
|
|
* **Superseded:** status names now use a private sparse registry, so the old
|
|
public `AKERR_MAX_ERR_VALUE` ceiling and its consumer ABI mismatch no longer
|
|
exist. `tests/err_maxval.c` covers arbitrary `int` values and registry
|
|
exhaustion.
|
|
* **Fixed:** the open-addressing probe mask (`& (AKERR_STATUS_NAME_SLOTS - 1)`)
|
|
could be mutated to `- 0` or `+ 1` — both of which index past the end of the
|
|
table — without any test noticing. `tests/err_maxval.c` only asserted that
|
|
*some* names registered before the table filled, which a collapsed probe
|
|
sequence still satisfies. It now requires a substantial number of entries and
|
|
reads every one of them back by its own distinct name, so a probe that
|
|
revisits slots fails on both counts.
|