# Mutation testing The unit tests tell us the library works. **Mutation testing tells us the tests work** — that they would actually fail if the library were broken. `scripts/mutation_test.py` deliberately breaks the library in small ways ("mutants"), one at a time, and runs the whole CTest suite against each broken copy: * if the tests **fail**, the mutant is **killed** — good, the suite caught it; * if the tests still **pass**, the mutant **survived** — a bug of that shape would slip through, so it points at a missing test. The **mutation score** is `killed / (killed + survived)`. A surviving mutant is a to-do item: write a test that distinguishes the mutant from the original. ## Running No third-party tools are required — just Python 3 and the normal cmake/ctest toolchain. The harness never touches your working tree; it copies the repo to a scratch directory and mutates the copy. ```sh # Default: mutate src/error.c and include/akerror.tmpl.h scripts/mutation_test.py # Faster: just the C source scripts/mutation_test.py --target src/error.c # See what would run without building anything scripts/mutation_test.py --target src/error.c --list # Gate CI: exit non-zero if the score drops below 90% scripts/mutation_test.py --threshold 90 ``` Via CMake (configures a build first if needed): ```sh cmake --build build --target mutation ``` Useful flags: `--timeout SECONDS` (per-suite build+test cap; a mutant that hangs is counted as killed), `--keep` (retain the scratch copy for debugging), `--work DIR` (use a specific scratch directory), `--junit FILE` (write a JUnit XML report — surviving mutants appear as failing test cases). ## CI reporting Both the unit tests and the mutation run emit JUnit XML that CI consumes: * `ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml"` — note the absolute path; `--output-junit` otherwise resolves relative to the test dir. * `scripts/mutation_test.py --junit mutation-junit.xml` `.gitea/workflows/ci.yaml` runs both and feeds the XML to `mikepenz/action-junit-report` (with `if: always()`, so results publish even when a gate fails). The reporter runs with `annotate_only: true`: Gitea does not implement the Checks API the action uses to create a check run, so creating one 404s (mikepenz/action-junit-report#23). `annotate_only` skips that call and the results surface via the job summary (`detailed_summary: true`) instead. The generated `*-junit.xml` files are git-ignored. ## Mutation operators Each mutant changes exactly one location by one of: | Tag | Operator | Example | |-----|--------------------------------|----------------------------------| | ROR | relational operator | `==` → `!=`, `<` → `<=`, `>=` → `>` | | LCR | logical connector | `&&` → `\|\|` | | BCR | boolean constant | `true` → `false` | | AOR | arithmetic / compound assign | `+` → `-`, `+=` → `-=` | | ICR | integer literal | `0` → `1`, `1` → `0` | | SDL | statement deletion | `err->refcount += 1;` → *(removed)* | Preprocessor control lines, comments, and the block of error-code / buffer-size `#define`s are skipped: mutating those produces equivalent or uninteresting mutants that only add noise. ## Interpreting survivors Not every survivor is a test gap — some mutants are **equivalent** (they don't change observable behaviour, e.g. resizing an internal scratch buffer). For each survivor, decide: 1. **Real gap** → add or strengthen a test in `tests/` so the mutant is killed, then re-run. 2. **Equivalent mutant** → no test can catch it; leave a note. If a specific line is a persistent source of equivalents, narrow the target with `--target` or extend the skip rules in `scripts/mutation_test.py`. Re-run after adding tests and confirm the score went up. ## Current status `src/error.c` scores 81.4% — 245 of 301 mutants killed (211 by a failing test, 24 by failing to compile, 10 by hanging the suite), 56 surviving. The CI gate is set to 65% for headroom. The ten timeout kills are all in the locking: deleting `akerr_mutex_init()` or the `akerr_initializing` re-entry guard deadlocks the very first test, which is the correct behaviour for a broken lock and is why the harness counts a hang as a kill. The remaining survivors are dominated by: * **Equivalent mutants** in `akerr_init`: deleting the `memset`/`NULL` setup of file-scope statics (`AKERR_ARRAY_ERROR`, `__akerr_last_ditch`, `__akerr_last_ignored`) changes nothing, because C already zero-initializes objects with static storage duration. `int oldid = 0;` → `1` is likewise dead: it is overwritten before use, and so is clearing `akerr_initializing` at the end of initialization — nothing reads that flag once the once-routine has returned. * **Lock acquisition** (`akerr_mutex_lock`/`unlock` deletions, and the `akerr_init()` call at the head of an entry point). These are the one category where a survivor does *not* mean the mutant is harmless. Removing a lock leaves a real race, and the assertions in `tests/err_threads_pool.c` only fire when the race actually loses: rebuilding the surviving mutant and running that test ten times caught it **four** times. The same mutant under `scripts/thread_test.sh` failed **five of five**, with no false positive on the unmutated library — but the mutation harness builds without sanitizers, so it never sees that. Deleting an `akerr_init()` call survives for a duller reason: something else has always initialized the library by the time that line runs. * **The default logger** (`vfprintf`, `va_end`, and the `return` in the no-stdlib branch): the other tests replace `akerr_log_method` with the in-process capturing logger, so nothing observes what the default one writes to a real stderr. Killing these needs a test that captures a child's stderr. The *handler* internals next to them are no longer in this category: `tests/err_unhandled_null.c` and `tests/err_exit_status.c` read a forked child's exit code, which kills every mutant in `akerr_exit()` and in `akerr_default_handler_unhandled_error()` — all twelve of them, including the `status < 0` → `status < 1` variant that only a test asserting `akerr_exit(0)` exits 0 can distinguish. * **Static assertions** (`akerr_assert_name_slots_pow2` and the occupancy cap it guards): a mutated compile-time assertion that still compiles has no runtime behavior to observe. Unkillable by construction — the assertion is itself the test, and `tests/err_maxval.c` covers the runtime consequence. * **Hash and probe details** in `akerr_status_slot`: dropping one of the multiply steps in `akerr_status_hash` leaves a worse but still correct hash, and probing backwards (`slot - 1u`) is an equally valid sequence over a power-of-two table. Both are behaviorally equivalent. * **The `capacity <= 0` guard** in `akerr_copy_string`, which is defensive: both call sites pass a positive constant. Findings surfaced by mutation testing: * **Open:** the harness builds every mutant with the default CMake options, so a mutant that only breaks under concurrency is judged by a suite running without ThreadSanitizer. Mutating under `-DAKERR_SANITIZE=thread` would close that, and needs a way to pass CMake options through to the mutant build. See "Mutation testing judges concurrency mutants without a sanitizer" in `TODO.md`. * **Superseded:** status names now use a private sparse registry, so the old public `AKERR_MAX_ERR_VALUE` ceiling and its consumer ABI mismatch no longer exist. `tests/err_maxval.c` covers arbitrary `int` values and registry exhaustion. * **Fixed:** the open-addressing probe mask (`& (AKERR_STATUS_NAME_SLOTS - 1)`) could be mutated to `- 0` or `+ 1` — both of which index past the end of the table — without any test noticing. `tests/err_maxval.c` only asserted that *some* names registered before the table filled, which a collapsed probe sequence still satisfies. It now requires a substantial number of entries and reads every one of them back by its own distinct name, so a probe that revisits slots fails on both counts.