Produce machine-readable results and surface them in the Gitea pipeline:
- ctest: run with --output-junit to write ctest-junit.xml. The path must be
absolute ("$(pwd)/...") because --output-junit otherwise resolves relative to
the --test-dir build directory.
- mutation_test.py: new --junit FILE option writes a JUnit report where each
mutant is a test case and a surviving mutant is a <failure> (so gaps show up
as failing tests).
- .gitea/workflows/ci.yaml: both jobs generate their XML and feed it to
mikepenz/action-junit-report with `if: always()`, so results publish even
when a gate fails. Mutation publishing is display-only (fail_on_failure:
false); the --threshold flag remains the gate.
- .gitignore: ignore the generated *-junit.xml artifacts.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
114 lines
4.8 KiB
Markdown
114 lines
4.8 KiB
Markdown
# Mutation testing
|
|
|
|
The unit tests tell us the library works. **Mutation testing tells us the tests
|
|
work** — that they would actually fail if the library were broken.
|
|
|
|
`scripts/mutation_test.py` deliberately breaks the library in small ways
|
|
("mutants"), one at a time, and runs the whole CTest suite against each broken
|
|
copy:
|
|
|
|
* if the tests **fail**, the mutant is **killed** — good, the suite caught it;
|
|
* if the tests still **pass**, the mutant **survived** — a bug of that shape
|
|
would slip through, so it points at a missing test.
|
|
|
|
The **mutation score** is `killed / (killed + survived)`. A surviving mutant is
|
|
a to-do item: write a test that distinguishes the mutant from the original.
|
|
|
|
## Running
|
|
|
|
No third-party tools are required — just Python 3 and the normal
|
|
cmake/ctest toolchain. The harness never touches your working tree; it copies
|
|
the repo to a scratch directory and mutates the copy.
|
|
|
|
```sh
|
|
# Default: mutate src/error.c and include/akerror.tmpl.h
|
|
scripts/mutation_test.py
|
|
|
|
# Faster: just the C source
|
|
scripts/mutation_test.py --target src/error.c
|
|
|
|
# See what would run without building anything
|
|
scripts/mutation_test.py --target src/error.c --list
|
|
|
|
# Gate CI: exit non-zero if the score drops below 90%
|
|
scripts/mutation_test.py --threshold 90
|
|
```
|
|
|
|
Via CMake (configures a build first if needed):
|
|
|
|
```sh
|
|
cmake --build build --target mutation
|
|
```
|
|
|
|
Useful flags: `--timeout SECONDS` (per-suite build+test cap; a mutant that
|
|
hangs is counted as killed), `--keep` (retain the scratch copy for debugging),
|
|
`--work DIR` (use a specific scratch directory), `--junit FILE` (write a JUnit
|
|
XML report — surviving mutants appear as failing test cases).
|
|
|
|
## CI reporting
|
|
|
|
Both the unit tests and the mutation run emit JUnit XML that CI consumes:
|
|
|
|
* `ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml"` — note the
|
|
absolute path; `--output-junit` otherwise resolves relative to the test dir.
|
|
* `scripts/mutation_test.py --junit mutation-junit.xml`
|
|
|
|
`.gitea/workflows/ci.yaml` runs both and feeds the XML to
|
|
`mikepenz/action-junit-report` (with `if: always()`, so results publish even
|
|
when a gate fails). The generated `*-junit.xml` files are git-ignored.
|
|
|
|
## Mutation operators
|
|
|
|
Each mutant changes exactly one location by one of:
|
|
|
|
| Tag | Operator | Example |
|
|
|-----|--------------------------------|----------------------------------|
|
|
| ROR | relational operator | `==` → `!=`, `<` → `<=`, `>=` → `>` |
|
|
| LCR | logical connector | `&&` → `\|\|` |
|
|
| BCR | boolean constant | `true` → `false` |
|
|
| AOR | arithmetic / compound assign | `+` → `-`, `+=` → `-=` |
|
|
| ICR | integer literal | `0` → `1`, `1` → `0` |
|
|
| SDL | statement deletion | `err->refcount += 1;` → *(removed)* |
|
|
|
|
Preprocessor control lines, comments, and the block of error-code / buffer-size
|
|
`#define`s are skipped: mutating those produces equivalent or uninteresting
|
|
mutants that only add noise.
|
|
|
|
## Interpreting survivors
|
|
|
|
Not every survivor is a test gap — some mutants are **equivalent** (they don't
|
|
change observable behaviour, e.g. resizing an internal scratch buffer). For each
|
|
survivor, decide:
|
|
|
|
1. **Real gap** → add or strengthen a test in `tests/` so the mutant is killed,
|
|
then re-run.
|
|
2. **Equivalent mutant** → no test can catch it; leave a note. If a specific
|
|
line is a persistent source of equivalents, narrow the target with
|
|
`--target` or extend the skip rules in `scripts/mutation_test.py`.
|
|
|
|
Re-run after adding tests and confirm the score went up.
|
|
|
|
## Current status
|
|
|
|
`src/error.c` scores ~74% (the CI gate is set to 65% for headroom). The
|
|
remaining survivors are dominated by:
|
|
|
|
* **Equivalent mutants** in `akerr_init`: deleting the `memset`/`NULL` setup of
|
|
file-scope statics (`AKERR_ARRAY_ERROR`, `__akerr_last_ditch`,
|
|
`__akerr_last_ignored`) changes nothing, because C already zero-initializes
|
|
objects with static storage duration. `int oldid = 0;` → `1` is likewise
|
|
dead: it is overwritten before use.
|
|
* **Default logger / handler internals** (`vfprintf`, `va_end`, the
|
|
`errctx == NULL` branch, `exit(1)`): killing these needs a subprocess-based
|
|
test that captures a child's stderr and exit code, rather than the in-process
|
|
capturing logger the other tests use.
|
|
|
|
Findings surfaced by mutation testing:
|
|
|
|
* **Fixed:** `AKERR_MAX_ERR_VALUE` was `AKERR_LAST_ERRNO_VALUE + 15`, below
|
|
`AKERR_NOT_IMPLEMENTED` (+16) and `AKERR_BADEXC` (+17). `akerr_name_for_status`
|
|
rejects any status `> AKERR_MAX_ERR_VALUE`, so those codes could never store or
|
|
return a name and the `akerr_name_for_status(AKERR_BADEXC, ...)` call in
|
|
`akerr_init` was dead code (which is why deleting it survived). The max is now
|
|
`+ 17`, and `tests/err_maxval.c` guards the invariant so it can't regress.
|