All checks were successful
The threshold had been 80 against src/stdlib.c alone. There are three more sources now and nobody had measured them, so that number was a guess carried forward. Measured: 72.3%, 188 of a 260-mutant sample from the 1701 the four sources generate. Gate set to 65 -- a ratchet with headroom for the runner and for the sample shifting as sources change, not a target. A sample rather than the whole set, because 1701 rebuilds and test runs is hours. --max-mutants samples by even index rather than at random, so the same 260 run every time and the gate stays reproducible; sampling all four files beats exhausting one of them, which is what this job did before. 72.3% against the 89.6% reported at 0.1.0 is a change in denominator, not a regression in the tests. That figure covered one 561-line file; this covers four totalling 1716 lines, and most of the added surface is argument validation whose mutants are frequently *equivalent* -- 12 of the 72 survivors are `errno = 0` deleted from a wrapper whose libc call always sets errno, which no test that could be written would catch. The README breaks all 72 down and says which are worth acting on; TODO.md 2.4 carries the three clusters that are. Two of them were real and are fixed here and in the previous commit: the right child's `depth + 1` in the depth-first walk, and aksl_tree_remove on an empty tree, which without its guard dereferences NULL. Neither had a test; both do now. That is what the harness is for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
184 lines
8.4 KiB
YAML
184 lines
8.4 KiB
YAML
name: libakstdlib CI Build
|
|
run-name: ${{ gitea.actor }} libakstdlib test
|
|
on: [push]
|
|
|
|
jobs:
|
|
cmake_build:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- run: echo "Triggered by ${{ gitea.event_name }} from ${{ gitea.repository }}@${{ gitea.ref }}. Building on ${{ runner.os }}."
|
|
- name: Check out repository code
|
|
uses: actions/checkout@v4
|
|
with:
|
|
# A top-level build uses deps/libakerror via add_subdirectory, so the
|
|
# submodule has to be present or the configure step fails outright.
|
|
submodules: recursive
|
|
- name: dependencies
|
|
run: |
|
|
sudo apt-get update -y
|
|
sudo apt-get install -y cmake gcc moreutils
|
|
# This step used to clone and install libakerror@main. That was two
|
|
# different libakerrors depending on where you built: the install went to
|
|
# /usr/local, while the build below is top-level and so compiles
|
|
# deps/libakerror at the pinned commit via add_subdirectory -- the
|
|
# installed one was never actually linked against. TODO.md 2.3.
|
|
#
|
|
# The submodule wins. It is the version this repository pins, tests and
|
|
# ships against, and a CI that tests a different one is testing something
|
|
# nobody runs. It is installed here as well as compiled, because an
|
|
# installed libakstdlib is not usable without it: akstdlibConfig.cmake
|
|
# calls find_dependency(akerror), and the top-level build pulls the
|
|
# submodule in EXCLUDE_FROM_ALL so `cmake --install` on this project
|
|
# installs only this project.
|
|
- name: install the pinned libakerror
|
|
run: |
|
|
cmake -S deps/libakerror -B build-akerror
|
|
cmake --build build-akerror
|
|
sudo cmake --install build-akerror
|
|
- name: build and test
|
|
run: |
|
|
# -DAKSL_WERROR=ON here and not locally: a warning that stops the build
|
|
# mid-thought teaches people to turn warnings off, but a warning that
|
|
# reaches main is one nobody will ever look at again.
|
|
cmake -S . -B build -DAKSL_WERROR=ON
|
|
cmake --build build
|
|
sudo cmake --install build
|
|
ctest --test-dir build --output-on-failure
|
|
# Build something against what was just installed. The suite links the
|
|
# build tree, so it says nothing about whether the *installed* package is
|
|
# usable -- and the two have come apart before: find_dependency(akerror)
|
|
# resolving is a property of the install, not of the build.
|
|
- name: consume the installed package
|
|
run: |
|
|
sudo ldconfig
|
|
cmake -S tests/consumer -B build-consumer
|
|
cmake --build build-consumer
|
|
./build-consumer/consumer
|
|
- run: echo "🍏 This job's status is ${{ job.status }}."
|
|
|
|
# The sanitizer build is its own job rather than a step on the one above,
|
|
# because three of the defects fixed in 0.2.0 only ever misbehaved under
|
|
# instrumentation and the tests that pin them are worth failing on their own.
|
|
sanitizers:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- name: Check out repository code
|
|
uses: actions/checkout@v4
|
|
with:
|
|
submodules: recursive
|
|
- name: dependencies
|
|
run: |
|
|
sudo apt-get update -y
|
|
sudo apt-get install -y cmake gcc
|
|
- name: build and test under ASan + UBSan
|
|
run: |
|
|
cmake -S . -B build-asan -DAKSL_SANITIZE=ON -DAKSL_WERROR=ON
|
|
cmake --build build-asan
|
|
ctest --test-dir build-asan --output-on-failure
|
|
- run: echo "🍏 This job's status is ${{ job.status }}."
|
|
|
|
coverage:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- name: Check out repository code
|
|
uses: actions/checkout@v4
|
|
with:
|
|
submodules: recursive
|
|
- name: dependencies
|
|
run: |
|
|
sudo apt-get update -y
|
|
sudo apt-get install -y cmake gcc python3
|
|
# scripts/coverage.py needs nothing but python3 and gcc's own gcov, so
|
|
# there is no lcov/gcovr to install here.
|
|
#
|
|
# The gate is a ratchet, not a target. Across all four sources the suite
|
|
# covers 99.5% of lines (1708/1716), 46.0% of branches and 100% of
|
|
# functions (154/154), so 90/40 fails on a real regression -- a test
|
|
# deleted, or new untested code added -- without tripping over rounding.
|
|
# The report is printed either way; the uncovered lines it lists are the
|
|
# missing tests.
|
|
#
|
|
# Eight lines are uncovered and every one is deliberate: four
|
|
# `HANDLE(e, AKERR_ITERATOR_BREAK)` macro artifacts, the size_t overflow
|
|
# guard in strbuf_reserve (reaching it needs a buffer near SIZE_MAX), and
|
|
# the short transfer with neither feof nor ferror set, which the standard
|
|
# permits and Linux never produces. README.md has the detail.
|
|
#
|
|
# Branch coverage sits far below line coverage because most branches are
|
|
# inside libakerror's FAIL_*/ATTEMPT/FINISH expansions -- pool exhaustion,
|
|
# stack-trace limits, akerr_valid_error_address failures -- which this
|
|
# library has no way to reach. Chasing them here would be testing
|
|
# libakerror's macros, which is libakerror's mutation suite's job: macros
|
|
# expand at the call site, so coverage cannot see them properly from
|
|
# either side. The gate stays at 40 for that reason and not for want of
|
|
# tests.
|
|
- name: coverage
|
|
run: |
|
|
cmake -S . -B build-coverage -DAKSL_COVERAGE=ON \
|
|
-DAKSL_COVERAGE_THRESHOLD=90 \
|
|
-DAKSL_COVERAGE_BRANCH_THRESHOLD=40
|
|
cmake --build build-coverage
|
|
ctest --test-dir build-coverage --output-on-failure
|
|
cat build-coverage/coverage-summary.txt
|
|
- run: echo "🍏 This job's status is ${{ job.status }}."
|
|
|
|
mutation_test:
|
|
runs-on: ubuntu-latest
|
|
steps:
|
|
- name: Check out repository code
|
|
uses: actions/checkout@v4
|
|
with:
|
|
# The harness copies the repo and builds the copy top-level, so it
|
|
# needs deps/libakerror present just like the main job does.
|
|
submodules: recursive
|
|
- name: dependencies
|
|
run: |
|
|
sudo apt-get update -y
|
|
sudo apt-get install -y cmake gcc moreutils python3
|
|
# Verify the tests actually catch bugs: break the library many ways and
|
|
# confirm the suite fails.
|
|
#
|
|
# A sample rather than the whole set. The four sources generate 1701
|
|
# mutants and each one is a full rebuild and re-run of the suite, which is
|
|
# hours. --max-mutants samples by even index, not at random, so the same
|
|
# 260 run every time and the gate is reproducible -- and sampling all four
|
|
# files beats exhausting one of them, which is what this job used to do.
|
|
#
|
|
# The threshold is a ratchet, not a quality bar. The measured score is
|
|
# 72.3% (188/260). 65 leaves headroom for the runner and for the sample
|
|
# shifting as sources change, while still failing on a real regression --
|
|
# a test deleted, or new untested code added.
|
|
#
|
|
# It is well below the 89.6% this job reported at 0.1.0, and that is a
|
|
# change in denominator rather than in the tests: that figure covered one
|
|
# 561-line file, this one covers four totalling 1716 lines, most of the
|
|
# new surface being argument validation whose mutants are frequently
|
|
# equivalent. `errno = 0` deleted from a wrapper whose libc call always
|
|
# sets errno cannot be distinguished by any test that could be written.
|
|
# The survivors worth acting on are named in TODO.md; raise the gate as
|
|
# they become assertions.
|
|
- name: mutation testing
|
|
run: |
|
|
python3 scripts/mutation_test.py \
|
|
--target src/stdlib.c \
|
|
--target src/string.c \
|
|
--target src/stream.c \
|
|
--target src/collections.c \
|
|
--max-mutants 260 \
|
|
--junit mutation-junit.xml \
|
|
--threshold 65
|
|
# Publish even when the threshold gate fails, so survivors are visible --
|
|
# each one is a missing test. Display-only (fail_on_failure: false); the
|
|
# --threshold above is the gate. annotate_only avoids the Checks API 404
|
|
# on Gitea (mikepenz/action-junit-report#23).
|
|
- name: publish mutation results
|
|
if: always()
|
|
uses: mikepenz/action-junit-report@v4
|
|
with:
|
|
report_paths: 'mutation-junit.xml'
|
|
annotate_only: true
|
|
detailed_summary: true
|
|
include_passed: true
|
|
fail_on_failure: 'false'
|
|
- run: echo "🍏 This job's status is ${{ job.status }}."
|