Files
libakstdlib/.gitea/workflows/ci.yaml
Andrew Kesterson f8425b8729
All checks were successful
libakstdlib CI Build / cmake_build (push) Successful in 2m53s
libakstdlib CI Build / sanitizers (push) Successful in 2m51s
libakstdlib CI Build / coverage (push) Successful in 2m43s
libakstdlib CI Build / mutation_test (push) Successful in 12m37s
Gate mutation testing on a measured score, not an inherited one
The threshold had been 80 against src/stdlib.c alone. There are three
more sources now and nobody had measured them, so that number was a
guess carried forward.

Measured: 72.3%, 188 of a 260-mutant sample from the 1701 the four
sources generate. Gate set to 65 -- a ratchet with headroom for the
runner and for the sample shifting as sources change, not a target.

A sample rather than the whole set, because 1701 rebuilds and test runs
is hours. --max-mutants samples by even index rather than at random, so
the same 260 run every time and the gate stays reproducible; sampling
all four files beats exhausting one of them, which is what this job did
before.

72.3% against the 89.6% reported at 0.1.0 is a change in denominator,
not a regression in the tests. That figure covered one 561-line file;
this covers four totalling 1716 lines, and most of the added surface is
argument validation whose mutants are frequently *equivalent* -- 12 of
the 72 survivors are `errno = 0` deleted from a wrapper whose libc call
always sets errno, which no test that could be written would catch. The
README breaks all 72 down and says which are worth acting on; TODO.md
2.4 carries the three clusters that are.

Two of them were real and are fixed here and in the previous commit: the
right child's `depth + 1` in the depth-first walk, and aksl_tree_remove
on an empty tree, which without its guard dereferences NULL. Neither had
a test; both do now. That is what the harness is for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:15:08 -04:00

184 lines
8.4 KiB
YAML

name: libakstdlib CI Build
run-name: ${{ gitea.actor }} libakstdlib test
on: [push]
jobs:
cmake_build:
runs-on: ubuntu-latest
steps:
- run: echo "Triggered by ${{ gitea.event_name }} from ${{ gitea.repository }}@${{ gitea.ref }}. Building on ${{ runner.os }}."
- name: Check out repository code
uses: actions/checkout@v4
with:
# A top-level build uses deps/libakerror via add_subdirectory, so the
# submodule has to be present or the configure step fails outright.
submodules: recursive
- name: dependencies
run: |
sudo apt-get update -y
sudo apt-get install -y cmake gcc moreutils
# This step used to clone and install libakerror@main. That was two
# different libakerrors depending on where you built: the install went to
# /usr/local, while the build below is top-level and so compiles
# deps/libakerror at the pinned commit via add_subdirectory -- the
# installed one was never actually linked against. TODO.md 2.3.
#
# The submodule wins. It is the version this repository pins, tests and
# ships against, and a CI that tests a different one is testing something
# nobody runs. It is installed here as well as compiled, because an
# installed libakstdlib is not usable without it: akstdlibConfig.cmake
# calls find_dependency(akerror), and the top-level build pulls the
# submodule in EXCLUDE_FROM_ALL so `cmake --install` on this project
# installs only this project.
- name: install the pinned libakerror
run: |
cmake -S deps/libakerror -B build-akerror
cmake --build build-akerror
sudo cmake --install build-akerror
- name: build and test
run: |
# -DAKSL_WERROR=ON here and not locally: a warning that stops the build
# mid-thought teaches people to turn warnings off, but a warning that
# reaches main is one nobody will ever look at again.
cmake -S . -B build -DAKSL_WERROR=ON
cmake --build build
sudo cmake --install build
ctest --test-dir build --output-on-failure
# Build something against what was just installed. The suite links the
# build tree, so it says nothing about whether the *installed* package is
# usable -- and the two have come apart before: find_dependency(akerror)
# resolving is a property of the install, not of the build.
- name: consume the installed package
run: |
sudo ldconfig
cmake -S tests/consumer -B build-consumer
cmake --build build-consumer
./build-consumer/consumer
- run: echo "🍏 This job's status is ${{ job.status }}."
# The sanitizer build is its own job rather than a step on the one above,
# because three of the defects fixed in 0.2.0 only ever misbehaved under
# instrumentation and the tests that pin them are worth failing on their own.
sanitizers:
runs-on: ubuntu-latest
steps:
- name: Check out repository code
uses: actions/checkout@v4
with:
submodules: recursive
- name: dependencies
run: |
sudo apt-get update -y
sudo apt-get install -y cmake gcc
- name: build and test under ASan + UBSan
run: |
cmake -S . -B build-asan -DAKSL_SANITIZE=ON -DAKSL_WERROR=ON
cmake --build build-asan
ctest --test-dir build-asan --output-on-failure
- run: echo "🍏 This job's status is ${{ job.status }}."
coverage:
runs-on: ubuntu-latest
steps:
- name: Check out repository code
uses: actions/checkout@v4
with:
submodules: recursive
- name: dependencies
run: |
sudo apt-get update -y
sudo apt-get install -y cmake gcc python3
# scripts/coverage.py needs nothing but python3 and gcc's own gcov, so
# there is no lcov/gcovr to install here.
#
# The gate is a ratchet, not a target. Across all four sources the suite
# covers 99.5% of lines (1708/1716), 46.0% of branches and 100% of
# functions (154/154), so 90/40 fails on a real regression -- a test
# deleted, or new untested code added -- without tripping over rounding.
# The report is printed either way; the uncovered lines it lists are the
# missing tests.
#
# Eight lines are uncovered and every one is deliberate: four
# `HANDLE(e, AKERR_ITERATOR_BREAK)` macro artifacts, the size_t overflow
# guard in strbuf_reserve (reaching it needs a buffer near SIZE_MAX), and
# the short transfer with neither feof nor ferror set, which the standard
# permits and Linux never produces. README.md has the detail.
#
# Branch coverage sits far below line coverage because most branches are
# inside libakerror's FAIL_*/ATTEMPT/FINISH expansions -- pool exhaustion,
# stack-trace limits, akerr_valid_error_address failures -- which this
# library has no way to reach. Chasing them here would be testing
# libakerror's macros, which is libakerror's mutation suite's job: macros
# expand at the call site, so coverage cannot see them properly from
# either side. The gate stays at 40 for that reason and not for want of
# tests.
- name: coverage
run: |
cmake -S . -B build-coverage -DAKSL_COVERAGE=ON \
-DAKSL_COVERAGE_THRESHOLD=90 \
-DAKSL_COVERAGE_BRANCH_THRESHOLD=40
cmake --build build-coverage
ctest --test-dir build-coverage --output-on-failure
cat build-coverage/coverage-summary.txt
- run: echo "🍏 This job's status is ${{ job.status }}."
mutation_test:
runs-on: ubuntu-latest
steps:
- name: Check out repository code
uses: actions/checkout@v4
with:
# The harness copies the repo and builds the copy top-level, so it
# needs deps/libakerror present just like the main job does.
submodules: recursive
- name: dependencies
run: |
sudo apt-get update -y
sudo apt-get install -y cmake gcc moreutils python3
# Verify the tests actually catch bugs: break the library many ways and
# confirm the suite fails.
#
# A sample rather than the whole set. The four sources generate 1701
# mutants and each one is a full rebuild and re-run of the suite, which is
# hours. --max-mutants samples by even index, not at random, so the same
# 260 run every time and the gate is reproducible -- and sampling all four
# files beats exhausting one of them, which is what this job used to do.
#
# The threshold is a ratchet, not a quality bar. The measured score is
# 72.3% (188/260). 65 leaves headroom for the runner and for the sample
# shifting as sources change, while still failing on a real regression --
# a test deleted, or new untested code added.
#
# It is well below the 89.6% this job reported at 0.1.0, and that is a
# change in denominator rather than in the tests: that figure covered one
# 561-line file, this one covers four totalling 1716 lines, most of the
# new surface being argument validation whose mutants are frequently
# equivalent. `errno = 0` deleted from a wrapper whose libc call always
# sets errno cannot be distinguished by any test that could be written.
# The survivors worth acting on are named in TODO.md; raise the gate as
# they become assertions.
- name: mutation testing
run: |
python3 scripts/mutation_test.py \
--target src/stdlib.c \
--target src/string.c \
--target src/stream.c \
--target src/collections.c \
--max-mutants 260 \
--junit mutation-junit.xml \
--threshold 65
# Publish even when the threshold gate fails, so survivors are visible --
# each one is a missing test. Display-only (fail_on_failure: false); the
# --threshold above is the gate. annotate_only avoids the Checks API 404
# on Gitea (mikepenz/action-junit-report#23).
- name: publish mutation results
if: always()
uses: mikepenz/action-junit-report@v4
with:
report_paths: 'mutation-junit.xml'
annotate_only: true
detailed_summary: true
include_passed: true
fail_on_failure: 'false'
- run: echo "🍏 This job's status is ${{ job.status }}."