39 Commits

Author SHA1 Message Date
5eaa956f50 Stop an unhandled error from exiting zero
Some checks failed
libakerror CI Build / cmake_build (push) Successful in 2m47s
libakerror CI Build / coverage (push) Successful in 2m48s
libakerror CI Build / thread_sanitizer (push) Failing after 2m49s
libakerror CI Build / mutation_test (push) Successful in 39m15s
An unhandled error could kill the process and still report success. The
default handler ended in exit(errctx->status), and an exit status is one
byte wide: the kernel keeps the low 8 bits of the argument and discards
the rest. Consumer statuses start at AKERR_FIRST_CONSUMER_STATUS (256),
so the first status any consumer can reserve exited 0 and a shell saw a
clean run. Status 300 exited 44, an unrelated error's code.

There is no wider exit() to reach for. _exit(), _Exit(), quick_exit()
and the raw exit_group syscall all truncate identically, and even
waitid(), whose si_status is a full int, reports the truncated value --
the truncation happened before the parent looked.

akerr_exit() now owns that mapping and the default handler calls it: 0
exits 0, 1 through 255 exit the status, and anything else exits
AKERR_EXIT_STATUS_UNREPRESENTABLE (125) rather than a low byte that is
either a lie or a claim of success. Only values that were already being
delivered wrong behave differently. Call it instead of exit() anywhere
you leave the process on a status; it is declared AKERR_NORETURN.

akerr_exit(0) exits 0, because 0 is this library's success status. That
is not a hole in the rule: PROCESS opens with case 0, which marks a zero
status handled, so a successful context never reaches FINISH_NORETURN's
call to the handler at all.

tests/err_exit_status.c drives one table through akerr_exit() and
through the default handler in forked children and requires identical
exit codes, so the handler cannot grow a mapping of its own. With the
clamp removed it fails with "akerr_exit(256) exited 0, want 125". The
full-width status was already reaching the log and still does, which the
same test asserts against the captured stack trace.

2.0.1. No ABI break: the soname stays libakerror.so.2 and nothing that
already existed changed shape. akerr_exit() is a new exported symbol, so
a consumer that starts calling it needs 2.0.1 at link time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 07:25:37 -04:00
756933c600 Record the mutation score and the concurrency mutants it misses
Some checks failed
libakerror CI Build / cmake_build (push) Successful in 2m48s
libakerror CI Build / coverage (push) Successful in 2m47s
libakerror CI Build / thread_sanitizer (push) Failing after 2m49s
libakerror CI Build / mutation_test (push) Successful in 38m22s
src/error.c now scores 81.2%: 238 of 293 mutants killed, 204 by a
failing test, 24 by failing to compile, and 10 by hanging the suite --
deleting akerr_mutex_init() or the akerr_initializing re-entry guard
deadlocks the first test, which is the right answer for a broken lock.

Lock deletions are the one survivor category where surviving does not
mean harmless, so measure it rather than assume: rebuilt, the surviving
"delete the pool lock" mutant fails tests/err_threads_pool.c in 4 of 10
runs and fails under scripts/thread_test.sh in 5 of 5, with no false
positive on the unmutated library. The property assertions alone are a
coin flip on a missing lock; the sanitizer run is what holds that line.
The harness builds mutants with default CMake options and so never sees
it -- TODO item 8.

Also warn that a sanitized test binary run by hand does not inherit the
halt_on_error CTest gives it, and will print a race and still exit 0.
That is how the 5-of-5 above first read as 2 of 5.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:45:20 -04:00
be24f80022 Make the error pool and status registry thread safe
Every entry point may now be called from any thread. akerr_init() runs
exactly once however many threads race into it, the pool hands each slot
to exactly one thread, and reservations, registrations and lookups are
serialized against each other.

One recursive lock covers both tables (src/lock.h, private). Recursive
because raising an error re-enters the library -- FAIL needs a pool slot
and a status name -- and single because two locks would mean an ordering
to get wrong. Registry bodies that use the early-returning FAIL_*_RETURN
macros are split into *_locked functions behind wrappers that take and
release the lock on one path; consumer callbacks are never called under
it.

This is an ABI break, hence 2.0.0 and SOVERSION 2:

- akerr_next_error() now returns a context that already holds its
  reference. Finding a free slot and claiming it has to be one operation,
  or two threads scanning at once are handed the same slot.
  ENSURE_ERROR_READY no longer increments.
- __akerr_last_ignored is thread-local, as is the last-ditch context used
  to report akerr_release_error(NULL).

The threading backend is chosen at configure time by AKERR_THREADS
(auto, pthread, none). auto fails the configure when it cannot find
POSIX threads rather than quietly building a library that reports itself
thread safe and is not. generrno.sh stamps the decision into the
generated header as AKERR_THREAD_SAFE, so a consumer cannot disagree
with the library about it.

Tests: err_threads_init, err_threads_pool and err_threads_registry
assert exclusive slot ownership, exactly one winner for a contested
range, and every registered name readable back under contention.
AKERR_SANITIZE builds the library and the tests with any sanitizer;
scripts/thread_test.sh runs the suite under ThreadSanitizer and CI runs
it. Removing the pool lock makes both the sanitizer and the plain
assertions fail, so the tests are not vacuous.

Documented in README.md and UPGRADING.md, including what this does not
cover: renaming a status while another thread looks it up, and which of
two simultaneous unhandled errors sets the exit status.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-31 08:31:22 -04:00
5ff87908e7 Use the library's own error idioms inside the library
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m48s
libakerror CI Build / coverage (push) Successful in 2m46s
libakerror CI Build / mutation_test (push) Successful in 16m9s
Four things in src/error.c did by hand what the macros already do, or
skipped checks the library would have caught for a consumer.

akerr_copy_string() returned void and validated only its capacity, while
writing through a caller-supplied pointer for a caller-supplied length.
It is now __akerr_copy_string() and raises: AKERR_NULLPOINTER for a NULL
destination or source, AKERR_VALUE for a capacity with no room for a
terminator. Both call sites PASS it, and the owner copy in
akerr_reserve_status_range() now gates the commit, so a failed copy
cannot leave a range claimed under an empty owner. It is exported under
the internal prefix rather than static so tests/err_copy_string.c can
drive those guards; nothing else can reach them.

__akerr_name_library_status() and the band reservation in akerr_init()
hand-rolled the log/handler/release sequence. Both now use
ATTEMPT/CATCH/PROCESS/FINISH_NORETURN. PASS does not fit: both sites are
void and have no caller to propagate to, so the terminal form of the same
idiom is the right one -- an unhandled failure prints its stack trace and
goes to akerr_handler_unhandled_error, which terminates, exactly as
before but without the bespoke plumbing. The legacy set path in
akerr_name_for_status() had the same shape and now handles its refusal
with HANDLE_DEFAULT, converting it to the "Unknown Error" sentinel.

Every remaining `if (x) { FAIL_RETURN }` in the registry is now
FAIL_ZERO_RETURN or FAIL_NONZERO_RETURN, and akerr_register_status_name()
checks both owner and name before passing either down --
akerr_store_status_name() reads a NULL owner as "caller did not identify
itself" for the legacy path, so a NULL arriving through the owned entry
point would have skipped the ownership check entirely.

New tests: err_copy_string (the guards above), err_library_status_fatal
(WILL_FAIL -- proves a refused library-status registration terminates).

Tests: ctest 31/31, mutation 80.7% (was 77.5%), line coverage 98.9%.
Branch coverage on src/error.c drops 64.5% -> 50.4%, just over its gate:
each FAIL_* site carries ~6 branch outcomes of error-construction
machinery that only run when that failure fires, and each PASS around a
call that cannot fail carries ~25, so added validation lowers the ratio
by construction. Recorded in TODO.md item 7.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 21:53:33 -04:00
64da04e83b Raise errors from the status registry instead of returning codes
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m47s
libakerror CI Build / coverage (push) Successful in 2m46s
libakerror CI Build / mutation_test (push) Successful in 14m56s
akerr_reserve_status_range() and akerr_register_status_name() returned
private int enumerations, which was the one place in the library where a
failure was not an akerr_ErrorContext *. They now return one like
everything else: NULL on success, and on refusal an error whose status is
a real code in the library's reserved band, so it can be CATCH-ed,
HANDLE-d, PASS-ed, or left to propagate into a stack trace. Both are
marked AKERR_NOIGNORE, so discarding the result warns at compile time.

AKERR_STATUS_RANGE_OK and AKERR_STATUS_NAME_OK are gone; the remaining
seven codes move into the AKERR_* offset span and get registered names.
AKERR_LAST_LIBRARY_STATUS replaces AKERR_BADEXC as the top of that span
in the reserved-band static assert and the exhaustiveness sweep.

The refusal detail that used to go straight to akerr_log_method now
travels in the error message, so a caller that handles the error decides
whether it is reported. The two-argument akerr_name_for_status() set path
is the exception: it returns a name and cannot raise, so it logs and
releases. akerr_init() likewise has no caller to raise into, so failing
to reserve its own band or name its own codes is logged and fatal --
that can only happen on a misconfigured build, and continuing would
degrade every later stack trace to "Unknown Error".

Move the 1.0.0 upgrade notice out of README.md into UPGRADING.md and
rewrite its return-code tables in terms of the statuses now raised.

Tests: ctest 29/29, coverage 97.5% line / 64.5% branch, mutation 77.5%
(was 77.6%; the new survivors are the fatal init path, which needs a
library built with an undersized name table -- TODO item 7).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 20:58:47 -04:00
ba2430bfa1 Reduce TODO.md to outstanding work
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m46s
libakerror CI Build / coverage (push) Successful in 2m45s
libakerror CI Build / mutation_test (push) Successful in 15m10s
The status-code ownership notes described work that has landed. That record
belongs in the commit that made the change and in the comments around the code
it constrains, not in a file whose purpose is naming what is left.

Keeps the six open items and the two unrelated pre-existing issues, and
promotes the missing sanitizer run to the top: mutation testing found an
out-of-bounds probe whose failure mode was a silent BSS write, which no
assertion-based test was positioned to catch.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 18:19:53 -04:00
f1283e21a3 Enforce status-code ownership and harden the name registry
Reservations were advisory bookkeeping: any component could name any status,
so the registry only detected declared-range overlap between components that
both opted in. Naming a status now requires a reservation.
akerr_register_status_name() checks that the range belongs to the caller, and
the legacy two-argument akerr_name_for_status() set path, which cannot
identify its caller, requires that some reservation covers the status. Every
refusal is logged and names the real owner, because a name that fails to
register degrades that code to "Unknown Error" in every later stack trace.

Fix a reservation made before the first PREPARE_ERROR being silently
discarded. akerr_init() clears the tables, so whichever component first
triggered it wiped an earlier reservation and the next component to claim the
same range was told it was free, producing exactly the undetected aliasing
the registry exists to prevent. Every registry entry point now calls
akerr_init(), which sets its guard before doing any work so those calls do
not recurse.

Replace the linear-scan name array with an open-addressed hash table, taking
lookup from O(n) to O(1) and raising usable capacity from 512 entries (366
free to consumers after errno registration) to 3072 (~2900 free). Both table
sizes are build-time overridable and applied PRIVATE: they live entirely in
src/error.c, so raising them cannot desynchronize a library from its
consumers the way AKERR_MAX_ERR_VALUE could. Exhausting either table is now
logged and returned to the caller rather than silently dropping the entry.
No dynamic allocation is introduced; both tables remain file-scope arrays,
and the library's undefined-symbol set gains only strcmp and strlen.

Register names for AKERR_EOF, AKERR_ITERATOR_BREAK and AKERR_NOT_IMPLEMENTED,
which had none and rendered as "Unknown Error" in every stack trace carrying
them. err_error_names.c now sweeps the whole AKERR_* offset span so a code
added without a name fails there instead of in production traces.

Add static assertions that the slot count is a power of two and that
AKERR_BADEXC stays inside the library's own 0-255 band, the latter guarding
against a host errno space large enough to push library codes into the range
consumers are told to allocate from.

Set a project version and soname (1.0.0 / libakerror.so.1) so a stale
installed library can no longer be silently paired with newer headers, and so
akerror.pc ships a real Version field instead of an empty one.

Mutation testing surfaced an out-of-bounds probe in the new table that the
suite did not catch: masking with SLOTS rather than SLOTS-1 indexes past the
array, and err_maxval.c asserted only that some names registered before the
table filled, which a collapsed probe sequence still satisfies. It now
requires a substantial entry count and reads every entry back by its own
distinct name.

Tests: 28/28 pass. Coverage 99.4% line / 86.8% branch. Mutation score for
src/error.c 74% -> 77.3%.

Compatibility: source and ABI break. AKERR_MAX_ERR_VALUE and the
__AKERR_ERROR_NAMES data symbol are gone, custom codes must move out of
0-255, and names must be registered against a reserved range. README.md
carries the migration steps.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 18:19:25 -04:00
11d21068df Add collision-safe status code registry
Replace the consumer-sized status-name array with private sparse storage and accept arbitrary integer status values. Add explicit range reservations with overlap diagnostics, reserve the library's 0-255 compatibility band, and harden pointer and string boundary handling.

Update regression coverage and document the required migration for custom status-code consumers.
2026-07-30 13:53:52 -04:00
0bb3a4d52c TODO.md
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m45s
libakerror CI Build / coverage (push) Successful in 2m45s
libakerror CI Build / mutation_test (push) Successful in 8m1s
2026-07-30 13:22:50 -04:00
539293cc1c Expand error release and handler test coverage
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m46s
libakerror CI Build / coverage (push) Successful in 2m45s
libakerror CI Build / mutation_test (push) Successful in 8m0s
2026-07-30 02:05:31 -04:00
9f0034a56e Add gcov code coverage to the test suite
scripts/coverage.py configures an instrumented build tree
(-DAKERR_COVERAGE=ON), runs the CTest suite in it, and reports merged
gcov line/branch/function coverage per library source. Like the mutation
harness it has no third-party dependencies and supports --threshold and
--junit; thresholds gate each file as well as the total so the generated
status-name table cannot mask a regression in src/error.c.

Only the library is instrumented. The public header's macros cannot be
measured this way -- GCC attributes an expanded macro to its call site,
so header logic would report as lines of the test that used it -- which
is what mutation testing against include/akerror.tmpl.h is for.

Coverage flags are applied per target rather than globally, so they do
not leak into the exported/installed target interface.

Current numbers for src/error.c are 94.0% line and 59.5% branch; the CI
gate is set to 90/50 to keep headroom, matching the convention used for
the mutation score threshold.

Tests run: ctest (23/23), cmake --build build --target coverage,
threshold gate verified failing at --threshold 99, cmake --install
checked for flag leakage, out-of-tree --build-dir checked.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-07-30 02:05:30 -04:00
0c0d81249f Avoid mutation target collisions when embedded
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m44s
libakerror CI Build / mutation_test (push) Successful in 7m38s
2026-07-29 18:01:58 -04:00
426efbb2d4 Add CLAUDE.md
Some checks failed
libakerror CI Build / cmake_build (push) Successful in 2m44s
libakerror CI Build / mutation_test (push) Has been cancelled
2026-07-29 17:42:56 -04:00
4ae1decde2 Expand error status test coverage
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m44s
libakerror CI Build / mutation_test (push) Successful in 7m46s
- Added explicit status validation across error handling tests.
- Added lifecycle and slot-leak checks to older tests.
- Improved unhandled-error propagation coverage.
- Added repository guidance and a test runner script.
- Verified all 23 CTest tests pass.

Co-Authored by Codex GPT 5.4
2026-07-29 17:12:52 -04:00
4212ff0b28 Fix format-string use of __FILE__/__func__ and name_for_status lower bound
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m48s
libakerror CI Build / mutation_test (push) Successful in 7m40s
Two hardening fixes flagged by the earlier review:

1. FAIL passed __FILE__ and __func__ directly as the snprintf format string.
   __FILE__ expands to a string literal that could contain a '%' (a build path
   under a directory with a percent sign), and __func__ is not a literal at all;
   either way snprintf would read nonexistent varargs. Pass them as "%s"
   arguments instead.

2. akerr_name_for_status guarded the upper bound but not the lower one, so a
   negative status indexed __AKERR_ERROR_NAMES[negative] -- an out-of-bounds
   read, or an out-of-bounds write when a name was supplied. Reject status < 0.

Regression tests err_format_string (uses #line to put a conversion specifier in
__FILE__) and err_name_bounds fail against the old code (verified) and pass now.
Full suite: 23/23, no warnings.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-28 10:52:29 -04:00
de13b290d4 Fix refcount leak and stack-trace buffer overflow
Two memory-safety bugs in the macro core:

1. Refcount leak. ENSURE_ERROR_READY incremented refcount on every FAIL/SUCCEED
   rather than only when it acquired a fresh context from the pool. A function
   that FAILed a context more than once and then propagated arrived at its
   caller with refcount 2; the caller released once, leaking the slot. After
   AKERR_MAX_ARRAY_ERROR leaks the pool is exhausted and the library exit(1)s.
   Move the increment inside the acquisition branch.

2. Stack-trace overflow. Each appended frame passed the full buffer length to
   snprintf instead of the space remaining, and advanced the cursor by
   snprintf's would-be return value, so a trace that filled the buffer wrote
   past the end of stacktracebuf and ran the cursor out of bounds. Add
   AKERR_STACKTRACE_APPEND, which bounds the write to the remaining space and
   clamps the cursor advance.

Regression tests err_refcount_double_fail and err_stacktrace_bounds fail against
the old code (verified) and pass now. Full suite: 21/21, no warnings.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-28 09:32:55 -04:00
3e24356f07 Document that CATCH/FAIL_*_BREAK must not be used inside a loop
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m47s
libakerror CI Build / mutation_test (push) Successful in 6m59s
These macros leave the ATTEMPT block with a C break, which only escapes the
innermost loop/switch. Nesting them in a loop inside an ATTEMPT lets the rest of
the block run with an error already pending. Document the correct patterns:
iterate with PASS / FAIL_*_RETURN (which return, not break), or move the loop
into a helper returning akerr_ErrorContext * and CATCH the single call. Also
note that merely extracting the loop into a function does not fix it if the
helper still wraps the loop in ATTEMPT.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-28 08:49:29 -04:00
8a026d3006 Include passed tests in the JUnit report summary
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m48s
libakerror CI Build / mutation_test (push) Successful in 7m13s
The reporter warned "No annotations found ... configure 'include_passed' as
'true'" because with annotate_only the summary only listed failures. Set
include_passed: true on both reporter steps so the job summary table lists the
passing tests too.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 21:15:13 -04:00
792e646957 Work around Gitea Checks API 404 in the JUnit reporter
Some checks failed
libakerror CI Build / cmake_build (push) Successful in 2m45s
libakerror CI Build / mutation_test (push) Has been cancelled
mikepenz/action-junit-report defaults to creating a check run via the Checks
API, which Gitea does not support -- the call 404s and the publish step fails
(mikepenz/action-junit-report#23). Set annotate_only: true on both reporter
steps to skip check creation, and detailed_summary: true so results still show
up in the job summary (which Gitea's runner does render).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 21:10:09 -04:00
43516c7e73 Emit JUnit XML from tests + mutation, consume it in CI
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 3m10s
libakerror CI Build / mutation_test (push) Successful in 6m58s
Produce machine-readable results and surface them in the Gitea pipeline:

- ctest: run with --output-junit to write ctest-junit.xml. The path must be
  absolute ("$(pwd)/...") because --output-junit otherwise resolves relative to
  the --test-dir build directory.
- mutation_test.py: new --junit FILE option writes a JUnit report where each
  mutant is a test case and a surviving mutant is a <failure> (so gaps show up
  as failing tests).
- .gitea/workflows/ci.yaml: both jobs generate their XML and feed it to
  mikepenz/action-junit-report with `if: always()`, so results publish even
  when a gate fails. Mutation publishing is display-only (fail_on_failure:
  false); the --threshold flag remains the gate.
- .gitignore: ignore the generated *-junit.xml artifacts.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 20:51:28 -04:00
536a269aad Derive err_maxval's code set from the header, not a hardcoded list
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m41s
libakerror CI Build / mutation_test (push) Successful in 6m52s
The previous err_maxval hardcoded the list of AKERR_* codes, which silently
rots the moment a code is added. Instead, parse the generated akerror.h at
runtime: discover every "#define AKERR_<NAME> (AKERR_LAST_ERRNO_VALUE + N)",
take the highest offset actually defined, and assert AKERR_MAX_ERR_VALUE covers
it. A compile-time cross-check ties the parsed ceiling to the compiled macro so
the test can't pass by reading a stale header.

CMake injects the header path as AKERR_GENERATED_HEADER. Verified: the test
fails (max_err_value >= highest_code) when pointed at a +15 header while
AKERR_BADEXC is +17, and now also strengthens header mutation coverage.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 17:26:32 -04:00
10f7203e8f Fix AKERR_MAX_ERR_VALUE to cover all AKERR_* codes
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m41s
libakerror CI Build / mutation_test (push) Successful in 6m54s
AKERR_MAX_ERR_VALUE was AKERR_LAST_ERRNO_VALUE + 15, but the highest defined
code, AKERR_BADEXC, is + 17 (AKERR_NOT_IMPLEMENTED is + 16). akerr_name_for_status
rejects any status above the max, so those codes could never have a registered
name and the AKERR_BADEXC registration in akerr_init was dead code -- a gap
found by mutation testing. Bump the max to + 17.

- err_maxval: new test asserting the reserved AKERR_* range exceeds the number
  of AKERR_* codes and that every code is individually indexable. Fails against
  the old + 15 value (verified), guarding against regression.
- err_error_names: now also checks AKERR_BADEXC's name, which the fix makes
  reachable.

Mutation score on src/error.c rises 71% -> 74%: the previously-dead BADEXC
registration and the name_for_status upper-bound check are now killable.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 17:17:26 -04:00
43f46dca64 Add mutation testing to validate the test suite
Introduce a self-contained mutation testing harness that verifies the unit
tests actually catch bugs: it makes small deliberate breakages to the library
(flip comparisons, delete statements, swap true/false, etc.), rebuilds, and
runs the whole CTest suite against each mutant. Tests that still pass reveal a
gap; tests that fail "kill" the mutant.

- scripts/mutation_test.py: the engine (stdlib only, no LLVM/clang deps).
  Operators ROR/LCR/BCR/AOR/ICR/SDL over src/error.c and the macro header.
  Mutates a scratch copy, never the working tree. Supports --target, --list,
  --max-mutants sampling, --threshold gating, --timeout.
- CMakeLists.txt: 'mutation' custom target (cmake --build build --target mutation).
- .gitea/workflows/ci.yaml: gated mutation job on src/error.c (threshold 65%).
- tests/MUTATION.md: how to run, interpret survivors, and known equivalents.

Close the real gaps the harness found in src/error.c (score 53% -> 71%):
- err_error_names: the AKERR_* codes have their names registered by akerr_init
- err_release_clears: releasing a context wipes it before reuse
- err_pool_exhaust: akerr_next_error returns NULL when the pool is full and
  always hands back the lowest free slot

Also surfaced (documented, not fixed): AKERR_MAX_ERR_VALUE (+15) is below
AKERR_NOT_IMPLEMENTED (+16) and AKERR_BADEXC (+17), so those codes can never
have a name registered.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 17:03:53 -04:00
e5f761662c Reformat new test files with Stroustrup style
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m38s
Apply emacs CC-mode "stroustrup" style (c-basic-offset 4, indent-tabs-mode t)
to the test files added in the previous commit, matching the existing house
style. Whitespace only; no behavioral change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 16:35:32 -04:00
4daa411f3f Expand test coverage for error-handling macros and error pool
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m42s
Add 11 CTest programs and a shared test helper covering gaps left by the
original four tests (which only exercised FAIL/CATCH/HANDLE and CLEANUP):

- err_capture.h: capturing akerr_log_method + NDEBUG-proof AKERR_CHECK so
  tests can assert on message/status/stacktrace content, not just exit codes
- err_success: clean nested return does not break/handle/leak
- err_pool_refcount: 100k raise->catch->handle cycles leak 0 pool slots
- err_handle_default / err_handle_group / err_handle_dispatch: handler routing
- err_pass / err_ignore / err_swallow: PASS, IGNORE, FINISH(e,false)
- err_break_variants: FAIL_*_BREAK and FAIL_*_RETURN
- err_errno: system errno name lookup + "Unknown Error" boundary
- err_custom_handler: override the unhandled-error hook, assert non-fatally

Register tests via a foreach loop in CMakeLists.txt. Ignore build/.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-27 16:18:17 -04:00
4fad0cec59 Add gitea workflow
All checks were successful
libakerror CI Build / cmake_build (push) Successful in 2m39s
2026-06-27 10:22:26 -04:00
5b1a6c2cc7 Remove spurious duplicate lines from the stacktrace. Only log the DETECT and FAIL points, everything else is noise. 2026-06-27 08:37:05 -04:00
675c60b5e0 Removed AKERR_HEAP, AKERR_REGISTRY, AKERR_BEHAVIOR, they belong to libakgl. Replaced AKERR_BEHAVIOR which was ambiguous with AKERR_BADEXC which is raised when a function falls to call SUCCEED_RETURN(). 2026-06-22 08:15:07 -04:00
93f5e93480 Add an error type for circular references 2026-06-02 17:12:14 -04:00
b5435041f2 VALID() wasn't resetting invalid error contexts back to NULL before atempting to ENSURE_READY() them 2026-05-24 19:41:32 -04:00
235033d633 VALID() wasn't properly handling NULL returns, leading to false positives 2026-05-24 19:14:35 -04:00
7700af06a1 Changes maximum stacktrace string values to better allow for inclusion of values up to PATH_MAX. I don't like it. The values are huge. Need a better more sensible way. 2026-05-24 09:50:21 -04:00
be2dba8416 Fix a bug in the new VALID() method, false is not true 2026-05-21 21:44:52 -04:00
03e9b8a96d Update docs, replace the missing VALID() macro to safeguard against misbehaving functions, add PASS, remove CATCH_AND_RETURN 2026-05-15 19:41:22 -04:00
768a235da4 Fix builds when in a submodule 2026-05-12 21:29:07 -04:00
4a09ca87fd Fix cmake duplicate targets 2026-05-12 16:56:49 -04:00
51b6b23b4c Make the build work from a subdirectory dependency 2026-05-12 16:44:06 -04:00
166a478e6a Add AKERR_EOF, properly call akerr_init_errno 2026-05-12 16:42:33 -04:00
0c29f5d69f Fix handling of CATCH() or FAIL() macros around functions that should return an (akerr_ErrorContext *) but return an invalid pointer (to something not in our exception array) 2026-05-06 12:36:56 -04:00
55 changed files with 5990 additions and 158 deletions

127
.gitea/workflows/ci.yaml Normal file
View File

@@ -0,0 +1,127 @@
name: libakerror CI Build
run-name: ${{ gitea.actor }} libakerror test
on: [push]
jobs:
cmake_build:
runs-on: ubuntu-latest
steps:
- run: echo "Triggered by ${{ gitea.event_name }} from ${{ gitea.repository }}@${{ gitea.ref }}. Building on ${{ runner.os }}."
- name: Check out repository code
uses: actions/checkout@v4
- name: dependencies
run: |
sudo apt-get update -y
sudo apt-get install -y cmake gcc moreutils
- name: build and install
run: |
mkdir installdir
cmake -S . -B build -DCMAKE_INSTALL_PREFIX=installdir
cmake --build build
cmake --install build
# --output-junit resolves relative to the test dir, so give an absolute
# path to land the report in the workspace root for the reporter below.
- name: test (JUnit)
run: ctest --test-dir build --output-on-failure --output-junit "$(pwd)/ctest-junit.xml"
# annotate_only: true skips creating a check run via the Checks API, which
# Gitea does not support and 404s on (mikepenz/action-junit-report#23).
# Results surface via the job summary instead.
- name: publish test results
if: always()
uses: mikepenz/action-junit-report@v4
with:
report_paths: 'ctest-junit.xml'
annotate_only: true
detailed_summary: true
include_passed: true
fail_on_failure: 'true'
- run: echo "🍏 This job's status is ${{ job.status }}."
coverage:
runs-on: ubuntu-latest
steps:
- name: Check out repository code
uses: actions/checkout@v4
- name: dependencies
run: |
sudo apt-get update -y
sudo apt-get install -y cmake gcc moreutils python3
# Run the suite against a gcov-instrumented build and gate on coverage of
# the library sources. Thresholds keep headroom below the current numbers
# (src/error.c: ~94% line, ~60% branch) and apply per file as well as to
# the total, so the generated status-name table cannot mask a regression.
- name: coverage
run: |
python3 scripts/coverage.py --junit coverage-junit.xml --threshold 90 --branch-threshold 50
# Publish even when the threshold gate fails, so gaps are visible.
# Display-only (fail_on_failure: false); the --threshold above is the gate.
# annotate_only avoids the Checks API 404 on Gitea (see note above).
- name: publish coverage results
if: always()
uses: mikepenz/action-junit-report@v4
with:
report_paths: 'coverage-junit.xml'
annotate_only: true
detailed_summary: true
include_passed: true
fail_on_failure: 'false'
- run: echo "🍏 This job's status is ${{ job.status }}."
thread_sanitizer:
runs-on: ubuntu-latest
steps:
- name: Check out repository code
uses: actions/checkout@v4
- name: dependencies
run: |
sudo apt-get update -y
sudo apt-get install -y cmake gcc moreutils
# The thread tests assert exclusive ownership of pool slots and reserved
# ranges, which is checkable without tooling and runs in the job above.
# This is the run that proves there is no data race underneath them.
# libtsan arrives with gcc (libgcc-N-dev depends on it); the script
# disables ASLR because TSan aborts on kernels with vm.mmap_rnd_bits > 28.
- name: thread sanitizer
run: |
scripts/thread_test.sh build/tsan --output-junit "$(pwd)/tsan-junit.xml"
- name: publish thread sanitizer results
if: always()
uses: mikepenz/action-junit-report@v4
with:
report_paths: 'tsan-junit.xml'
annotate_only: true
detailed_summary: true
include_passed: true
fail_on_failure: 'true'
- run: echo "🍏 This job's status is ${{ job.status }}."
mutation_test:
runs-on: ubuntu-latest
steps:
- name: Check out repository code
uses: actions/checkout@v4
- name: dependencies
run: |
sudo apt-get update -y
sudo apt-get install -y cmake gcc moreutils python3
# Verify the tests actually catch bugs: break the library many ways and
# confirm the suite fails. Gated on src/error.c (fast, deterministic).
# The threshold keeps headroom below the current score for equivalent
# mutants; see tests/MUTATION.md. Run the full default target locally for
# deeper (slower) coverage including the macro header.
- name: mutation testing
run: |
python3 scripts/mutation_test.py --target src/error.c --junit mutation-junit.xml --threshold 65
# Publish even when the threshold gate fails, so survivors are visible.
# Display-only (fail_on_failure: false); the --threshold above is the gate.
# annotate_only avoids the Checks API 404 on Gitea (see note above).
- name: publish mutation results
if: always()
uses: mikepenz/action-junit-report@v4
with:
report_paths: 'mutation-junit.xml'
annotate_only: true
detailed_summary: true
include_passed: true
fail_on_failure: 'false'
- run: echo "🍏 This job's status is ${{ job.status }}."

3
.gitignore vendored Normal file
View File

@@ -0,0 +1,3 @@
include/akerror.h
build/
*-junit.xml

136
AGENTS.md Normal file
View File

@@ -0,0 +1,136 @@
# Repository Guidelines
## Project Structure & Module Organization
This is a small C library built with CMake. Core implementation lives in
`src/error.c`. `src/lock.h` is private: it selects the threading backend
(pthread or none) and is not installed. The public header is generated at build
time from `include/akerror.tmpl.h` by `scripts/generrno.sh`, which also
generates `src/errno.c` under the build directory and stamps in whether the
build is thread safe. CMake package and pkg-config templates are in `cmake/` and
`akerror.pc.in`. Tests are one-file C programs in `tests/`; shared test helpers
live beside them, such as `tests/err_capture.h` and `tests/err_threads.h`.
## Build, Test, and Development Commands
Use an out-of-tree build:
```sh
cmake -S . -B build
cmake --build build
ctest --test-dir build --output-on-failure
```
`cmake -S . -B build` configures the build and generates `akerror.h`.
`cmake --build build` compiles `libakerror` and test executables. `ctest`
runs the registered unit tests. To emit CI-style results, use:
```sh
ctest --test-dir build --output-on-failure --output-junit "$(pwd)/ctest-junit.xml"
```
The suite includes thread tests. The run that proves there is no data race under
them is ThreadSanitizer:
```sh
scripts/thread_test.sh
```
which configures `build/tsan` with `-DAKERR_SANITIZE=thread`, builds the library
*and* the tests with it, and runs CTest with ASLR disabled (TSan aborts with an
"unexpected memory mapping" on kernels with `vm.mmap_rnd_bits` above 28).
`AKERR_SANITIZE` takes any sanitizer list, so `-DAKERR_SANITIZE=address,undefined`
works the same way. CTest gives each test `halt_on_error=1`, so a report fails
the test rather than being printed and passed over — running a sanitized test
binary by hand does **not** inherit that. Set it yourself
(`TSAN_OPTIONS=halt_on_error=1 ./build/tsan/test_err_threads_pool`) or the
binary can print a race and still exit 0.
Mutation testing is available through:
```sh
cmake --build build --target mutation
scripts/mutation_test.py --target src/error.c --threshold 65
```
Code coverage is available through:
```sh
cmake --build build --target coverage
scripts/coverage.py --threshold 90 --branch-threshold 50
```
`scripts/coverage.py` configures its own instrumented build tree (default
`build/coverage`, `-DAKERR_COVERAGE=ON`), runs the CTest suite there, and
reports gcov line/branch coverage per library source. Thresholds gate each
file as well as the total. Coverage measures the library's own sources only --
`src/error.c`, `src/lock.h` and the generated `src/errno.c`; the public header's
macros expand at their call sites, so mutation testing is what checks those.
To build without threads (no locking, no thread-local storage, thread tests not
registered), configure with `-DAKERR_THREADS=none`. `auto` is the default and
fails the configure rather than falling back.
## Coding Style & Naming Conventions
Use C99-compatible C and follow the surrounding style. Functions and types use
the `akerr_` prefix; macros and constants use `AKERR_` or all-caps macro names
such as `PREPARE_ERROR`. Keep generated-code changes in templates or generator
scripts, not in build outputs. Preserve concise comments for invariants,
macro constraints, and non-obvious error lifecycle behavior.
**Locking.** One recursive lock (`akerr_state_lock`, see `src/lock.h`) covers
both the error pool and the status registry. A function named `*_locked` is
called with that lock already held; a function without the suffix takes it.
Never take it in a body written with the `FAIL_*_RETURN` macros — those return
from the middle of the function and would skip the unlock. Split it instead: the
locked body does the work, and a thin wrapper takes the lock, calls it, and
releases it on the single return path. Do not call consumer code
(`akerr_log_method`, `akerr_handler_unhandled_error`) while holding it.
**Exiting.** Never call `exit()` with a status value. Call `akerr_exit()`, which
owns the one mapping from an akerr status to an exit code: an exit status is a
byte, and every consumer status starts at 256, so passing a status through
`exit()` silently truncates it — status 256 exits 0 and reports success. The
only exits that bypass it are the two that have no status to map, both in
`src/error.c`: `ENSURE_ERROR_READY`'s pool-exhaustion abort and the NULL-context
case in `akerr_default_handler_unhandled_error()`. This rule applies to test
programs too, except where the test's whole point is to observe the raw
truncation.
## Testing Guidelines
Add tests as `tests/err_<behavior>.c`. Register each new test in the
`AKERR_TESTS` list in `CMakeLists.txt`; tests that intentionally abort must
also be listed in `AKERR_WILL_FAIL_TESTS`. Prefer focused executable tests that
return zero on success and use existing helpers such as `AKERR_CHECK`. Run the
CTest suite before submitting changes, and run mutation testing and coverage
when changing core control-flow, reference counting, stack-trace, or handler
behavior.
Tests that drive the library from several threads go in the `AKERR_THREAD_SAFE`
branch of the `AKERR_TESTS` list — an `-DAKERR_THREADS=none` build has no
threading to test — and use `tests/err_threads.h`, which runs a body on
`AKERR_TEST_THREADS` threads that meet at a barrier first.
Count failures per thread with `AKERR_TCHECK` rather than returning early: a
thread that abandons its work leaves the others holding pool slots and turns one
failure into a cascade. Anything the test shares between its own threads must go
through `__atomic` builtins — a race in the test is still a race, and
ThreadSanitizer cannot tell you whose it is. The capturing logger in
`err_capture.h` is single-threaded (shared buffer, shared length); use
`akerr_thread_logger` instead. Run `scripts/thread_test.sh` when changing
anything that touches the pool, the registry, initialization, or the lock.
## Commit & Pull Request Guidelines
Recent commits use short, imperative, sentence-case subjects, for example
`Fix refcount leak and stack-trace buffer overflow`. Keep commits focused and
describe the observable behavior changed. Pull requests should include a brief
summary, tests run, and any compatibility impact for public macros, generated
headers, installation paths, or CMake/pkg-config consumers.
## Agent-Specific Instructions
Do not overwrite uncommitted user changes. Avoid editing generated files in
`build/`; update `include/akerror.tmpl.h`, `src/error.c`, CMake files, tests,
or scripts instead.

1
CLAUDE.md Normal file
View File

@@ -0,0 +1 @@
See @AGENTS.md

View File

@@ -1,59 +1,315 @@
cmake_minimum_required(VERSION 3.10)
project(akerror LANGUAGES C)
# 1.0.0 replaced the consumer-sized __AKERR_ERROR_NAMES array with private
# storage. 2.0.0 makes the library thread safe, which is a second ABI break in
# the same places: __akerr_last_ignored became thread-local storage, and
# ENSURE_ERROR_READY no longer takes the pool reference that akerr_next_error()
# now takes for it. Consumer code compiled against a 1.x header would
# double-count every reference. Hence the major bump and the SOVERSION, so a
# stale installed libakerror.so cannot be silently paired with new headers.
# 2.0.1 fixes the unhandled-error exit code, which reported success for any
# status whose low byte was zero. It adds akerr_exit() but breaks nothing: the
# soname is unchanged and no existing entry point changed shape.
project(akerror VERSION 2.0.1 LANGUAGES C)
include(GNUInstallDirs)
include(CMakePackageConfigHelpers)
include(CTest)
set(AKERR_USE_STDLIB 1 CACHE BOOL "Use the C standard library")
set(AKERR_COVERAGE 0 CACHE BOOL "Instrument the build with gcov coverage counters")
set(AKERR_SANITIZE "" CACHE STRING
"Sanitizers to build the library and tests with, e.g. thread or address,undefined")
# Threading backend for the library's global state (the error pool and the
# status registry). "auto" takes POSIX threads when they exist and fails the
# configure when they do not: a build that silently fell back to no locking
# would produce a library that reports itself thread safe and is not. Say
# -DAKERR_THREADS=none to mean it on purpose.
set(AKERR_THREADS "auto" CACHE STRING "Threading backend: auto, pthread, or none")
set_property(CACHE AKERR_THREADS PROPERTY STRINGS auto pthread none)
if(AKERR_THREADS STREQUAL "auto" OR AKERR_THREADS STREQUAL "pthread")
set(THREADS_PREFER_PTHREAD_FLAG ON)
find_package(Threads)
if(CMAKE_USE_PTHREADS_INIT)
set(AKERR_THREAD_SAFE 1)
elseif(AKERR_THREADS STREQUAL "pthread")
message(FATAL_ERROR
"-DAKERR_THREADS=pthread was requested but no POSIX thread library "
"was found.")
else()
message(FATAL_ERROR
"No POSIX thread library was found. libakerror serializes its "
"global state with a recursive pthread mutex; without one it "
"cannot be thread safe. Configure with -DAKERR_THREADS=none to "
"build a deliberately single-threaded library instead.")
endif()
elseif(AKERR_THREADS STREQUAL "none")
set(AKERR_THREAD_SAFE 0)
else()
message(FATAL_ERROR
"AKERR_THREADS must be auto, pthread or none, not '${AKERR_THREADS}'")
endif()
# Size of the private status-name hash table. Must be a power of two; usable
# capacity is 75% of it (src/error.c asserts both). The host's errno list
# consumes part of that at akerr_init() time, so the remainder is what all
# consumer libraries in the process share. These are applied PRIVATE on purpose:
# the table lives entirely in src/error.c, so raising them never changes
# anything a consumer can see. That is what makes them safe to tune, unlike the
# AKERR_MAX_ERR_VALUE they replaced.
set(AKERR_STATUS_NAME_SLOTS 4096 CACHE STRING
"Slots in the status-name table (power of two; 75% usable)")
set(AKERR_MAX_RESERVED_STATUS_RANGES 64 CACHE STRING
"Maximum number of status ranges that may be reserved")
set(akerror_install_cmakedir "${CMAKE_INSTALL_LIBDIR}/cmake/akerror")
set(SCRIPT ${CMAKE_SOURCE_DIR}/scripts/generrno.sh)
set(INFILE ${CMAKE_SOURCE_DIR}/include/akerror.tmpl.h)
set(OUTFILES ${CMAKE_SOURCE_DIR}/src/errno.c ${CMAKE_SOURCE_DIR}/include/akerror.h)
# Coverage instrumentation. Applied per target (not globally) so it never leaks
# into the exported/installed target interface. Only the library is
# instrumented: the tests are the thing doing the covering, and the public
# header's macros cannot be measured this way at all -- GCC attributes an
# expanded macro to its call site, so header logic shows up as test-file lines.
# Coverage of those macros is what mutation testing (--target
# include/akerror.tmpl.h) is for.
if(AKERR_COVERAGE)
if(NOT CMAKE_C_COMPILER_ID MATCHES "GNU|Clang")
message(FATAL_ERROR
"AKERR_COVERAGE requires GCC or Clang, not ${CMAKE_C_COMPILER_ID}")
endif()
# -O0 keeps line counts attributable; no inlining or code motion.
set(AKERR_COVERAGE_FLAGS --coverage -O0 -g)
if(CMAKE_C_COMPILER_ID STREQUAL "GNU")
# Record absolute source paths so gcov resolves sources built from the
# generated directory (GCC 8+; harmless to check).
include(CheckCCompilerFlag)
check_c_compiler_flag(-fprofile-abs-path AKERR_HAVE_PROFILE_ABS_PATH)
if(AKERR_HAVE_PROFILE_ABS_PATH)
list(APPEND AKERR_COVERAGE_FLAGS -fprofile-abs-path)
endif()
endif()
endif()
# Add coverage compile/link flags to one target, if coverage is enabled.
function(akerr_instrument_for_coverage _target)
if(AKERR_COVERAGE)
target_compile_options(${_target} PRIVATE ${AKERR_COVERAGE_FLAGS})
set_property(TARGET ${_target} APPEND_STRING
PROPERTY LINK_FLAGS " --coverage")
endif()
endfunction()
# Sanitizers. Unlike coverage these go on the tests as well as the library:
# ThreadSanitizer only sees a race if every thread that touches the memory was
# compiled with it, and the threads live in the test programs.
# cmake -S . -B build/tsan -DAKERR_SANITIZE=thread
if(AKERR_SANITIZE AND NOT CMAKE_C_COMPILER_ID MATCHES "GNU|Clang")
message(FATAL_ERROR
"AKERR_SANITIZE requires GCC or Clang, not ${CMAKE_C_COMPILER_ID}")
endif()
function(akerr_instrument_for_sanitizers _target)
if(AKERR_SANITIZE)
target_compile_options(${_target} PRIVATE
-fsanitize=${AKERR_SANITIZE}
-fno-omit-frame-pointer -g -O1)
set_property(TARGET ${_target} APPEND_STRING
PROPERTY LINK_FLAGS " -fsanitize=${AKERR_SANITIZE}")
endif()
endfunction()
set(SCRIPT ${CMAKE_CURRENT_SOURCE_DIR}/scripts/generrno.sh)
set(INFILE ${CMAKE_CURRENT_SOURCE_DIR}/include/akerror.tmpl.h)
set(GENERATED_DIR ${CMAKE_CURRENT_BINARY_DIR}/generated)
set(GENERATED_ERRNO_C ${GENERATED_DIR}/src/errno.c)
set(GENERATED_AKERROR_H ${GENERATED_DIR}/include/akerror.h)
# The threading decision is stamped into the generated header, so the header has
# to be regenerated when it changes. Makefile generators compare timestamps
# rather than command lines, so carry the value through a file: configure_file
# rewrites it only when the content differs, which is exactly the trigger we
# want and no trigger at all on an unchanged reconfigure.
set(GENERATED_THREAD_STAMP ${CMAKE_CURRENT_BINARY_DIR}/akerr_thread_safe.stamp)
configure_file(cmake/thread_safe.stamp.in ${GENERATED_THREAD_STAMP} @ONLY)
add_custom_command(
OUTPUT ${OUTFILES}
COMMAND ${CMAKE_COMMAND} -E make_directory ${CMAKE_SOURCE_DIR}
COMMAND /usr/bin/env bash ${SCRIPT} ${CMAKE_SOURCE_DIR}
DEPENDS ${SCRIPT} ${INFILE}
OUTPUT ${GENERATED_ERRNO_C} ${GENERATED_AKERROR_H}
COMMAND ${CMAKE_COMMAND} -E make_directory ${GENERATED_DIR}
COMMAND /usr/bin/env bash
${SCRIPT}
${CMAKE_CURRENT_SOURCE_DIR}
${GENERATED_DIR}
${AKERR_THREAD_SAFE}
DEPENDS ${SCRIPT} ${INFILE} ${GENERATED_THREAD_STAMP}
VERBATIM
)
set_source_files_properties(src/errno.c PROPERTIES GENERATED TRUE)
add_library(akerror SHARED
src/error.c
${GENERATED_ERRNO_C}
)
add_custom_target(generrno DEPENDS src/errno.c)
target_include_directories(akerror PUBLIC
$<BUILD_INTERFACE:${GENERATED_DIR}/include>
$<INSTALL_INTERFACE:${CMAKE_INSTALL_INCLUDEDIR}>
)
find_package(PkgConfig REQUIRED)
add_library(akerror STATIC
src/error.c
src/errno.c
)
add_dependencies(akerror generrno)
add_library(akerror::akerror ALIAS akerror)
# The threading backend is PRIVATE: src/lock.h is not installed, so which
# primitive the library locks with is invisible to a consumer. What a consumer
# does see -- whether the library locks at all -- travels in the generated
# header instead, where it cannot disagree with this build.
if(AKERR_THREAD_SAFE)
set(AKERR_THREADS_DEFINE AKERR_THREADS_PTHREAD=1)
else()
set(AKERR_THREADS_DEFINE AKERR_THREADS_NONE=1)
endif()
target_compile_definitions(akerror
PUBLIC AKERR_USE_STDLIB=${AKERR_USE_STDLIB}
PRIVATE AKERR_STATUS_NAME_SLOTS=${AKERR_STATUS_NAME_SLOTS}
PRIVATE AKERR_MAX_RESERVED_STATUS_RANGES=${AKERR_MAX_RESERVED_STATUS_RANGES}
PRIVATE ${AKERR_THREADS_DEFINE}
)
add_executable(test_err_catch tests/err_catch.c)
add_executable(test_err_cleanup tests/err_cleanup.c)
add_executable(test_err_trace tests/err_trace.c)
add_executable(test_err_unhandled tests/err_unhandled.c)
add_test(NAME err_catch COMMAND test_err_catch)
add_test(NAME err_cleanup COMMAND test_err_cleanup)
add_test(NAME err_trace COMMAND test_err_trace)
add_test(NAME err_unhandled COMMAND test_err_unhandled)
if(AKERR_THREAD_SAFE)
target_link_libraries(akerror PRIVATE Threads::Threads)
endif()
# Specify include directories for the library's headers (if applicable)
target_include_directories(akerror PUBLIC
$<BUILD_INTERFACE:${CMAKE_CURRENT_SOURCE_DIR}/include>
$<INSTALL_INTERFACE:${CMAKE_INSTALL_INCLUDEDIR}/>
set_target_properties(akerror PROPERTIES
VERSION ${PROJECT_VERSION}
SOVERSION ${PROJECT_VERSION_MAJOR}
)
target_link_libraries(test_err_catch PRIVATE akerror)
target_link_libraries(test_err_cleanup PRIVATE akerror)
target_link_libraries(test_err_trace PRIVATE akerror)
target_link_libraries(test_err_unhandled PRIVATE akerror)
akerr_instrument_for_coverage(akerror)
akerr_instrument_for_sanitizers(akerror)
# Each test is one source file in tests/ built into test_<name> and registered
# as CTest <name>. Tests expected to abort (unhandled error / contract
# violation) go in AKERR_WILL_FAIL_TESTS; all others must exit 0.
set(AKERR_TESTS
err_catch
err_cleanup
err_trace
err_improper_closure
err_success
err_pool_refcount
err_handle_default
err_handle_group
err_handle_dispatch
err_pass
err_ignore
err_swallow
err_errno
err_break_variants
err_custom_handler
err_error_names
err_release_clears
err_pool_exhaust
err_maxval
err_name_ownership
err_registry_init_order
err_status_exception
err_copy_string
err_library_status_fatal
err_refcount_double_fail
err_stacktrace_bounds
err_name_bounds
err_format_string
err_unhandled_null
err_exit_status
err_release_null
err_release_refcount
)
# These drive the library from many threads at once. They are worth running on
# their own -- they assert exclusive ownership of pool slots and of reserved
# ranges, which is checkable without a sanitizer -- but the run that proves the
# absence of a race is the one under -DAKERR_SANITIZE=thread.
if(AKERR_THREAD_SAFE)
list(APPEND AKERR_TESTS
err_threads_init
err_threads_pool
err_threads_registry
)
endif()
set(AKERR_WILL_FAIL_TESTS
err_trace
err_improper_closure
err_library_status_fatal
)
foreach(_test IN LISTS AKERR_TESTS)
add_executable(test_${_test} tests/${_test}.c)
target_include_directories(test_${_test} PRIVATE ${CMAKE_CURRENT_SOURCE_DIR}/tests)
target_link_libraries(test_${_test} PRIVATE akerror)
if(AKERR_THREAD_SAFE)
target_link_libraries(test_${_test} PRIVATE Threads::Threads)
endif()
akerr_instrument_for_sanitizers(test_${_test})
add_test(NAME ${_test} COMMAND test_${_test})
# A sanitizer report is a test failure. Without halt_on_error the runtime
# prints and continues, which leaves a race to be noticed in the log by
# somebody reading it -- and under a race storm the reporting itself is slow
# enough to look like a hang.
if(AKERR_SANITIZE)
set_tests_properties(${_test} PROPERTIES ENVIRONMENT
"TSAN_OPTIONS=halt_on_error=1;ASAN_OPTIONS=halt_on_error=1;UBSAN_OPTIONS=halt_on_error=1:print_stacktrace=1")
endif()
endforeach()
set_tests_properties(
${AKERR_WILL_FAIL_TESTS}
PROPERTIES WILL_FAIL TRUE
)
# Coverage and mutation testing are meta-checks on the test suite itself, and
# both rebuild and re-run the whole suite, so they are manual targets rather
# than CTest tests.
find_package(Python3 COMPONENTS Interpreter)
if(Python3_FOUND)
# Code coverage: which library lines/branches the CTest suite reaches.
# cmake --build build --target coverage
# The script configures and drives its own instrumented build tree (under
# ${CMAKE_BINARY_DIR}/coverage) so this build's binaries and its coverage
# counters can never be stale or half-instrumented. Reports via gcov.
add_custom_target(coverage
COMMAND ${Python3_EXECUTABLE}
${CMAKE_CURRENT_SOURCE_DIR}/scripts/coverage.py
--source-root ${CMAKE_CURRENT_SOURCE_DIR}
--build-dir ${CMAKE_CURRENT_BINARY_DIR}/coverage
--cmake ${CMAKE_COMMAND}
--ctest ${CMAKE_CTEST_COMMAND}
WORKING_DIRECTORY ${CMAKE_CURRENT_SOURCE_DIR}
USES_TERMINAL
COMMENT "Running the test suite instrumented for coverage"
)
# Mutation testing: break the library in small ways and confirm the test
# suite notices.
# cmake --build build --target mutation
# When embedded in another project, use a namespaced target to avoid
# collisions with mutation targets provided by sibling dependencies.
if(CMAKE_SOURCE_DIR STREQUAL CMAKE_CURRENT_SOURCE_DIR)
set(AKERR_MUTATION_TARGET mutation)
else()
set(AKERR_MUTATION_TARGET akerror_mutation)
endif()
add_custom_target(${AKERR_MUTATION_TARGET}
COMMAND ${Python3_EXECUTABLE}
${CMAKE_CURRENT_SOURCE_DIR}/scripts/mutation_test.py
--source-root ${CMAKE_CURRENT_SOURCE_DIR}
WORKING_DIRECTORY ${CMAKE_CURRENT_SOURCE_DIR}
USES_TERMINAL
COMMENT "Running mutation tests (breaks the library, expects tests to fail)"
)
endif()
set(main_lib_dest "lib/my_library-${MY_LIBRARY_VERSION}")
install(TARGETS akerror EXPORT akerror DESTINATION "lib/")
install(TARGETS akerror
EXPORT akerrorTargets
ARCHIVE DESTINATION ${CMAKE_INSTALL_LIBDIR}
@@ -62,13 +318,10 @@ install(TARGETS akerror
INCLUDES DESTINATION ${CMAKE_INSTALL_INCLUDEDIR}
)
install(EXPORT akerror FILE akerrorTargets.cmake DESTINATION ${CMAKE_INSTALL_LIBDIR}/cmake/akerror)
install(FILES "include/akerror.h" DESTINATION "include/")
install(FILES ${GENERATED_AKERROR_H} DESTINATION "include/")
install(FILES ${CMAKE_CURRENT_BINARY_DIR}/akerror.pc DESTINATION "lib/pkgconfig/")
install(EXPORT akerror
install(EXPORT akerrorTargets
FILE akerrorTargets.cmake
NAMESPACE akerror::
DESTINATION ${akerror_install_cmakedir}

312
README.md
View File

@@ -2,6 +2,23 @@
This library provides a TRY/CATCH style exception handling mechanism for C.
![build badge](https://source.starfort.tech/andrew/libakerror/actions/workflows/ci.yaml/badge.svg?branch=main)
## Upgrading
2.0.1 fixes an unhandled error killing the process and still reporting success:
the exit code was the status truncated to a byte, and every consumer status
starts at 256. Use `akerr_exit()` instead of `exit()` — see
[Exit status](#exit-status). No ABI break.
2.0.0 makes the library thread safe. That is an ABI break — `__akerr_last_ignored`
became thread-local storage and the pool now takes its own reference — so
everything built against a 1.x header must be rebuilt. 1.0.0 replaced the
consumer-sized status-name array with a private, ownership-enforced registry.
See [UPGRADING.md](UPGRADING.md) for all three, what was removed, how to migrate,
the capacity limits and how to raise them, and the thread-safety rules.
# Why?
There is nothing wrong with C as it is. This library does not claim to fix some problem with C.
@@ -62,14 +79,132 @@ Any function which uses the `PREPARE_ERROR` macro should have a return type of `
The library uses integer values to specify error codes inside of its context. These integer return codes are defined in `akerror.h` in the form of `AKERR_xxxxx` where `xxxxx` is the name of the error code in question. See `akerror.h` for a list of defined errors and their descriptions.
You can define additional error types by defining additional `AKERR_xxxxx` values. Error values up to 255 are reserved by the library (`AKERR_xxxxx` begins where `errno` stops), so please begin your error values at 256. When you add additional error codes, you need to define `-DAKERR_MAX_ERR_VALUE=n` to the compiler, where `n` is the maximum error code you have defined. If you define custom error codes, `AKERR_MAX_ERR_VALUE` must be >= 256 or the compiler will throw an error.
You can define additional error types as integer constants. Values 0 through 255
are reserved by libakerror (the host's errno values plus the `AKERR_*` codes);
consumers allocate codes starting at `AKERR_FIRST_CONSUMER_STATUS` (256). Status
names are stored sparsely, so any `int` is a legal status and no compile
definition is needed to use large values. Note that no consumer status can be a
process exit code — see [Exit status](#exit-status).
Define a human-friendly name for the error with the `akerr_name_for_status` method somewhere in your code's initialization before the error may be used:
Every library that may coexist in one process must reserve its range during
initialization. `akerr_reserve_status_range()` and
`akerr_register_status_name()` report failure the way everything else in this
library does — they return `akerr_ErrorContext *`, and they are marked
`AKERR_NOIGNORE` — so a collision is an exception you can `CATCH`, `HANDLE`, or
`PASS` up out of your initialization:
```c
akerr_name_for_status(129, "Some Error Code Description")
#define MYLIB_OWNER "my-library"
akerr_ErrorContext AKERR_NOIGNORE *mylib_init(void)
{
PREPARE_ERROR(errctx);
/* Another component owning part of the range propagates to our caller. */
PASS(errctx, akerr_reserve_status_range(256, 16, MYLIB_OWNER));
/* Then name each code, quoting the owner you reserved with. */
PASS(errctx, akerr_register_status_name(MYLIB_OWNER, 256,
"Some Error Code Description"));
SUCCEED_RETURN(errctx);
}
```
Reservations are process-local, fixed-capacity, and idempotent when the same
owner repeats the exact same range. They preserve compile-time integer constants
so `HANDLE` still works, since `case` labels require them.
Naming a status outside your reservation raises `AKERR_STATUS_NAME_FOREIGN`, and
naming one nobody reserved raises `AKERR_STATUS_NAME_UNRESERVED`; a colliding
reservation raises `AKERR_STATUS_RANGE_OVERLAP`, with a message naming the real
owner. See [UPGRADING.md](UPGRADING.md) for the full list of statuses these two
functions raise, the capacity limits and how to raise them, and thread-safety
rules.
# Thread safety
The library is thread safe as built by default. Every entry point may be called
from any thread at any time, including the first one: `akerr_init()` runs
exactly once no matter how many threads race into it.
What that covers:
* **The error pool.** Finding a free slot in `AKERR_ARRAY_ERROR` and taking its
reference is one operation under a lock, so two threads can never be handed
the same context. A context is then owned by the thread that raised it, all
the way through `CATCH`, `HANDLE`, and release.
* **The status registry.** Reservations and name registrations are serialized
against each other and against lookups. Two threads reserving the same range
cannot both win — exactly one gets `NULL` and the other gets
`AKERR_STATUS_RANGE_OVERLAP` naming the winner.
* **Per-thread state.** The context behind `IGNORE` (`__akerr_last_ignored`) and
the last-ditch context used to report `akerr_release_error(NULL)` are
thread-local, so one thread's ignored error is never another's.
What it does not cover, and cannot:
* **Sharing one error context between threads.** The library hands a context to
one thread; passing it to another is your synchronization to do.
* **`akerr_log_method` and `akerr_handler_unhandled_error`.** Set them during
startup, before you spawn threads. They are read on every error and the
library never writes them after initialization, so setting one while other
threads are raising errors is a race the library cannot mediate.
* **Renaming a status that other threads are looking up.**
`akerr_name_for_status(status, NULL)` returns a pointer into the registry,
valid for the life of the process; registering a *second* name for the same
status overwrites that buffer in place. Register names during initialization.
Registering a *new* status concurrently is fine.
* **Which unhandled error terminates the process.** An error that reaches
`FINISH_NORETURN` unhandled prints its stack trace and calls
`akerr_handler_unhandled_error`, which by default calls `akerr_exit()`. Each
thread's trace is whole — the buffer belongs to its context, and each line is
one call to `akerr_log_method` — but if two threads get there at the same
instant, both traces print and the exit status is whichever one won.
There is one lock, it is recursive, and it covers both the pool and the
registry. That means error construction is serialized across threads: raising an
error is the exceptional path, and correctness there is worth more than
throughput. A program that raises errors on its hot path will feel it.
`AKERR_THREAD_SAFE` in the generated header is `1` for a thread-safe build, so a
consumer can check what it linked against:
```c
#if AKERR_THREAD_SAFE
/* ... start worker threads ... */
#endif
```
## Building single threaded
The threading backend is chosen when libakerror is configured. `auto` (the
default) takes POSIX threads, and **fails the configure** if it cannot find
them rather than quietly producing a library that says it is thread safe and is
not. To mean it:
```sh
cmake -S . -B build -DAKERR_THREADS=none
```
That builds with no locking and no thread-local storage, stamps
`AKERR_THREAD_SAFE 0` into the header, and calling the library from more than
one thread is then undefined.
## Proving it
The thread tests (`tests/err_threads_*.c`) assert the properties above directly:
exclusive ownership of pool slots, exactly one winner for a contested range,
every registered name readable back. They run in the normal suite. The run that
proves the *absence* of a data race underneath them is ThreadSanitizer:
```sh
scripts/thread_test.sh
```
which configures `build/tsan` with `-DAKERR_SANITIZE=thread`, builds the library
and every test with it, and runs the suite. Under that build a sanitizer report
fails the test rather than being printed and passed over.
# Installation
@@ -89,14 +224,19 @@ The build process relies upon `scripts/generrno.sh` which performs the following
## Dependencies
This library depends upon `stdlib`. If you don't want to link against stdlib, you must modify the library code to include headers and link against a library that provides the following:
This library depends upon `stdlib`, and upon POSIX threads unless it is built
with `-DAKERR_THREADS=none` (see "Thread safety" above). If you don't want to link against stdlib, you must modify the library code to include headers and link against a library that provides the following:
- `memset` function
- `strncpy` function
- `strlen` function
- `strcmp` function
- `sprintf` function
- `exit` function
- `bool` type
- `NULL` type
- `INT_MAX` constant
- `PATH_MAX` constant
... then you can compile it thusly:
@@ -137,6 +277,14 @@ pkg_check_modules(akerror REQUIRED akerror)
target_link_libraries(YOUR_TARGET PRIVATE akerror::akerror)
```
Using this project as a submodule with cmake:
```cmake
add_subdirectory(deps/libakerror EXCLUDE_FROM_ALL)
target_link_libraries(YOUR_PROJECT PRIVATE akerror::akerror)
```
## (Optional) Configuring the logging function
@@ -200,6 +348,8 @@ ATTEMPT {
This will assign the return value of the function in question to the akerr_ErrorContext previously prepared in the current scope. If the function returns an akerr_ErrorContext that indicates any type of error, the `ATTEMPT` block is immediately exited, and the `CLEANUP` block begins.
(One caveat: because this exit is implemented with a C `break`, `CATCH` must not be used inside a loop within the `ATTEMPT` block — see the section "Important: do not use CATCH or FAIL_*_BREAK inside a loop" below.)
## Setting errors from functions or expressions returning integer
For functions that return integer, such as logical comparisons or most standard library functions, use the `FAIL_ZERO_BREAK` and `FAIL_NONZERO_BREAK` macros. These macros allow you to capture an integer return code from an expression or function and set an error code in the current context based off that return.
@@ -222,6 +372,71 @@ ATTEMPT {
When either of these two macros are used, the `ATTEMPT` block is immediately exited, and the `CLEANUP` block begins.
## Important: do not use CATCH or FAIL_*_BREAK inside a loop
`CATCH`, `FAIL_ZERO_BREAK`, `FAIL_NONZERO_BREAK`, and `FAIL_BREAK` leave the `ATTEMPT` block by executing a C `break` statement. In C, `break` only exits the *innermost* enclosing `for`, `while`, `do`, or `switch`. Therefore **these macros must not be used inside a loop (or a nested `switch`) that is itself inside an `ATTEMPT` block.** If you do, the `break` escapes only the loop — not the `ATTEMPT` — and the rest of the `ATTEMPT` body then runs with an error already pending:
```c
ATTEMPT {
for ( int i = 0; i < n; i++ ) {
CATCH(errctx, process(items[i])); // WRONG: break exits the for loop, not the ATTEMPT
}
// ... this code still executes, with errctx already in an error state ...
} CLEANUP {
} PROCESS(errctx) {
} FINISH(errctx, true);
```
Note that moving the loop into a helper function does **not** fix this on its own — if the helper still wraps the loop in an `ATTEMPT` and uses `CATCH`/`FAIL_*_BREAK` inside it, it has the exact same problem. The fix is to iterate with `return`-based macros, which are unaffected by loop nesting. Use one of the two patterns below.
**Pattern 1 — use `PASS` (or a `FAIL_*_RETURN` macro) inside the loop.** These exit the *enclosing function* with a `return` rather than a `break`, so loop nesting is irrelevant. Use this when the loop should stop and propagate on the first error:
```c
akerr_ErrorContext AKERR_NOIGNORE *process_all(Item *items, int n)
{
PREPARE_ERROR(errctx);
for ( int i = 0; i < n; i++ ) {
PASS(errctx, process(items[i])); // returns from process_all on the first error
}
SUCCEED_RETURN(errctx);
}
```
**Pattern 2 — move the loop into a helper and `CATCH` the single call.** When you need a `CLEANUP` block or want to `HANDLE` the error locally, put the loop in its own `akerr_ErrorContext *`-returning function (written per Pattern 1) and `CATCH` that one call. The `CATCH` is then not inside a loop, so its `break` scopes to the `ATTEMPT` correctly:
```c
ATTEMPT {
CATCH(errctx, process_all(items, n)); // a single CATCH, not looped
} CLEANUP {
// ... always runs ...
} PROCESS(errctx) {
} HANDLE(errctx, AKERR_VALUE) {
// ... handle a failure from any iteration ...
} FINISH(errctx, true);
```
# Passing errors
Sometimes you can't actually do anything about the errors that come out of a given method, but you want that error to be propagated back up the call chain, and to be properly reported. If this is your goal, you can avoid using a `ATTEMPT ... FINISH` block, and simply use the `PASS` macro.
```
PREPARE_ERROR(e);
PASS(e, some_method_that_may_fail());
SUCCEED_RETURN(e);
```
This does the same thing as this, but with less code:
```
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, some_method_that_may_fail());
} CLEANUP {
} PROCESS(e) {
} FINISH(e, true);
SUCCEED_RETURN(e);
```
# Handling errors
Inside of the `PROCESS { ... }` block, you must handle any errors that occurred during the `ATTEMPT { ... }` block. You do this with `HANDLE`, `HANDLE_GROUP`, and `HANDLE_DEFAULT`.
@@ -310,20 +525,24 @@ FAIL_NONZERO_RETURN(errctx, strcmp("not", "equal"), AKERR_VALUE, "Strings are no
# Uncaught errors
## Misbehaving methods
Any function which returns `akerr_ErrorContext *` and completes successfully MUST call `SUCCEED_RETURN(errctx)`. Failure to do this may result in an invalid `akerr_ErrorContext *` being returned, which will cause an `AKERR_BEHAVIOR` error to be triggered from your code.
## Ensuring that all error codes are captured
Any function which returns `akerr_ErrorContext *` should also be marked with `ERROR_NOIGNORE`.
Any function which returns `akerr_ErrorContext *` should also be marked with `AKERROR_NOIGNORE`.
```c
akerr_ErrorContext ERROR_NOIGNORE *f(...);
akerr_ErrorContext AKERROR_NOIGNORE *f(...);
```
This will cause a compile-time error if the return value of such a function is not used. "Used" here means assigned to a variable - it does not necessarily mean that the value is checked. However assuming that such functions are called inside of `ATTEMPT { ... }` blocks, it is safe to assume that such returns will be caught with `CATCH(...)`; therefore this error is a generally effective safeguard against careless coding where errors are not checked.
Beware that `ERROR_NOIGNORE` is not a failsafe - it implements the `warn_unused_result` mechanic. By design users may explicitly ignore an error code from a function marked with `warn_unused_result` by explicitly casting the return to `void`.
Beware that `AKERROR_NOIGNORE` is not a failsafe - it implements the `warn_unused_result` mechanic. By design users may explicitly ignore an error code from a function marked with `warn_unused_result` by explicitly casting the return to `void`.
```c
#define ERROR_NOIGNORE __attribute__((warn_unused_result))
#define AKERROR_NOIGNORE __attribute__((warn_unused_result))
```
## Stack Traces
@@ -354,3 +573,80 @@ From bottom to top, we have:
* Above that, the `FINISH()` macro in `func2()` which detected an unhandled error and passed it out of the function
* Above that, a reference to the line where the `FAIL()` macro set the error code and provided the message which is printed here
## Exit status
**Never call `exit()` with an akerr status. Call `akerr_exit()`.**
```c
void akerr_exit(int status);
```
This applies everywhere you are leaving the process on account of a status, not
just in an unhandled-error handler: a CLI's top-level `HANDLE` block, an
initialization routine that cannot continue, a `main()` that ends by reporting
the status it finished with. One function owns the mapping so that a given
status produces the same exit code no matter which of your exits it left by.
The mapping exists because a process exit status is one byte wide. `exit()`
accepts an `int` and the kernel keeps the low 8 bits of it — `_exit()`,
`_Exit()`, `quick_exit()` and the raw `exit_group` syscall all behave
identically, and even `waitid()`, whose `si_status` is a full `int`, sees the
truncated value because the truncation happened before the parent looked. There
is no wider `exit()` to reach for.
That leaves 0 through 255 as the only statuses an exit code can carry, and
consumer statuses start at `AKERR_FIRST_CONSUMER_STATUS` (256) — so *no* consumer
status can be an exit code:
```
status exit code
0 0 (success)
1 .. 255 the status
negative, or > 255 AKERR_EXIT_STATUS_UNREPRESENTABLE (125)
```
Passing the low byte instead would have made status 256 exit 0 and report
success to the shell, and status 300 exit 44 — an unrelated error's code. 125 is
the conventional "the tool itself failed" status; 126, 127 and 128+*n* already
belong to the shell.
Status 0 exits 0, because 0 is this library's success status. That is not a hole
in the rule that an unhandled error never exits 0: `PROCESS` opens with `case
0`, which marks a zero status handled, so a successful context cannot reach
`FINISH_NORETURN`'s call to the handler at all.
Above 255, the exit code tells you the process died of an error, not which one.
125 is inside the library's reserved band and so is also some host's `errno`, and
every status above 255 collapses onto it. **The stack trace is what identifies
the error** — it is printed before the handler runs, carrying the status at full
width along with its registered name.
### Replacing the handler
After the trace is printed, `FINISH_NORETURN` calls
`akerr_handler_unhandled_error`. The default implementation,
`akerr_default_handler_unhandled_error()`, hands `errctx->status` to
`akerr_exit()` (a NULL context exits 1). Replace it if you need something else
to happen first — a core dump, a crash reporter, a flush — and finish by calling
`akerr_exit()` so the exit code still means what it means everywhere else:
```c
static void mylib_handler(akerr_ErrorContext *e)
{
mylib_flush_telemetry();
if ( e == NULL ) {
akerr_exit(AKERR_API);
}
akerr_exit(e->status);
}
akerr_handler_unhandled_error = &mylib_handler;
```
`akerr_exit()` is declared `AKERR_NORETURN`, so the compiler knows a handler
ending in one of those calls is complete rather than falling off the end.
Set the handler once, before you start any threads. `tests/err_custom_handler.c`
installs one that does not exit at all, which is how the test suite asserts on
unhandled errors without dying.

122
TODO.md Normal file
View File

@@ -0,0 +1,122 @@
# TODO
Working notes for `libakerror`. Outstanding items only.
## 1. Only ThreadSanitizer is wired into CI, not ASan/UBSan
`AKERR_SANITIZE` builds the library and the tests with any sanitizer list, and
CI runs `-DAKERR_SANITIZE=thread` through `scripts/thread_test.sh`. Nothing runs
`address,undefined` yet, and that is the one that covers the original
motivation: mutation testing caught an out-of-bounds probe in the status-name
hash table that the suite could not, because the failure mode was a write into
adjacent BSS, which does not crash. Sharpening one test closed that instance;
ASan would catch the whole class regardless of how sharp the assertions are.
The machinery is in place — this is one more job in
`.gitea/workflows/ci.yaml` running
`cmake -S . -B build/asan -DAKERR_SANITIZE=address,undefined`. Left separate
because ASan and TSan cannot be combined in one build.
## 2. `HANDLE`-level status aliasing is still undetectable
Two components can compile the same integer into a `case` label without ever
reserving a range or registering a name, and nothing sees it. Ownership
enforcement covers *naming*, which is the part the library mediates; the `case`
label never reaches it.
Closing this needs the `if`/`else if` handler ladder — rewriting
`PROCESS`/`HANDLE`/`HANDLE_GROUP`/`HANDLE_DEFAULT`/`FINISH` so status matching
is not restricted to integer constant expressions. That would also allow
matching on ranges or predicates, and would let a handler resolve a code through
its owner. It touches the most load-bearing code in the library and every
consumer's error handling at once, so it wants its own change.
Note it would *not* by itself fix the "don't use `CATCH` or `FAIL_*_BREAK`
inside a loop" hazard: that comes from exiting via `break`, not from `switch`.
## 3. No registry introspection
There is no way to ask who owns a status, or to enumerate reservations. The
"coordinate ranges at the dependency-stack level" advice in UPGRADING.md
therefore has no tooling behind it.
A read-only accessor plus a dump through `akerr_log_method` would let a startup
self-check or a CI job print the whole map for a linked stack. Cheap, additive,
and the natural next step for multi-component adoption.
## 4. No way to release a reservation
A plugin host that `dlopen`s many distinct plugins over a process lifetime
accumulates ranges until the table fills. Reloading the *same* plugin is fine —
an identical repeat by the same owner is idempotent.
## 5. Renaming a status is not safe against a concurrent lookup
`akerr_name_for_status(status, NULL)` returns a pointer into the registry rather
than a copy, which is what makes it usable from inside `FAIL` — it needs no
buffer and no error context of its own. Registering a *second* name for a status
that already has one (`tests/err_name_ownership.c` covers that it is allowed)
overwrites that buffer in place, so a thread reading the name at that moment can
see a torn string. Every other registry operation is serialized; this one cannot
be, because the reader is outside the lock by the time it reads the characters.
Documented in README.md and UPGRADING.md as "register names during
initialization". Closing it properly means making a registered name immutable —
either refusing a rename outright (a behavior change, and
`tests/err_name_ownership.c` asserts the current contract), or copying names
into a bump-allocated arena and publishing the pointer with a release store, so
a rename allocates new storage instead of rewriting live storage. The arena is
the better answer; it costs a second capacity limit and its exhaustion path.
## 6. Deprecate the two-argument name-registration path
`akerr_name_for_status(status, name)` cannot identify its caller, so it can only
check that *some* reservation covers the status, not that the caller owns it. It
exists for migration. Once consumers have moved to
`akerr_register_status_name()`, make the set path a no-op or remove it and leave
`akerr_name_for_status()` as pure lookup.
## 7. `akerr_init()`'s own reservation failure is untested
`tests/err_library_status_fatal.c` covers the terminal path in
`__akerr_name_library_status()` by naming a status the library does not own. The
band reservation in `akerr_init()` has no such handle: it can only fail in a
build whose tables are too small for the library's own entries, and both sizes
are `PRIVATE` to the library target, so a test executable cannot set them.
Closing it means a second library target built with tiny tables plus a
`WILL_FAIL` test linked against it. Nothing in the CMake does that yet: the
sanitizer and coverage options vary the *flags* of the one library target, not
its compile definitions.
Related: branch coverage on `src/error.c` now sits just above its 50% gate.
Every `FAIL_*` site carries about six branch outcomes of error-construction
machinery (`ENSURE_ERROR_READY`, `AKERR_STACKTRACE_APPEND`) that only run when
that specific failure fires, and every `PASS` site around a call that cannot
fail carries about twenty-five. Validating more inputs therefore lowers the
ratio by construction. Before adding defensive checks, expect to add a test that
drives them, as `tests/err_copy_string.c` does.
## 8. Mutation testing judges concurrency mutants without a sanitizer
`scripts/mutation_test.py` configures each mutant build with the default CMake
options, so a mutant that only breaks under concurrency is judged by a suite
running without ThreadSanitizer. Deleting the pool's `akerr_mutex_lock()` call
survives the run even though it is a real race: rebuilt and run directly, that
mutant fails `tests/err_threads_pool.c` in 4 of 10 runs, and fails under
`scripts/thread_test.sh` in 5 of 5. So 81.2% is a floor for that category, not a
verdict.
Closing it means a `--cmake-arg` passthrough on the harness so the mutant build
can be configured with `-DAKERR_SANITIZE=thread`. The whole run then costs a
TSan-instrumented suite per mutant (roughly 6s instead of 0.4s), so it belongs
behind a flag rather than in the default target or in CI.
## Unrelated pre-existing issues
- The `AKERR_USE_STDLIB=OFF` build does not compile at all: `bool`, `PATH_MAX`
and `NULL` are used unconditionally but only included under the stdlib branch.
The README's dependency list states what a replacement must provide, but the
header still needs its includes untangled for that configuration to work.
- `CMakeLists.txt` sets `main_lib_dest` from `MY_LIBRARY_VERSION`, which is never
defined and never read. Dead line.

347
UPGRADING.md Normal file
View File

@@ -0,0 +1,347 @@
# Bug fix: unhandled-error exit status (2.0.1)
An unhandled error could kill the process and still report success.
`akerr_default_handler_unhandled_error()` ended in `exit(errctx->status)`, and a
process exit status is one byte wide — the kernel keeps the low 8 bits of the
argument and discards the rest. Consumer statuses start at
`AKERR_FIRST_CONSUMER_STATUS` (256), so **the first status any consumer can
reserve exited 0**, and a shell or supervisor watching `$?` saw a clean run.
Status 300 exited 44, which is some unrelated error's code. No status a consumer
owns could ever come out of `$?` intact, and there is no wider `exit()` to reach
for: `_exit()`, `_Exit()`, `quick_exit()` and the raw `exit_group` syscall all
truncate identically, and even `waitid()`, whose `si_status` is a full `int`,
reports the truncated value.
The mapping now lives in one place, `akerr_exit()`, which the default handler
calls:
```
status exit code
0 0 (success)
1 .. 255 the status
negative, or > 255 AKERR_EXIT_STATUS_UNREPRESENTABLE (125)
```
Statuses 0 through 255 are unchanged, which covers every `errno` and every
`AKERR_*` code. Only the values that were already being delivered wrong behave
differently, and they now exit 125 instead of a truncated byte.
**What you should change.** Call `akerr_exit()` instead of `exit()` anywhere you
leave the process on an akerr status — your own unhandled-error handler, a
top-level `HANDLE` block, an init routine that cannot continue — so one mapping
covers every exit. It is declared `AKERR_NORETURN`. If you were reading a
consumer status out of `$?`, you were never getting it: read the stack trace,
which carries the status at full width along with its registered name, or
install a handler that maps your own statuses into a byte. See "Exit status" in
[README.md](README.md).
No ABI break. The soname stays `libakerror.so.2` and nothing you already call
changed shape. `akerr_exit()` is a new exported symbol, so a consumer that
starts calling it needs 2.0.1 or later at link time.
# Upgrade notice: thread safety (2.0.0)
2.0.0 makes the library thread safe. Every entry point may be called from any
thread, `akerr_init()` runs exactly once however many threads race into it, and
the error pool and the status registry are serialized.
This is an ABI break. Rebuild libakerror and everything that includes its
header; the soname moved to `libakerror.so.2` so the two cannot be mixed by
accident.
What moved at the ABI:
* `__akerr_last_ignored` is thread-local storage. An ignored error is a fact
about the thread that ignored it, and one shared slot had two threads
overwriting each other's. The `IGNORE` macro expands at *your* call site, so
your objects reference the symbol under whichever storage model your header
said — which is why this cannot be mixed.
* `akerr_next_error()` returns a context that **already holds one reference**.
Finding a free slot and claiming it has to be one operation under the pool
lock, or two threads scanning at once are handed the same slot.
`ENSURE_ERROR_READY` therefore no longer increments the count. Code compiled
against a 1.x header and linked against 2.x would count every reference twice
and never return a slot to the pool.
What changed in the API: nothing you call, unless you call `akerr_next_error()`
yourself. If you do, you now own a reference and must release it — the same
thing you were doing already if you were using the context for anything.
New build options, both on libakerror itself:
| Option | Default | Meaning |
| --- | --- | --- |
| `AKERR_THREADS` | `auto` | Threading backend: `auto`, `pthread`, or `none` |
| `AKERR_SANITIZE` | (empty) | Sanitizers for the library and its tests, e.g. `thread` |
`auto` takes POSIX threads and **fails the configure** when it cannot find them,
rather than quietly building a library that reports itself thread safe and is
not. `-DAKERR_THREADS=none` is how you say you meant it: no locking, no
thread-local storage, undefined if you then use more than one thread.
The generated header records which one you built, as `AKERR_THREAD_SAFE` (`1` or
`0`), so a consumer can test what it linked against and cannot disagree with the
library about it.
## What thread safety here does and does not mean
Safe from any thread, with no coordination on your part:
* Raising, catching, handling, passing, ignoring and releasing errors.
* `akerr_reserve_status_range()` and `akerr_register_status_name()`. Two threads
reserving overlapping ranges cannot both succeed: one gets `NULL`, the other
gets `AKERR_STATUS_RANGE_OVERLAP` naming the winner.
* `akerr_name_for_status(status, NULL)` lookups, concurrently with each other
and with registrations of *other* statuses.
* `akerr_init()`, from any number of threads at once.
Still yours to coordinate:
* **One error context is owned by one thread.** The library hands it to the
thread that raised it. Handing it to another thread is your synchronization.
* **`akerr_log_method` and `akerr_handler_unhandled_error`** are read on every
error and written by nobody but you. Set them during startup, before spawning.
* **Renaming a status while another thread looks it up.**
`akerr_name_for_status()` returns a pointer into the registry — stable for the
life of the process, which is what makes it usable from a stack trace — and
registering a second name for the same status overwrites that buffer in place.
Register names during initialization. This is the one registry operation the
lock cannot make safe, because the reader is outside the lock by the time it
reads the string.
* **Which unhandled error wins.** Two threads reaching `FINISH_NORETURN` with
unhandled errors at the same instant both print a complete stack trace (the
buffer belongs to the context, and each line is a single `akerr_log_method`
call) and both call `akerr_handler_unhandled_error`. The process exits with
whichever status got there first.
## Cost
One recursive lock covers both the pool and the registry, so error
*construction* is serialized process-wide. Errors are the exceptional path and
correctness there is worth more than throughput, but a program that raises
errors in a hot loop will feel it.
The per-thread last-ditch context is a whole `akerr_ErrorContext` (tens of
kilobytes) in thread-local storage, allocated per thread on first use of the
library from that thread.
## Proving it
`tests/err_threads_init.c`, `tests/err_threads_pool.c` and
`tests/err_threads_registry.c` assert the properties directly and run in the
normal suite. The run that proves there is no data race underneath them is
ThreadSanitizer:
```sh
scripts/thread_test.sh
```
# Upgrade notice: custom status codes (1.0.0)
Version 1.0.0 replaces the consumer-sized status-name array with a private
registry, and makes status-code ownership explicit and enforced. This is a
source and ABI break. The library now carries a version and an soname
(`libakerror.so.1`), so a stale installed library can no longer be silently
paired with a newer header — but anything built against a pre-1.0.0 header must
be rebuilt.
What was removed:
* `AKERR_MAX_ERR_VALUE` — the registry is sparse and accepts any `int`, so
consumers no longer size it. Delete every compile definition and source
reference. A stale `-DAKERR_MAX_ERR_VALUE=...` is now harmless but useless.
* `__AKERR_ERROR_NAMES` — the name table is private to the library. Code that
touched this data symbol was using an undocumented interface; use
`akerr_name_for_status()`.
* `AKERR_STATUS_RANGE_OK` and `AKERR_STATUS_NAME_OK` — the registry functions no
longer return an `int`. Success is a `NULL` `akerr_ErrorContext *`, like every
other function in the library. See "The registry raises errors" below.
To migrate:
1. Rebuild libakerror and every dependent library against the new header.
2. Move custom status codes out of the reserved `0``255` band. Use fixed
integer constants beginning at `AKERR_FIRST_CONSUMER_STATUS` (256) rather
than offsets from `AKERR_LAST_ERRNO_VALUE`, so a libc that grows an errno
cannot move your codes.
3. Assign a distinct range to every library that may coexist in one process, and
coordinate those ranges at the application or dependency-stack level.
4. Reserve the complete range with `akerr_reserve_status_range()` during
initialization, and treat the error it returns like any other error: `CATCH`
it, `PASS` it, or let it propagate out of your init function.
5. Register names with `akerr_register_status_name()`, passing the same owner
string you reserved with. **You can no longer name a status you have not
reserved** — see "Ownership is enforced" below.
For example:
```c
enum {
MYLIB_ERR_BASE = AKERR_FIRST_CONSUMER_STATUS, /* 256 */
MYLIB_ERR_PARSE = MYLIB_ERR_BASE,
MYLIB_ERR_STORAGE,
MYLIB_ERR_LIMIT
};
#define MYLIB_OWNER "mylib"
akerr_ErrorContext AKERR_NOIGNORE *mylib_init(void)
{
PREPARE_ERROR(errctx);
/* Any collision propagates out of mylib_init() to the caller. */
PASS(errctx, akerr_reserve_status_range(MYLIB_ERR_BASE,
MYLIB_ERR_LIMIT - MYLIB_ERR_BASE,
MYLIB_OWNER));
PASS(errctx, akerr_register_status_name(MYLIB_OWNER, MYLIB_ERR_PARSE,
"Parse Error"));
PASS(errctx, akerr_register_status_name(MYLIB_OWNER, MYLIB_ERR_STORAGE,
"Storage Error"));
SUCCEED_RETURN(errctx);
}
```
If your `init` cannot return an `akerr_ErrorContext *` — a C API with an `int`
return, say — catch the error and convert it. Close with `FINISH_NORETURN`
rather than `FINISH`: nothing can propagate out of a function that does not
return an error context, and `FINISH(errctx, false)` in an `int`-returning
function compiles the (dead) propagation branch anyway, which warns.
```c
int mylib_init(void)
{
PREPARE_ERROR(errctx);
int rc = MYLIB_INIT_OK;
ATTEMPT {
CATCH(errctx, akerr_reserve_status_range(MYLIB_ERR_BASE,
MYLIB_ERR_LIMIT - MYLIB_ERR_BASE,
MYLIB_OWNER));
} CLEANUP {
} PROCESS(errctx) {
} HANDLE_DEFAULT(errctx) {
LOG_ERROR_WITH_MESSAGE(errctx, "mylib could not claim its status range");
rc = MYLIB_INIT_FAILED; /* another component owns part of the range */
} FINISH_NORETURN(errctx);
return rc;
}
```
You do not need to call `akerr_init()` first. Every registry entry point calls
it for you, so a library that reserves its range before anything else in the
process has touched libakerror keeps that reservation. (Before 1.0.0 this was
silently destructive: `akerr_init()` cleared the tables, so a reservation made
too early was discarded and the next component to claim the same range was told
it was free.)
## Ownership is enforced
Reserving a range is no longer advisory bookkeeping. A status may only be named
from inside a reservation:
* `akerr_register_status_name(owner, status, name)` requires that `status` fall
in a range reserved by `owner`. Naming another component's status fails with
`AKERR_STATUS_NAME_FOREIGN`, and naming a status nobody reserved fails with
`AKERR_STATUS_NAME_UNRESERVED`.
* The two-argument `akerr_name_for_status(status, name)` set path still works,
but it cannot identify its caller, so it can only require that *some*
reservation covers the status. Prefer the owned form: it is the one that
catches a component writing into a range that is not its own.
Every refusal is reported, because a name that fails to register degrades that
code to `"Unknown Error"` in every later stack trace.
This detects *name* collisions. It cannot detect two components compiling the
same integer into a `HANDLE` label without ever registering a name, so every
co-resident library should still reserve its range.
## The registry raises errors
`akerr_reserve_status_range()` and `akerr_register_status_name()` return
`akerr_ErrorContext *`, exactly like the rest of the library. They are marked
`AKERR_NOIGNORE`, so discarding the result is a compile-time warning, and an
error you catch but do not handle propagates out of your init function instead
of leaving you with a range you do not actually own.
```c
ATTEMPT {
CATCH(errctx, akerr_reserve_status_range(256, 16, MYLIB_OWNER));
} CLEANUP {
} PROCESS(errctx) {
} HANDLE(errctx, AKERR_STATUS_RANGE_OVERLAP) {
/* Another component owns part of it -- the message names which. */
} FINISH(errctx, true);
```
`akerr_reserve_status_range()` raises:
| Status | Meaning |
| --- | --- |
| — (returns `NULL`) | Range reserved, or an identical range was already reserved by the same owner |
| `AKERR_STATUS_RANGE_OVERLAP` | Part of the range is owned by someone else (the message names the owner) |
| `AKERR_STATUS_RANGE_FULL` | No reservation slots remain |
| `AKERR_STATUS_RANGE_INVALID` | `count < 1`, NULL/empty/over-long owner, or the range overflows `int` |
`akerr_register_status_name()` raises:
| Status | Meaning |
| --- | --- |
| — (returns `NULL`) | Name registered |
| `AKERR_STATUS_NAME_UNRESERVED` | No reservation contains this status |
| `AKERR_STATUS_NAME_FOREIGN` | The status is in a range owned by someone else |
| `AKERR_STATUS_NAME_FULL` | The name registry is full |
| `AKERR_STATUS_NAME_INVALID` | NULL/empty/over-long owner, or a NULL name |
These are ordinary status codes inside the library's reserved band, so they can
be matched with `HANDLE`, grouped with `HANDLE_GROUP`, and they print with a
name and a stack trace when they go unhandled.
`akerr_name_for_status(status, name)` is the one exception: it returns a name
rather than an error context, so it cannot raise. A refused registration through
that path is reported through `akerr_log_method` and reads back as
`"Unknown Error"`.
Repeating an identical reservation for the same owner is a no-op. A *subset* or
*superset* of your own range is not — it raises `AKERR_STATUS_RANGE_OVERLAP`.
Reserve the whole range in one call.
The library holds itself to the same rule at startup: if `akerr_init()` cannot
reserve its own `0``255` band or name one of its own codes, it logs the failure
and terminates the program. That can only happen in a misconfigured build (a
name table too small for the library's own entries), and the alternative is a
process whose stack traces silently read `"Unknown Error"` for built-in codes.
## Capacity
Both tables are fixed size, allocated in BSS, and never grow:
| Limit | Default | Build-time override |
| --- | --- | --- |
| Status names | 3072 usable (4096 slots, 75% load) | `-DAKERR_STATUS_NAME_SLOTS=<power of two>` |
| Reserved ranges | 64 | `-DAKERR_MAX_RESERVED_STATUS_RANGES=<n>` |
`akerr_init()` consumes one name slot per host errno value plus one per
`AKERR_*` code — around 150 on glibc, leaving roughly 2900 for all consumers in
the process to share. That figure varies with the host libc, so treat it as
approximate rather than a budget to fill.
Both overrides are CMake cache variables and apply `PRIVATE` to the library
target. The tables live entirely inside `src/error.c`, so raising them changes
nothing a consumer can observe — unlike the `AKERR_MAX_ERR_VALUE` they replaced,
where a mismatch between the library and its consumers corrupted memory. Set
them when configuring libakerror itself:
```sh
cmake -S . -B build -DAKERR_STATUS_NAME_SLOTS=16384
```
Exhausting either table raises an error to the caller; it is never silent.
## Thread safety
As of 2.0.0 the registry is serialized: reservations, registrations and lookups
are all safe to call concurrently. See "What thread safety here does and does
not mean" at the top of this document for the two things that are still yours to
coordinate.

View File

@@ -0,0 +1 @@
@AKERR_THREAD_SAFE@

View File

@@ -6,38 +6,132 @@
#include <stdbool.h>
#include <string.h>
#include <stdio.h>
#include <limits.h>
#endif
#define AKERR_MAX_ERROR_CONTEXT_STRING_LENGTH 1024
/*
* Threading.
*
* scripts/generrno.sh stamps this value in at build time from the AKERR_THREADS
* build option, the same way it stamps AKERR_LAST_ERRNO_VALUE. It is generated
* rather than defined by the consumer on purpose: whether the library
* serializes its global state and whether __akerr_last_ignored is a
* thread-local are the same decision, and a consumer that disagreed with the
* library about it would link against a differently shaped symbol.
*
* 1 The error pool and the status registry are mutex protected, and the
* per-thread state below is thread local. Every entry point may be called
* from any thread. See "Thread safety" in README.md for what that does and
* does not cover.
* 0 The library was built -DAKERR_THREADS=none for a single-threaded
* process: no locking, no thread-local storage, and calling it from more
* than one thread is undefined.
*
* Consumers can test it: #if AKERR_THREAD_SAFE.
*/
#define AKERR_THREAD_SAFE AKERR_THREAD_SAFE_SED
#if AKERR_THREAD_SAFE == 1
#if defined(__GNUC__) || defined(__clang__)
#define AKERR_THREAD_LOCAL __thread
#elif defined(__STDC_VERSION__) && __STDC_VERSION__ >= 201112L
#define AKERR_THREAD_LOCAL _Thread_local
#elif defined(_MSC_VER)
#define AKERR_THREAD_LOCAL __declspec(thread)
#else
#error "libakerror was built thread safe, but this compiler has no thread-local storage specifier that akerror.h knows about. Rebuild libakerror with -DAKERR_THREADS=none, or add the spelling here."
#endif
#else
#define AKERR_THREAD_LOCAL
#endif
// FIXME: This is huge now. It used to be 1000 bytes, then I wanted to report errors
// related to filesystem paths, which made it grow beyond PATH_MAX, then I started
// reporting messages including 2 file paths (PATH_MAX * 2), so now to make the compiler warnings
// shut up, it's enormous (PATH_MAX*3).
#define AKERR_MAX_ERROR_CONTEXT_STRING_LENGTH 12384
#define AKERR_MAX_ERROR_NAME_LENGTH 64
#define AKERR_MAX_ERROR_FNAME_LENGTH 256
#define AKERR_MAX_ERROR_FNAME_LENGTH PATH_MAX
#define AKERR_MAX_ERROR_FUNCTION_LENGTH 128
#define AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH 2048
#define AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH (AKERR_MAX_ERROR_CONTEXT_STRING_LENGTH + AKERR_MAX_ERROR_NAME_LENGTH + AKERR_MAX_ERROR_FNAME_LENGTH + AKERR_MAX_ERROR_FUNCTION_LENGTH + 16)
#define AKERR_LAST_ERRNO_VALUE AKERR_LAST_ERRNO_VALUE_SED
#define AKERR_NULLPOINTER (AKERR_LAST_ERRNO_VALUE + 1)
#define AKERR_OUTOFBOUNDS (AKERR_LAST_ERRNO_VALUE + 2)
#define AKERR_API (AKERR_LAST_ERRNO_VALUE + 3)
#define AKERR_ATTRIBUTE (AKERR_LAST_ERRNO_VALUE + 4)
#define AKERR_TYPE (AKERR_LAST_ERRNO_VALUE + 5)
#define AKERR_KEY (AKERR_LAST_ERRNO_VALUE + 6)
#define AKERR_HEAP (AKERR_LAST_ERRNO_VALUE + 7)
#define AKERR_INDEX (AKERR_LAST_ERRNO_VALUE + 8)
#define AKERR_FORMAT (AKERR_LAST_ERRNO_VALUE + 9)
#define AKERR_IO (AKERR_LAST_ERRNO_VALUE + 10)
#define AKERR_REGISTRY (AKERR_LAST_ERRNO_VALUE + 11)
#define AKERR_VALUE (AKERR_LAST_ERRNO_VALUE + 12)
#define AKERR_BEHAVIOR (AKERR_LAST_ERRNO_VALUE + 13)
#define AKERR_RELATIONSHIP (AKERR_LAST_ERRNO_VALUE + 14)
#define AKERR_NULLPOINTER (AKERR_LAST_ERRNO_VALUE + 1) /** A pointer had a NULL value where such was not permissible */
#define AKERR_OUTOFBOUNDS (AKERR_LAST_ERRNO_VALUE + 2) /** Attempt to access a datastructure outside of bounds */
#define AKERR_API (AKERR_LAST_ERRNO_VALUE + 3) /** An otherwise unspecified API contract has been violated */
#define AKERR_ATTRIBUTE (AKERR_LAST_ERRNO_VALUE + 4) /** Relates to accessing of attributes on objects */
#define AKERR_TYPE (AKERR_LAST_ERRNO_VALUE + 5) /** An object had the incorrect type */
#define AKERR_KEY (AKERR_LAST_ERRNO_VALUE + 6) /** A key was either invalid for or not present in a map */
#define AKERR_INDEX (AKERR_LAST_ERRNO_VALUE + 8) /** An error occurred when attempting to index an indexable datastructure (other than out of bounds) */
#define AKERR_FORMAT (AKERR_LAST_ERRNO_VALUE + 9) /** An error occurred in the formatting of an object (usually a string) */
#define AKERR_IO (AKERR_LAST_ERRNO_VALUE + 10) /** An unspecified IO error occurred. */
#define AKERR_VALUE (AKERR_LAST_ERRNO_VALUE + 11) /** A provided value was invalid */
#define AKERR_RELATIONSHIP (AKERR_LAST_ERRNO_VALUE + 12) /** An error occurred in establishing, maintaining or severing a relationship between two objects */
#define AKERR_EOF (AKERR_LAST_ERRNO_VALUE + 13) /** The end of a stream or file has been encountered */
#define AKERR_CIRCULAR_REFERENCE (AKERR_LAST_ERRNO_VALUE + 14) /** Indicates that a circular reference has been found in a linked list */
#define AKERR_ITERATOR_BREAK (AKERR_LAST_ERRNO_VALUE + 15) /** Used to prematurely end an iteration cycle (such as when searching a graph and the desired node has been found) */
#define AKERR_NOT_IMPLEMENTED (AKERR_LAST_ERRNO_VALUE + 16) /** A method was called that is defined but not currently implemented */
#define AKERR_BADEXC (AKERR_LAST_ERRNO_VALUE + 17) /** The libakerr library was given an akerr_ErrorContext to parse that did not come from AKERR_ARRAY_ERROR (likely an uninitialized pointer) */
#ifndef AKERR_MAX_ERR_VALUE
#define AKERR_MAX_ERR_VALUE (AKERR_LAST_ERRNO_VALUE + 14)
#elif AKERR_MAX_ERR_VALUE < 256
#error user-defined AKERR_MAX_ERR_VALUE must be >= 256
#endif
/*
* Registry failures. These are ordinary status codes, not a private return
* enumeration: akerr_reserve_status_range() and akerr_register_status_name()
* return akerr_ErrorContext * like everything else in this library, so a refused
* reservation can be CATCH-ed, HANDLE-d, PASS-ed, or left unhandled to produce a
* stack trace and stop the program. Success returns NULL.
*/
#define AKERR_STATUS_RANGE_OVERLAP (AKERR_LAST_ERRNO_VALUE + 18) /** Some part of the range is already owned by someone else */
#define AKERR_STATUS_RANGE_FULL (AKERR_LAST_ERRNO_VALUE + 19) /** No reservation slots remain (see AKERR_MAX_RESERVED_STATUS_RANGES) */
#define AKERR_STATUS_RANGE_INVALID (AKERR_LAST_ERRNO_VALUE + 20) /** Bad count, bad owner string, or the range overflows int */
#define AKERR_STATUS_NAME_UNRESERVED (AKERR_LAST_ERRNO_VALUE + 21) /** No owner has reserved a range containing this status */
#define AKERR_STATUS_NAME_FOREIGN (AKERR_LAST_ERRNO_VALUE + 22) /** The status lies in a range reserved by a different owner */
#define AKERR_STATUS_NAME_FULL (AKERR_LAST_ERRNO_VALUE + 23) /** The name registry is full (raise AKERR_STATUS_NAME_SLOTS) */
#define AKERR_STATUS_NAME_INVALID (AKERR_LAST_ERRNO_VALUE + 24) /** NULL/empty/over-long owner, or a NULL name */
extern char __AKERR_ERROR_NAMES[AKERR_MAX_ERR_VALUE+1][AKERR_MAX_ERROR_NAME_LENGTH];
/* The last status the library defines for itself. Everything from
* AKERR_LAST_ERRNO_VALUE + 1 through here must have a registered name. */
#define AKERR_LAST_LIBRARY_STATUS AKERR_STATUS_NAME_INVALID
/*
* Status values 0 through 255 are reserved by libakerror at akerr_init() time:
* the host's errno values plus the AKERR_* codes above. Consumers allocate from
* 256 upwards. Reserving any part of this band fails with
* AKERR_STATUS_RANGE_OVERLAP naming AKERR_LIBRARY_OWNER.
*/
#define AKERR_LIBRARY_OWNER "libakerror"
#define AKERR_RESERVED_STATUS_COUNT 256
#define AKERR_FIRST_CONSUMER_STATUS AKERR_RESERVED_STATUS_COUNT
/*
* The library reserves status values 0 through 255 for itself (see akerr_init),
* which must contain every AKERR_* code above. AKERR_LAST_ERRNO_VALUE is
* derived from the host's errno list at build time, so on a platform with an
* unusually large errno space these codes could escape the band and collide
* with consumer codes allocated at 256. Fail the build instead.
*/
typedef char akerr_assert_codes_within_reserved_band[(AKERR_LAST_LIBRARY_STATUS < 256) ? 1 : -1];
/*
* A process exit status is one byte wide. exit() takes an int, but the kernel
* keeps only the low 8 bits of it and throws the rest away, so 0 through 255
* are the only statuses that can also be exit codes -- and every consumer status
* begins at AKERR_FIRST_CONSUMER_STATUS (256). No wider variant exists to reach
* for: _exit(), _Exit(), quick_exit() and the raw exit_group syscall all
* truncate identically, and even waitid()'s int-wide si_status reports the
* truncated value, because the truncation happened before the parent looked.
*
* akerr_exit() therefore substitutes AKERR_EXIT_STATUS_UNREPRESENTABLE for any
* status it cannot deliver intact, rather than passing the low byte -- status
* 256 would exit 0 and report success. 125 is the conventional "the tool itself
* failed" code (126, 127 and 128+n belong to the shell). It is inside the
* library's reserved band, so it is also some host's errno: the exit code says
* only that the process died of an error, and the stack trace carries the real
* status.
*/
#define AKERR_EXIT_STATUS_MAX 255
#define AKERR_EXIT_STATUS_UNREPRESENTABLE 125
#define AKERR_MAX_ARRAY_ERROR 128
@@ -58,25 +152,107 @@ typedef struct
} akerr_ErrorContext;
#define AKERR_NOIGNORE __attribute__((warn_unused_result))
/* akerr_exit() does not come back, and the compiler should know it: a handler
* whose last statement is a call to it is complete, not falling off the end. */
#define AKERR_NORETURN __attribute__((noreturn))
typedef void (*akerr_ErrorUnhandledErrorHandler)(akerr_ErrorContext *errctx);
typedef void (*akerr_ErrorLogFunction)(const char *f, ...);
extern akerr_ErrorContext AKERR_ARRAY_ERROR[AKERR_MAX_ARRAY_ERROR];
/*
* Set these before starting threads. They are read on every error and written
* by nothing but your own code, so changing one while other threads are raising
* errors is a data race the library cannot mediate.
*/
extern akerr_ErrorUnhandledErrorHandler akerr_handler_unhandled_error;
extern akerr_ErrorLogFunction akerr_log_method;
extern akerr_ErrorContext *__akerr_last_ignored;
extern akerr_ErrorContext *__akerr_last_prepared;
extern int __akerr_last_attempt_id;
/*
* The error IGNORE() last swallowed, per thread: an ignored error is a fact
* about the thread that ignored it, and one shared slot would have two threads
* overwriting each other's. Thread local only when AKERR_THREAD_SAFE is 1.
*/
extern AKERR_THREAD_LOCAL akerr_ErrorContext *__akerr_last_ignored;
akerr_ErrorContext AKERR_NOIGNORE *akerr_release_error(akerr_ErrorContext *ptr);
/*
* Check a context out of the pool. The returned context already carries one
* reference: finding a free slot and claiming it is a single operation under
* the pool lock, because two threads scanning at once would otherwise be handed
* the same slot. Release it with akerr_release_error() (or let RELEASE_ERROR,
* SUCCEED_RETURN or FINISH do it for you). Returns NULL when every slot is
* checked out.
*/
akerr_ErrorContext AKERR_NOIGNORE *akerr_next_error();
/*
* Look up (name == NULL) or register (name != NULL) the display name for a
* status. Registration succeeds only if some owner has reserved a range
* containing `status`; prefer akerr_register_status_name(), which also checks
* that the range belongs to you and raises an error saying why a registration
* was refused. This entry point cannot return an error context, so a refusal
* here is reported through akerr_log_method and reads back as "Unknown Error".
* Never returns NULL -- an unregistered status reads back as "Unknown Error".
*/
char *akerr_name_for_status(int status, char *name);
/*
* Register a display name for a status you own. `owner` must match the owner
* string passed to akerr_reserve_status_range() for the range containing
* `status`. Returns NULL on success, or an error context whose status is one of
* the AKERR_STATUS_NAME_* codes above -- CATCH it, HANDLE it, or let it
* propagate.
*/
akerr_ErrorContext AKERR_NOIGNORE *akerr_register_status_name(const char *owner, int status, const char *name);
/*
* Claim `count` status values starting at `first_status` for `owner`. Repeating
* an identical reservation for the same owner is a no-op; any other collision is
* refused. Returns NULL on success, or an error context whose status is one of
* the AKERR_STATUS_RANGE_* codes above. Treat any error as an initialization
* failure: either handle it or let it propagate out of your init function.
*/
akerr_ErrorContext AKERR_NOIGNORE *akerr_reserve_status_range(int first_status, int count, const char *owner);
void akerr_init();
/*
* Terminate the process, reporting `status`. Use this instead of exit()
* anywhere you are leaving on account of an akerr status -- an unhandled-error
* handler of your own, a CLI's top-level HANDLE block, an init routine that
* cannot continue -- so that every exit out of the library's status space maps
* the same way.
*
* Exits with `status` when 0 <= status <= AKERR_EXIT_STATUS_MAX, and with
* AKERR_EXIT_STATUS_UNREPRESENTABLE otherwise (see above). Status 0 exits 0:
* zero is this library's success status, and passing it here says the program
* finished, not that it failed with a code that got lost.
*/
void AKERR_NORETURN akerr_exit(int status);
/*
* The default akerr_handler_unhandled_error: logs nothing further -- the stack
* trace has already been printed by the time it runs -- and hands `ptr->status`
* to akerr_exit(), or exits 1 when `ptr` is NULL. Replace it if you need a
* different mapping, and call akerr_exit() from your replacement.
*/
void akerr_default_handler_unhandled_error(akerr_ErrorContext *ptr);
void akerr_default_logger(const char *f, ...);
int akerr_valid_error_address(akerr_ErrorContext *ptr);
/* defined in src/errno.c which is built dynamically at build time from system errno definitions */
void akerr_init_errno(void);
/*
* Internal. Names a status in the library's own reserved band on behalf of
* akerr_init() and the generated errno table, which have no caller to raise
* into: a failure here is logged and terminates the program. Not part of the
* consumer API -- use akerr_register_status_name(), which raises instead.
*/
void __akerr_name_library_status(int status, const char *name);
/*
* Internal. Bounded string copy into a fixed buffer, always NUL-terminated.
* Raises AKERR_NULLPOINTER for a NULL destination or source and AKERR_VALUE for
* a capacity that leaves no room for a terminator. Exported so the library's
* own tests can drive those guards; not part of the consumer API.
*/
akerr_ErrorContext AKERR_NOIGNORE *__akerr_copy_string(char *destination, int capacity,
const char *source);
#define LOG_ERROR_WITH_MESSAGE(__err_context, __err_message) \
akerr_log_method("%s%s:%s:%d: %s %d (%s): %s", (char *)&__err_context->stacktracebuf, (char *)__FILE__, (char *)__func__, __LINE__, __err_message, __err_context->status, akerr_name_for_status(__err_context->status, NULL), __err_context->message); \
@@ -93,6 +269,11 @@ void akerr_init_errno(void);
akerr_init(); \
akerr_ErrorContext __attribute__ ((unused)) *__err_context = NULL;
/*
* akerr_next_error() hands back a context that already holds one reference --
* it has to, or a second thread could be given the same slot between the scan
* and the increment. There is nothing to increment here.
*/
#define ENSURE_ERROR_READY(__err_context) \
if ( __err_context == NULL ) { \
__err_context = akerr_next_error(); \
@@ -100,10 +281,31 @@ void akerr_init_errno(void);
akerr_log_method("%s:%s:%d: Unable to pull an error context from the array!", __FILE__, (char *)__func__, __LINE__); \
exit(1); \
} \
} \
__err_context->refcount += 1;
}
/*
/*
* Append a formatted line to the error's stack-trace buffer, bounded by the
* space that remains so a deep propagation chain cannot write past the end of
* stacktracebuf. snprintf reports the length it *would* have written, which on
* truncation exceeds what it actually wrote, so the cursor advance is clamped
* to the remaining space.
*/
#define AKERR_STACKTRACE_APPEND(__err_context, ...) \
do { \
char *__akerr_stb = (char *)__err_context->stacktracebuf; \
size_t __akerr_used = (size_t)(__err_context->stacktracebufptr - __akerr_stb); \
if ( __akerr_used < AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH ) { \
size_t __akerr_rem = AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH - __akerr_used; \
int __akerr_n = snprintf(__err_context->stacktracebufptr, __akerr_rem, __VA_ARGS__); \
if ( __akerr_n < 0 ) { \
__akerr_n = 0; \
} \
__err_context->stacktracebufptr += ((size_t)__akerr_n < __akerr_rem) \
? (size_t)__akerr_n : (__akerr_rem - 1); \
} \
} while ( 0 )
/*
* Failure and success methods for functions that return akerr_ErrorContext *
*/
@@ -158,11 +360,11 @@ void akerr_init_errno(void);
#define FAIL(__err_context, __err, __message, ...) \
ENSURE_ERROR_READY(__err_context); \
__err_context->status = __err; \
snprintf((char *)__err_context->fname, AKERR_MAX_ERROR_FNAME_LENGTH, __FILE__); \
snprintf((char *)__err_context->function, AKERR_MAX_ERROR_FUNCTION_LENGTH, __func__); \
snprintf((char *)__err_context->fname, AKERR_MAX_ERROR_FNAME_LENGTH, "%s", __FILE__); \
snprintf((char *)__err_context->function, AKERR_MAX_ERROR_FUNCTION_LENGTH, "%s", __func__); \
__err_context->lineno = __LINE__; \
snprintf((char *)__err_context->message, AKERR_MAX_ERROR_CONTEXT_STRING_LENGTH, __message, ## __VA_ARGS__); \
__err_context->stacktracebufptr += snprintf(__err_context->stacktracebufptr, AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH, "%s:%s:%d: %d (%s) : %s\n", (char *)__err_context->fname, (char *)__err_context->function, __err_context->lineno, __err_context->status, akerr_name_for_status(__err_context->status, NULL), (__err_context->message == NULL ? "" : __err_context->message));
AKERR_STACKTRACE_APPEND(__err_context, "%s:%s:%d: %d (%s) : %s\n", (char *)__err_context->fname, (char *)__err_context->function, __err_context->lineno, __err_context->status, akerr_name_for_status(__err_context->status, NULL), (__err_context->message == NULL ? "" : __err_context->message));
#define SUCCEED(__err_context) \
@@ -177,12 +379,18 @@ void akerr_init_errno(void);
switch ( 0 ) { \
case 0: \
#define VALID(__err_context, __stmt) \
__stmt; \
if ( akerr_valid_error_address(__err_context) == 0 ) { \
__err_context = NULL; \
FAIL(__err_context, AKERR_BADEXC, "Received (akerr_ErrorContext *) from an invalid memory region. (Did the method finish without calling SUCCEED_RETURN?)"); \
}
#define DETECT(__err_context, __stmt) \
__stmt; \
VALID(__err_context, __stmt); \
if ( __err_context != NULL ) { \
__err_context->stacktracebufptr += snprintf(__err_context->stacktracebufptr, AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH, "%s:%s:%d: Detected error %d from array (refcount %d)\n", (char *)__FILE__, (char *)__func__, __LINE__, __err_context->arrayid, __err_context->refcount); \
if ( __err_context->status != 0 ) { \
__err_context->stacktracebufptr += snprintf(__err_context->stacktracebufptr, AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH, "%s:%s:%d\n", (char *)__FILE__, (char *)__func__, __LINE__); \
AKERR_STACKTRACE_APPEND(__err_context, "%s:%s:%d\n", (char *)__FILE__, (char *)__func__, __LINE__); \
break; \
} \
}
@@ -190,6 +398,13 @@ void akerr_init_errno(void);
#define CATCH(__err_context, __stmt) \
DETECT(__err_context, __err_context = __stmt);
#define PASS(__err_context, __stmt) \
switch ( 0 ) { \
case 0: \
DETECT(__err_context, __err_context = __stmt); \
} \
FINISH_LOGIC(__err_context, true);
#define IGNORE(__stmt) \
__akerr_last_ignored = __stmt; \
if ( __akerr_last_ignored != NULL ) { \
@@ -222,18 +437,17 @@ void akerr_init_errno(void);
__err_context->stacktracebufptr = (char *)&__err_context->stacktracebuf; \
__err_context->handled = true;
#define FINISH_LOGIC(__err_context, __pass_up) \
if ( __err_context != NULL ) { \
if ( __err_context->handled == false && __pass_up == true ) { \
return __err_context; \
} \
} \
#define FINISH(__err_context, __pass_up) \
}; \
}; \
if ( __err_context != NULL ) { \
__err_context->stacktracebufptr += snprintf(__err_context->stacktracebufptr, AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH, "%s:%s:%d\n", (char *)__FILE__, (char *)__func__, __LINE__); \
if ( __err_context->handled == false && __pass_up == true && __err_context->arrayid > 0 ) { \
return __err_context; \
} else if ( __err_context->handled == false ) { \
LOG_ERROR_WITH_MESSAGE(__err_context, "Unhandled Error"); \
akerr_handler_unhandled_error(__err_context); \
} \
} \
FINISH_LOGIC(__err_context, __pass_up) \
RELEASE_ERROR(__err_context);
#define FINISH_NORETURN(__err_context) \
@@ -247,12 +461,4 @@ void akerr_init_errno(void);
} \
RELEASE_ERROR(__err_context);
#define CATCH_AND_RETURN(__err_context, __stmt) \
ATTEMPT { \
CATCH(__err_context, __stmt); \
} CLEANUP { \
} PROCESS(__err_context) { \
} FINISH(__err_context, true);
#endif // _AKERR_H_

407
scripts/coverage.py Executable file
View File

@@ -0,0 +1,407 @@
#!/usr/bin/env python3
"""
Code coverage harness for libakerror.
Coverage measures which parts of the library the CTest suite actually executes.
It is the complement to mutation testing (scripts/mutation_test.py): coverage
finds code the tests never reach, mutation testing finds code the tests reach
but do not really check.
What is measured is the library's own translation units: src/error.c and the
generated src/errno.c. The macros in the public header are deliberately not
measured -- GCC attributes an expanded macro to its call site, so header logic
would be reported as lines of the test that used it. Use mutation testing
(--target include/akerror.tmpl.h) to check how well those macros are tested.
This harness has no third-party dependencies -- just gcov, which ships with the
compiler, plus the project's normal cmake/ctest toolchain. By default it
configures its own instrumented build directory so the ordinary build tree is
left alone and coverage counters can never be stale.
Usage:
scripts/coverage.py [options]
--source-root DIR repo root (default: parent of this script's dir)
--build-dir DIR instrumented build dir (default: <root>/build/coverage)
--no-configure reuse the build dir as-is (do not cmake/build)
--no-run report existing counters (do not reset and run ctest)
--threshold PCT exit non-zero if line coverage < PCT (default: 0 = off)
--branch-threshold P exit non-zero if branch coverage < P (default: 0 = off)
--junit PATH write a JUnit XML report to this path
--max-uncovered N uncovered lines to list per file (default: 40, 0 = all)
--exclude SUBSTR skip reported paths containing SUBSTR; repeatable
--gcov PROG gcov program (default: $GCOV or "gcov")
"""
import argparse
import collections
import json
import os
import subprocess
import sys
# --------------------------------------------------------------------------- #
# Running the instrumented suite
# --------------------------------------------------------------------------- #
def configure_and_build(cmake, root, build, jobs):
"""Configure an instrumented build dir and build everything in it."""
cmds = [
[cmake, "-S", root, "-B", build, "-DAKERR_COVERAGE=ON"],
[cmake, "--build", build] + (["--parallel", str(jobs)] if jobs else []),
]
for cmd in cmds:
print("+ " + " ".join(cmd))
if subprocess.call(cmd) != 0:
return False
return True
def reset_counters(build):
"""Delete accumulated .gcda files so each run reports one suite run."""
removed = 0
for dirpath, _dirs, files in os.walk(build):
for name in files:
if name.endswith(".gcda"):
os.unlink(os.path.join(dirpath, name))
removed += 1
if removed:
print(f"Reset {removed} coverage data file(s).")
def run_ctest(build, ctest):
"""Run the suite. Returns True if every test passed.
Coverage counters are flushed at exit(), and the tests that are expected to
fail exit via the library's unhandled-error handler (exit(), not abort()),
so WILL_FAIL tests still contribute their counters.
"""
cmd = [ctest, "--output-on-failure"]
print("+ " + " ".join(cmd) + f" (in {build})")
return subprocess.call(cmd, cwd=build) == 0
# --------------------------------------------------------------------------- #
# Collecting gcov data
# --------------------------------------------------------------------------- #
class FileCov:
"""Merged coverage for one source file, across every object that built it.
A file compiled into more than one object -- or a header included by several
translation units -- is reported once, with counts summed. Branches are
merged by (line, index within line), so differing expansions of the same
line contribute the union of their branches.
"""
def __init__(self):
self.lines = collections.Counter() # lineno -> execution count
self.branches = collections.Counter() # (lineno, idx) -> taken count
self.funcs = collections.Counter() # (name, start_line) -> count
def merge(self, entry):
for ln in entry.get("lines", []):
no = ln["line_number"]
self.lines[no] += ln.get("count", 0)
for idx, br in enumerate(ln.get("branches", [])):
if br.get("throw"):
continue
self.branches[(no, idx)] += br.get("count", 0)
for fn in entry.get("functions", []):
key = (fn.get("name", "?"), fn.get("start_line", 0))
self.funcs[key] += fn.get("execution_count", 0)
@staticmethod
def _ratio(counter):
total = len(counter)
hit = sum(1 for v in counter.values() if v > 0)
return hit, total
def line_stats(self):
return self._ratio(self.lines)
def branch_stats(self):
return self._ratio(self.branches)
def func_stats(self):
return self._ratio(self.funcs)
def uncovered_lines(self):
return sorted(no for no, count in self.lines.items() if count == 0)
def find_notes(build):
"""Every .gcno in the build tree: one per instrumented translation unit."""
notes = []
for dirpath, _dirs, files in os.walk(build):
for name in files:
if name.endswith(".gcno"):
notes.append(os.path.join(dirpath, name))
return sorted(notes)
def gcov_json(gcov, note, cwd):
"""Run gcov on one .gcno and return its parsed JSON, or None on failure.
--stdout keeps gcov from littering .gcov files in the build tree. A .gcno
with no matching .gcda still reports, with all counts zero, which is the
correct answer for a translation unit no test executed.
"""
cmd = [gcov, "--branch-probabilities", "--json-format", "--stdout", note]
proc = subprocess.Popen(cmd, cwd=cwd, stdout=subprocess.PIPE,
stderr=subprocess.PIPE)
out, err = proc.communicate()
if proc.returncode != 0:
sys.stderr.write(f"warning: {' '.join(cmd)} failed:\n"
f"{err.decode(errors='replace')}")
return None
try:
return json.loads(out.decode(errors="replace"))
except ValueError as exc:
sys.stderr.write(f"warning: unparseable gcov output for {note}: {exc}\n")
return None
def display_path(path, build, root):
"""Label a source: build-relative for generated files, else repo-relative.
Returns None for anything outside both trees (system headers, toolchain
internals). Generated sources are labelled relative to the build dir so the
report and its thresholds do not change with the build dir's location.
"""
for base in (build, root):
if path.startswith(base + os.sep):
return os.path.relpath(path, base)
return None
def collect(gcov, build, root, excludes):
"""Merge gcov data for every instrumented unit into {display path: FileCov}."""
covs = {}
for note in find_notes(build):
data = gcov_json(gcov, note, build)
if data is None:
continue
# Paths in the report are relative to the directory the unit was
# compiled in, which gcov records in the notes file.
compile_dir = data.get("current_working_directory") or build
for entry in data.get("files", []):
path = entry.get("file", "")
if not os.path.isabs(path):
path = os.path.join(compile_dir, path)
display = display_path(os.path.realpath(path), build, root)
if display is None:
continue # system headers, toolchain internals
if any(x in display for x in excludes):
continue
covs.setdefault(display, FileCov()).merge(entry)
return covs
# --------------------------------------------------------------------------- #
# Reporting
# --------------------------------------------------------------------------- #
def pct(hit, total):
return 100.0 * hit / total if total else 100.0
def fmt_ratio(hit, total):
if not total:
return f"{'-':>9} -"
return f"{hit:>4}/{total:<4} {pct(hit, total):5.1f}%"
def compress(numbers):
"""[1,2,3,7,9,10] -> '1-3, 7, 9-10' for readable uncovered-line lists."""
out, start, prev = [], None, None
for n in numbers:
if start is None:
start = prev = n
elif n == prev + 1:
prev = n
else:
out.append(f"{start}" if start == prev else f"{start}-{prev}")
start = prev = n
if start is not None:
out.append(f"{start}" if start == prev else f"{start}-{prev}")
return ", ".join(out)
def report(covs, max_uncovered):
"""Print the per-file table and uncovered detail; return overall percentages."""
width = max([len(p) for p in covs] + [len("TOTAL")])
print("\n" + "=" * 72)
print("CODE COVERAGE SUMMARY")
print("=" * 72)
print(f" {'FILE':<{width}} {'LINES':^15} {'BRANCHES':^15} FUNCS")
totals = [0, 0, 0, 0, 0, 0] # lines hit/total, branches hit/total, funcs
for path in sorted(covs):
cov = covs[path]
lh, lt = cov.line_stats()
bh, bt = cov.branch_stats()
fh, ft = cov.func_stats()
for i, v in enumerate((lh, lt, bh, bt, fh, ft)):
totals[i] += v
print(f" {path:<{width}} {fmt_ratio(lh, lt)} {fmt_ratio(bh, bt)} "
f"{fh}/{ft}")
lh, lt, bh, bt, fh, ft = totals
print(f" {'-' * width} {'-' * 15} {'-' * 15} -----")
print(f" {'TOTAL':<{width}} {fmt_ratio(lh, lt)} {fmt_ratio(bh, bt)} "
f"{fh}/{ft}")
for path in sorted(covs):
missing = covs[path].uncovered_lines()
if not missing:
continue
shown = missing if not max_uncovered else missing[:max_uncovered]
more = "" if len(shown) == len(missing) else \
f" ... (+{len(missing) - len(shown)} more)"
print(f"\n uncovered in {path} ({len(missing)} line(s)):")
print(f" {compress(shown)}{more}")
return pct(lh, lt), pct(bh, bt)
def _xml_escape(text):
return (str(text).replace("&", "&amp;").replace("<", "&lt;")
.replace(">", "&gt;").replace('"', "&quot;"))
def write_junit(path, covs, line_threshold, branch_threshold):
"""One <testcase> per file per metric; below-threshold is a <failure>."""
cases = []
for f in sorted(covs):
cov = covs[f]
cases.append((f, "lines", cov.line_stats(), line_threshold))
cases.append((f, "branches", cov.branch_stats(), branch_threshold))
for metric, thr, stats in (("lines", line_threshold,
[c.line_stats() for c in covs.values()]),
("branches", branch_threshold,
[c.branch_stats() for c in covs.values()])):
cases.append(("TOTAL", metric,
(sum(h for h, _t in stats), sum(t for _h, t in stats)),
thr))
fails = sum(1 for _f, _m, (h, t), thr in cases
if t and thr > 0 and pct(h, t) < thr)
out = ['<?xml version="1.0" encoding="UTF-8"?>',
f'<testsuites name="coverage" tests="{len(cases)}" '
f'failures="{fails}">',
f' <testsuite name="coverage" tests="{len(cases)}" '
f'failures="{fails}">']
for f, metric, (hit, total), thr in cases:
name = _xml_escape(f"{f} {metric}")
detail = _xml_escape(f"{hit}/{total} ({pct(hit, total):.1f}%)"
if total else "no data")
out.append(f' <testcase name="{name}" '
f'classname="coverage.{_xml_escape(metric)}" time="0">')
if total and thr > 0 and pct(hit, total) < thr:
out.append(f' <failure message="{metric} coverage '
f'{pct(hit, total):.1f}% &lt; threshold {thr:.1f}%">'
f'{detail}</failure>')
else:
out.append(f' <system-out>{detail}</system-out>')
out.append(' </testcase>')
out.append(' </testsuite>')
out.append('</testsuites>')
with open(path, "w") as fh:
fh.write("\n".join(out) + "\n")
def main():
# Line-buffer stdout so progress is visible live under CI / the cmake target.
try:
sys.stdout.reconfigure(line_buffering=True)
except (AttributeError, ValueError):
pass
here = os.path.dirname(os.path.abspath(__file__))
default_root = os.path.dirname(here)
ap = argparse.ArgumentParser(description="Code coverage for libakerror")
ap.add_argument("--source-root", default=default_root)
ap.add_argument("--build-dir", default=None)
ap.add_argument("--configure", dest="configure", action="store_true",
default=True)
ap.add_argument("--no-configure", dest="configure", action="store_false")
ap.add_argument("--run", dest="run", action="store_true", default=True)
ap.add_argument("--no-run", dest="run", action="store_false")
ap.add_argument("--threshold", type=float, default=0.0,
help="fail if line coverage is below this percentage")
ap.add_argument("--branch-threshold", type=float, default=0.0,
help="fail if branch coverage is below this percentage")
ap.add_argument("--junit", default=None,
help="write a JUnit XML report to this path")
ap.add_argument("--max-uncovered", type=int, default=40)
ap.add_argument("--exclude", action="append", default=None,
help="skip reported paths containing this substring")
ap.add_argument("--gcov", default=os.environ.get("GCOV", "gcov"))
ap.add_argument("--cmake", default=os.environ.get("CMAKE", "cmake"))
ap.add_argument("--ctest", default=os.environ.get("CTEST", "ctest"))
ap.add_argument("-j", "--jobs", type=int, default=0)
args = ap.parse_args()
root = os.path.realpath(args.source_root)
build = os.path.realpath(args.build_dir or os.path.join(root, "build",
"coverage"))
# The tests exercise the library; their own source is not what we measure.
excludes = args.exclude if args.exclude is not None else ["tests/"]
if args.configure:
if not configure_and_build(args.cmake, root, build, args.jobs):
sys.stderr.write("\nInstrumented build FAILED; aborting.\n")
return 2
elif not os.path.isdir(build):
sys.stderr.write(f"No build dir at {build} (drop --no-configure).\n")
return 2
tests_ok = True
if args.run:
reset_counters(build)
tests_ok = run_ctest(build, args.ctest)
if not tests_ok:
sys.stderr.write("\nwarning: some tests FAILED; coverage below is "
"still reported, but the run is not green.\n")
covs = collect(args.gcov, build, root, excludes)
if not covs:
sys.stderr.write("No coverage data found. Was the build instrumented "
"(-DAKERR_COVERAGE=ON) and the suite run?\n")
return 2
line_pct, branch_pct = report(covs, args.max_uncovered)
if args.junit:
junit_path = os.path.abspath(args.junit)
write_junit(junit_path, covs, args.threshold, args.branch_threshold)
print(f"\nJUnit report written to: {junit_path}")
rc = 0
if not tests_ok:
print("\nFAIL: the CTest suite did not pass.")
rc = 1
# Thresholds gate every file as well as the total: a small, well-covered
# file (the generated status-name table) must not mask a regression in a
# bigger one. Files with no branches at all are not branch-gated.
checks = [("total", "line", line_pct, args.threshold),
("total", "branch", branch_pct, args.branch_threshold)]
for path in sorted(covs):
lh, lt = covs[path].line_stats()
bh, bt = covs[path].branch_stats()
checks.append((path, "line", pct(lh, lt), args.threshold))
if bt:
checks.append((path, "branch", pct(bh, bt), args.branch_threshold))
for path, metric, value, threshold in checks:
if threshold > 0 and value < threshold:
print(f"\nFAIL: {path} {metric} coverage {value:.1f}% < threshold "
f"{threshold:.1f}%")
rc = 1
return rc
if __name__ == "__main__":
sys.exit(main())

View File

@@ -1,16 +1,43 @@
#!/bin/bash
srcdir=$1
rm -f ${srcdir}/src/errno.c
echo "#include <akerror.h>" >> ${srcdir}/src/errno.c
echo "#include <errno.h>" >> ${srcdir}/src/errno.c
echo "void akerr_init_errno(void) {" >> ${srcdir}/src/errno.c
outdir=$2
# 1 when the library was configured with a threading backend, 0 for
# -DAKERR_THREADS=none. Stamped into the header so a consumer cannot disagree
# with the library about whether it locks and whether its per-thread state is
# thread local. Defaults to 1 for a hand-run of this script.
thread_safe=${3:-1}
if [ "${thread_safe}" != "0" ] && [ "${thread_safe}" != "1" ]; then
echo "$0: thread-safe argument must be 0 or 1, got '${thread_safe}'" >&2
exit 1
fi
mkdir -p ${outdir}/src
mkdir -p ${outdir}/include
rm -f ${outdir}/src/errno.c
echo "#include <akerror.h>" >> ${outdir}/src/errno.c
echo "#include <errno.h>" >> ${outdir}/src/errno.c
cat >> ${outdir}/src/errno.c <<'EOF'
/*
* These names belong to the library's own reserved band, and this runs from
* akerr_init(), which has no caller to raise into -- so it goes through
* __akerr_name_library_status(), which reports a refusal and terminates rather
* than leaving every later stack trace to print "Unknown Error" for an errno.
* Keeping the branch in src/error.c also keeps this generated file free of
* control flow no test can reach.
*/
EOF
echo "void akerr_init_errno(void) {" >> ${outdir}/src/errno.c
maxval=$(errno --list | cut -d ' ' -f 2 | sort -g | tail -n 1)
errno --list | while read LINE; do
define=$(echo "$LINE" | cut -d ' ' -f 1);
value=$(echo "$LINE" | cut -d ' ' -f 2);
desc=$(echo "$LINE" | cut -d ' ' -f 3-);
echo " akerr_name_for_status(${define}, \"${desc}\");" >> ${srcdir}/src/errno.c ;
echo " __akerr_name_library_status(${define}, \"${desc}\");" >> ${outdir}/src/errno.c ;
done;
echo "}" >> ${srcdir}/src/errno.c
sed "s/#define AKERR_LAST_ERRNO_VALUE .*/#define AKERR_LAST_ERRNO_VALUE ${maxval}/" ${srcdir}/include/akerror.tmpl.h > ${srcdir}/include/akerror.h
echo "}" >> ${outdir}/src/errno.c
sed -e "s/#define AKERR_LAST_ERRNO_VALUE .*/#define AKERR_LAST_ERRNO_VALUE ${maxval}/" \
-e "s/#define AKERR_THREAD_SAFE .*/#define AKERR_THREAD_SAFE ${thread_safe}/" \
${srcdir}/include/akerror.tmpl.h > ${outdir}/include/akerror.h

436
scripts/mutation_test.py Executable file
View File

@@ -0,0 +1,436 @@
#!/usr/bin/env python3
"""
Mutation testing harness for libakerror.
Mutation testing measures how good the test suite is at catching bugs. It works
by making many small, deliberate breakages ("mutants") to the library source --
flipping a comparison, deleting a statement, swapping true/false -- and then
running the whole CTest suite against each one. If the tests fail, the mutant is
"killed" (good: the tests noticed the bug). If the tests still pass, the mutant
"survived" (bad: a real bug of that shape would slip through unnoticed).
The mutation score is killed / (killed + survived). Surviving mutants are printed
with file:line and the exact change so they can be turned into new test cases.
This harness has no third-party dependencies (Python stdlib + the project's
normal cmake/ctest toolchain). It never mutates the real working tree: it copies
the repo to a scratch directory and mutates there.
Usage:
scripts/mutation_test.py [options]
--source-root DIR repo root to copy (default: parent of this script's dir)
--target FILE source file to mutate, relative to root; repeatable.
Default: src/error.c and include/akerror.tmpl.h
--work DIR scratch dir for the mutated copy (default: a temp dir)
--timeout SECONDS per-suite ctest timeout (default: 120)
--threshold PCT exit non-zero if mutation score < PCT (default: 0 = off)
--list only list the mutants that would be run, then exit
--keep keep the scratch working copy on exit (for debugging)
-j N (reserved) currently runs sequentially
"""
import argparse
import os
import re
import shutil
import subprocess
import sys
import tempfile
# --------------------------------------------------------------------------- #
# Mutation operators
#
# Each operator yields zero or more (start, end, replacement) edits for a single
# line of source. The driver applies exactly one edit per mutant so every mutant
# differs from the original by one localized change.
# --------------------------------------------------------------------------- #
# Relational operator replacement: map each operator to the alternatives that
# meaningfully change behaviour (not merely the strict negation).
_REL = {
"==": ["!="],
"!=": ["=="],
"<=": ["<", "=="],
">=": [">", "=="],
"<": ["<=", ">"],
">": [">=", "<"],
}
# Match a relational operator that is NOT part of ->, <<, >>, =>, <=, >=, ==, !=
# unless we intend it. We tokenize the two-char operators first, then single.
_REL_TWO = re.compile(r"(==|!=|<=|>=)")
_REL_ONE = re.compile(r"(?<![-<>=!+])([<>])(?![=<>])")
_LOGICAL = {"&&": "||", "||": "&&"}
_LOG_RE = re.compile(r"(&&|\|\|)")
_BOOL = {"true": "false", "false": "true"}
_BOOL_RE = re.compile(r"\b(true|false)\b")
# Arithmetic / compound-assignment on whitespace-delimited operands only, to
# avoid touching ++, --, ->, unary signs, or pointer/format punctuation.
_ARITH_RE = re.compile(r"(?<=\s)([+\-])(?=\s)")
_ARITH = {"+": "-", "-": "+"}
_COMPOUND_RE = re.compile(r"(\+=|-=)")
_COMPOUND = {"+=": "-=", "-=": "+="}
# Integer literal replacement: 0 <-> 1 (word-bounded, not inside identifiers or
# larger numbers, not a float).
_INT_RE = re.compile(r"(?<![\w.])([01])(?![\w.])")
_INT = {"0": "1", "1": "0"}
def _op_edits(line):
"""Yield (tag, start, end, replacement) for every candidate mutation."""
# Relational (two-char first so we don't split them with the one-char pass)
for m in _REL_TWO.finditer(line):
for alt in _REL[m.group(1)]:
yield ("ROR", m.start(1), m.end(1), alt)
for m in _REL_ONE.finditer(line):
for alt in _REL[m.group(1)]:
yield ("ROR", m.start(1), m.end(1), alt)
for m in _LOG_RE.finditer(line):
yield ("LCR", m.start(1), m.end(1), _LOGICAL[m.group(1)])
for m in _BOOL_RE.finditer(line):
yield ("BCR", m.start(1), m.end(1), _BOOL[m.group(1)])
for m in _COMPOUND_RE.finditer(line):
yield ("AOR", m.start(1), m.end(1), _COMPOUND[m.group(1)])
for m in _ARITH_RE.finditer(line):
yield ("AOR", m.start(1), m.end(1), _ARITH[m.group(1)])
for m in _INT_RE.finditer(line):
yield ("ICR", m.start(1), m.end(1), _INT[m.group(1)])
# Statement-deletion: neutralize a whole statement. We only delete statements
# that are safe to drop without guaranteeing a compile error, so a surviving
# deletion is a genuine test gap rather than compiler noise.
_STMT_DELETABLE = re.compile(
r"""^\s*(
break |
return\b[^;]* |
[A-Za-z_][-\w>().\[\]* ]*\s*=\s*[^;]* | # assignments
[A-Za-z_][\w]*\s*\([^;]*\) # bare function calls
)\s*;\s*(\\?)\s*$""",
re.VERBOSE,
)
# --------------------------------------------------------------------------- #
# Deciding which lines are eligible to mutate
# --------------------------------------------------------------------------- #
# Skip preprocessor control and the block of constant/error-code #defines in the
# template header: mutating buffer sizes or renumbering error codes produces
# equivalent or uninteresting mutants that swamp the signal.
_SKIP_LINE = re.compile(
r"""^\s*(
\#\s*(include|ifn?def|ifdef|if|elif|else|endif|error|pragma|undef) |
\#\s*define\s+AKERR_(MAX|LAST|NULLPOINTER|OUTOFBOUNDS|API|ATTRIBUTE|
TYPE|KEY|INDEX|FORMAT|IO|VALUE|RELATIONSHIP|EOF|CIRCULAR_REFERENCE|
ITERATOR_BREAK|NOT_IMPLEMENTED|BADEXC|NOIGNORE|USE_STDLIB)\b |
\* | // # comment bodies / line comments
)""",
re.VERBOSE,
)
def _is_comment_or_blank(line):
s = line.strip()
return (not s) or s.startswith("//") or s.startswith("/*") or s.startswith("*")
def eligible(line):
if _is_comment_or_blank(line):
return False
if _SKIP_LINE.match(line):
return False
return True
class Mutant:
__slots__ = ("path", "lineno", "op", "before", "after", "col")
def __init__(self, path, lineno, op, before, after, col):
self.path = path
self.lineno = lineno
self.op = op
self.before = before
self.after = after
self.col = col
def describe(self):
return (f"{self.path}:{self.lineno} [{self.op}] "
f"col{self.col}: {self.before.strip()} -> {self.after.strip()}")
def generate_mutants(root, rel_target):
"""Enumerate all mutants for one target file."""
abspath = os.path.join(root, rel_target)
with open(abspath, "r") as fh:
lines = fh.readlines()
mutants = []
for i, line in enumerate(lines, start=1):
if not eligible(line):
continue
# substitution operators
seen = set()
for tag, s, e, repl in _op_edits(line):
key = (s, e, repl)
if key in seen:
continue
seen.add(key)
mutated = line[:s] + repl + line[e:]
if mutated == line:
continue
mutants.append(Mutant(rel_target, i, tag, line, mutated, s))
# statement deletion
m = _STMT_DELETABLE.match(line)
if m:
indent = line[: len(line) - len(line.lstrip())]
cont = "\\" if line.rstrip().endswith("\\") else ""
deleted = f"{indent}/* mutant: deleted */ {cont}\n" if cont else f"{indent};\n"
mutants.append(Mutant(rel_target, i, "SDL", line, deleted, 0))
return mutants
# --------------------------------------------------------------------------- #
# Build / test orchestration against a scratch copy
# --------------------------------------------------------------------------- #
class Runner:
def __init__(self, work, timeout):
self.work = work
self.build = os.path.join(work, "build")
self.timeout = timeout
def _run(self, cmd, timeout=None):
return subprocess.run(
cmd, cwd=self.work, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
timeout=timeout,
)
def configure(self):
r = self._run(["cmake", "-S", ".", "-B", "build"], timeout=self.timeout)
return r.returncode == 0, r.stdout
def build_and_test(self):
"""Return ('killed-compile' | 'killed-test' | 'killed-timeout' | 'survived')."""
try:
b = self._run(["cmake", "--build", "build"], timeout=self.timeout)
except subprocess.TimeoutExpired:
return "killed-timeout"
if b.returncode != 0:
return "killed-compile"
try:
t = subprocess.run(
["ctest", "--test-dir", "build", "--output-on-failure",
"--stop-on-failure"],
cwd=self.work, stdout=subprocess.PIPE, stderr=subprocess.STDOUT,
timeout=self.timeout,
)
except subprocess.TimeoutExpired:
return "killed-timeout"
return "survived" if t.returncode == 0 else "killed-test"
def _xml_escape(s):
return (s.replace("&", "&amp;").replace("<", "&lt;").replace(">", "&gt;")
.replace('"', "&quot;"))
def write_junit(path, records, targets):
"""Write a JUnit XML report. One <testcase> per mutant; a surviving mutant
is a <failure> (test-suite gap), a killed mutant is a passing case."""
by_file = {t: [] for t in targets}
for m, result in records:
by_file.setdefault(m.path, []).append((m, result))
total = len(records)
total_fail = sum(1 for _, r in records if r == "survived")
out = ['<?xml version="1.0" encoding="UTF-8"?>',
f'<testsuites name="mutation" tests="{total}" failures="{total_fail}">']
for f, items in by_file.items():
if not items:
continue
fails = sum(1 for _, r in items if r == "survived")
out.append(f' <testsuite name="mutation:{_xml_escape(f)}" '
f'tests="{len(items)}" failures="{fails}">')
for m, result in items:
name = _xml_escape(m.describe())
cls = "mutation." + _xml_escape(m.path)
if result == "survived":
detail = _xml_escape(f"{m.before.strip()} -> {m.after.strip()}")
out.append(f' <testcase name="{name}" classname="{cls}" time="0">')
out.append(f' <failure message="survived mutant '
f'({_xml_escape(m.op)} at {_xml_escape(m.path)}:'
f'{m.lineno})">{detail}</failure>')
out.append(' </testcase>')
else:
out.append(f' <testcase name="{name}" classname="{cls}" '
f'time="0"><system-out>{_xml_escape(result)}'
'</system-out></testcase>')
out.append(' </testsuite>')
out.append('</testsuites>')
with open(path, "w") as fh:
fh.write("\n".join(out) + "\n")
def copy_tree(src, dst):
ignore = shutil.ignore_patterns("build", ".git", "*.o", "*.so", "*~",
"#*#", "*.iso", "*.png")
shutil.copytree(src, dst, ignore=ignore, symlinks=True)
def read_lines(path):
with open(path) as fh:
return fh.readlines()
def write_lines(path, lines):
with open(path, "w") as fh:
fh.writelines(lines)
def main():
# Line-buffer stdout so progress is visible live under CI / the cmake target.
try:
sys.stdout.reconfigure(line_buffering=True)
except (AttributeError, ValueError):
pass
here = os.path.dirname(os.path.abspath(__file__))
default_root = os.path.dirname(here)
ap = argparse.ArgumentParser(description="Mutation testing for libakerror")
ap.add_argument("--source-root", default=default_root)
ap.add_argument("--target", action="append", default=None)
ap.add_argument("--work", default=None)
ap.add_argument("--timeout", type=int, default=120)
ap.add_argument("--threshold", type=float, default=0.0)
ap.add_argument("--junit", default=None,
help="write a JUnit XML report to this path")
ap.add_argument("--max-mutants", type=int, default=0,
help="cap the run at N evenly-sampled mutants (0 = all)")
ap.add_argument("--list", action="store_true")
ap.add_argument("--keep", action="store_true")
ap.add_argument("-j", type=int, default=1)
args = ap.parse_args()
root = os.path.abspath(args.source_root)
targets = args.target or ["src/error.c", "include/akerror.tmpl.h"]
# Enumerate mutants from the pristine sources.
all_mutants = []
for t in targets:
all_mutants.extend(generate_mutants(root, t))
print(f"Generated {len(all_mutants)} mutants across {len(targets)} file(s):")
for t in targets:
n = sum(1 for m in all_mutants if m.path == t)
print(f" {t}: {n}")
# Optional even-strided sampling to bound run time (CI / smoke tests).
if args.max_mutants and len(all_mutants) > args.max_mutants:
step = len(all_mutants) / args.max_mutants
sampled = [all_mutants[int(i * step)] for i in range(args.max_mutants)]
print(f"Sampling {len(sampled)} of {len(all_mutants)} mutants "
f"(--max-mutants {args.max_mutants}).")
all_mutants = sampled
if args.list:
for m in all_mutants:
print(" " + m.describe())
return 0
if not all_mutants:
print("No mutants generated; nothing to do.")
return 0
# Scratch working copy.
work_parent = args.work or tempfile.mkdtemp(prefix="akerr_mut_")
work = os.path.join(work_parent, "src") if args.work else work_parent
if os.path.exists(work):
shutil.rmtree(work)
print(f"\nCopying sources to scratch dir: {work}")
copy_tree(root, work)
runner = Runner(work, args.timeout)
print("Configuring baseline ...")
ok, out = runner.configure()
if not ok:
sys.stderr.write(out.decode(errors="replace"))
sys.stderr.write("\nBaseline configure FAILED; aborting.\n")
return 2
print("Verifying baseline is green (no mutation) ...")
baseline = runner.build_and_test()
if baseline != "survived":
sys.stderr.write(f"Baseline is not green ({baseline}); aborting. "
"Fix the suite before mutation testing.\n")
return 2
print("Baseline OK.\n")
# Group mutants by file so we mutate one file at a time and restore it.
killed = {"killed-compile": 0, "killed-test": 0, "killed-timeout": 0}
survivors = []
records = []
total = len(all_mutants)
# Cache pristine contents per target.
pristine = {t: read_lines(os.path.join(work, t)) for t in targets}
for idx, m in enumerate(all_mutants, start=1):
tgt_abs = os.path.join(work, m.path)
lines = list(pristine[m.path])
lines[m.lineno - 1] = m.after
write_lines(tgt_abs, lines)
try:
result = runner.build_and_test()
finally:
write_lines(tgt_abs, pristine[m.path]) # always restore
records.append((m, result))
if result == "survived":
survivors.append(m)
mark = "SURVIVED"
else:
killed[result] += 1
mark = result.upper()
print(f"[{idx}/{total}] {mark:16} {m.describe()}")
total_killed = sum(killed.values())
score = 100.0 * total_killed / total if total else 100.0
print("\n" + "=" * 72)
print("MUTATION TESTING SUMMARY")
print("=" * 72)
print(f" total mutants : {total}")
print(f" killed (test) : {killed['killed-test']}")
print(f" killed (compile): {killed['killed-compile']}")
print(f" killed (timeout): {killed['killed-timeout']}")
print(f" survived : {len(survivors)}")
print(f" mutation score : {score:.1f}%")
if survivors:
print("\nSurviving mutants (test-suite gaps -- turn these into tests):")
for m in survivors:
print(" " + m.describe())
if args.junit:
junit_path = os.path.abspath(args.junit)
write_junit(junit_path, records, targets)
print(f"\nJUnit report written to: {junit_path}")
if not args.keep and not args.work:
shutil.rmtree(work_parent, ignore_errors=True)
else:
print(f"\nScratch working copy kept at: {work}")
if args.threshold > 0 and score < args.threshold:
print(f"\nFAIL: mutation score {score:.1f}% < threshold {args.threshold:.1f}%")
return 1
return 0
if __name__ == "__main__":
sys.exit(main())

45
scripts/thread_test.sh Executable file
View File

@@ -0,0 +1,45 @@
#!/bin/bash
#
# Run the test suite under ThreadSanitizer.
#
# The thread tests assert properties -- exclusive ownership of pool slots and of
# reserved status ranges -- that hold or fail without any tooling. This is the
# run that proves the absence of a data race underneath them, so it is the one
# that has to be easy to type.
#
# Usage: scripts/thread_test.sh [build directory]
set -o errexit
set -o nounset
set -o pipefail
BUILD_DIR=${1:-build/tsan}
SANITIZE=${AKERR_SANITIZE:-thread}
# Work from the repository root whatever directory this was invoked from, so a
# relative build directory always lands in the same place.
cd "$(dirname "$0")/.."
for tool in cmake ctest; do
if ! command -v "${tool}" >/dev/null 2>&1; then
echo "$0: ${tool} is required and was not found" >&2
exit 1
fi
done
# ThreadSanitizer maps its shadow memory at fixed addresses and aborts with
# "FATAL: ThreadSanitizer: unexpected memory mapping" on kernels configured for
# more ASLR entropy than it allows (vm.mmap_rnd_bits > 28, the default on
# several recent distributions). Running with ASLR disabled sidesteps it without
# needing root, which "sysctl -w vm.mmap_rnd_bits=28" would.
RUNNER=()
if command -v setarch >/dev/null 2>&1; then
RUNNER=(setarch -R)
else
echo "$0: setarch not found; if ThreadSanitizer aborts with an unexpected" \
"memory mapping, lower vm.mmap_rnd_bits to 28" >&2
fi
cmake -S . -B "${BUILD_DIR}" -DAKERR_SANITIZE="${SANITIZE}"
cmake --build "${BUILD_DIR}"
"${RUNNER[@]}" ctest --test-dir "${BUILD_DIR}" --output-on-failure "${@:2}"

View File

@@ -1,19 +1,149 @@
#include "akerror.h"
#include "lock.h"
#if defined(AKERR_USE_STDLIB) && AKERR_USE_STDLIB == 1
#include <stdlib.h>
#include <stdarg.h>
#include <stdio.h>
#endif // AKERR_USE_STDLIB
akerr_ErrorContext __akerr_last_ditch;
akerr_ErrorContext *__akerr_last_ignored;
/*
* Per-thread state.
*
* The last-ditch context is where a failure gets reported when there is no pool
* slot to report it from -- akerr_release_error(NULL). One shared copy would
* have two threads formatting a message into the same buffer, so each thread
* gets its own. Thread-local storage is zero initialized, which is why
* akerr_last_ditch_context() below sets the stack-trace cursor lazily rather
* than akerr_init() setting it for everyone: akerr_init() runs on one thread
* and cannot reach the others' copies.
*
* It is not small (an akerr_ErrorContext is tens of kilobytes), but the storage
* is allocated per thread only when that thread first touches the library's
* thread-local block, and the alternative is a shared buffer that two threads
* can be writing at once.
*/
static AKERR_THREAD_LOCAL akerr_ErrorContext __akerr_last_ditch;
AKERR_THREAD_LOCAL akerr_ErrorContext *__akerr_last_ignored;
akerr_ErrorUnhandledErrorHandler akerr_handler_unhandled_error;
akerr_ErrorLogFunction akerr_log_method = NULL;
char __AKERR_ERROR_NAMES[AKERR_MAX_ERR_VALUE+1][AKERR_MAX_ERROR_NAME_LENGTH];
/*
* One recursive lock over both the error pool and the status registry. See
* src/lock.h for why it is one lock, and why it is recursive.
*
* Everything below that touches either table does so through a function whose
* name ends in _locked, called from a wrapper that takes the lock and releases
* it on the single return path. The wrappers exist because the locked bodies
* are written with the FAIL_*_RETURN macros, which return from the middle of a
* function -- so those bodies cannot be the ones holding the lock.
*/
static akerr_Mutex akerr_state_lock;
/*
* Status-name registry.
*
* Storage is an open-addressed hash table keyed by status value. Both sizes are
* private to this translation unit -- they are deliberately NOT in the public
* header, because a consumer-visible table bound is exactly the ABI hazard this
* registry replaced. Overriding them changes only this file, so a library and
* its consumers can never disagree about the layout.
*
* The table is never resized or rehashed, so a pointer handed out by
* akerr_name_for_status() stays valid for the life of the process. Entries are
* never removed, so probing needs no tombstones.
*/
#ifndef AKERR_STATUS_NAME_SLOTS
#define AKERR_STATUS_NAME_SLOTS 4096
#endif
#ifndef AKERR_MAX_RESERVED_STATUS_RANGES
#define AKERR_MAX_RESERVED_STATUS_RANGES 64
#endif
/* Probing masks with SLOTS-1, so the slot count must be a power of two. This is
* the C99-portable spelling of a static assertion (a negative array bound). */
typedef char akerr_assert_name_slots_pow2[
(AKERR_STATUS_NAME_SLOTS > 0 &&
(AKERR_STATUS_NAME_SLOTS & (AKERR_STATUS_NAME_SLOTS - 1)) == 0) ? 1 : -1];
/* Cap occupancy at 75% so linear probing always meets an empty slot. */
#define AKERR_MAX_REGISTERED_STATUS_NAMES \
(AKERR_STATUS_NAME_SLOTS - (AKERR_STATUS_NAME_SLOTS / 4))
#define AKERR_MAX_STATUS_RANGE_OWNER_LENGTH 64
typedef struct
{
int status;
int used;
char name[AKERR_MAX_ERROR_NAME_LENGTH];
} akerr_StatusName;
typedef struct
{
int first;
int last;
char owner[AKERR_MAX_STATUS_RANGE_OWNER_LENGTH];
} akerr_StatusRange;
static akerr_StatusName akerr_status_names[AKERR_STATUS_NAME_SLOTS];
static int akerr_status_name_count;
static akerr_StatusRange akerr_status_ranges[AKERR_MAX_RESERVED_STATUS_RANGES];
static int akerr_status_range_count;
akerr_ErrorContext AKERR_ARRAY_ERROR[AKERR_MAX_ARRAY_ERROR];
/*
* Bounded copy into a fixed buffer. Every argument is checked: this writes
* through a caller-supplied pointer for a caller-supplied length, so a NULL or
* a non-positive capacity here is a memory error waiting to happen, not
* something to absorb and return from quietly.
*
* Exported under the internal __akerr_ prefix rather than kept static so that
* tests/err_copy_string.c can reach these guards. Both in-library callers
* validate their arguments first, so nothing else can drive them.
*/
akerr_ErrorContext *__akerr_copy_string(char *destination, int capacity,
const char *source)
{
PREPARE_ERROR(errctx);
FAIL_NONZERO_RETURN(errctx, (destination == NULL || source == NULL),
AKERR_NULLPOINTER,
"__akerr_copy_string got a NULL %s",
destination == NULL ? "destination buffer" : "source string");
FAIL_NONZERO_RETURN(errctx, (capacity <= 0), AKERR_VALUE,
"__akerr_copy_string got a capacity of %d; a buffer must "
"have room for at least a terminator", capacity);
strncpy(destination, source, (size_t)capacity - 1);
destination[capacity - 1] = '\0';
SUCCEED_RETURN(errctx);
}
/*
* Compare against each element address rather than testing the address range.
* A range test also accepts pointers into the *interior* of an element, which
* would then be treated as the head of an akerr_ErrorContext and written
* through. Keep this an element-wise scan; it is not a missed optimization.
*
* Takes no lock: the addresses of the pool slots are fixed for the life of the
* process, and nothing here reads a slot's contents.
*/
int akerr_valid_error_address(akerr_ErrorContext *ptr)
{
// Is this within the memory region occupied by AKERR_ARRAY_ERROR?
if ( ptr == NULL ) {
return 1;
}
for ( int i = 0; i < AKERR_MAX_ARRAY_ERROR; i++ ) {
if ( ptr == &AKERR_ARRAY_ERROR[i] ) {
return 1;
}
}
return 0;
}
void akerr_default_logger(const char *fmt, ...)
{
#if defined(AKERR_USE_STDLIB) && AKERR_USE_STDLIB == 1
@@ -27,68 +157,237 @@ void akerr_default_logger(const char *fmt, ...)
#endif
}
void akerr_init()
/*
* The library naming its own codes.
*
* akerr_init() returns void and runs before any consumer frame exists, so there
* is nothing to PASS an error to: this *is* the top of the stack. FINISH_NORETURN
* is the library's idiom for that position -- the same one main() uses -- so an
* unhandled failure prints its stack trace and goes to
* akerr_handler_unhandled_error, which terminates.
*
* That is fatal on purpose. The library can only fail to name its own status
* codes if the build is misconfigured -- a name table too small to hold even
* the library's own entries, or a reservation that did not take -- and the
* consequence of continuing is every later stack trace in the process printing
* "Unknown Error" for a code the library defines. That is a startup defect, and
* it is far cheaper to see it at init than to debug it from a degraded trace.
*
* The generated errno table calls this rather than registering names directly,
* so that all of this control flow lives here in one place.
*/
void __akerr_name_library_status(int status, const char *name)
{
static int inited = 0;
if ( inited == 0 ) {
for (int i = 0; i < AKERR_MAX_ARRAY_ERROR; i++ ) {
memset((void *)&AKERR_ARRAY_ERROR[i], 0x00, sizeof(akerr_ErrorContext));
AKERR_ARRAY_ERROR[i].arrayid = i;
AKERR_ARRAY_ERROR[i].stacktracebufptr = (char *)&AKERR_ARRAY_ERROR[i].stacktracebuf;
}
__akerr_last_ignored = NULL;
PREPARE_ERROR(errctx);
ATTEMPT {
CATCH(errctx, akerr_register_status_name(AKERR_LIBRARY_OWNER, status, name));
} CLEANUP {
} PROCESS(errctx) {
} FINISH_NORETURN(errctx);
}
/*
* The calling thread's last-ditch context.
*
* Thread-local storage starts zeroed, so a NULL stack-trace cursor means this
* thread has not used its copy yet. Checking the cursor rather than a separate
* flag keeps the whole thing self-describing: the cursor is the one field that
* must not be zero for the context to be usable at all.
*/
static akerr_ErrorContext *akerr_last_ditch_context(void)
{
if ( __akerr_last_ditch.stacktracebufptr == NULL ) {
memset((void *)&__akerr_last_ditch, 0x00, sizeof(akerr_ErrorContext));
__akerr_last_ditch.stacktracebufptr = (char *)&__akerr_last_ditch.stacktracebuf;
if ( akerr_log_method == NULL ) {
akerr_log_method = &akerr_default_logger;
}
akerr_handler_unhandled_error = &akerr_default_handler_unhandled_error;
memset((void *)&__AKERR_ERROR_NAMES[0], 0x00, ((AKERR_MAX_ERR_VALUE+1) * AKERR_MAX_ERROR_NAME_LENGTH));
akerr_name_for_status(AKERR_NULLPOINTER, "Null Pointer Error");
akerr_name_for_status(AKERR_OUTOFBOUNDS, "Out Of Bounds Error");
akerr_name_for_status(AKERR_API, "API Error");
akerr_name_for_status(AKERR_ATTRIBUTE, "Attribute Error");
akerr_name_for_status(AKERR_TYPE, "Type Error");
akerr_name_for_status(AKERR_KEY, "Key Error");
akerr_name_for_status(AKERR_HEAP, "Heap Error");
akerr_name_for_status(AKERR_INDEX, "Index Error");
akerr_name_for_status(AKERR_FORMAT, "Format Error");
akerr_name_for_status(AKERR_IO, "Input Output Error");
akerr_name_for_status(AKERR_REGISTRY, "Registry Error");
akerr_name_for_status(AKERR_VALUE, "Value Error");
akerr_name_for_status(AKERR_BEHAVIOR, "Behavior Error");
akerr_name_for_status(AKERR_RELATIONSHIP, "Relationship Error");
inited = 1;
}
return &__akerr_last_ditch;
}
static AKERR_THREAD_LOCAL int akerr_initializing;
static akerr_Once akerr_state_once = AKERR_ONCE_INIT;
/*
* Runs exactly once per process, under akerr_once(). Everything it calls --
* the registry entry points, and the pool underneath them -- calls akerr_init()
* itself, so the first thing it does is raise the re-entry flag; see
* akerr_init() below.
*
* The lock is initialized before anything that could take it. That ordering is
* the reason this work lives in a once-routine instead of a
* check-a-flag-and-go: a second thread arriving while this one is still
* populating the tables must block until they are complete, not walk them.
*/
static void akerr_init_state(void)
{
akerr_mutex_init(&akerr_state_lock);
akerr_initializing = 1;
for (int i = 0; i < AKERR_MAX_ARRAY_ERROR; i++ ) {
memset((void *)&AKERR_ARRAY_ERROR[i], 0x00, sizeof(akerr_ErrorContext));
AKERR_ARRAY_ERROR[i].arrayid = i;
AKERR_ARRAY_ERROR[i].stacktracebufptr = (char *)&AKERR_ARRAY_ERROR[i].stacktracebuf;
}
__akerr_last_ignored = NULL;
(void)akerr_last_ditch_context();
if ( akerr_log_method == NULL ) {
akerr_log_method = &akerr_default_logger;
}
akerr_handler_unhandled_error = &akerr_default_handler_unhandled_error;
memset((void *)&akerr_status_names[0], 0x00, sizeof(akerr_status_names));
memset((void *)&akerr_status_ranges[0], 0x00, sizeof(akerr_status_ranges));
akerr_status_name_count = 0;
akerr_status_range_count = 0;
/* errno and AKERR_* values are the library-owned compatibility band.
* This must precede every registration below: naming a status is only
* permitted inside a reserved range. Terminal for the same reason as
* __akerr_name_library_status(), and handled the same way: without this
* band the library owns nothing, so none of the names below could
* register either. */
PREPARE_ERROR(errctx);
ATTEMPT {
CATCH(errctx, akerr_reserve_status_range(0, AKERR_RESERVED_STATUS_COUNT,
AKERR_LIBRARY_OWNER));
} CLEANUP {
} PROCESS(errctx) {
} FINISH_NORETURN(errctx);
/* Every AKERR_* code gets a name; tests/err_error_names.c asserts the
* list is exhaustive so a new code cannot be added without one. */
__akerr_name_library_status(AKERR_NULLPOINTER, "Null Pointer Error");
__akerr_name_library_status(AKERR_OUTOFBOUNDS, "Out Of Bounds Error");
__akerr_name_library_status(AKERR_API, "API Error");
__akerr_name_library_status(AKERR_ATTRIBUTE, "Attribute Error");
__akerr_name_library_status(AKERR_TYPE, "Type Error");
__akerr_name_library_status(AKERR_KEY, "Key Error");
__akerr_name_library_status(AKERR_INDEX, "Index Error");
__akerr_name_library_status(AKERR_FORMAT, "Format Error");
__akerr_name_library_status(AKERR_IO, "Input Output Error");
__akerr_name_library_status(AKERR_VALUE, "Value Error");
__akerr_name_library_status(AKERR_RELATIONSHIP, "Relationship Error");
__akerr_name_library_status(AKERR_EOF, "End Of File");
__akerr_name_library_status(AKERR_CIRCULAR_REFERENCE, "Circular Reference Error");
__akerr_name_library_status(AKERR_ITERATOR_BREAK, "Iterator Break");
__akerr_name_library_status(AKERR_NOT_IMPLEMENTED, "Not Implemented");
__akerr_name_library_status(AKERR_BADEXC, "Invalid akerr_ErrorContext");
__akerr_name_library_status(AKERR_STATUS_RANGE_OVERLAP, "Status Range Overlap");
__akerr_name_library_status(AKERR_STATUS_RANGE_FULL, "Status Range Table Full");
__akerr_name_library_status(AKERR_STATUS_RANGE_INVALID, "Invalid Status Range");
__akerr_name_library_status(AKERR_STATUS_NAME_UNRESERVED, "Unreserved Status Name");
__akerr_name_library_status(AKERR_STATUS_NAME_FOREIGN, "Foreign Status Name");
__akerr_name_library_status(AKERR_STATUS_NAME_FULL, "Status Name Registry Full");
__akerr_name_library_status(AKERR_STATUS_NAME_INVALID, "Invalid Status Name");
#if (defined(AKERR_USE_STDLIB) && AKERR_USE_STDLIB == 1) || (!defined(AKERR_USE_STDLIB))
akerr_init_errno();
#endif
akerr_initializing = 0;
}
/*
* Idempotent, and safe to call from any thread at any time. Every public entry
* point calls it -- so that a consumer reserving its range before anything else
* touches the library cannot have that reservation wiped by a later first-use
* of the pool -- which means it is also on the path of everything
* akerr_init_state() itself calls.
*
* The re-entry guard is thread local, and has to be: only the thread running
* the once-routine may skip past it. A second thread arriving mid-initialization
* must block inside akerr_once() until the tables are complete, which a shared
* flag set at the top of initialization would have let it walk right past.
*/
void akerr_init()
{
if ( akerr_initializing != 0 ) {
return;
}
akerr_once(&akerr_state_once, &akerr_init_state);
}
/*
* Every way out of the library's status space goes through here, so that a
* status becomes an exit code exactly one way no matter who is leaving.
*
* Only 0 through AKERR_EXIT_STATUS_MAX survive the trip -- see the note on
* AKERR_EXIT_STATUS_UNREPRESENTABLE in the header for why there is no wider
* exit() to reach for. Everything else exits with that sentinel instead of its
* low byte, because the low byte is either a lie (status 300 exiting 44, which
* is some other error's code) or a disaster (status 256, the first status a
* consumer can own, exiting 0 and telling the shell the program succeeded).
*
* Status 0 exits 0, because 0 is this library's success status and an exit code
* of 0 is what success is called out here. What keeps an *unhandled* error from
* exiting 0 is not this function: PROCESS opens with `case 0`, which marks a
* zero status handled, so a successful context can never reach
* FINISH_NORETURN's call to the handler in the first place.
*
* Nothing is logged here. Callers arrive from a position that has already
* reported -- FINISH_NORETURN logs the stack trace, carrying the status at full
* width, before it calls the handler -- and a second line naming a number the
* trace already gave would only invite the reader to trust the exit code.
*/
void akerr_exit(int status)
{
if ( status < 0 || status > AKERR_EXIT_STATUS_MAX ) {
exit(AKERR_EXIT_STATUS_UNREPRESENTABLE);
}
exit(status);
}
/*
* The last stop for an unhandled error. A handler invoked with no error at all
* has no status to report and nothing akerr_exit() could map, so it exits 1
* directly: a plain generic failure.
*/
void akerr_default_handler_unhandled_error(akerr_ErrorContext *errctx)
{
if ( errctx == NULL ) {
exit(1);
}
exit(errctx->status);
if ( errctx == NULL ) {
exit(1);
}
akerr_exit(errctx->status);
}
/*
* Claim the lowest free slot. Finding it and taking the reference are one
* operation under the lock: a scan that returned an unclaimed slot would hand
* the same one to every thread that scanned before the first of them got around
* to incrementing the count.
*/
akerr_ErrorContext *akerr_next_error()
{
akerr_ErrorContext *found = (akerr_ErrorContext *)NULL;
akerr_init();
akerr_mutex_lock(&akerr_state_lock);
for (int i = 0; i < AKERR_MAX_ARRAY_ERROR; i++ ) {
if ( AKERR_ARRAY_ERROR[i].refcount == 0 ) {
return &AKERR_ARRAY_ERROR[i];
found = &AKERR_ARRAY_ERROR[i];
found->refcount = 1;
break;
}
}
return (akerr_ErrorContext *)NULL;
akerr_mutex_unlock(&akerr_state_lock);
return found;
}
/*
* The wipe returns the slot to the pool, so it and the decrement that triggers
* it are one operation under the lock. Otherwise a thread that saw the count
* reach zero could be handed the slot by akerr_next_error() and start writing
* its error into it while the releasing thread was still memsetting it.
*/
akerr_ErrorContext *akerr_release_error(akerr_ErrorContext *err)
{
int oldid = 0;
akerr_ErrorContext *remaining = err;
akerr_init();
if ( err == NULL ) {
akerr_ErrorContext *errctx = &__akerr_last_ditch;
akerr_ErrorContext *errctx = akerr_last_ditch_context();
FAIL_RETURN(errctx, AKERR_NULLPOINTER, "akerr_release_error got NULL context pointer");
}
akerr_mutex_lock(&akerr_state_lock);
if ( err->refcount > 0 ) {
err->refcount -= 1;
}
@@ -97,21 +396,314 @@ akerr_ErrorContext *akerr_release_error(akerr_ErrorContext *err)
memset(err, 0x00, sizeof(akerr_ErrorContext));
err->stacktracebufptr = (char *)&err->stacktracebuf;
err->arrayid = oldid;
return NULL;
remaining = NULL;
}
return err;
akerr_mutex_unlock(&akerr_state_lock);
return remaining;
}
// returns or sets the name for the given status.
// Call with name = NULL to retrieve a status.
/*
* Scatter the status across the table. Status values are typically dense runs
* (errno 1..N, then a library's block at its base), which linear probing on the
* raw value would pile into one cluster, so mix the bits first.
*/
static unsigned akerr_status_hash(int status)
{
unsigned h = (unsigned)status;
h ^= h >> 16;
h *= 0x85ebca6bu;
h ^= h >> 13;
h *= 0xc2b2ae35u;
h ^= h >> 16;
return h & (unsigned)(AKERR_STATUS_NAME_SLOTS - 1);
}
/*
* Find the slot holding `status`. With create != 0, claim a free slot for it if
* it is not present yet. Returns NULL when the status is absent and either no
* slot was requested or the registry is full. Caller holds akerr_state_lock.
*
* The `& (AKERR_STATUS_NAME_SLOTS - 1)` below is load-bearing and fails
* silently: off by one in either direction and the probe indexes past
* akerr_status_names, writing into whatever BSS follows rather than crashing.
* A test that only counts how many names registered before the table filled
* cannot see that -- a probe sequence collapsed to two slots still registers
* "some" names. Any change to the probe sequence, the occupancy cap, or the
* power-of-two assumption needs a test that reads every entry back by its own
* distinct value; tests/err_maxval.c does.
*/
static akerr_StatusName *akerr_status_slot(int status, int create)
{
unsigned slot = akerr_status_hash(status);
for ( int probe = 0; probe < AKERR_STATUS_NAME_SLOTS; probe++ ) {
akerr_StatusName *entry = &akerr_status_names[slot];
if ( entry->used == 0 ) {
if ( create == 0 ||
akerr_status_name_count >= AKERR_MAX_REGISTERED_STATUS_NAMES ) {
return NULL;
}
entry->used = 1;
entry->status = status;
entry->name[0] = '\0';
akerr_status_name_count++;
return entry;
}
if ( entry->status == status ) {
return entry;
}
slot = (slot + 1u) & (unsigned)(AKERR_STATUS_NAME_SLOTS - 1);
}
/* Unreachable: occupancy is capped below the slot count, so the probe above
* always meets a free slot. Present so a future change to that cap cannot
* turn this into a runaway loop. */
return NULL;
}
/* The reservation covering `status`, or NULL if nobody has claimed it. Caller
* holds akerr_state_lock. */
static akerr_StatusRange *akerr_range_for_status(int status)
{
for ( int i = 0; i < akerr_status_range_count; i++ ) {
if ( status >= akerr_status_ranges[i].first &&
status <= akerr_status_ranges[i].last ) {
return &akerr_status_ranges[i];
}
}
return NULL;
}
/*
* Shared body of both registration entry points. A NULL owner means the caller
* did not identify itself (the legacy two-argument akerr_name_for_status path):
* the status must still lie inside *some* reservation, but we cannot check that
* it is the caller's. Every refusal raises an error -- a name that silently
* fails to register degrades into "Unknown Error" in stack traces, which is
* exactly the kind of quiet loss this registry exists to prevent -- so the
* message carries everything a caller needs to see in a stack trace.
*
* Caller holds akerr_state_lock. The FAIL_* macros below re-enter the library
* to build their error -- a pool slot from akerr_next_error(), and a status
* name for the stack trace -- and that re-entry is why the lock is recursive.
*/
static akerr_ErrorContext AKERR_NOIGNORE *akerr_store_status_name_locked(const char *owner,
int status,
const char *name)
{
akerr_StatusRange *range;
akerr_StatusName *entry;
PREPARE_ERROR(errctx);
FAIL_NONZERO_RETURN(errctx, (name == NULL), AKERR_STATUS_NAME_INVALID,
"Refusing to name status %d for %s: the name is NULL",
status, owner == NULL ? "an unnamed caller" : owner);
FAIL_NONZERO_RETURN(errctx,
(owner != NULL && ( owner[0] == '\0' ||
strlen(owner) >= AKERR_MAX_STATUS_RANGE_OWNER_LENGTH )),
AKERR_STATUS_NAME_INVALID,
"Refusing to name status %d (\"%s\"): the owner string "
"is empty or longer than %d characters",
status, name, AKERR_MAX_STATUS_RANGE_OWNER_LENGTH - 1);
range = akerr_range_for_status(status);
FAIL_ZERO_RETURN(errctx, range, AKERR_STATUS_NAME_UNRESERVED,
"Refusing to name status %d (\"%s\") for %s: no reserved "
"range contains it. Call akerr_reserve_status_range() first.",
status, name, owner == NULL ? "an unnamed caller" : owner);
FAIL_NONZERO_RETURN(errctx,
(owner != NULL && strcmp(owner, range->owner) != 0),
AKERR_STATUS_NAME_FOREIGN,
"Refusing to name status %d (\"%s\") for %s: that status "
"is in range %d..%d owned by %s.",
status, name, owner,
range->first, range->last, range->owner);
entry = akerr_status_slot(status, 1);
FAIL_ZERO_RETURN(errctx, entry, AKERR_STATUS_NAME_FULL,
"Status name registry is full (%d entries); dropping name "
"\"%s\" for status %d. Rebuild libakerror with a larger "
"AKERR_STATUS_NAME_SLOTS.",
AKERR_MAX_REGISTERED_STATUS_NAMES, name, status);
PASS(errctx, __akerr_copy_string(entry->name, AKERR_MAX_ERROR_NAME_LENGTH, name));
SUCCEED_RETURN(errctx);
}
/*
* Register a name for a status inside a range the caller reserved. Both strings
* are checked here rather than only inside akerr_store_status_name_locked(): the
* store accepts a NULL owner for the legacy akerr_name_for_status() path, so a
* NULL arriving through *this* entry point would be read as "caller did not
* identify itself" and skip the ownership check entirely.
*
* Caller holds akerr_state_lock.
*/
static akerr_ErrorContext AKERR_NOIGNORE *akerr_register_status_name_locked(const char *owner,
int status,
const char *name)
{
PREPARE_ERROR(errctx);
FAIL_NONZERO_RETURN(errctx, (owner == NULL), AKERR_STATUS_NAME_INVALID,
"Refusing to name status %d: the owner string is NULL. "
"Pass the same owner you reserved the range with.",
status);
FAIL_NONZERO_RETURN(errctx, (name == NULL), AKERR_STATUS_NAME_INVALID,
"Refusing to name status %d for %s: the name is NULL",
status, owner);
PASS(errctx, akerr_store_status_name_locked(owner, status, name));
SUCCEED_RETURN(errctx);
}
akerr_ErrorContext *akerr_register_status_name(const char *owner, int status, const char *name)
{
akerr_ErrorContext *errctx;
akerr_init();
akerr_mutex_lock(&akerr_state_lock);
errctx = akerr_register_status_name_locked(owner, status, name);
akerr_mutex_unlock(&akerr_state_lock);
return errctx;
}
/*
* Return or set a name. Status magnitude is unrelated to storage size.
*
* The set path is the legacy two-argument form. It returns a name, so it cannot
* hand an error back to its caller and cannot raise: it handles the refusal
* here, converting it to the "Unknown Error" sentinel the way any function that
* must return a value converts a caught error into one.
* akerr_register_status_name() is the form that raises.
*
* The lookup path (name == NULL) deliberately stays clear of all of this. FAIL
* calls it to render a status into a stack trace, so it must not itself need an
* error context.
*
* The name is returned by pointer into the registry, which never resizes and
* never removes an entry, so the pointer is good for the life of the process.
* Its *contents* are stable as long as nobody registers a second name for the
* same status: a rename overwrites the buffer in place, and a lookup on another
* thread can be reading it. Register names during initialization -- renaming a
* live status while other threads run is the one registry operation the lock
* cannot make safe, because the reader is outside it by then.
*/
static akerr_ErrorContext AKERR_NOIGNORE *akerr_store_status_name(const char *owner,
int status,
const char *name)
{
akerr_ErrorContext *errctx;
akerr_mutex_lock(&akerr_state_lock);
errctx = akerr_store_status_name_locked(owner, status, name);
akerr_mutex_unlock(&akerr_state_lock);
return errctx;
}
char *akerr_name_for_status(int status, char *name)
{
if ( status > AKERR_MAX_ERR_VALUE ) {
return "Unknown Error";
}
akerr_StatusName *entry;
char *found = "Unknown Error";
akerr_init();
if ( name != NULL ) {
strncpy((char *)&__AKERR_ERROR_NAMES[status], name, AKERR_MAX_ERROR_NAME_LENGTH);
PREPARE_ERROR(errctx);
int refused = 0;
/* The store takes and releases the lock itself, so the handler below
* calls akerr_log_method -- consumer code, which may do anything at all
* including calling back into this library -- without holding it. */
ATTEMPT {
CATCH(errctx, akerr_store_status_name(NULL, status, name));
} CLEANUP {
} PROCESS(errctx) {
} HANDLE_DEFAULT(errctx) {
LOG_ERROR_WITH_MESSAGE(errctx, "** REFUSED STATUS NAME **");
refused = 1;
} FINISH_NORETURN(errctx);
if ( refused != 0 ) {
return "Unknown Error";
}
}
return (char *)&__AKERR_ERROR_NAMES[status];
akerr_mutex_lock(&akerr_state_lock);
entry = akerr_status_slot(status, 0);
if ( entry != NULL ) {
found = entry->name;
}
akerr_mutex_unlock(&akerr_state_lock);
return found;
}
/* Reserve an inclusive status interval and reject collisions. Caller holds
* akerr_state_lock: the overlap scan and the entry that follows it are one
* decision, so two threads claiming overlapping ranges at once must not be able
* to both find the table clear. */
static akerr_ErrorContext AKERR_NOIGNORE *akerr_reserve_status_range_locked(int first_status,
int count,
const char *owner)
{
int last_status;
PREPARE_ERROR(errctx);
FAIL_NONZERO_RETURN(errctx,
(count <= 0 || owner == NULL || owner[0] == '\0' ||
strlen(owner) >= AKERR_MAX_STATUS_RANGE_OWNER_LENGTH ||
first_status > INT_MAX - (count - 1)),
AKERR_STATUS_RANGE_INVALID,
"Invalid status range reservation: %d status values from "
"%d for %s (count must be positive, the owner string "
"non-empty and shorter than %d characters, and the range "
"must not overflow int)",
count, first_status, owner == NULL ? "(null)" : owner,
AKERR_MAX_STATUS_RANGE_OWNER_LENGTH);
last_status = first_status + count - 1;
for ( int i = 0; i < akerr_status_range_count; i++ ) {
if ( first_status <= akerr_status_ranges[i].last &&
last_status >= akerr_status_ranges[i].first ) {
if ( first_status == akerr_status_ranges[i].first &&
last_status == akerr_status_ranges[i].last &&
strcmp(owner, akerr_status_ranges[i].owner) == 0 ) {
SUCCEED_RETURN(errctx);
}
FAIL_RETURN(errctx, AKERR_STATUS_RANGE_OVERLAP,
"Status range %d..%d requested by %s overlaps %d..%d "
"owned by %s",
first_status, last_status, owner,
akerr_status_ranges[i].first,
akerr_status_ranges[i].last,
akerr_status_ranges[i].owner);
}
}
FAIL_NONZERO_RETURN(errctx,
(akerr_status_range_count == AKERR_MAX_RESERVED_STATUS_RANGES),
AKERR_STATUS_RANGE_FULL,
"Status range table is full (%d ranges); refusing %d..%d "
"for %s.",
AKERR_MAX_RESERVED_STATUS_RANGES,
first_status, last_status, owner);
/* The owner copy is what commits the entry, so the count only advances
* once it has succeeded -- a half-written reservation would claim the
* range under an empty owner nobody could ever match. */
akerr_status_ranges[akerr_status_range_count].first = first_status;
akerr_status_ranges[akerr_status_range_count].last = last_status;
PASS(errctx, __akerr_copy_string(akerr_status_ranges[akerr_status_range_count].owner,
AKERR_MAX_STATUS_RANGE_OWNER_LENGTH, owner));
akerr_status_range_count++;
SUCCEED_RETURN(errctx);
}
akerr_ErrorContext *akerr_reserve_status_range(int first_status, int count, const char *owner)
{
akerr_ErrorContext *errctx;
akerr_init();
akerr_mutex_lock(&akerr_state_lock);
errctx = akerr_reserve_status_range_locked(first_status, count, owner);
akerr_mutex_unlock(&akerr_state_lock);
return errctx;
}

121
src/lock.h Normal file
View File

@@ -0,0 +1,121 @@
#ifndef _AKERR_LOCK_H_
#define _AKERR_LOCK_H_
/*
* Serialization for the library's process-global state: the error pool
* (AKERR_ARRAY_ERROR) and the status registry. Private to the library -- none
* of this appears in the installed header, so the backend is not part of the
* ABI and can be changed without touching a consumer.
*
* The backend is chosen at configure time by the AKERR_THREADS build option and
* never by autodetection here. A build that quietly decided it did not need
* locking is exactly the failure this has to prevent: it would produce a
* library that reports itself thread safe and is not.
*
* AKERR_THREADS_PTHREAD POSIX threads.
* AKERR_THREADS_NONE No locking at all, for a build that has declared
* itself single threaded (-DAKERR_THREADS=none).
*
* One lock covers both tables, and it is recursive. Both are deliberate:
*
* - Raising an error re-enters the library. FAIL() calls
* akerr_name_for_status() to render the status into the stack trace and
* ENSURE_ERROR_READY() to check a context out of the pool, so a refusal
* raised from inside a locked registry operation takes the lock again on
* the same thread. A non-recursive mutex deadlocks there.
* - With a single lock there is no lock ordering to get wrong, and no way for
* a future caller to acquire the pool and the registry in the opposite
* order from this file.
*
* The cost is that error *construction* is serialized across threads. Errors
* are the exceptional path; correctness is worth more there than throughput.
*/
/*
* PTHREAD_MUTEX_RECURSIVE is XSI, so glibc hides it under a strict -std=c99
* without _XOPEN_SOURCE. No feature-test macro is defined here, because the
* public header already needs the same one for PATH_MAX: a build strict enough
* to lose one has already lost the other. Build with -D_XOPEN_SOURCE=700 if you
* need strict C99.
*/
#if defined(AKERR_THREADS_PTHREAD) && AKERR_THREADS_PTHREAD == 1
#include <pthread.h>
#include <stdlib.h>
typedef pthread_mutex_t akerr_Mutex;
typedef pthread_once_t akerr_Once;
#define AKERR_ONCE_INIT PTHREAD_ONCE_INIT
/*
* Terminal on failure. There is no error context to raise into: the pool one
* would come from is the thing this lock protects, and every path that could
* report the failure needs the lock to do it. A process whose error library
* silently stopped locking is worse than one that stops here.
*/
static void akerr_mutex_init(akerr_Mutex *mutex)
{
pthread_mutexattr_t attr;
if ( pthread_mutexattr_init(&attr) != 0 ||
pthread_mutexattr_settype(&attr, PTHREAD_MUTEX_RECURSIVE) != 0 ||
pthread_mutex_init(mutex, &attr) != 0 ) {
abort();
}
pthread_mutexattr_destroy(&attr);
}
static void akerr_mutex_lock(akerr_Mutex *mutex)
{
pthread_mutex_lock(mutex);
}
static void akerr_mutex_unlock(akerr_Mutex *mutex)
{
pthread_mutex_unlock(mutex);
}
static void akerr_once(akerr_Once *once, void (*routine)(void))
{
pthread_once(once, routine);
}
#elif defined(AKERR_THREADS_NONE) && AKERR_THREADS_NONE == 1
typedef char akerr_Mutex;
typedef int akerr_Once;
#define AKERR_ONCE_INIT 0
static void akerr_mutex_init(akerr_Mutex *mutex)
{
(void)mutex;
}
static void akerr_mutex_lock(akerr_Mutex *mutex)
{
(void)mutex;
}
static void akerr_mutex_unlock(akerr_Mutex *mutex)
{
(void)mutex;
}
/*
* The flag is raised before the routine runs, so a routine that calls back into
* akerr_init() sees initialization already in progress and does not recurse --
* the same short-circuit the pthread backend gets from akerr_initializing.
*/
static void akerr_once(akerr_Once *once, void (*routine)(void))
{
if ( *once == 0 ) {
*once = 1;
routine();
}
}
#else
#error "No threading backend selected. Build libakerror through its CMake, which defines AKERR_THREADS_PTHREAD or AKERR_THREADS_NONE from the AKERR_THREADS option."
#endif
#endif // _AKERR_LOCK_H_

6
test.sh Normal file
View File

@@ -0,0 +1,6 @@
cmake -S . -B build
cmake --build build
ctest --test-dir build --output-on-failure --output-junit "$(pwd)/ctest-junit.xml"
scripts/thread_test.sh build/tsan --output-junit "$(pwd)/tsan-junit.xml"
python3 scripts/mutation_test.py --target src/error.c --junit mutation-junit.xml --threshold 65
python3 scripts/coverage.py --junit coverage-junit.xml --threshold 90 --branch-threshold 50

160
tests/MUTATION.md Normal file
View File

@@ -0,0 +1,160 @@
# Mutation testing
The unit tests tell us the library works. **Mutation testing tells us the tests
work** — that they would actually fail if the library were broken.
`scripts/mutation_test.py` deliberately breaks the library in small ways
("mutants"), one at a time, and runs the whole CTest suite against each broken
copy:
* if the tests **fail**, the mutant is **killed** — good, the suite caught it;
* if the tests still **pass**, the mutant **survived** — a bug of that shape
would slip through, so it points at a missing test.
The **mutation score** is `killed / (killed + survived)`. A surviving mutant is
a to-do item: write a test that distinguishes the mutant from the original.
## Running
No third-party tools are required — just Python 3 and the normal
cmake/ctest toolchain. The harness never touches your working tree; it copies
the repo to a scratch directory and mutates the copy.
```sh
# Default: mutate src/error.c and include/akerror.tmpl.h
scripts/mutation_test.py
# Faster: just the C source
scripts/mutation_test.py --target src/error.c
# See what would run without building anything
scripts/mutation_test.py --target src/error.c --list
# Gate CI: exit non-zero if the score drops below 90%
scripts/mutation_test.py --threshold 90
```
Via CMake (configures a build first if needed):
```sh
cmake --build build --target mutation
```
Useful flags: `--timeout SECONDS` (per-suite build+test cap; a mutant that
hangs is counted as killed), `--keep` (retain the scratch copy for debugging),
`--work DIR` (use a specific scratch directory), `--junit FILE` (write a JUnit
XML report — surviving mutants appear as failing test cases).
## CI reporting
Both the unit tests and the mutation run emit JUnit XML that CI consumes:
* `ctest --test-dir build --output-junit "$(pwd)/ctest-junit.xml"` — note the
absolute path; `--output-junit` otherwise resolves relative to the test dir.
* `scripts/mutation_test.py --junit mutation-junit.xml`
`.gitea/workflows/ci.yaml` runs both and feeds the XML to
`mikepenz/action-junit-report` (with `if: always()`, so results publish even
when a gate fails). The reporter runs with `annotate_only: true`: Gitea does not
implement the Checks API the action uses to create a check run, so creating one
404s (mikepenz/action-junit-report#23). `annotate_only` skips that call and the
results surface via the job summary (`detailed_summary: true`) instead. The
generated `*-junit.xml` files are git-ignored.
## Mutation operators
Each mutant changes exactly one location by one of:
| Tag | Operator | Example |
|-----|--------------------------------|----------------------------------|
| ROR | relational operator | `==``!=`, `<``<=`, `>=``>` |
| LCR | logical connector | `&&``\|\|` |
| BCR | boolean constant | `true``false` |
| AOR | arithmetic / compound assign | `+``-`, `+=``-=` |
| ICR | integer literal | `0``1`, `1``0` |
| SDL | statement deletion | `err->refcount += 1;`*(removed)* |
Preprocessor control lines, comments, and the block of error-code / buffer-size
`#define`s are skipped: mutating those produces equivalent or uninteresting
mutants that only add noise.
## Interpreting survivors
Not every survivor is a test gap — some mutants are **equivalent** (they don't
change observable behaviour, e.g. resizing an internal scratch buffer). For each
survivor, decide:
1. **Real gap** → add or strengthen a test in `tests/` so the mutant is killed,
then re-run.
2. **Equivalent mutant** → no test can catch it; leave a note. If a specific
line is a persistent source of equivalents, narrow the target with
`--target` or extend the skip rules in `scripts/mutation_test.py`.
Re-run after adding tests and confirm the score went up.
## Current status
`src/error.c` scores 81.2% — 238 of 293 mutants killed (204 by a failing test,
24 by failing to compile, 10 by hanging the suite), 55 surviving. The CI gate is
set to 65% for headroom.
The ten timeout kills are all in the locking: deleting `akerr_mutex_init()` or
the `akerr_initializing` re-entry guard deadlocks the very first test, which is
the correct behaviour for a broken lock and is why the harness counts a hang as
a kill.
The remaining survivors are dominated by:
* **Equivalent mutants** in `akerr_init`: deleting the `memset`/`NULL` setup of
file-scope statics (`AKERR_ARRAY_ERROR`, `__akerr_last_ditch`,
`__akerr_last_ignored`) changes nothing, because C already zero-initializes
objects with static storage duration. `int oldid = 0;``1` is likewise
dead: it is overwritten before use, and so is clearing `akerr_initializing`
at the end of initialization — nothing reads that flag once the once-routine
has returned.
* **Lock acquisition** (`akerr_mutex_lock`/`unlock` deletions, and the
`akerr_init()` call at the head of an entry point). These are the one category
where a survivor does *not* mean the mutant is harmless. Removing a lock
leaves a real race, and the assertions in `tests/err_threads_pool.c` only fire
when the race actually loses: rebuilding the surviving mutant and running that
test ten times caught it **four** times. The same mutant under
`scripts/thread_test.sh` failed **five of five**, with no false positive on
the unmutated library — but the mutation harness builds without sanitizers, so
it never sees that. Deleting an `akerr_init()` call survives for a duller
reason: something else has always initialized the library by the time that
line runs.
* **Default logger / handler internals** (`vfprintf`, `va_end`, the
`errctx == NULL` branch, `exit(1)`): killing these needs a subprocess-based
test that captures a child's stderr and exit code, rather than the in-process
capturing logger the other tests use.
* **Static assertions** (`akerr_assert_name_slots_pow2` and the occupancy cap
it guards): a mutated compile-time assertion that still compiles has no
runtime behavior to observe. Unkillable by construction — the assertion is
itself the test, and `tests/err_maxval.c` covers the runtime consequence.
* **Hash and probe details** in `akerr_status_slot`: dropping one of the
multiply steps in `akerr_status_hash` leaves a worse but still correct hash,
and probing backwards (`slot - 1u`) is an equally valid sequence over a
power-of-two table. Both are behaviorally equivalent.
* **The `capacity <= 0` guard** in `akerr_copy_string`, which is defensive: both
call sites pass a positive constant.
Findings surfaced by mutation testing:
* **Open:** the harness builds every mutant with the default CMake options, so a
mutant that only breaks under concurrency is judged by a suite running without
ThreadSanitizer. Mutating under `-DAKERR_SANITIZE=thread` would close that,
and needs a way to pass CMake options through to the mutant build. See
"Mutation testing judges concurrency mutants without a sanitizer" in
`TODO.md`.
* **Superseded:** status names now use a private sparse registry, so the old
public `AKERR_MAX_ERR_VALUE` ceiling and its consumer ABI mismatch no longer
exist. `tests/err_maxval.c` covers arbitrary `int` values and registry
exhaustion.
* **Fixed:** the open-addressing probe mask (`& (AKERR_STATUS_NAME_SLOTS - 1)`)
could be mutated to `- 0` or `+ 1` — both of which index past the end of the
table — without any test noticing. `tests/err_maxval.c` only asserted that
*some* names registered before the table filled, which a collapsed probe
sequence still satisfies. It now requires a substantial number of entries and
reads every one of them back by its own distinct name, so a probe that
revisits slots fails on both counts.

View File

@@ -0,0 +1,92 @@
#include "akerror.h"
#include "err_capture.h"
/*
* Cover the conditional break/return failure macros:
* FAIL_ZERO_BREAK / FAIL_NONZERO_BREAK / FAIL_BREAK (inside ATTEMPT)
* FAIL_ZERO_RETURN / FAIL_NONZERO_RETURN / FAIL_RETURN (direct return)
* Each is checked in both its failing and its non-failing (pass-through) state.
*/
akerr_ErrorContext *zero_break(int x)
{
PREPARE_ERROR(e);
ATTEMPT {
FAIL_ZERO_BREAK(e, x, AKERR_VALUE, "x was zero");
} CLEANUP {
} PROCESS(e) {
} FINISH(e, true);
SUCCEED_RETURN(e);
}
akerr_ErrorContext *nonzero_break(int x)
{
PREPARE_ERROR(e);
ATTEMPT {
FAIL_NONZERO_BREAK(e, x, AKERR_INDEX, "x was nonzero");
} CLEANUP {
} PROCESS(e) {
} FINISH(e, true);
SUCCEED_RETURN(e);
}
akerr_ErrorContext *always_break(void)
{
PREPARE_ERROR(e);
ATTEMPT {
FAIL_BREAK(e, AKERR_IO, "always");
} CLEANUP {
} PROCESS(e) {
} FINISH(e, true);
SUCCEED_RETURN(e);
}
akerr_ErrorContext *zero_return(int x)
{
PREPARE_ERROR(e);
FAIL_ZERO_RETURN(e, x, AKERR_KEY, "x was zero");
SUCCEED_RETURN(e);
}
akerr_ErrorContext *nonzero_return(int x)
{
PREPARE_ERROR(e);
FAIL_NONZERO_RETURN(e, x, AKERR_TYPE, "x was nonzero");
SUCCEED_RETURN(e);
}
int main(void)
{
akerr_capture_install();
akerr_init();
akerr_ErrorContext *r;
r = zero_break(0);
AKERR_CHECK_STATUS(r, AKERR_VALUE);
r = akerr_release_error(r);
AKERR_CHECK(zero_break(7) == NULL);
r = nonzero_break(7);
AKERR_CHECK_STATUS(r, AKERR_INDEX);
r = akerr_release_error(r);
AKERR_CHECK(nonzero_break(0) == NULL);
r = always_break();
AKERR_CHECK_STATUS(r, AKERR_IO);
r = akerr_release_error(r);
r = zero_return(0);
AKERR_CHECK_STATUS(r, AKERR_KEY);
r = akerr_release_error(r);
AKERR_CHECK(zero_return(7) == NULL);
r = nonzero_return(7);
AKERR_CHECK_STATUS(r, AKERR_TYPE);
r = akerr_release_error(r);
AKERR_CHECK(nonzero_return(0) == NULL);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_break_variants ok\n");
return 0;
}

131
tests/err_capture.h Normal file
View File

@@ -0,0 +1,131 @@
#ifndef AKERR_TEST_CAPTURE_H
#define AKERR_TEST_CAPTURE_H
/*
* Shared test helpers for libakerror.
*
* Installs a capturing implementation of akerr_log_method so tests can assert
* on the *content* of log/stacktrace output (messages, status codes, error
* names) instead of relying solely on process exit codes.
*
* Also provides AKERR_CHECK(), a NDEBUG-proof assertion that fails the test by
* returning non-zero from main() (unlike assert(), which is compiled out in
* release builds and would silently turn a test into a no-op).
*/
#include "akerror.h"
#include <stdarg.h>
#include <stdio.h>
#include <string.h>
#define AKERR_CAPTURE_BUFSZ 65536
static char akerr_capture_buf[AKERR_CAPTURE_BUFSZ];
static size_t akerr_capture_len = 0;
static void __attribute__((unused)) akerr_capture_logger(const char *fmt, ...)
{
va_list ap;
va_start(ap, fmt);
int n = vsnprintf(akerr_capture_buf + akerr_capture_len,
AKERR_CAPTURE_BUFSZ - akerr_capture_len, fmt, ap);
va_end(ap);
if ( n > 0 ) {
akerr_capture_len += (size_t)n;
if ( akerr_capture_len >= AKERR_CAPTURE_BUFSZ ) {
akerr_capture_len = AKERR_CAPTURE_BUFSZ - 1;
}
}
}
static void __attribute__((unused)) akerr_capture_reset(void)
{
akerr_capture_len = 0;
akerr_capture_buf[0] = '\0';
}
/*
* Install the capturing logger. akerr_init() only assigns a default logger when
* akerr_log_method is NULL, and it is idempotent, so calling this either before
* or after the first PREPARE_ERROR keeps our logger in place.
*/
static void __attribute__((unused)) akerr_capture_install(void)
{
akerr_capture_reset();
akerr_log_method = &akerr_capture_logger;
}
/* Count array slots currently checked out of the pool (refcount != 0). */
static int __attribute__((unused)) akerr_slots_in_use(void)
{
int n = 0;
for ( int i = 0; i < AKERR_MAX_ARRAY_ERROR; i++ ) {
if ( AKERR_ARRAY_ERROR[i].refcount != 0 ) {
n++;
}
}
return n;
}
#define AKERR_CHECK(cond) \
do { \
if ( !(cond) ) { \
fprintf(stderr, "CHECK FAILED: %s at %s:%d\n", \
#cond, __FILE__, __LINE__); \
return 1; \
} \
} while ( 0 )
#define AKERR_CHECK_STATUS(errctx, expected_status) \
do { \
AKERR_CHECK((errctx) != NULL); \
AKERR_CHECK((errctx)->status == (expected_status)); \
} while ( 0 )
/*
* Helpers for functions that report failure by returning akerr_ErrorContext *.
* Both release the context they consume, so a test that makes thousands of
* failing calls cannot exhaust the pool. AKERR_CHECK_RAISES keeps a copy of the
* message for AKERR_CHECK_MESSAGE_CONTAINS, since the context is gone by then.
*/
static char __attribute__((unused)) akerr_last_message[AKERR_MAX_ERROR_CONTEXT_STRING_LENGTH];
#define AKERR_CHECK_SUCCEEDS(expr) \
do { \
akerr_ErrorContext *__akerr_result = (expr); \
if ( __akerr_result != NULL ) { \
fprintf(stderr, "UNEXPECTED ERROR from %s: %d (%s): %s" \
" at %s:%d\n", #expr, __akerr_result->status, \
akerr_name_for_status(__akerr_result->status, NULL), \
__akerr_result->message, __FILE__, __LINE__); \
RELEASE_ERROR(__akerr_result); \
return 1; \
} \
} while ( 0 )
#define AKERR_CHECK_RAISES(expr, expected_status) \
do { \
akerr_ErrorContext *__akerr_result = (expr); \
AKERR_CHECK(__akerr_result != NULL); \
snprintf(akerr_last_message, sizeof(akerr_last_message), "%s", \
__akerr_result->message); \
if ( __akerr_result->status != (expected_status) ) { \
fprintf(stderr, "WRONG STATUS from %s: got %d, want %s" \
" at %s:%d\n", #expr, __akerr_result->status, \
#expected_status, __FILE__, __LINE__); \
RELEASE_ERROR(__akerr_result); \
return 1; \
} \
RELEASE_ERROR(__akerr_result); \
AKERR_CHECK(__akerr_result == NULL); \
} while ( 0 )
#define AKERR_CHECK_MESSAGE_CONTAINS(needle) \
AKERR_CHECK(strstr(akerr_last_message, (needle)) != NULL)
#define AKERR_CHECK_CONTAINS(needle) \
AKERR_CHECK(strstr(akerr_capture_buf, (needle)) != NULL)
#define AKERR_CHECK_NOT_CONTAINS(needle) \
AKERR_CHECK(strstr(akerr_capture_buf, (needle)) == NULL)
#endif // AKERR_TEST_CAPTURE_H

View File

@@ -1,4 +1,5 @@
#include "akerror.h"
#include "err_capture.h"
akerr_ErrorContext *func2(void)
{
@@ -31,6 +32,11 @@ int main(void)
} CLEANUP {
} PROCESS(errctx) {
} HANDLE(errctx, AKERR_NULLPOINTER) {
akerr_log_method("Caught exception");
AKERR_CHECK_STATUS(errctx, AKERR_NULLPOINTER);
akerr_log_method("Caught exception");
} FINISH_NORETURN(errctx);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_catch ok\n");
return 0;
}

View File

@@ -1,4 +1,5 @@
#include "akerror.h"
#include "err_capture.h"
int x;
@@ -34,10 +35,15 @@ int main(void)
} CLEANUP {
} PROCESS(errctx) {
} HANDLE(errctx, AKERR_NULLPOINTER) {
AKERR_CHECK_STATUS(errctx, AKERR_NULLPOINTER);
if ( x == 0 ) {
fprintf(stderr, "Cleanup works\n");
return 0;
akerr_log_method("Cleanup works\n");
} else {
return 1;
}
return 1;
} FINISH_NORETURN(errctx);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_cleanup ok\n");
return 0;
}

58
tests/err_copy_string.c Normal file
View File

@@ -0,0 +1,58 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* __akerr_copy_string() is the only place in the library that writes through a
* caller-supplied pointer for a caller-supplied length, so it validates every
* argument and raises rather than returning quietly. Both in-library callers
* check their arguments before calling it, so this test is what drives those
* guards -- without it they are unreachable code that no build ever exercises.
*/
int main(void)
{
char buf[8];
akerr_capture_install();
akerr_init();
/* The happy path: bounded, and always terminated. */
memset(buf, 'x', sizeof(buf));
AKERR_CHECK_SUCCEEDS(__akerr_copy_string(buf, (int)sizeof(buf), "abc"));
AKERR_CHECK(strcmp(buf, "abc") == 0);
/* A source longer than the buffer is truncated, never overrun. */
memset(buf, 'x', sizeof(buf));
AKERR_CHECK_SUCCEEDS(__akerr_copy_string(buf, (int)sizeof(buf), "abcdefghijkl"));
AKERR_CHECK(strlen(buf) == sizeof(buf) - 1);
AKERR_CHECK(buf[sizeof(buf) - 1] == '\0');
AKERR_CHECK(strcmp(buf, "abcdefg") == 0);
/* A capacity of exactly one holds nothing but the terminator. */
memset(buf, 'x', sizeof(buf));
AKERR_CHECK_SUCCEEDS(__akerr_copy_string(buf, 1, "abc"));
AKERR_CHECK(buf[0] == '\0');
AKERR_CHECK(buf[1] == 'x'); /* and wrote nothing past its capacity */
/* NULL pointers raise instead of faulting, and the message says which. */
AKERR_CHECK_RAISES(__akerr_copy_string(NULL, (int)sizeof(buf), "abc"),
AKERR_NULLPOINTER);
AKERR_CHECK_MESSAGE_CONTAINS("destination");
AKERR_CHECK_RAISES(__akerr_copy_string(buf, (int)sizeof(buf), NULL),
AKERR_NULLPOINTER);
AKERR_CHECK_MESSAGE_CONTAINS("source");
/* A capacity with no room for a terminator is a value error, not a write. */
memset(buf, 'x', sizeof(buf));
AKERR_CHECK_RAISES(__akerr_copy_string(buf, 0, "abc"), AKERR_VALUE);
AKERR_CHECK_MESSAGE_CONTAINS("capacity of 0");
AKERR_CHECK_RAISES(__akerr_copy_string(buf, -1, "abc"), AKERR_VALUE);
AKERR_CHECK(buf[0] == 'x'); /* nothing was written */
/* Each refusal handed its context back to the pool. */
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_copy_string ok\n");
return 0;
}

View File

@@ -0,0 +1,45 @@
#include "akerror.h"
#include "err_capture.h"
/*
* The unhandled-error hook (akerr_handler_unhandled_error) is overridable.
* Install a non-fatal handler so an unhandled error can be asserted directly
* (status, invocation) instead of relying on process death / WILL_FAIL.
*/
static int custom_fired = 0;
static int custom_status = 0;
static void my_handler(akerr_ErrorContext *e)
{
custom_fired = 1;
custom_status = (e != NULL) ? e->status : -1;
/* deliberately does NOT exit() */
}
akerr_ErrorContext *boom(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_TYPE, "unhandled on purpose");
}
int main(void)
{
akerr_capture_install();
akerr_init(); /* sets the default handler... */
akerr_handler_unhandled_error = &my_handler; /* ...which we then override */
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, boom());
} CLEANUP {
} PROCESS(e) {
/* no HANDLE for AKERR_TYPE -> stays unhandled */
} FINISH_NORETURN(e);
AKERR_CHECK(custom_fired == 1);
AKERR_CHECK(custom_status == AKERR_TYPE);
AKERR_CHECK_CONTAINS("Unhandled Error");
fprintf(stderr, "err_custom_handler ok\n");
return 0;
}

48
tests/err_errno.c Normal file
View File

@@ -0,0 +1,48 @@
#include "akerror.h"
#include "err_capture.h"
#include <errno.h>
/*
* The library imports system errno codes and their descriptions at build time
* (scripts/generrno.sh -> akerr_init_errno). Verify:
* - a system errno (EACCES) has a registered, non-empty name;
* - an unregistered status returns the "Unknown Error" sentinel;
* - a system errno can be raised, propagated and handled like any AKERR_* code.
*/
static int handled = 0;
akerr_ErrorContext *boom(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, EACCES, "permission denied");
}
int main(void)
{
akerr_capture_install();
akerr_init();
char *nm = akerr_name_for_status(EACCES, NULL);
AKERR_CHECK(nm != NULL);
AKERR_CHECK(nm[0] != '\0');
AKERR_CHECK(strcmp(nm, "Unknown Error") != 0);
AKERR_CHECK(strcmp(akerr_name_for_status(1000000, NULL),
"Unknown Error") == 0);
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, boom());
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, EACCES) {
AKERR_CHECK_STATUS(e, EACCES);
handled = 1;
} FINISH_NORETURN(e);
AKERR_CHECK(handled == 1);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_errno ok (EACCES name: \"%s\")\n", nm);
return 0;
}

83
tests/err_error_names.c Normal file
View File

@@ -0,0 +1,83 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* akerr_init() registers a human-readable name for each library error code.
* Verify the names are actually installed (mutation testing showed the
* registration calls could be deleted without any test noticing).
*
* This list must stay exhaustive. AKERR_EOF, AKERR_ITERATOR_BREAK and
* AKERR_NOT_IMPLEMENTED were previously valid codes with no registered name, so
* they rendered as "Unknown Error" in every stack trace that carried them --
* the same class of silent gap that a too-small AKERR_MAX_ERR_VALUE used to
* cause. The sweep below walks the whole AKERR_* offset span so a newly added
* code without a name fails here rather than showing up in production traces.
*/
static const struct {
int code;
const char *name;
} expected[] = {
{ AKERR_NULLPOINTER, "Null Pointer Error" },
{ AKERR_OUTOFBOUNDS, "Out Of Bounds Error" },
{ AKERR_API, "API Error" },
{ AKERR_ATTRIBUTE, "Attribute Error" },
{ AKERR_TYPE, "Type Error" },
{ AKERR_KEY, "Key Error" },
{ AKERR_INDEX, "Index Error" },
{ AKERR_FORMAT, "Format Error" },
{ AKERR_IO, "Input Output Error" },
{ AKERR_VALUE, "Value Error" },
{ AKERR_RELATIONSHIP, "Relationship Error" },
{ AKERR_EOF, "End Of File" },
{ AKERR_CIRCULAR_REFERENCE, "Circular Reference Error" },
{ AKERR_ITERATOR_BREAK, "Iterator Break" },
{ AKERR_NOT_IMPLEMENTED, "Not Implemented" },
{ AKERR_BADEXC, "Invalid akerr_ErrorContext" },
{ AKERR_STATUS_RANGE_OVERLAP, "Status Range Overlap" },
{ AKERR_STATUS_RANGE_FULL, "Status Range Table Full" },
{ AKERR_STATUS_RANGE_INVALID, "Invalid Status Range" },
{ AKERR_STATUS_NAME_UNRESERVED, "Unreserved Status Name" },
{ AKERR_STATUS_NAME_FOREIGN, "Foreign Status Name" },
{ AKERR_STATUS_NAME_FULL, "Status Name Registry Full" },
{ AKERR_STATUS_NAME_INVALID, "Invalid Status Name" },
};
int main(void)
{
akerr_init();
for ( unsigned i = 0; i < sizeof(expected) / sizeof(expected[0]); i++ ) {
char *nm = akerr_name_for_status(expected[i].code, NULL);
AKERR_CHECK(nm != NULL);
AKERR_CHECK(strcmp(nm, expected[i].name) == 0);
}
/*
* Every value in the library's own offset span must resolve to a real name.
* AKERR_LAST_ERRNO_VALUE + 7 is the one deliberate hole (a removed code);
* anything else nameless is a code someone added without registering it.
*/
for ( int offset = 1;
offset <= AKERR_LAST_LIBRARY_STATUS - AKERR_LAST_ERRNO_VALUE;
offset++ ) {
int code = AKERR_LAST_ERRNO_VALUE + offset;
char *nm = akerr_name_for_status(code, NULL);
if ( offset == 7 ) {
AKERR_CHECK(strcmp(nm, "Unknown Error") == 0);
continue;
}
if ( strcmp(nm, "Unknown Error") == 0 || nm[0] == '\0' ) {
fprintf(stderr, "AKERR_LAST_ERRNO_VALUE + %d (%d) has no name\n",
offset, code);
return 1;
}
}
/* Every AKERR_* code must sit inside the band the library reserves. */
AKERR_CHECK(AKERR_LAST_LIBRARY_STATUS < AKERR_FIRST_CONSUMER_STATUS);
fprintf(stderr, "err_error_names ok\n");
return 0;
}

206
tests/err_exit_status.c Normal file
View File

@@ -0,0 +1,206 @@
#include "akerror.h"
#include "err_capture.h"
#include <unistd.h>
#include <sys/wait.h>
/*
* An unhandled error must never leave the process looking like a success.
*
* The default handler used to exit(errctx->status) unconditionally, and an exit
* status is one byte wide: status 256 -- AKERR_FIRST_CONSUMER_STATUS, the very
* first code any consumer can reserve -- exited 0 and told the shell the
* program succeeded. Status 300 exited 44, which is some unrelated error's code.
*
* akerr_exit() now owns that mapping, and the default handler is one of its
* callers. The same table therefore drives both: a status must produce the same
* exit code whether a consumer calls akerr_exit() from their own handler or
* lets the library's handler run.
*
* akerr_exit(0) exits 0 -- zero is the library's success status. The thing that
* keeps an unhandled error off that path is PROCESS's `case 0`, which marks a
* zero status handled before FINISH_NORETURN can reach the handler, and the
* last three cases assert that, plus that the status an exit code could not
* carry is still recoverable from the stack trace.
*
* Neither exit path returns, so those cases run in forked children.
*/
#define TEST_OWNER "err_exit_status"
#define TEST_STATUS AKERR_FIRST_CONSUMER_STATUS
#define TEST_STATUS_NAME "Consumer Status Two Fifty Six"
static const struct {
int status;
int expect;
} exit_cases[] = {
/* status expected exit */
{ 0, 0 }, /* Zero is the library's success status, and exit code 0 is what that is called out here */
{ 1, 1 }, /* Lowest status an exit code can carry */
{ AKERR_VALUE, AKERR_VALUE }, /* An ordinary library status, delivered intact */
{ AKERR_EXIT_STATUS_MAX, AKERR_EXIT_STATUS_MAX }, /* Highest status an exit code can carry */
{ AKERR_FIRST_CONSUMER_STATUS, AKERR_EXIT_STATUS_UNREPRESENTABLE }, /* 256: low byte 0, the case that used to exit success */
{ 300, AKERR_EXIT_STATUS_UNREPRESENTABLE }, /* Low byte 44 would alias an unrelated status */
{ 65536, AKERR_EXIT_STATUS_UNREPRESENTABLE }, /* Low byte 0 again, further out */
{ -1, AKERR_EXIT_STATUS_UNREPRESENTABLE }, /* Negative: low byte 255 */
};
/*
* Run body in a child and return its exit status, or -1 if it did not exit
* normally. The 99 sentinel catches a body that returns instead of terminating.
*/
static int child_exit_status(void (*body)(void))
{
pid_t pid = fork();
if ( pid == 0 ) {
body();
_exit(99);
}
int status = 0;
if ( pid < 0 || waitpid(pid, &status, 0) != pid || !WIFEXITED(status) ) {
return -1;
}
return WEXITSTATUS(status);
}
/* The status the next child leaves with. Set before each fork. */
static int pending_status;
/* The way a consumer's own handler is expected to leave. */
static void run_akerr_exit(void)
{
akerr_exit(pending_status);
}
/* The way the library leaves when nothing handled the error. */
static void run_default_handler(void)
{
akerr_ErrorContext *slot = akerr_next_error();
if ( slot == NULL ) {
_exit(98);
}
slot->status = pending_status;
akerr_default_handler_unhandled_error(slot);
}
static akerr_ErrorContext AKERR_NOIGNORE *raise_consumer_error(void)
{
PREPARE_ERROR(errctx);
FAIL_RETURN(errctx, TEST_STATUS, "consumer status, deliberately unhandled");
}
static akerr_ErrorContext AKERR_NOIGNORE *raise_nothing(void)
{
PREPARE_ERROR(errctx);
SUCCEED_RETURN(errctx);
}
/* A full propagation to the top of the stack, with nothing handling it. */
static void unhandled_consumer_error(void)
{
PREPARE_ERROR(errctx);
ATTEMPT {
CATCH(errctx, raise_consumer_error());
} CLEANUP {
} PROCESS(errctx) {
/* no HANDLE for TEST_STATUS -> stays unhandled */
} FINISH_NORETURN(errctx);
}
/*
* Reached through FINISH_NORETURN rather than by calling the handler directly,
* so the child exercises the path a consumer actually takes. The handler is set
* explicitly because an earlier case in this process may have replaced it.
*/
static void unhandled_consumer_error_fatal(void)
{
akerr_handler_unhandled_error = &akerr_default_handler_unhandled_error;
unhandled_consumer_error();
}
static int trace_fired = -2;
static void nonfatal_handler(akerr_ErrorContext *e)
{
trace_fired = (e != NULL) ? e->status : -1;
}
/*
* The same shape as unhandled_consumer_error(), except nothing fails. PROCESS
* opens with `case 0`, so a zero status is handled and the handler must not
* run -- which is what keeps akerr_exit(0) exiting 0 from being a hole in the
* "an unhandled error never exits 0" rule.
*/
static void successful_operation(void)
{
PREPARE_ERROR(errctx);
ATTEMPT {
CATCH(errctx, raise_nothing());
} CLEANUP {
} PROCESS(errctx) {
} FINISH_NORETURN(errctx);
}
int main(void)
{
akerr_capture_install();
akerr_init();
/* The three outcomes have to stay distinguishable from each other. */
AKERR_CHECK(AKERR_EXIT_STATUS_UNREPRESENTABLE != 0);
AKERR_CHECK(AKERR_EXIT_STATUS_UNREPRESENTABLE != 1);
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(TEST_STATUS, 1, TEST_OWNER));
AKERR_CHECK_SUCCEEDS(akerr_register_status_name(TEST_OWNER, TEST_STATUS,
TEST_STATUS_NAME));
for ( size_t i = 0; i < sizeof(exit_cases) / sizeof(exit_cases[0]); i++ ) {
pending_status = exit_cases[i].status;
int direct = child_exit_status(&run_akerr_exit);
if ( direct != exit_cases[i].expect ) {
fprintf(stderr, "akerr_exit(%d) exited %d, want %d\n",
exit_cases[i].status, direct, exit_cases[i].expect);
return 1;
}
/* The handler must not carry a mapping of its own. */
int handled = child_exit_status(&run_default_handler);
if ( handled != direct ) {
fprintf(stderr, "default handler on status %d exited %d,"
" but akerr_exit(%d) exited %d\n",
exit_cases[i].status, handled, exit_cases[i].status, direct);
return 1;
}
}
/* The NULL-context exit is asserted by tests/err_unhandled_null.c. */
/* End to end: an unhandled consumer error kills the process non-zero. */
AKERR_CHECK(child_exit_status(&unhandled_consumer_error_fatal)
== AKERR_EXIT_STATUS_UNREPRESENTABLE);
/*
* And the status the exit code could not carry is in the trace. Run with a
* handler that returns so the assertions happen in this process, where the
* captured log lives.
*/
akerr_handler_unhandled_error = &nonfatal_handler;
akerr_capture_reset();
unhandled_consumer_error();
AKERR_CHECK(trace_fired == TEST_STATUS);
AKERR_CHECK_CONTAINS("Unhandled Error");
AKERR_CHECK_CONTAINS("256");
AKERR_CHECK_CONTAINS(TEST_STATUS_NAME);
/* A zero status is handled by PROCESS and never reaches the handler. */
trace_fired = -2;
akerr_capture_reset();
successful_operation();
AKERR_CHECK(trace_fired == -2);
AKERR_CHECK_NOT_CONTAINS("Unhandled Error");
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_exit_status ok\n");
return 0;
}

34
tests/err_format_string.c Normal file
View File

@@ -0,0 +1,34 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* FAIL records the source file and function names with snprintf. Those names
* must be passed as %s ARGUMENTS, not used as the format string -- otherwise a
* path containing a printf conversion (say a build directory with a '%') is
* interpreted as a format and reads nonexistent varargs (undefined behavior).
*
* #line lets us make __FILE__ contain a conversion specifier; the stored name
* must come back verbatim.
*/
#line 1 "pct%dname.c"
akerr_ErrorContext *raise_with_percent_in_filename(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_VALUE, "boom");
}
#line 22 "tests/err_format_string.c"
int main(void)
{
akerr_init();
akerr_ErrorContext *e = raise_with_percent_in_filename();
AKERR_CHECK(e != NULL);
AKERR_CHECK(strcmp(e->fname, "pct%dname.c") == 0);
e = akerr_release_error(e);
fprintf(stderr, "err_format_string ok\n");
return 0;
}

View File

@@ -0,0 +1,41 @@
#include "akerror.h"
#include "err_capture.h"
/*
* An error whose status has no matching HANDLE block must fall through to
* HANDLE_DEFAULT, and the non-matching HANDLE body must not run.
*/
static int specific_fired = 0;
static int default_fired = 0;
static int default_status = 0;
akerr_ErrorContext *boom(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_TYPE, "wrong type");
}
int main(void)
{
akerr_capture_install();
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, boom());
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, AKERR_KEY) {
specific_fired = 1; /* must NOT run: error is AKERR_TYPE */
} HANDLE_DEFAULT(e) {
default_fired = 1;
default_status = e->status;
} FINISH_NORETURN(e);
AKERR_CHECK(specific_fired == 0);
AKERR_CHECK(default_fired == 1);
AKERR_CHECK(default_status == AKERR_TYPE);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_handle_default ok\n");
return 0;
}

View File

@@ -0,0 +1,45 @@
#include "akerror.h"
#include "err_capture.h"
/*
* With several distinct HANDLE blocks, only the one matching the raised status
* may run. Raise the middle code and confirm exact dispatch.
*/
static int a_fired = 0;
static int b_fired = 0;
static int c_fired = 0;
static int b_status = 0;
akerr_ErrorContext *boom(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_TYPE, "the middle one");
}
int main(void)
{
akerr_capture_install();
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, boom());
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, AKERR_KEY) {
a_fired = 1;
} HANDLE(e, AKERR_TYPE) {
b_fired = 1;
b_status = e->status;
} HANDLE(e, AKERR_IO) {
c_fired = 1;
} FINISH_NORETURN(e);
AKERR_CHECK(a_fired == 0);
AKERR_CHECK(b_fired == 1);
AKERR_CHECK(b_status == AKERR_TYPE);
AKERR_CHECK(c_fired == 0);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_handle_dispatch ok\n");
return 0;
}

63
tests/err_handle_group.c Normal file
View File

@@ -0,0 +1,63 @@
#include "akerror.h"
#include "err_capture.h"
/*
* HANDLE_GROUP lets several status codes share one handler body via
* case-fallthrough. The first member of the group uses HANDLE (which emits the
* leading break that terminates the previous case); each additional member
* uses HANDLE_GROUP; the shared body follows the last member. Verify that two
* different status codes both reach the shared body.
*/
static int group_fired = 0;
static int other_fired = 0;
static int group_status = 0;
static int other_status = 0;
akerr_ErrorContext *boom(int status)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, status, "boom %d", status);
}
akerr_ErrorContext *run(int status)
{
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, boom(status));
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, AKERR_KEY)
HANDLE_GROUP(e, AKERR_INDEX) {
group_fired++;
group_status = e->status;
} HANDLE(e, AKERR_IO) {
other_fired++;
other_status = e->status;
} FINISH(e, false);
return e;
}
int main(void)
{
akerr_capture_install();
run(AKERR_KEY);
AKERR_CHECK(group_fired == 1);
AKERR_CHECK(group_status == AKERR_KEY);
AKERR_CHECK(other_fired == 0);
run(AKERR_INDEX);
AKERR_CHECK(group_fired == 2);
AKERR_CHECK(group_status == AKERR_INDEX);
AKERR_CHECK(other_fired == 0);
run(AKERR_IO);
AKERR_CHECK(group_fired == 2);
AKERR_CHECK(other_fired == 1);
AKERR_CHECK(other_status == AKERR_IO);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_handle_group ok\n");
return 0;
}

34
tests/err_ignore.c Normal file
View File

@@ -0,0 +1,34 @@
#include "akerror.h"
#include "err_capture.h"
/*
* IGNORE deliberately swallows an error: it records the context in
* __akerr_last_ignored, logs it with an "IGNORED ERROR" marker, and lets
* execution continue.
*/
akerr_ErrorContext *boom(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_VALUE, "this error is ignored on purpose");
}
int main(void)
{
akerr_capture_install();
int reached_after_ignore = 0;
PREPARE_ERROR(e);
(void)e;
IGNORE(boom());
reached_after_ignore = 1;
AKERR_CHECK(__akerr_last_ignored != NULL);
AKERR_CHECK(__akerr_last_ignored->status == AKERR_VALUE);
AKERR_CHECK(reached_after_ignore == 1);
AKERR_CHECK_CONTAINS("IGNORED ERROR");
AKERR_CHECK_CONTAINS("this error is ignored on purpose");
fprintf(stderr, "err_ignore ok\n");
return 0;
}

View File

@@ -0,0 +1,22 @@
#include "akerror.h"
#include <stdio.h>
akerr_ErrorContext AKERR_NOIGNORE *improper_closure(void)
{
PREPARE_ERROR(errctx);
ATTEMPT {
} CLEANUP {
} PROCESS(errctx) {
} FINISH(errctx, true);
fprintf(stderr, "Improperly returning from improper_closure\n");
}
int main(void)
{
PREPARE_ERROR(errctx);
ATTEMPT {
CATCH(errctx, improper_closure());
} CLEANUP {
} PROCESS(errctx) {
} FINISH_NORETURN(errctx);
}

View File

@@ -0,0 +1,31 @@
#include "akerror.h"
#include "err_capture.h"
/*
* The library naming one of its own codes is not allowed to fail quietly.
* __akerr_name_library_status() runs from akerr_init() and from the generated
* errno table, neither of which has a caller to raise into, so a refusal there
* goes through FINISH_NORETURN: stack trace, then akerr_handler_unhandled_error,
* which terminates the process.
*
* In a correct build that can only happen with a name table too small to hold
* the library's own entries, which no test can configure (the slot count is
* PRIVATE to the library target). Calling the helper for a status the library
* does not own reaches the same refusal, so this test covers the terminal path
* itself.
*
* Registered in AKERR_WILL_FAIL_TESTS: reaching the end of main() means the
* failure was swallowed, and that is the bug this test exists to catch.
*/
int main(void)
{
akerr_init();
/* Nobody has reserved 9999, so this registration is refused. */
__akerr_name_library_status(9999, "Not The Library's To Name");
fprintf(stderr, "err_library_status_fatal: a refused library-status "
"registration did NOT terminate\n");
return 0;
}

165
tests/err_maxval.c Normal file
View File

@@ -0,0 +1,165 @@
#include "akerror.h"
#include "err_capture.h"
#include <limits.h>
#include <string.h>
/*
* Status magnitude is no longer coupled to a public array bound: any int is a
* legal status, and storage is a private sparse registry. What bounds the
* registry now is its *capacity*, not the value of the largest code.
*
* Covers: arbitrary int status values, name truncation, range reservation
* semantics (overlap, idempotency, endpoints, validation, overflow), and both
* capacity limits -- the range table and the name table -- each of which must
* raise rather than dropping the registration quietly.
*
* Every refusal is an error context the caller owns, so each check below also
* asserts the context returns to the pool; a leak here would exhaust the
* 128-slot pool long before these loops finish.
*/
int main(void)
{
akerr_capture_install();
akerr_init();
/* Any int is a legal status, at either extreme of the range. */
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(INT_MIN, 1, "min-owner"));
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(INT_MAX, 1, "max-owner"));
AKERR_CHECK_SUCCEEDS(akerr_register_status_name("max-owner", INT_MAX, "Maximum Status"));
AKERR_CHECK_SUCCEEDS(akerr_register_status_name("min-owner", INT_MIN, "Minimum Status"));
AKERR_CHECK(strcmp(akerr_name_for_status(INT_MAX, NULL), "Maximum Status") == 0);
AKERR_CHECK(strcmp(akerr_name_for_status(INT_MIN, NULL), "Minimum Status") == 0);
/* A name longer than the buffer is truncated and always terminated. */
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(1000000, 1, "trunc"));
const char *long_name =
"abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ0123456789-extra";
AKERR_CHECK_SUCCEEDS(akerr_register_status_name("trunc", 1000000, long_name));
char *stored = akerr_name_for_status(1000000, NULL);
AKERR_CHECK(strlen(stored) == AKERR_MAX_ERROR_NAME_LENGTH - 1);
AKERR_CHECK(stored[AKERR_MAX_ERROR_NAME_LENGTH - 1] == '\0');
/* Reservation: overlap detection, and idempotency for an exact repeat. */
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(256, 16, "component-a"));
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(256, 16, "component-a"));
AKERR_CHECK_RAISES(akerr_reserve_status_range(260, 2, "component-b"),
AKERR_STATUS_RANGE_OVERLAP);
AKERR_CHECK_MESSAGE_CONTAINS("component-a");
/* The library's own 0..255 band is reserved and cannot be encroached on. */
AKERR_CHECK_RAISES(akerr_reserve_status_range(255, 1, "component-b"),
AKERR_STATUS_RANGE_OVERLAP);
AKERR_CHECK_MESSAGE_CONTAINS(AKERR_LIBRARY_OWNER);
/* An exact repeat by the owner is still a no-op. */
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(0, AKERR_RESERVED_STATUS_COUNT,
AKERR_LIBRARY_OWNER));
/* Argument validation. */
AKERR_CHECK_RAISES(akerr_reserve_status_range(INT_MAX, 2, "overflow"),
AKERR_STATUS_RANGE_INVALID);
AKERR_CHECK_RAISES(akerr_reserve_status_range(300, 0, "empty"),
AKERR_STATUS_RANGE_INVALID);
AKERR_CHECK_MESSAGE_CONTAINS("empty"); /* the message names the caller */
AKERR_CHECK_RAISES(akerr_reserve_status_range(300, -1, "negative"),
AKERR_STATUS_RANGE_INVALID);
AKERR_CHECK_RAISES(akerr_reserve_status_range(300, 1, NULL),
AKERR_STATUS_RANGE_INVALID);
AKERR_CHECK_RAISES(akerr_reserve_status_range(300, 1, ""),
AKERR_STATUS_RANGE_INVALID);
/* Owner strings: 63 chars fit, 64 do not, and neither does anything past. */
char owner63[AKERR_MAX_ERROR_NAME_LENGTH];
char owner64[AKERR_MAX_ERROR_NAME_LENGTH + 1];
char owner70[AKERR_MAX_ERROR_NAME_LENGTH + 7];
memset(owner63, 'a', sizeof(owner63) - 1);
owner63[sizeof(owner63) - 1] = '\0';
memset(owner64, 'b', sizeof(owner64) - 1);
owner64[sizeof(owner64) - 1] = '\0';
memset(owner70, 'c', sizeof(owner70) - 1);
owner70[sizeof(owner70) - 1] = '\0';
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(400, 1, owner63));
AKERR_CHECK_RAISES(akerr_reserve_status_range(401, 1, owner64),
AKERR_STATUS_RANGE_INVALID);
AKERR_CHECK_RAISES(akerr_reserve_status_range(402, 1, owner70),
AKERR_STATUS_RANGE_INVALID);
/* Partial overlaps at either endpoint, and a same-range different owner. */
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(500, 2, "endpoint"));
AKERR_CHECK_RAISES(akerr_reserve_status_range(499, 2, "left"),
AKERR_STATUS_RANGE_OVERLAP);
AKERR_CHECK_RAISES(akerr_reserve_status_range(500, 1, "endpoint"),
AKERR_STATUS_RANGE_OVERLAP); /* subset, not an exact repeat */
AKERR_CHECK_RAISES(akerr_reserve_status_range(501, 1, "endpoint"),
AKERR_STATUS_RANGE_OVERLAP);
AKERR_CHECK_RAISES(akerr_reserve_status_range(500, 2, "other"),
AKERR_STATUS_RANGE_OVERLAP); /* same range, wrong owner */
/* Claim room for the name-exhaustion sweep before filling the range table. */
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(2000000, 100000, "fill"));
/*
* Range table capacity. The limit is private to src/error.c on purpose, so
* discover it by filling rather than by hardcoding it here.
*/
int ranges_added = 0;
akerr_ErrorContext *range_err = NULL;
for ( int i = 0; i < 100000; i++ ) {
range_err = akerr_reserve_status_range(1000 + (i * 2), 1, "pad");
if ( range_err != NULL ) {
break;
}
ranges_added++;
}
AKERR_CHECK(ranges_added > 0);
AKERR_CHECK_STATUS(range_err, AKERR_STATUS_RANGE_FULL);
AKERR_CHECK(strstr(range_err->message, "range table is full") != NULL);
RELEASE_ERROR(range_err);
AKERR_CHECK(range_err == NULL);
/*
* Name table capacity. Exhaustion must be reported, not silent: a dropped
* name degrades every future stack trace for that code to "Unknown Error".
*/
int full_at = -1;
for ( int i = 0; i < 100000; i++ ) {
char name[32];
snprintf(name, sizeof(name), "Filled %d", i);
akerr_ErrorContext *name_err =
akerr_register_status_name("fill", 2000000 + i, name);
if ( name_err != NULL ) {
AKERR_CHECK_STATUS(name_err, AKERR_STATUS_NAME_FULL);
AKERR_CHECK(strstr(name_err->message, "registry is full") != NULL);
AKERR_CHECK(strstr(name_err->message, "AKERR_STATUS_NAME_SLOTS") != NULL);
RELEASE_ERROR(name_err);
AKERR_CHECK(name_err == NULL);
full_at = i;
break;
}
}
/*
* The table must actually hold everything it accepted. A probe sequence
* that revisits slots instead of walking the table -- e.g. masking with
* SLOTS rather than SLOTS-1 -- both collapses the usable capacity and
* loses earlier entries, and each check below catches it independently.
* The floor assumes at least the default table size (4096 slots).
*/
AKERR_CHECK(full_at > 256);
for ( int i = 0; i < full_at; i++ ) {
char expected[32];
snprintf(expected, sizeof(expected), "Filled %d", i);
AKERR_CHECK(strcmp(akerr_name_for_status(2000000 + i, NULL), expected) == 0);
}
/* A dropped name reads back as the sentinel, and earlier ones survive. */
AKERR_CHECK(strcmp(akerr_name_for_status(2000000 + full_at, NULL),
"Unknown Error") == 0);
AKERR_CHECK(strcmp(akerr_name_for_status(INT_MIN, NULL), "Minimum Status") == 0);
/* Thousands of refusals later, every context went back to the pool. */
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_maxval ok (%d consumer names before full)\n", full_at);
return 0;
}

24
tests/err_name_bounds.c Normal file
View File

@@ -0,0 +1,24 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/* Unregistered status values return the sentinel regardless of magnitude. */
int main(void)
{
akerr_init();
/* Below range. */
AKERR_CHECK(strcmp(akerr_name_for_status(-1, NULL), "Unknown Error") == 0);
AKERR_CHECK(strcmp(akerr_name_for_status(-9999, NULL), "Unknown Error") == 0);
AKERR_CHECK(strcmp(akerr_name_for_status(1000000, NULL),
"Unknown Error") == 0);
/* A valid code must still resolve to its real name. */
AKERR_CHECK(strcmp(akerr_name_for_status(AKERR_NULLPOINTER, NULL),
"Null Pointer Error") == 0);
fprintf(stderr, "err_name_bounds ok\n");
return 0;
}

126
tests/err_name_ownership.c Normal file
View File

@@ -0,0 +1,126 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* Reserving a range used to be pure bookkeeping: akerr_name_for_status() would
* name any status for any caller, so two components could still register names
* for the same code -- and HANDLE the same code -- with nothing detecting it.
* Reservation only caught components that both opted in AND declared ranges
* that happened to overlap.
*
* Naming a status is now permitted only inside a reservation:
* - akerr_register_status_name() requires the range to belong to the caller,
* and raises AKERR_STATUS_NAME_* when it does not;
* - the legacy two-argument akerr_name_for_status() set path cannot identify
* its caller, so it can only require that *some* reservation covers the
* status -- still enough to stop a code nobody claimed. It returns a name
* rather than an error context, so its refusals are logged instead.
* Either way the refusal is visible, because a name that fails to register
* degrades the status to "Unknown Error" in every later stack trace.
*/
int main(void)
{
akerr_capture_install();
akerr_init();
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(256, 16, "lib-a"));
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(512, 16, "lib-b"));
/* The owner of a range may name statuses inside it. */
AKERR_CHECK_SUCCEEDS(akerr_register_status_name("lib-a", 256, "A Parse Error"));
AKERR_CHECK_SUCCEEDS(akerr_register_status_name("lib-a", 271, "A Last Error"));
AKERR_CHECK(strcmp(akerr_name_for_status(256, NULL), "A Parse Error") == 0);
/* Naming another owner's status is refused and names the real owner. */
AKERR_CHECK_RAISES(akerr_register_status_name("lib-b", 256, "B Hijack"),
AKERR_STATUS_NAME_FOREIGN);
AKERR_CHECK_MESSAGE_CONTAINS("lib-a");
AKERR_CHECK_MESSAGE_CONTAINS("lib-b");
AKERR_CHECK(strcmp(akerr_name_for_status(256, NULL), "A Parse Error") == 0);
/* Including the library's own reserved band. */
AKERR_CHECK_RAISES(akerr_register_status_name("lib-b", AKERR_VALUE, "B Value"),
AKERR_STATUS_NAME_FOREIGN);
AKERR_CHECK_MESSAGE_CONTAINS(AKERR_LIBRARY_OWNER);
AKERR_CHECK(strcmp(akerr_name_for_status(AKERR_VALUE, NULL), "Value Error") == 0);
/* A status nobody reserved cannot be named through either entry point. */
AKERR_CHECK_RAISES(akerr_register_status_name("lib-a", 9999, "Unclaimed"),
AKERR_STATUS_NAME_UNRESERVED);
AKERR_CHECK_MESSAGE_CONTAINS("no reserved range");
AKERR_CHECK(strcmp(akerr_name_for_status(9999, NULL), "Unknown Error") == 0);
/* The legacy path has no caller to raise into, so it logs the refusal.
* It also cannot name the caller, and must say so rather than printing a
* stray owner. */
akerr_capture_reset();
AKERR_CHECK(strcmp(akerr_name_for_status(9999, "Unclaimed Legacy"),
"Unknown Error") == 0);
AKERR_CHECK_CONTAINS("no reserved range");
AKERR_CHECK_CONTAINS("an unnamed caller");
AKERR_CHECK_CONTAINS("REFUSED STATUS NAME");
AKERR_CHECK(strcmp(akerr_name_for_status(9999, NULL), "Unknown Error") == 0);
/* ... and hands the context it raised back to the pool. */
AKERR_CHECK(akerr_slots_in_use() == 0);
/* Boundaries: just outside lib-a's range is not lib-a's to name. */
AKERR_CHECK_RAISES(akerr_register_status_name("lib-a", 255, "Below"),
AKERR_STATUS_NAME_FOREIGN);
AKERR_CHECK_RAISES(akerr_register_status_name("lib-a", 272, "Above"),
AKERR_STATUS_NAME_UNRESERVED);
/* The legacy set path still works inside any reservation. */
AKERR_CHECK(strcmp(akerr_name_for_status(513, "B Legacy"), "B Legacy") == 0);
AKERR_CHECK(strcmp(akerr_name_for_status(513, NULL), "B Legacy") == 0);
/* Re-registering your own status overwrites the name. */
AKERR_CHECK_SUCCEEDS(akerr_register_status_name("lib-a", 256, "A Renamed"));
AKERR_CHECK(strcmp(akerr_name_for_status(256, NULL), "A Renamed") == 0);
/*
* Argument validation. Each message must identify the caller it refused,
* since that message is the whole report a consumer gets.
*/
AKERR_CHECK_RAISES(akerr_register_status_name(NULL, 257, "No Owner"),
AKERR_STATUS_NAME_INVALID);
AKERR_CHECK_MESSAGE_CONTAINS("257");
AKERR_CHECK_RAISES(akerr_register_status_name("", 257, "Empty Owner"),
AKERR_STATUS_NAME_INVALID);
AKERR_CHECK_MESSAGE_CONTAINS("Empty Owner");
AKERR_CHECK_RAISES(akerr_register_status_name("lib-a", 257, NULL),
AKERR_STATUS_NAME_INVALID);
AKERR_CHECK_MESSAGE_CONTAINS("lib-a");
/* Over-long owner strings: 63 characters fit, 64 and beyond do not. */
char owner63[64];
char owner64[65];
char owner70[71];
memset(owner63, 'a', sizeof(owner63) - 1);
owner63[sizeof(owner63) - 1] = '\0';
memset(owner64, 'b', sizeof(owner64) - 1);
owner64[sizeof(owner64) - 1] = '\0';
memset(owner70, 'c', sizeof(owner70) - 1);
owner70[sizeof(owner70) - 1] = '\0';
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(600, 1, owner63));
AKERR_CHECK_SUCCEEDS(akerr_register_status_name(owner63, 600, "Long Owner"));
AKERR_CHECK_RAISES(akerr_register_status_name(owner64, 600, "Too Long"),
AKERR_STATUS_NAME_INVALID);
AKERR_CHECK_MESSAGE_CONTAINS("63");
AKERR_CHECK_RAISES(akerr_register_status_name(owner70, 600, "Far Too Long"),
AKERR_STATUS_NAME_INVALID);
AKERR_CHECK(strcmp(akerr_name_for_status(600, NULL), "Long Owner") == 0);
/* A refused registration must not consume a slot or leave a partial entry. */
AKERR_CHECK(strcmp(akerr_name_for_status(257, NULL), "Unknown Error") == 0);
/* Lookup is unaffected by ownership -- anyone may read any name. */
AKERR_CHECK(strcmp(akerr_name_for_status(271, NULL), "A Last Error") == 0);
/* Every refusal above released its context. */
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_name_ownership ok\n");
return 0;
}

47
tests/err_pass.c Normal file
View File

@@ -0,0 +1,47 @@
#include "akerror.h"
#include "err_capture.h"
/*
* PASS bubbles an error up to the caller without a local ATTEMPT/PROCESS block:
* if the wrapped call fails, PASS returns the context from the current function.
* Verify the error reaches main and that the code after PASS is skipped on
* failure.
*/
static int reached_after_pass = 0;
akerr_ErrorContext *inner(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_IO, "inner failed");
}
akerr_ErrorContext *outer(void)
{
PREPARE_ERROR(e);
PASS(e, inner());
reached_after_pass = 1; /* must NOT run: PASS returned already */
SUCCEED_RETURN(e);
}
int main(void)
{
akerr_capture_install();
int handled = 0;
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, outer());
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, AKERR_IO) {
AKERR_CHECK_STATUS(e, AKERR_IO);
handled = 1;
} FINISH_NORETURN(e);
AKERR_CHECK(handled == 1);
AKERR_CHECK(reached_after_pass == 0);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_pass ok\n");
return 0;
}

40
tests/err_pool_exhaust.c Normal file
View File

@@ -0,0 +1,40 @@
#include "akerror.h"
#include "err_capture.h"
/*
* The error pool is a fixed array of AKERR_MAX_ARRAY_ERROR slots. When every
* slot is checked out, akerr_next_error() must return NULL rather than run off
* the end of the array; and it must always hand back the lowest free slot.
* Mutation testing showed both the terminating "return NULL" and the scan
* bounds could be broken without any test noticing.
*/
int main(void)
{
akerr_init();
akerr_ErrorContext *slots[AKERR_MAX_ARRAY_ERROR];
/* Check out every slot. Each arrives holding its own reference, so the
* next request cannot be handed the same one. */
for ( int i = 0; i < AKERR_MAX_ARRAY_ERROR; i++ ) {
slots[i] = akerr_next_error();
AKERR_CHECK(slots[i] != NULL);
AKERR_CHECK(slots[i]->refcount == 1);
}
/* Pool is fully exhausted: the next request must fail cleanly. */
AKERR_CHECK(akerr_next_error() == NULL);
/* Free exactly the first slot; the scan must find and return it. */
slots[0]->refcount = 0;
AKERR_CHECK(akerr_next_error() == slots[0]);
/* Tidy up. */
for ( int i = 0; i < AKERR_MAX_ARRAY_ERROR; i++ ) {
slots[i]->refcount = 0;
}
fprintf(stderr, "err_pool_exhaust ok\n");
return 0;
}

60
tests/err_pool_refcount.c Normal file
View File

@@ -0,0 +1,60 @@
#include "akerror.h"
#include "err_capture.h"
/*
* Pool-hygiene regression test. The error pool is a fixed 128-slot static
* array; a single leaked reference per error would silently exhaust it. Run a
* large number of full raise -> catch -> handle cycles and confirm every slot
* is returned to the pool (refcount 0) and the pool can still hand out
* contexts afterwards.
*/
#define ITERATIONS 100000
static int handled_status = 0;
akerr_ErrorContext *boom(void)
{
PREPARE_ERROR(e);
ATTEMPT {
FAIL(e, AKERR_VALUE, "boom %d", 1);
} CLEANUP {
} PROCESS(e) {
} FINISH(e, true);
SUCCEED_RETURN(e);
}
/* One full raise -> catch -> handle cycle. Returns NULL (context released). */
akerr_ErrorContext *one_cycle(void)
{
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, boom());
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, AKERR_VALUE) {
handled_status = e->status;
} FINISH(e, false);
return e;
}
int main(void)
{
akerr_capture_install();
akerr_init();
AKERR_CHECK(akerr_slots_in_use() == 0);
for ( int iter = 0; iter < ITERATIONS; iter++ ) {
handled_status = 0;
(void)one_cycle();
AKERR_CHECK(handled_status == AKERR_VALUE);
}
AKERR_CHECK(akerr_slots_in_use() == 0);
akerr_ErrorContext *probe = akerr_next_error();
AKERR_CHECK(probe != NULL);
fprintf(stderr, "err_pool_refcount ok (%d cycles, 0 leaked)\n", ITERATIONS);
return 0;
}

View File

@@ -0,0 +1,56 @@
#include "akerror.h"
#include "err_capture.h"
/*
* Regression test for the refcount leak: ENSURE_ERROR_READY must increment
* refcount only when it *acquires* a fresh context, not on every FAIL/SUCCEED.
* A function that calls FAIL more than once on the same context and then
* propagates used to arrive at the caller with refcount 2; the caller released
* once, leaking the slot. After enough leaks the pool is exhausted and the
* library exit(1)s.
*/
static int handled_status = 0;
akerr_ErrorContext *validate(void)
{
PREPARE_ERROR(e);
ATTEMPT {
FAIL(e, AKERR_VALUE, "condition 1 failed"); /* acquires the context */
FAIL(e, AKERR_KEY, "condition 2 failed"); /* must NOT re-acquire it */
} CLEANUP {
} PROCESS(e) {
} FINISH(e, true); /* unhandled -> propagate */
SUCCEED_RETURN(e);
}
/* One raise -> catch -> handle cycle; returns NULL (context released). */
akerr_ErrorContext *one_cycle(void)
{
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, validate());
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, AKERR_KEY) {
handled_status = e->status;
} HANDLE(e, AKERR_VALUE) {
} FINISH(e, false);
return e;
}
int main(void)
{
akerr_init();
AKERR_CHECK(akerr_slots_in_use() == 0);
for ( int i = 0; i < 32; i++ ) {
handled_status = 0;
(void)one_cycle();
AKERR_CHECK(handled_status == AKERR_KEY);
}
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_refcount_double_fail ok\n");
return 0;
}

View File

@@ -0,0 +1,55 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* A library may reserve its status range from its own init() before anything in
* the process has raised an error, i.e. before akerr_init() has run. That used
* to be silently destructive: akerr_init() clears the range and name tables, so
* whichever component first triggered it (via PREPARE_ERROR) wiped the earlier
* reservation, and the *next* component to claim the same range was told OK --
* producing exactly the undetected aliasing the registry exists to prevent.
*
* Every public registry entry point now calls akerr_init() itself, so the
* tables are only ever cleared before the first reservation, never after one.
*
* Note this test must not call akerr_init() or PREPARE_ERROR first -- the
* uninitialized entry is the whole point.
*/
int main(void)
{
akerr_capture_install();
/* Cold call: no akerr_init(), no PREPARE_ERROR anywhere yet. */
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(256, 16, "early-lib"));
AKERR_CHECK_SUCCEEDS(akerr_register_status_name("early-lib", 256, "Early Error"));
/* Something else now uses the library for the first time. */
akerr_init();
PREPARE_ERROR(e);
(void)e;
/* The early reservation and its name must both have survived. */
AKERR_CHECK(strcmp(akerr_name_for_status(256, NULL), "Early Error") == 0);
AKERR_CHECK_RAISES(akerr_reserve_status_range(256, 16, "late-lib"),
AKERR_STATUS_RANGE_OVERLAP);
AKERR_CHECK_MESSAGE_CONTAINS("early-lib");
AKERR_CHECK_RAISES(akerr_reserve_status_range(260, 2, "late-lib"),
AKERR_STATUS_RANGE_OVERLAP);
/* The library's own initialization still happened exactly once. */
AKERR_CHECK(strcmp(akerr_name_for_status(AKERR_NULLPOINTER, NULL),
"Null Pointer Error") == 0);
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(0, AKERR_RESERVED_STATUS_COUNT,
AKERR_LIBRARY_OWNER));
/* An identical repeat by the original owner is still idempotent. */
AKERR_CHECK_SUCCEEDS(akerr_reserve_status_range(256, 16, "early-lib"));
/* Raising from a cold registry must not strand a pool slot either. */
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_registry_init_order ok\n");
return 0;
}

View File

@@ -0,0 +1,45 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* Releasing an error context back to the pool must wipe it, so the next caller
* that checks it out never sees stale status/message/stacktrace from a previous
* error. Mutation testing showed the clearing memset in akerr_release_error
* could be deleted without any test noticing.
*/
akerr_ErrorContext *boom(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_VALUE, "stale dirty message that must not survive");
}
int main(void)
{
akerr_capture_install();
akerr_init();
/* Raise and fully handle an error; FINISH_NORETURN releases it to the pool. */
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, boom());
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, AKERR_VALUE) {
AKERR_CHECK_STATUS(e, AKERR_VALUE);
} FINISH_NORETURN(e);
AKERR_CHECK(e == NULL);
/* The next context handed out is the slot we just released: it must be clean. */
akerr_ErrorContext *slot = akerr_next_error();
AKERR_CHECK(slot != NULL);
AKERR_CHECK(slot->status == 0);
AKERR_CHECK(slot->message[0] == '\0');
AKERR_CHECK(slot->stacktracebuf[0] == '\0');
AKERR_CHECK(strstr(slot->message, "stale dirty message") == NULL);
fprintf(stderr, "err_release_clears ok\n");
return 0;
}

43
tests/err_release_null.c Normal file
View File

@@ -0,0 +1,43 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* akerr_release_error(NULL) is an API contract violation the library reports
* rather than crashes on: it raises AKERR_NULLPOINTER against the internal
* last-ditch context and hands that back, so the caller gets a describable
* error instead of a dereferenced NULL. Nothing exercised that path.
*
* The last-ditch context deliberately lives outside AKERR_ARRAY_ERROR so
* reporting this failure cannot consume a pool slot -- which is exactly what
* akerr_valid_error_address() reports on, and what this test checks.
*/
int main(void)
{
akerr_capture_install();
/* akerr_init() sets up the last-ditch context's stacktrace cursor. */
akerr_init();
AKERR_CHECK(akerr_slots_in_use() == 0);
akerr_ErrorContext *ret = akerr_release_error(NULL);
AKERR_CHECK(ret != NULL);
AKERR_CHECK(ret->status == AKERR_NULLPOINTER);
AKERR_CHECK(strstr(ret->message, "NULL context pointer") != NULL);
/* Reported through FAIL, so the stack trace names the error too. */
AKERR_CHECK(strstr(ret->stacktracebuf, "Null Pointer Error") != NULL);
/* Not a pool slot, and no slot was checked out to report the failure. */
AKERR_CHECK(akerr_valid_error_address(ret) == 0);
AKERR_CHECK(akerr_slots_in_use() == 0);
/* The last-ditch context is a singleton: the same one comes back. */
akerr_ErrorContext *again = akerr_release_error(NULL);
AKERR_CHECK(again == ret);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_release_null ok\n");
return 0;
}

View File

@@ -0,0 +1,73 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* akerr_release_error() drops one reference and only recycles the slot when the
* last one goes away. Two of its refcount edges had no test:
*
* - refcount > 1: the release must decrement and return the context intact,
* not wipe it out from under the reference that is still held.
* - refcount == 0: releasing a context nobody holds must not underflow the
* count; it wipes and returns NULL like any other fully-released slot.
*
* The macro API never produces a refcount above 1 (akerr_next_error() takes the
* one reference a fresh slot gets), so this test sets the count directly to
* model a caller that took an extra reference, and clears it to model a slot
* nobody holds.
*/
int main(void)
{
akerr_capture_install();
akerr_init();
AKERR_CHECK(akerr_slots_in_use() == 0);
/* Check out a slot and give it two holders. */
akerr_ErrorContext *held = akerr_next_error();
AKERR_CHECK(held != NULL);
int slotid = held->arrayid;
held->refcount = 2;
held->status = AKERR_VALUE;
snprintf((char *)held->message, AKERR_MAX_ERROR_CONTEXT_STRING_LENGTH,
"still referenced");
/* First release: one holder left, so the context survives untouched. */
akerr_ErrorContext *ret = akerr_release_error(held);
AKERR_CHECK(ret == held);
AKERR_CHECK(held->refcount == 1);
AKERR_CHECK(held->status == AKERR_VALUE);
AKERR_CHECK(strcmp(held->message, "still referenced") == 0);
AKERR_CHECK(akerr_slots_in_use() == 1);
/* Second release: last holder gone, so the slot is wiped and recycled. */
ret = akerr_release_error(held);
AKERR_CHECK(ret == NULL);
AKERR_CHECK(held->refcount == 0);
AKERR_CHECK(held->status == 0);
AKERR_CHECK(held->message[0] == '\0');
/* The wipe must preserve the slot's identity and stacktrace cursor. */
AKERR_CHECK(held->arrayid == slotid);
AKERR_CHECK(held->stacktracebufptr == (char *)&held->stacktracebuf);
AKERR_CHECK(akerr_slots_in_use() == 0);
/*
* Releasing an unheld slot: refcount is already 0, so there is nothing to
* decrement and the count must not go negative.
*/
akerr_ErrorContext *unheld = akerr_next_error();
AKERR_CHECK(unheld != NULL);
AKERR_CHECK(unheld->refcount == 1);
unheld->refcount = 0;
ret = akerr_release_error(unheld);
AKERR_CHECK(ret == NULL);
AKERR_CHECK(unheld->refcount == 0);
AKERR_CHECK(akerr_slots_in_use() == 0);
/* The pool is still healthy afterwards. */
akerr_ErrorContext *probe = akerr_next_error();
AKERR_CHECK(probe != NULL);
fprintf(stderr, "err_release_refcount ok\n");
return 0;
}

View File

@@ -0,0 +1,54 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* Regression test for the stack-trace buffer overflow. Each frame appended a
* line with snprintf, but passed the *full* buffer length as the size rather
* than the space remaining, and advanced the cursor by snprintf's would-be
* return value. A trace that filled the buffer therefore wrote past the end of
* stacktracebuf and ran the cursor out of bounds.
*
* We place a context in a struct with a guard region right after it, position
* the trace cursor near the end of the buffer, append one more frame, and
* require that nothing was written past the buffer and the cursor stayed in
* bounds.
*/
static struct {
akerr_ErrorContext ctx;
unsigned char guard[512];
} probe;
akerr_ErrorContext *append_frame(akerr_ErrorContext *e)
{
FAIL_RETURN(e, AKERR_VALUE,
"an error message long enough to overflow a nearly full stack trace buffer");
}
int main(void)
{
akerr_init();
memset(&probe, 0x00, sizeof(probe));
memset(probe.guard, 0xAA, sizeof(probe.guard));
akerr_ErrorContext *e = &probe.ctx;
e->refcount = 1;
/* Two bytes short of full: any real frame would overflow the old code. */
e->stacktracebufptr = probe.ctx.stacktracebuf
+ AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH - 2;
(void)append_frame(e);
/* Nothing may have been written past the end of stacktracebuf. */
for ( unsigned i = 0; i < sizeof(probe.guard); i++ ) {
AKERR_CHECK(probe.guard[i] == 0xAA);
}
/* The cursor must remain within the buffer. */
AKERR_CHECK(e->stacktracebufptr
<= probe.ctx.stacktracebuf + AKERR_MAX_ERROR_STACKTRACE_BUF_LENGTH);
fprintf(stderr, "err_stacktrace_bounds ok\n");
return 0;
}

View File

@@ -0,0 +1,100 @@
#include "akerror.h"
#include "err_capture.h"
#include <string.h>
/*
* The registry entry points report failure the same way every other function in
* this library does: they return akerr_ErrorContext *. A refused reservation is
* therefore an ordinary error, and everything that works on an ordinary error
* must work on it -- CATCH, HANDLE, PASS, propagation to the caller, and the
* stack trace an unhandled one prints.
*
* This is what distinguishes the current design from the integer return codes
* it replaced: a consumer that ignores the result gets a compiler warning
* (AKERR_NOIGNORE), and a consumer that catches it but does not handle it gets
* the error propagated out of its init function rather than a silently
* unreserved range.
*/
#define TEST_OWNER "exception-test"
/* A consumer init function in the shape the documentation recommends. */
static akerr_ErrorContext AKERR_NOIGNORE *component_init(int base, const char *owner)
{
PREPARE_ERROR(errctx);
ATTEMPT {
CATCH(errctx, akerr_reserve_status_range(base, 4, owner));
CATCH(errctx, akerr_register_status_name(owner, base, "Component Error"));
} CLEANUP {
} PROCESS(errctx) {
} FINISH(errctx, true);
SUCCEED_RETURN(errctx);
}
int main(void)
{
akerr_capture_install();
akerr_init();
PREPARE_ERROR(errctx);
/* A successful reservation raises nothing at all. */
AKERR_CHECK_SUCCEEDS(component_init(256, TEST_OWNER));
AKERR_CHECK(strcmp(akerr_name_for_status(256, NULL), "Component Error") == 0);
AKERR_CHECK(akerr_slots_in_use() == 0);
/*
* A collision propagates out of the component's init and is caught here.
* HANDLE proves the status is a real, distinct exception a consumer can
* dispatch on -- not an opaque nonzero int.
*/
int handled_overlap = 0;
ATTEMPT {
CATCH(errctx, component_init(258, "other-component"));
} CLEANUP {
} PROCESS(errctx) {
} HANDLE(errctx, AKERR_STATUS_RANGE_OVERLAP) {
handled_overlap = 1;
} HANDLE_DEFAULT(errctx) {
handled_overlap = -1;
} FINISH_NORETURN(errctx);
AKERR_CHECK(handled_overlap == 1);
AKERR_CHECK(akerr_slots_in_use() == 0);
/* The name-side refusals dispatch the same way. */
int handled_foreign = 0;
ATTEMPT {
CATCH(errctx, akerr_register_status_name("interloper", 256, "Hijack"));
} CLEANUP {
} PROCESS(errctx) {
} HANDLE(errctx, AKERR_STATUS_NAME_FOREIGN) {
handled_foreign = 1;
} FINISH_NORETURN(errctx);
AKERR_CHECK(handled_foreign == 1);
AKERR_CHECK(strcmp(akerr_name_for_status(256, NULL), "Component Error") == 0);
/*
* An unhandled refusal carries a stack trace naming the real owner, so the
* report a consumer gets is the library's normal one.
*/
akerr_capture_reset();
ATTEMPT {
CATCH(errctx, akerr_reserve_status_range(AKERR_VALUE, 1, "encroacher"));
} CLEANUP {
} PROCESS(errctx) {
} HANDLE_DEFAULT(errctx) {
LOG_ERROR_WITH_MESSAGE(errctx, "Reservation refused");
} FINISH_NORETURN(errctx);
AKERR_CHECK_CONTAINS("Reservation refused");
AKERR_CHECK_CONTAINS("Status Range Overlap");
AKERR_CHECK_CONTAINS(AKERR_LIBRARY_OWNER);
AKERR_CHECK_CONTAINS("encroacher");
AKERR_CHECK_CONTAINS("error.c");
/* Nothing above stranded a pool slot. */
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_status_exception ok\n");
return 0;
}

44
tests/err_success.c Normal file
View File

@@ -0,0 +1,44 @@
#include "akerror.h"
#include "err_capture.h"
/*
* The success path: a nested call that returns cleanly (NULL) must not break
* out of the caller's ATTEMPT block, must not enter any handler, and must not
* consume a slot from the error pool.
*/
static int default_fired = 0;
akerr_ErrorContext *ok_func(void)
{
PREPARE_ERROR(e);
ATTEMPT {
} CLEANUP {
} PROCESS(e) {
} FINISH(e, true);
SUCCEED_RETURN(e);
}
int main(void)
{
akerr_capture_install();
akerr_init();
int reached_after_catch = 0;
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, ok_func());
reached_after_catch = 1; /* must run: no break on success */
} CLEANUP {
} PROCESS(e) {
} HANDLE_DEFAULT(e) {
default_fired = 1; /* must NOT run */
} FINISH_NORETURN(e);
AKERR_CHECK(reached_after_catch == 1);
AKERR_CHECK(default_fired == 0);
AKERR_CHECK(e == NULL);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_success ok\n");
return 0;
}

48
tests/err_swallow.c Normal file
View File

@@ -0,0 +1,48 @@
#include "akerror.h"
#include "err_capture.h"
/*
* FINISH(ctx, false) closes an ATTEMPT block without propagating: an unhandled
* error is dropped (not returned, not passed to the unhandled handler) and the
* context is released back to the pool. Because we use FINISH (not
* FINISH_NORETURN) the process must continue normally and exit 0.
*/
static int wrong_handler_fired = 0;
static int swallowed_status = 0;
akerr_ErrorContext *boom(void)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_VALUE, "unhandled but swallowed");
}
/* FINISH(e, false) belongs in a context-returning function. On the swallow
* path e is released to NULL, so this returns NULL to its caller. */
akerr_ErrorContext *swallow_it(void)
{
PREPARE_ERROR(e);
ATTEMPT {
CATCH(e, boom());
} CLEANUP {
swallowed_status = (e != NULL) ? e->status : 0;
} PROCESS(e) {
} HANDLE(e, AKERR_KEY) {
wrong_handler_fired = 1; /* does not match AKERR_VALUE */
} FINISH(e, false);
return e; /* NULL: released even though unhandled */
}
int main(void)
{
akerr_capture_install();
akerr_ErrorContext *res = swallow_it();
AKERR_CHECK(wrong_handler_fired == 0);
AKERR_CHECK(swallowed_status == AKERR_VALUE);
AKERR_CHECK(res == NULL); /* released even though unhandled */
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_swallow ok\n");
return 0;
}

149
tests/err_threads.h Normal file
View File

@@ -0,0 +1,149 @@
#ifndef AKERR_TEST_THREADS_H
#define AKERR_TEST_THREADS_H
/*
* Shared helpers for the thread-safety tests.
*
* These tests are checkable on their own -- they assert exclusive ownership of
* pool slots and of reserved ranges, which is a property, not a symptom -- but
* the run that proves the absence of a data race is the one under
* ThreadSanitizer:
*
* cmake -S . -B build/tsan -DAKERR_SANITIZE=thread
* cmake --build build/tsan
* ctest --test-dir build/tsan --output-on-failure
*
* Everything shared between the threads here is either read-only after the
* threads start, or touched through __atomic builtins. Anything else would be a
* race in the *test*, and TSan cannot tell whose bug it is reporting.
*
* A test body runs on AKERR_TEST_THREADS threads that meet at a barrier first,
* so they arrive at the library together instead of in start-up order. Failed
* checks are counted per thread rather than returned early: a thread that
* abandoned its work would leave the others holding pool slots and turn one
* failure into a cascade of unrelated ones.
*/
#include "akerror.h"
#include <pthread.h>
#include <stdio.h>
#define AKERR_TEST_THREADS 8
typedef struct
{
int id; /* 1-based: 0 means "no thread" below */
int failures;
pthread_barrier_t *barrier;
} akerr_ThreadArg;
#define AKERR_TCHECK(__arg, __cond) \
do { \
if ( !(__cond) ) { \
fprintf(stderr, "CHECK FAILED (thread %d): %s at %s:%d\n", \
(__arg)->id, #__cond, __FILE__, __LINE__); \
(__arg)->failures += 1; \
} \
} while ( 0 )
/*
* A logger that counts instead of printing. The capturing logger in
* err_capture.h appends to a shared buffer with a shared length, which is a
* data race the moment two threads log at once; these tests need a logger that
* is safe to install before spawning and still shows that something was
* reported.
*/
static int akerr_thread_log_count;
static void __attribute__((unused)) akerr_thread_logger(const char *fmt, ...)
{
(void)fmt;
__atomic_fetch_add(&akerr_thread_log_count, 1, __ATOMIC_RELAXED);
}
static int __attribute__((unused)) akerr_thread_logs(void)
{
return __atomic_load_n(&akerr_thread_log_count, __ATOMIC_RELAXED);
}
/*
* Independent bookkeeping of who holds which pool slot. The library's own
* refcount says a slot is checked out; this says which thread it was checked
* out to, which is the part a racing akerr_next_error() would get wrong by
* handing one slot to two threads at once.
*/
static int akerr_slot_owner[AKERR_MAX_ARRAY_ERROR];
/* Returns 0 on success, or the id of the thread that already holds the slot. */
static int __attribute__((unused)) akerr_slot_claim(int slot, int id)
{
int unowned = 0;
if ( __atomic_compare_exchange_n(&akerr_slot_owner[slot], &unowned, id, 0,
__ATOMIC_ACQ_REL, __ATOMIC_ACQUIRE) ) {
return 0;
}
return unowned;
}
static int __attribute__((unused)) akerr_slot_holder(int slot)
{
return __atomic_load_n(&akerr_slot_owner[slot], __ATOMIC_ACQUIRE);
}
/*
* Give the slot up *before* releasing the error context. The other order hands
* the slot back to the pool while this thread still claims it, and the next
* thread to be given it reports a violation that is the test's fault.
*/
static void __attribute__((unused)) akerr_slot_drop(int slot)
{
__atomic_store_n(&akerr_slot_owner[slot], 0, __ATOMIC_RELEASE);
}
/*
* Run `body` on AKERR_TEST_THREADS threads and return the total number of
* failed checks. A thread that cannot be created is itself a failure, but the
* ones already running still get joined.
*/
static int __attribute__((unused)) akerr_run_threads(void *(*body)(void *))
{
pthread_t threads[AKERR_TEST_THREADS];
akerr_ThreadArg args[AKERR_TEST_THREADS];
pthread_barrier_t barrier;
int started = 0;
int failures = 0;
if ( pthread_barrier_init(&barrier, NULL, AKERR_TEST_THREADS) != 0 ) {
fprintf(stderr, "CHECK FAILED: pthread_barrier_init at %s:%d\n",
__FILE__, __LINE__);
return 1;
}
for ( int i = 0; i < AKERR_TEST_THREADS; i++ ) {
args[i].id = i + 1;
args[i].failures = 0;
args[i].barrier = &barrier;
if ( pthread_create(&threads[i], NULL, body, &args[i]) != 0 ) {
fprintf(stderr, "CHECK FAILED: pthread_create for thread %d"
" at %s:%d\n", i + 1, __FILE__, __LINE__);
failures++;
break;
}
started++;
}
/* An unstarted thread never reaches the barrier, so the started ones would
* wait for it forever. Nothing to do but say so before hanging is diagnosed
* as a deadlock in the library. */
if ( started != AKERR_TEST_THREADS ) {
fprintf(stderr, "only %d of %d threads started; the barrier will not"
" release\n", started, AKERR_TEST_THREADS);
}
for ( int i = 0; i < started; i++ ) {
pthread_join(threads[i], NULL);
failures += args[i].failures;
}
pthread_barrier_destroy(&barrier);
return failures;
}
#endif // AKERR_TEST_THREADS_H

124
tests/err_threads_init.c Normal file
View File

@@ -0,0 +1,124 @@
#include "akerror.h"
#include "err_capture.h"
#include "err_threads.h"
#include <errno.h>
#include <string.h>
/*
* akerr_init() runs exactly once, no matter how many threads reach the library
* at the same instant.
*
* This is the hardest of the three to get right, because initialization is
* re-entrant: akerr_init() reserves the library's own status band and names its
* own codes, and every one of those calls goes through a public entry point
* that calls akerr_init() again. The guard against that recursion has to be
* per-thread, or a second thread arriving mid-initialization would see the flag
* the first thread raised on its way in and walk tables that are still being
* built.
*
* Nothing in main() touches the library before the threads start, so the race
* is real: whichever thread wins does the initializing, and the rest must block
* until it is finished rather than proceed on half-built tables.
*
* What proves it ran once rather than several times: each thread reserves its
* own status range as its first act. A second pass through akerr_init() would
* memset the range table, so a reservation made by a thread that raced ahead
* would silently vanish -- exactly the failure the pre-1.0.0 library had. After
* the join, every thread's reservation must still be there, still attributed to
* that thread.
*/
static int thread_range_base(int id)
{
return 400000 + (id * 16);
}
static void *init_racer(void *raw)
{
akerr_ThreadArg *arg = raw;
akerr_ErrorContext *e;
char owner[32];
char name[32];
int base = thread_range_base(arg->id);
snprintf(owner, sizeof(owner), "init-thread-%d", arg->id);
snprintf(name, sizeof(name), "Thread %d Error", arg->id);
pthread_barrier_wait(arg->barrier);
/* First touch of the library from this thread, and for one of them the
* first touch in the process. */
e = akerr_reserve_status_range(base, 16, owner);
AKERR_TCHECK(arg, e == NULL);
RELEASE_ERROR(e);
e = akerr_register_status_name(owner, base, name);
AKERR_TCHECK(arg, e == NULL);
RELEASE_ERROR(e);
/* The library's own band was reserved by whichever thread initialized, and
* every thread must see it as taken -- including the one that did it. A
* second initialization would have wiped the reservation and let this
* through. */
e = akerr_reserve_status_range(0, AKERR_RESERVED_STATUS_COUNT, owner);
AKERR_TCHECK(arg, e != NULL);
if ( e != NULL ) {
AKERR_TCHECK(arg, e->status == AKERR_STATUS_RANGE_OVERLAP);
AKERR_TCHECK(arg, strstr(e->message, AKERR_LIBRARY_OWNER) != NULL);
}
RELEASE_ERROR(e);
/* The name tables are complete as seen from every thread: the library's own
* codes, the generated errno names, and this thread's own registration. */
AKERR_TCHECK(arg, strcmp(akerr_name_for_status(AKERR_VALUE, NULL),
"Value Error") == 0);
AKERR_TCHECK(arg, strcmp(akerr_name_for_status(AKERR_NULLPOINTER, NULL),
"Null Pointer Error") == 0);
AKERR_TCHECK(arg, strcmp(akerr_name_for_status(EACCES, NULL),
"Unknown Error") != 0);
AKERR_TCHECK(arg, strcmp(akerr_name_for_status(base, NULL), name) == 0);
/* And an error raised from this thread renders with a name, which is the
* whole point of the tables being complete. */
PREPARE_ERROR(errctx);
ATTEMPT {
FAIL_BREAK(errctx, AKERR_TYPE, "raised during init race by thread %d",
arg->id);
} CLEANUP {
} PROCESS(errctx) {
} HANDLE(errctx, AKERR_TYPE) {
AKERR_TCHECK(arg, strstr(errctx->stacktracebuf, "Type Error") != NULL);
} FINISH_NORETURN(errctx);
return NULL;
}
int main(void)
{
/* Installed before anything initializes: akerr_init() only supplies the
* default logger when this is still NULL. */
akerr_log_method = &akerr_thread_logger;
int failures = akerr_run_threads(&init_racer);
AKERR_CHECK(failures == 0);
/* Every thread's reservation survived the race, under its own owner. */
for ( int id = 1; id <= AKERR_TEST_THREADS; id++ ) {
char owner[32];
char name[32];
int base = thread_range_base(id);
snprintf(owner, sizeof(owner), "init-thread-%d", id);
snprintf(name, sizeof(name), "Thread %d Error", id);
AKERR_CHECK(strcmp(akerr_name_for_status(base, NULL), name) == 0);
AKERR_CHECK_RAISES(akerr_reserve_status_range(base, 16, "verifier"),
AKERR_STATUS_RANGE_OVERLAP);
AKERR_CHECK_MESSAGE_CONTAINS(owner);
}
/* Nothing leaked a pool slot on the way through. */
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_threads_init ok (%d threads raced initialization)\n",
AKERR_TEST_THREADS);
return 0;
}

138
tests/err_threads_pool.c Normal file
View File

@@ -0,0 +1,138 @@
#include "akerror.h"
#include "err_capture.h"
#include "err_threads.h"
#include <string.h>
/*
* The error pool under contention.
*
* AKERR_ARRAY_ERROR is a fixed 128-slot array shared by every thread in the
* process, and a slot is checked out by finding refcount == 0 and taking a
* reference. Those two steps have to be one operation: a scan that returned an
* unclaimed slot would hand the same one to every thread that scanned before
* the first of them incremented the count, and each would then format its own
* error into the same buffers. That failure is invisible to a single-threaded
* test and produces a garbled message rather than a crash, so this test asserts
* exclusivity directly.
*
* akerr_slot_owner[] (see err_threads.h) is the test's own record of who holds
* which slot, kept with atomics. Every check-out claims its slot and every
* release drops the claim; a slot handed to two threads at once is caught by
* the claim failing, whether or not the resulting message is garbled.
*
* Each thread also asserts that the error it raised is the error it handles,
* message and all. That is the same property from the other end: a context
* cannot be exclusively ours if another thread's text shows up in it.
*/
#define ITERATIONS 2000
/* Raise an error and take ownership of whatever slot it came from. */
static akerr_ErrorContext AKERR_NOIGNORE *boom(akerr_ThreadArg *arg)
{
PREPARE_ERROR(e);
FAIL(e, AKERR_VALUE, "raised by thread %d", arg->id);
AKERR_TCHECK(arg, akerr_slot_claim(e->arrayid, arg->id) == 0);
return e;
}
static akerr_ErrorContext AKERR_NOIGNORE *ignorable(akerr_ThreadArg *arg)
{
PREPARE_ERROR(e);
FAIL_RETURN(e, AKERR_IO, "ignored by thread %d", arg->id);
}
/* One raise -> catch -> handle cycle, exclusively owned from end to end. */
static void one_cycle(akerr_ThreadArg *arg)
{
char expected[64];
PREPARE_ERROR(e);
snprintf(expected, sizeof(expected), "raised by thread %d", arg->id);
ATTEMPT {
CATCH(e, boom(arg));
} CLEANUP {
} PROCESS(e) {
/* case 0: the error we just raised came back clean, which can only
* mean another thread wrote over this context. */
int error_was_lost = 1;
AKERR_TCHECK(arg, error_was_lost == 0);
} HANDLE(e, AKERR_VALUE) {
AKERR_TCHECK(arg, akerr_slot_holder(e->arrayid) == arg->id);
AKERR_TCHECK(arg, strcmp(e->message, expected) == 0);
AKERR_TCHECK(arg, strstr(e->stacktracebuf, expected) != NULL);
/* Give the slot up before FINISH releases the context: the other order
* hands it back to the pool while this thread still claims it. */
akerr_slot_drop(e->arrayid);
} FINISH_NORETURN(e);
}
/* The same property against the raw pool API, with no macros in between. */
static void one_checkout(akerr_ThreadArg *arg)
{
akerr_ErrorContext *e = akerr_next_error();
AKERR_TCHECK(arg, e != NULL);
if ( e == NULL ) {
return;
}
/* The context arrives already holding its reference. */
AKERR_TCHECK(arg, e->refcount == 1);
AKERR_TCHECK(arg, akerr_slot_claim(e->arrayid, arg->id) == 0);
AKERR_TCHECK(arg, akerr_slot_holder(e->arrayid) == arg->id);
akerr_slot_drop(e->arrayid);
RELEASE_ERROR(e);
AKERR_TCHECK(arg, e == NULL);
}
static void *pool_body(void *raw)
{
akerr_ThreadArg *arg = raw;
char expected[64];
snprintf(expected, sizeof(expected), "ignored by thread %d", arg->id);
pthread_barrier_wait(arg->barrier);
for ( int i = 0; i < ITERATIONS; i++ ) {
one_cycle(arg);
one_checkout(arg);
}
/* An ignored error is a fact about the thread that ignored it: each thread
* must see its own, not the last one any thread swallowed. */
IGNORE(ignorable(arg));
AKERR_TCHECK(arg, __akerr_last_ignored != NULL);
if ( __akerr_last_ignored != NULL ) {
AKERR_TCHECK(arg, __akerr_last_ignored->status == AKERR_IO);
AKERR_TCHECK(arg, strcmp(__akerr_last_ignored->message, expected) == 0);
}
/* IGNORE keeps the reference by design; hand it back so the pool is empty
* at the end of the test. */
RELEASE_ERROR(__akerr_last_ignored);
return NULL;
}
int main(void)
{
akerr_log_method = &akerr_thread_logger;
akerr_init();
AKERR_CHECK(akerr_slots_in_use() == 0);
int failures = akerr_run_threads(&pool_body);
AKERR_CHECK(failures == 0);
/* Every context went back to the pool, and every claim was dropped. */
AKERR_CHECK(akerr_slots_in_use() == 0);
for ( int i = 0; i < AKERR_MAX_ARRAY_ERROR; i++ ) {
AKERR_CHECK(akerr_slot_holder(i) == 0);
}
/* Each thread's IGNORE reported through the log method. */
AKERR_CHECK(akerr_thread_logs() >= AKERR_TEST_THREADS);
fprintf(stderr, "err_threads_pool ok (%d threads x %d cycles)\n",
AKERR_TEST_THREADS, ITERATIONS);
return 0;
}

View File

@@ -0,0 +1,137 @@
#include "akerror.h"
#include "err_capture.h"
#include "err_threads.h"
#include <string.h>
/*
* The status registry under contention.
*
* Two properties, and they fail differently:
*
* - A reservation is a decision, not a write. The overlap scan and the entry
* that follows it have to be one operation, or two threads claiming the
* same range both find the table clear and both believe they own it. That
* is silent: neither gets an error, and the collision surfaces much later
* as one component's status rendering under another's name. The contested
* range below is claimed by every thread at once and exactly one may win.
*
* - The name table is an open-addressed hash table with linear probing. A
* concurrent insert that another thread's probe walks through -- an entry
* half claimed, a count incremented before the slot was marked used -- loses
* names or writes outside the table. So every thread registers a block of
* names and reads each one back while the others are still writing, and
* interleaves lookups of a name nobody is touching.
*
* Ownership enforcement has to hold under contention too: after every thread
* has reserved, each one tries to name a status inside its neighbour's range
* and must be refused. The barrier before that is what makes the expected
* refusal exactly AKERR_STATUS_NAME_FOREIGN rather than sometimes
* AKERR_STATUS_NAME_UNRESERVED, which is a real distinction and not just test
* tidiness: FOREIGN means the registry knew who owned it.
*/
#define NAMES_PER_THREAD 64
#define CONTESTED_FIRST 900000
#define CONTESTED_COUNT 64
static int contested_winners;
static int thread_range_base(int id)
{
return 500000 + (id * 1000);
}
static void *registry_body(void *raw)
{
akerr_ThreadArg *arg = raw;
akerr_ErrorContext *e;
char owner[32];
int base = thread_range_base(arg->id);
int victim = thread_range_base((arg->id % AKERR_TEST_THREADS) + 1);
snprintf(owner, sizeof(owner), "registry-%d", arg->id);
pthread_barrier_wait(arg->barrier);
/* One range, every thread, distinct owners. Exactly one may come back
* successful; the rest must be told who won. */
e = akerr_reserve_status_range(CONTESTED_FIRST, CONTESTED_COUNT, owner);
if ( e == NULL ) {
__atomic_fetch_add(&contested_winners, 1, __ATOMIC_RELAXED);
} else {
AKERR_TCHECK(arg, e->status == AKERR_STATUS_RANGE_OVERLAP);
AKERR_TCHECK(arg, strstr(e->message, "registry-") != NULL);
RELEASE_ERROR(e);
}
/* This thread's own range, which nobody contests. */
e = akerr_reserve_status_range(base, NAMES_PER_THREAD, owner);
AKERR_TCHECK(arg, e == NULL);
RELEASE_ERROR(e);
for ( int i = 0; i < NAMES_PER_THREAD; i++ ) {
char name[48];
snprintf(name, sizeof(name), "registry-%d name %d", arg->id, i);
e = akerr_register_status_name(owner, base + i, name);
AKERR_TCHECK(arg, e == NULL);
RELEASE_ERROR(e);
/* Read it back while the other threads are still inserting. */
AKERR_TCHECK(arg, strcmp(akerr_name_for_status(base + i, NULL), name) == 0);
/* And an entry nobody is touching: a probe sequence that a concurrent
* insert walked off loses names that were already there. */
AKERR_TCHECK(arg, strcmp(akerr_name_for_status(AKERR_VALUE, NULL),
"Value Error") == 0);
}
/* Everyone has reserved by the time anyone tries to trespass. */
pthread_barrier_wait(arg->barrier);
e = akerr_register_status_name(owner, victim, "Hijack");
AKERR_TCHECK(arg, e != NULL);
if ( e != NULL ) {
AKERR_TCHECK(arg, e->status == AKERR_STATUS_NAME_FOREIGN);
}
RELEASE_ERROR(e);
return NULL;
}
int main(void)
{
akerr_log_method = &akerr_thread_logger;
akerr_init();
AKERR_CHECK(akerr_slots_in_use() == 0);
int failures = akerr_run_threads(&registry_body);
AKERR_CHECK(failures == 0);
/* The contested range went to exactly one owner. */
AKERR_CHECK(contested_winners == 1);
/* Every name every thread registered is present and correct: nothing was
* dropped, overwritten, or attributed to the wrong thread. */
for ( int id = 1; id <= AKERR_TEST_THREADS; id++ ) {
int base = thread_range_base(id);
for ( int i = 0; i < NAMES_PER_THREAD; i++ ) {
char expected[48];
snprintf(expected, sizeof(expected), "registry-%d name %d", id, i);
AKERR_CHECK(strcmp(akerr_name_for_status(base + i, NULL), expected) == 0);
}
/* The trespass attempt did not leave a name behind. */
AKERR_CHECK(strcmp(akerr_name_for_status(base, NULL), "Hijack") != 0);
}
/* The library's own entries survived every one of those inserts. */
AKERR_CHECK(strcmp(akerr_name_for_status(AKERR_VALUE, NULL), "Value Error") == 0);
AKERR_CHECK(strcmp(akerr_name_for_status(AKERR_STATUS_NAME_FOREIGN, NULL),
"Foreign Status Name") == 0);
/* Thousands of refusals and registrations later, the pool is empty. */
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_threads_registry ok (%d threads x %d names)\n",
AKERR_TEST_THREADS, NAMES_PER_THREAD);
return 0;
}

View File

@@ -1,4 +1,10 @@
#include "akerror.h"
#include <stdlib.h>
static void expect_unhandled_nullpointer(akerr_ErrorContext *errctx)
{
exit((errctx != NULL && errctx->status == AKERR_NULLPOINTER) ? 1 : 0);
}
akerr_ErrorContext *func2(void)
{
@@ -25,6 +31,9 @@ akerr_ErrorContext *func1(void)
int main(void)
{
akerr_init();
akerr_handler_unhandled_error = &expect_unhandled_nullpointer;
PREPARE_ERROR(errctx);
ATTEMPT {
CATCH(errctx, func1());

View File

@@ -0,0 +1,64 @@
#include "akerror.h"
#include "err_capture.h"
#include <unistd.h>
#include <sys/wait.h>
/*
* The default unhandled-error handler is the library's last stop: it exits the
* process. Both of its exits were untested -- exit(1) for a NULL context (a
* handler invoked with no error at all) and, for a real one, the status handed
* to akerr_exit(). tests/err_exit_status.c covers what akerr_exit() does with a
* status; this covers that the handler reaches it, and the NULL case, which
* never gets that far.
*
* The handler never returns, so each case runs in a forked child and the test
* asserts the exact exit status. That is stricter than a WILL_FAIL test, which
* would pass on any non-zero exit, including one caused by an unrelated bug.
*/
/*
* Run the default handler on ctx in a child process and return the child's exit
* status, or -1 if it did not exit normally. The _exit() sentinel catches a
* handler that returns instead of terminating.
*/
static int handler_exit_status(akerr_ErrorContext *ctx)
{
pid_t pid = fork();
if ( pid == 0 ) {
akerr_default_handler_unhandled_error(ctx);
_exit(99);
}
int status = 0;
if ( pid < 0 || waitpid(pid, &status, 0) != pid || !WIFEXITED(status) ) {
return -1;
}
return WEXITSTATUS(status);
}
int main(void)
{
akerr_capture_install();
akerr_init();
/* No context to report: the handler has nothing to exit with but failure. */
AKERR_CHECK(handler_exit_status(NULL) == 1);
/*
* With a context, the status becomes the exit code. waitpid only reports
* its low 8 bits, which is where AKERR_VALUE (144) lands, and it is neither
* 0 nor the 1 used for the NULL case, so the two exits stay distinguishable.
*/
akerr_ErrorContext *slot = akerr_next_error();
AKERR_CHECK(slot != NULL);
slot->refcount = 1;
slot->status = AKERR_VALUE;
AKERR_CHECK((AKERR_VALUE & 0xff) != 0 && (AKERR_VALUE & 0xff) != 1);
AKERR_CHECK(handler_exit_status(slot) == (AKERR_VALUE & 0xff));
slot = akerr_release_error(slot);
AKERR_CHECK(slot == NULL);
AKERR_CHECK(akerr_slots_in_use() == 0);
fprintf(stderr, "err_unhandled_null ok\n");
return 0;
}