Files
libakerror/docs/thread-safety.md
Andrew Kesterson 5695061130
Some checks failed
libakerror CI Build / cmake_build (push) Successful in 2m47s
libakerror CI Build / coverage (push) Successful in 2m48s
libakerror CI Build / thread_sanitizer (push) Failing after 2m49s
libakerror CI Build / mutation_test (push) Successful in 39m56s
Split the README reference material into docs/
The README was 794 lines: the summary, the design rationale, the whole macro
reference, the threading contract, the build internals and the exit-status
specification in one file. It is now 178 lines -- summary, installation,
quickstart, and an index -- and the reference material lives in docs/, one file
per topic: architecture, usage, status-codes, uncaught-errors, exit-status,
thread-safety, building.

The prose moved as written. Inbound references followed it: UPGRADING.md,
TODO.md, include/akerror.tmpl.h and tests/err_threads_handoff.c now name the
docs/ file that owns the text they cite, and AGENTS.md says where new
documentation goes so the README does not grow back.

Five factual errors fixed in the moved text:

- Both NULL-pointer examples inverted their test. FAIL_ZERO_* fails when the
  expression is zero, so `(somePointer == NULL)` failed on a *valid* pointer.
  They now read `(somePointer != NULL)`.
- AKERROR_NOIGNORE, four times including the #define, is AKERR_NOIGNORE.
- FINISH_NORExbTURN is FINISH_NORETURN.
- "functiions" is "functions".
- The architecture link pointed at include/akerror.h, which is generated and
  not in the tree; it points at include/akerror.tmpl.h.

The quickstart is new text. It compiles under -Wall -Wextra -Werror and was run
through all three of its paths: handled usage error exits 0, unhandled IO error
prints a trace and exits with the status, success exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-01 16:56:07 -04:00

226 lines
11 KiB
Markdown

# Thread safety
The library is thread safe as built by default. Every entry point may be called
from any thread at any time, including the first one: `akerr_init()` runs
exactly once no matter how many threads race into it.
What that covers:
* **The error pool.** Finding a free slot in `AKERR_ARRAY_ERROR` and taking its
reference is one operation under a lock, so two threads can never be handed
the same context. A context is then owned by exactly one thread at a time, all
the way through `CATCH`, `HANDLE`, and release.
* **The status registry.** Reservations and name registrations are serialized
against each other and against lookups. Two threads reserving the same range
cannot both win — exactly one gets `NULL` and the other gets
`AKERR_STATUS_RANGE_OVERLAP` naming the winner.
* **Per-thread state.** The context behind `IGNORE` (`__akerr_last_ignored`) and
the last-ditch context used to report `akerr_release_error(NULL)` are
thread-local, so one thread's ignored error is never another's.
* **Handing a context from one thread to another.** A context is not thread
state — it lives in `AKERR_ARRAY_ERROR`, which is process-global — so it
outlives the thread that raised it. The reference count is the only field the
library reads across threads, and it is only ever touched under the pool lock,
so `akerr_release_error()` does not care which thread checked the slot out.
Raise on a worker, queue it, let the worker exit; the collector still has a
whole error with a whole stack trace. See
[Handing an error to another thread](#handing-an-error-to-another-thread).
What it does not cover, and cannot:
* **Two threads inside one context at the same time.** A context has one owner,
and only one. `FAIL` rewrites the message, `HANDLE` rewinds the stack-trace
cursor, and none of that is locked — two threads in one context splice their
messages together and truncate each other's trace, without crashing. Handing a
context *from* one thread *to* another is a different thing, and is supported.
* **`akerr_log_method` and `akerr_handler_unhandled_error`.** Set them during
startup, before you spawn threads. They are read on every error and the
library never writes them after initialization, so setting one while other
threads are raising errors is a race the library cannot mediate.
* **Renaming a status that other threads are looking up.**
`akerr_name_for_status(status, NULL)` returns a pointer into the registry,
valid for the life of the process; registering a *second* name for the same
status overwrites that buffer in place. Register names during initialization.
Registering a *new* status concurrently is fine.
* **Which unhandled error terminates the process.** An error that reaches
`FINISH_NORETURN` unhandled prints its stack trace and calls
`akerr_handler_unhandled_error`, which by default calls `akerr_exit()`. Each
thread's trace is whole — the buffer belongs to its context, and each line is
one call to `akerr_log_method` — but if two threads get there at the same
instant, both traces print and the exit status is whichever one won.
There is one lock, it is recursive, and it covers both the pool and the
registry. That means error construction is serialized across threads: raising an
error is the exceptional path, and correctness there is worth more than
throughput. A program that raises errors on its hot path will feel it.
`AKERR_THREAD_SAFE` in the generated header is `1` for a thread-safe build, so a
consumer can check what it linked against:
```c
#if AKERR_THREAD_SAFE
/* ... start worker threads ... */
#endif
```
## Handing an error to another thread
**Transfer** an error context; never **share** one. One thread owns it at a time,
and ownership moves in a single step: the thread giving it up stops touching it
in the same act that makes it visible to the thread taking it over.
This works because a context is not thread state. It lives in
`AKERR_ARRAY_ERROR`, which is process-global, so a context outlives the thread
that raised it — a worker can raise an error, queue it, and exit, and the
collector still has a whole error with a whole stack trace. Nothing in the
library reads a context through any pointer but the one its current owner handed
in. The one field it reads across threads is the reference count, and that is
only ever touched under the pool lock, so `akerr_release_error()` does not care
which thread checked the slot out. Release it wherever it ended up.
**Always** hand it over through something that synchronizes: a mutex, a
condition variable, `pthread_join`, or an acquire/release atomic. The content of
a context is written with no lock at all — that is deliberate, error
construction should not pay for a lock it does not need — so the handoff itself
is what publishes those writes. Push the pointer through a relaxed atomic or a
plain global and the receiver can read a half-written message, on a machine you
did not test on.
**Never** touch a context after you have given it away. Not a `->status`, not a
log line. The receiver may be inside `HANDLE` rewinding the trace cursor, or
inside `akerr_release_error()` memsetting the slot.
**Release it exactly once.** A context released twice from a stale pointer takes
the refcount-zero branch a second time and wipes the slot again — which by then
holds somebody else's live error. Nothing crashes; a different thread's error
just quietly goes blank.
```c
/* Producer. Owns the context until queue_push() returns, and not after. */
static void report_unit_failure(int unit)
{
PREPARE_ERROR(e);
FAIL(e, MYLIB_UNIT_FAILED, "unit %d stopped responding", unit);
/* The last thing this thread does to it, so the trace records the crossing.
The buffer belongs to the context, so it travels with it. */
AKERR_STACKTRACE_APPEND(e, "%s:%s:%d: queued for the collector\n",
__FILE__, __func__, __LINE__);
queue_push(e); /* takes the queue mutex; `e` is not ours after this */
}
/* Collector. Owns it from the moment queue_pop() returns. */
static void collect_one(akerr_ErrorContext *e)
{
ATTEMPT {
} CLEANUP {
} PROCESS(e) {
} HANDLE(e, MYLIB_UNIT_FAILED) {
restart_unit(e);
} HANDLE_DEFAULT(e) {
LOG_ERROR_WITH_MESSAGE(e, "collector: unrecognized failure");
} FINISH_NORETURN(e);
}
static void *collector(void *unused)
{
akerr_ErrorContext *e;
(void)unused;
while ( (e = queue_pop()) != NULL ) {
collect_one(e);
}
return NULL;
}
```
Four things about the receiving side:
* **Declare a plain pointer, not `PREPARE_ERROR`.** That macro *declares* a
fresh context variable set to `NULL`; it cannot adopt one. For the same
reason, never `CATCH` into the variable holding a received context — `CATCH`
assigns over it, and the slot you were handed is gone.
* **In a `void` helper, use `FINISH_NORETURN`.** `FINISH_LOGIC` decides whether
to propagate at run time, so the compiler still *parses* its
`return __err_context` even when the second argument is the literal `false`.
`FINISH(e, false)` in a function returning void therefore draws
`warning: 'return' with a value, in function returning void` from gcc — a
constraint violation, and a build failure under `-Werror`. To propagate
inside the collector's own call stack, give the
helper an `akerr_ErrorContext *` return and use `FINISH(e, true)` as usual —
just never let an error propagate out of the thread body itself, whose
`void *` return nobody reads. `PASS` has the same problem for the same reason:
in a thread body it compiles, hands the pointer back as a `void *`, and leaks
the slot.
* **`FINISH_NORETURN(e)` on a received context still terminates the process.**
That is right — an unhandled error is unhandled wherever it was raised — but
it is now the collector's thread deciding the exit status, not the raiser's.
* **Give the handler blocks their own function.** `ATTEMPT` is a `switch`, and a
`break` inside one written directly in a loop leaves the `switch`, not the
loop. Same hazard as
[do not use CATCH or FAIL_*_BREAK inside a loop](usage.md#important-do-not-use-catch-or-fail__break-inside-a-loop).
**Always** bound the queue. A queued error is a checked-out pool slot, and there
are `AKERR_MAX_ARRAY_ERROR` (128) of them in the entire process. Size *queue
depth + producers with an error in flight* well under that. When the pool runs
dry the library logs and calls `exit(1)` from inside `FAIL` — there is no slot
left to raise the failure *from*, which is exactly why a collector that stops
draining takes the process with it.
### If you need to keep it as well as report it
There is no copy or retain call, on purpose: a context is a pool slot, and two
owners of one slot is the thing this whole section exists to prevent. So either
read out what you want to keep — `status` is an `int`, `message` and
`stacktracebuf` are ordinary NUL-terminated strings you can `snprintf` into your
own, much smaller, record — or raise two errors and hand one of them over.
**Never** copy an `akerr_ErrorContext` by assignment and keep the copy:
```c
akerr_ErrorContext snapshot = *failed; /* looks fine. is not. */
```
`stacktracebufptr` points into the context's *own* `stacktracebuf`, so after that
assignment `snapshot`'s cursor still points into `failed`'s buffer.
`LOG_ERROR(&snapshot)` reads the array and prints correctly, so it looks healthy
— and then the first `AKERR_STACKTRACE_APPEND(&snapshot, ...)` writes into a
pool slot you no longer own, arbitrarily far from the copy. `arrayid` has the
same shape: it is restored after the wipe, so a copied id makes the destination
impersonate the source's slot forever. And the copy is not a pool address, so
`akerr_valid_error_address()` rejects it, every `CATCH` on it becomes
`AKERR_BADEXC`, and releasing it memsets your own storage while the real slot
stays checked out for the life of the process.
## Building single threaded
The threading backend is chosen when libakerror is configured. `auto` (the
default) takes POSIX threads, and **fails the configure** if it cannot find
them rather than quietly producing a library that says it is thread safe and is
not. To mean it:
```sh
cmake -S . -B build -DAKERR_THREADS=none
```
That builds with no locking and no thread-local storage, stamps
`AKERR_THREAD_SAFE 0` into the header, and calling the library from more than
one thread is then undefined.
## Proving it
The thread tests (`tests/err_threads_*.c`) assert the properties above directly:
exclusive ownership of pool slots, exactly one winner for a contested range,
every registered name readable back, and — in `err_threads_handoff.c` — an error
raised on one thread arriving whole on another and released there. They run in
the normal suite. The run that
proves the *absence* of a data race underneath them is ThreadSanitizer:
```sh
scripts/thread_test.sh
```
which configures `build/tsan` with `-DAKERR_SANITIZE=thread`, builds the library
and every test with it, and runs the suite. Under that build a sanitizer report
fails the test rather than being printed and passed over.