_validate carried the whole matrix in one function (cc C/13, over the
branch quality gate). Splitting it by what is actually being checked —
shared, sqlite-only, postgres-only — puts every piece at A/B.
Ordering is the part that had to survive: the chain's order is the error
messages' priority, so a run with several bad flags still reports the
same one it did before. The --vacuum/--apply pairing therefore stays in
the shared step ahead of the backend branch, where it was; it is a "do
not rewrite the whole file when you only meant to look" rule, which
holds before the question of which backend a flag belongs to.
No behavior change: all nine parser.error strings are byte-identical and
in the same order, and the eight usage-error cases pass unmodified.
The library only ever SELECTs/INSERTs into llm_calls (D15), so expiring
rows has to live outside it — holding DELETE would contradict the
REVOKE UPDATE, DELETE the deployment template recommends.
tools/telemetry_retention.py is dry-run by default and prints the row
count, the created_at window and the tenant_id spread so an operator can
tell whether the rows about to go are the intended ones. The Postgres
branch refuses partitioned targets with exit code 3 (DETACH/DROP
PARTITION is O(1); DELETE is not) and otherwise deletes in per-batch
transactions. Missing asyncpg exits 2 rather than degrading quietly:
this is an ops tool, and a silent "0 rows" reads as "already clean".
Exit codes are the contract with the scheduler, so argparse errors were
moved off 2 (now 1) to keep "bad flags" distinguishable from "cannot
reach the database".
The Postgres cases run against the real instance in throwaway schemas —
never public.llm_calls — and the batch case asserts the shared table's
row count is unchanged, so a search_path that failed to apply lands as a
red test instead of a deletion.