Records the human decision to test against the real gateway with real
data instead of building a mock gateway, six execution scenarios with
fault-source mixing, hard invariants scored from our own telemetry,
and the corpus inventory pulled from the lab server.
Classify empty completions as transient per human ruling (fixes flaky
real-gateway smoke and prevents caching empty responses), rename the
factory injection parameter gate to breaker per the frozen design,
rewrite the probe-entry cleanup without except BaseException, declare
python-dotenv explicitly, add a mid-backoff cancellation test, and
record all implementation errata in the design and architecture docs.
Includes config aggregation for multi-source env keys, from_env and
from_settings factories with explicit shared-backend injection,
gather_bounded, top-level exports, tightened import-linter layers with
the gate removed from the Makefile, and the finalized .env.example.
Also note that the seven architecture amendments predate the plan,
tighten T12 to independent implementation layers, and add fallback
guidance for read-only reference protocol imports.
Correct the section 6.1 reason enum against actual CHS code, make
retry_after_s a non-optional float, add structured_data to LLMResponse
extensions, finalize the structured parameter type, document SSE
missing-done semantics and watchdog thinking-token liveness, add the
StructuredMW layer to the default onion order, and register
middleware/structured.py in the module map.
Per 2026-07-20 human decisions: multi-source switching/cooldown ships
with M1 (retry-loop signatures freeze there), both GovDoc and
Video-Tree must pass the M1 onboarding smoke test, and Redis
integration tests use the lab's remote Redis with isolated namespaces.
Define the five-step ladder (native-schema prevention, repair, shape
validation, bounded feedback re-ask, ResultInvalidError) as D14;
rewrite section 7.9 accordingly, give the structured parameter its
three-tier semantics in the chat() signature, and scope the bounded
re-ask outside transport retry and breaker counting.
ruff check prints 'All checks passed!' on success, which the non-empty
output test counted as one problem and blocked every commit. Use
--quiet --output-format concise so success is silent and each
diagnostic is exactly one line.
Record tenacity vs self-built evaluation as D13 in ARCHITECTURE.md
(self-built wins: per-attempt orchestration, cancellation guarantees,
supply-chain discipline). Add research-wiki/ROADMAP.md expanding the
milestone table with ordering rationale and exit criteria.