39 Commits

Author SHA1 Message Date
iomgaa 17dcff41c3 Merge branch 'feat/issue-10-error-body-retention' 2026-08-16 23:34:20 -04:00
iomgaa fa4a7e220b test: make the 429 red line actually catch its violation
Verifier mutation test: flipping _translate_429 to parse the summary
left all 824 tests green. The padding was one long string value, so the
cut landed inside it - and head-and-tail retention kept the trailing
error object, leaving the summary parseable. Many keys put the cut
between structural tokens, where the summary stops being valid JSON.
Mutation now fails as it should. Also splits OCR 429 out on its own.
2026-08-16 06:50:50 -04:00
iomgaa 658086e2c0 docs: release 1.2.0 and unpin downstream from the 1.1 series
Issue #10 Task 6. The install pin moves from ==1.1.* to >=1.2,<2 - left
alone, everyone following the README would have stayed silently on
1.1.2 without this fix and without a warning. Telemetry field count
re-measured via inspect.signature: still 22.
2026-08-16 06:22:13 -04:00
iomgaa 9dada0be9d test: prove a rejected call's reason reaches the telemetry table
Issue #10 Task 5, the acceptance claim. Before the fix this asserted
against 'qwen_1 请求被拒: 400' and failed on the first substring - which
is exactly what the downstream batch was left with. Uses the real body
from the issue, and checks the trailing code too, since a head-only cut
would drop the one field you quote when chasing the provider.
2026-08-16 06:17:10 -04:00
iomgaa a3f4cc323f feat: keep the gateway's words on the OCR branches too
Issue #10 Task 4: the OCR side said only 'HTTP 404'. The issue reported
the chat path, but the batch that lost its 400 was reading tables - the
same blind spot, one transport over. Reuses the shared summarizer.
2026-08-16 06:14:07 -04:00
iomgaa 0edb9d397a feat: keep the gateway's words on every non-2xx chat branch
Issue #10 Task 3: five branches each built their own message, so adding
the summary would have meant five copies. Table-driven classification
composes it in one place instead, and the 429 split still parses the
untruncated body - reading the summary would demote an oversized
insufficient_quota to a plain rate limit and stop force_open.
2026-08-16 06:07:29 -04:00
iomgaa 484900d300 feat: add the single summarizer for HTTP error bodies
Issue #10 Task 2: head-and-tail rather than a head-only cut, because the
code and request_id that let you chase the provider sit at the very end
of a JSON error body. Cap 2048 follows k8s client-go for the same job.
2026-08-16 06:03:26 -04:00
iomgaa e302247022 feat: let every gateway error carry what the gateway said
Issue #10 Task 1: a rejected call's reason had nowhere to live. The
field goes on the base class because these errors all come from one HTTP
response - which class it is and what the peer said are orthogonal.
2026-08-16 06:01:38 -04:00
iomgaa 1489aab95d docs: add the missing imports to the plan's key interfaces
Codex review: the code blocks reference httpx and PolyGatewayError, but
neither module imports them today. A zero-context implementer copying
them verbatim would stall on F821.
2026-08-16 05:57:07 -04:00
iomgaa c2dd4a1cf4 docs: plan the implementation for issue #10 2026-08-16 05:50:37 -04:00
iomgaa 1801289277 docs: mark the issue #10 design approved 2026-08-16 05:34:46 -04:00
iomgaa 7462cad166 docs: widen the body cap to 2048 and keep the tail
The 500-char head-only rule came from a single sample. k8s client-go
caps the same thing at 2048; reprlib keeps head and tail because the
text is meant to be read. Gateway error bodies are JSON whose code and
request_id sit at the very end, so a head-only cut drops exactly what
you need to chase the provider. Version pinned at 1.2.0, which forces
the README install pin off ==1.1.*.
2026-08-16 05:24:30 -04:00
iomgaa 3cbe8aab91 docs: register the issue #10 design in the research wiki 2026-08-16 05:12:14 -04:00
iomgaa 707f8f7317 docs: pin the truncation rule to arithmetic after Codex review
"Truncate at cap and append the ellipsis" admits both 501 and 500 total
length; the two would desync test assertions from the telemetry length
promise. Cap is now the total including the marker.
2026-08-16 05:09:06 -04:00
iomgaa 10fbc5441e docs: design how the gateway's refusal survives the transport layer
Issue #10: the 400 body dies in _status_to_error, and telemetry only
writes str(exc), so adding a field alone would not make the refusal
queryable after the fact. Design keeps the summary in both the message
and a new base-class body_text, across every non-2xx branch and both
transports.
2026-08-16 05:03:33 -04:00
iomgaa 114fc8b1b3 Merge branch 'chore/packaging-metadata' 2026-08-07 21:48:15 -04:00
iomgaa 4f1ab21562 chore: give the package page a body and repo links
1.1.2 went out with an empty description on the registry page: without a
readme field there is no long_description, and twine only warns about
that -- it does not block the upload. Add readme and project.urls, and
record the wider lesson in the release procedure: a release is done when
the pages a downstream user actually opens look right, not when the local
steps go green. Also adds the missing step for creating a Release, which
is why the Releases page sat empty through eight tags.
2026-08-07 21:48:15 -04:00
iomgaa 2be89c47d8 Merge branch 'fix/issue-9-telemetry-ddl-probe' 2026-08-07 11:23:38 -04:00
iomgaa 7c60199680 chore: release 1.1.2
README first, per the release procedure: the 1.1.* pin still covers this
version, the 22-field count re-checked with inspect.signature, and the
telemetry row now states that an existing table needs no schema CREATE
privilege.
2026-08-07 11:23:29 -04:00
iomgaa 2e028d38f2 fix: probe for the telemetry table before creating it
PostgreSQL checks the schema CREATE privilege before the IF NOT EXISTS
existence test, so an account with only table-level INSERT was denied on
CREATE TABLE IF NOT EXISTS even though the table was right there and
writable. The denial set _failed and the whole recorder went no-op for
the process lifetime, silently: 150+ calls downstream lost their latency,
token and cost rows with nothing but one warning to show for it.

The probe is the direct fix. The larger fix is the criterion: structural
degradation now means "provably cannot write" (pool creation failed, or
the table is absent and cannot be created), not "something threw during
init" -- a probe or acquire failure just skips the row and retries on the
next call.

SQLite stays as it is on purpose. Measured: it short-circuits the
statement at parse time, so it passes even under another connection's
EXCLUSIVE lock or on a read-only file. A probe there would buy nothing;
the docstring now says so to keep symmetry-minded future edits away.
2026-08-07 11:21:33 -04:00
iomgaa c2e9f5396c docs: write down the release procedure that keeps getting skipped
Bumping the version is not releasing. 1.0.6 and 1.1.0 both got a version
bump and a changelog entry but were never uploaded, so the registry sat at
1.0.5 and downstream could not install any of those fixes.

The ordering matters in one non-obvious way: README has to be correct
before the build, because sdist freezes whatever is there at that moment.
That is exactly how 1.1.1 shipped with a stale README. The install pin is
called out by name since it is the easiest line to forget and the most
damaging to leave wrong.

Also records where the Gitea token actually lives — tea's config, not
.pypirc — after that misreading led to a wrong "no credentials" claim.
2026-08-06 12:44:06 -04:00
iomgaa d2cb8770df docs: correct the telemetry field count and document backpressure
Three README drifts, none of them about issue #8 alone:

- the telemetry row said 18 fields; record_llm_call takes 22 (verified by
  inspect.signature). ports.py claimed 20 in its own docstring, so both
  records of the same fact were stale.
- LLMResponse.cost is hardcoded to None on the chat path (retry.py:519).
  Cost only ever reaches telemetry. The old doc site named this as a known
  trap, so the capability row now says it outright.
- backpressure had no row at all, which is what issue #8 was about.
2026-08-06 12:33:45 -04:00
iomgaa 80bc94c42d docs: point the install pin at 1.1.x
The pin still said ==1.0.*, which caps downstream at 1.0.5 and hides
every fix since. Now that 1.1.1 is actually in the registry the pin can
move; it was missed when 1.1.0 was tagged.
2026-08-06 12:24:44 -04:00
iomgaa bb69ecf0da Merge branch 'feat/issue-8-stall-budget'
stall 判定改为非生产性等待口径(issue #8)并发布 1.1.1。
2026-08-06 12:12:42 -04:00
iomgaa 014fc2bfa7 chore: release 1.1.1
Patch rather than minor: the error surface is unchanged and no public
signature moved. What downstream must notice is timing, not types — the
worst-case call duration rises to roughly max_attempts * timeout_s now
that the retry budget actually applies.

Pre-release review caught an overreaching promise in the changelog entry:
the 429 bound holds only when the stall verdict can fire at all, i.e. when
the whole scope has no progress. The verdict is a conjunction, so a call
does not die while other calls in the scope are still producing — by
design — which leaves no hard per-call ceiling in that case. That property
predates this fix and is now stated with its precondition instead of as an
unconditional guarantee.

The wiki sync in the release checklist is a no-op again: the doc site has
been down since 2026-08-02 and its landing page names CHANGELOG.md as the
version source of truth, which this commit updates.
2026-08-06 11:57:06 -04:00
iomgaa f3e06eac89 chore: register the issue #8 design and plan in the research wiki
Registration pages carry the chosen approach, why the split is by "which
budget the time consumes", the five rejected alternatives with reasons,
and the 3.6 correction found during independent verification.
2026-08-06 11:09:37 -04:00
iomgaa a0a5cf7ecc fix: return 429 attempt time to the stall budget
Independent verification found the first cut had swapped one bug for a
worse one. The budgets were split by "did we send a request", so a 429
attempt counted as productive — but 429 is exempt from the retry budget,
so its time burned neither budget. Against a queueing gateway that holds
the request for the full timeout before answering 429, a call could hang
for 301 attempts / 25.2 hours, measured, versus 301 seconds before the
change.

The split is now by which budget the time consumes: time that burns
max_attempts is excluded from stall, time that does not (429 attempts
included) belongs to stall. Measured again: back to one attempt / 301s.

Only the chat loop needs this — embedding and ocr count 429 against
max_attempts unconditionally, so the gap never existed there. The stall
verdict moved into _stalled(), which both call sites had duplicated, to
keep __call__ under the complexity gate.
2026-08-06 10:55:51 -04:00
iomgaa bc4683d1f5 test: make the per-call clock invariant actually testable
The concurrency case used two RetryMW instances, so instance-level sharing
was hidden by object isolation and a clock promoted to an instance
attribute passed all seven cases. Both cases now reuse one mw, and a new
one idles past the window between two calls on that instance — the shape
that would expose _entered_at pinned to process start. Mutation-checked:
promoting the clock fails the new case.
2026-08-06 10:36:40 -04:00
iomgaa 3645e574d3 docs: record the stall metering change in architecture and changelog
ARCHITECTURE.md 7.3 now carries the new metering and notes that the G6
ttft guard became conservative redundancy. The changelog entry leads with
what downstream must act on: the worst-case call duration rises to
max_attempts * timeout_s, and any STALL_WINDOW_S that was inflated to work
around this can go back to the default.
2026-08-06 10:18:11 -04:00
iomgaa d05114e895 docs: align the stall window comments with the new metering
_validate_stall still guards stall_window_s >= max ttft_timeout_s, but its
stated reason no longer holds: TTFT waiting is productive time and never
reaches the stall account. The check is harmless and stays, so the
docstring now says why it is kept rather than implying a live hazard.
.env.example dropped the "must be >= max TTFT" advice for what the window
actually measures.
2026-08-06 10:09:06 -04:00
iomgaa 0477d9534b fix: apply the non-productive stall budget to the ocr loop
Same failure path as the embedding loop: one timed-out attempt drains the
wall-clock window, and the next round without a runnable source declares
the scope dead in _on_no_runnable. All three governance loops now meter
stall the same way.
2026-08-06 09:51:30 -04:00
iomgaa 6d0f3c9044 fix: apply the non-productive stall budget to the embedding loop
The embedding loop shares the wall-clock entered_at and the same stall
verdict, so it failed the same way through a different path: one timed-out
attempt, then any round with no runnable source, and _on_no_runnable
declared the scope dead. Issue #8 only recorded the chat path; the
regression test pins this one.
2026-08-06 09:42:07 -04:00
iomgaa 02c3d06ec6 fix: bill only non-productive waiting against the chat stall budget
Issue #8: with timeout_s >= stall_window_s a single timed-out request
exhausted the stall window before the second attempt was even dispatched,
so LLM_MAX_RETRIES never applied and the whole scope was declared dead.

Root cause is that real attempts and non-productive waiting charged the
same wall clock, while the stall budget is the smaller of the two. The new
StallClock subtracts attempt time from the stall account, leaving the two
budgets orthogonal: attempts bill max_attempts, waiting bills
stall_window_s. The dual-condition verdict, the inf semantics of
progress_age_s, the 429 exemption and the error surface are untouched.

The productive boundary is _attempt itself, telemetry included, so a slow
recorder cannot push a call into a stalled verdict.
2026-08-06 09:20:21 -04:00
iomgaa 573e505a4b docs: add the implementation plan for issue #8
Six tasks: StallClock plus the chat loop, then embedding, ocr, the config
comments, the full-suite regression with doc sync, and independent
verification. Codex review raised four points, all confirmed and folded in:
a stale line reference in the fidelity section, explicit cancellation
acceptance for T2/T3 (the new attempting() wrapper now wraps their existing
cancel paths), a telemetry-boundary test pinning the design's claim that
telemetry jitter must not feed the stall verdict, and concrete test
construction for the embedding/ocr regressions.
2026-08-06 09:09:31 -04:00
iomgaa ce2dda7d45 docs: sharpen the productive-time boundary after Codex review
Two internal-consistency fixes from the independent design review:
the 429 saturation argument wrongly claimed exponential backoff growth
(429 skips the retry budget, so max(fails, 1) pins the delay to the base
tier), and "productive" was defined as waiting on the response while the
StallClock actually wraps all of _attempt. The boundary is now stated as
_attempt itself, including per-attempt accounting and telemetry, with the
rationale that telemetry jitter must not participate in the stall verdict.
2026-08-06 08:16:07 -04:00
iomgaa bfe423ddf8 docs: bill only non-productive waiting against the stall budget
Issue #8: a single request that burns its full timeout_s also exhausts
stall_window_s, so the retry budget silently never applies. Root cause is
that both budgets charge the same wall-clock time. The design makes the two
budgets orthogonal — real attempts bill the retry budget, everything else
bills the stall budget — which drops the timeout_s / stall_window_s coupling
instead of guarding it with an assembly-time check.
2026-08-06 08:01:07 -04:00
iomgaa 9c2824ce8a fix: admit governance backend failures into scope-level unavailability (issue #7)
A fail-closed limiter or breaker backend means the scope cannot emit a
single request, yet GovernanceBackendError sat directly under
PolyGatewayError. A caller writing only `except GatewayUnavailableError`
dropped it into the catch-all branch, so a Redis blip burned a backlog's
business failure budget into the dead letter queue over a fault a restart
would clear. It now inherits GatewayUnavailableError with a
governance_backend_down reason and a 5 second retry_after_s.

The two unknown-source sites split out into SourceNotConfiguredError,
deliberately outside the retryable family: a misconfigured source name
must burn its budget and surface rather than retry forever in silence.

README now states which errors reach callers and which the retry loop
absorbs. TransientError and SourceDeadError read like caller contracts but
never arrive, and a downstream project wrote a whole design section on
that false premise before checking the source.

Independent verification caught the split not actually holding on the only
path production uses, and caught the fix for that opening a second hole on
the accounting path. Both are fixed and pinned by tests that go through
the wrappers rather than the private methods underneath them.
2026-08-06 06:55:38 -04:00
iomgaa 5853c3f8ff fix: keep the accounting path degrading after the wrapper change
Letting SourceNotConfiguredError through the gate wrappers opened a hole
the recheck caught: _record_quietly only degrades GovernanceBackendError,
so an assembly defect raised from the accounting side would now escape and
destroy a response from a call that had already genuinely succeeded. That
inverts the exact invariant _record_quietly exists to hold.

Widening _record_quietly is the right fix rather than narrowing the
wrappers, because that layer degrades by what the path is (accounting, the
call is already done) rather than by which error type shows up. Narrowing
would have left 4 of 9 wrapper methods as exceptions to a rule nobody can
remember.

No backend raises it from an accounting method today, so this is a
guardrail for whoever adds source-name validation to a breaker backend.

The stub that first reported this green was wrong: its record_success
lacked count_attempt, so it raised TypeError and the wrapper relabeled it.
Fixed signature, then the test failed as it should have.

Also finishes the three-to-five leak path correction across the four
remaining spots, including the wiki summary card that indexes this design.
2026-08-06 06:39:52 -04:00
iomgaa a57a5cea72 fix: let assembly defects pierce the gate wrappers
Independent verification caught that the split shipped in the previous
commit did not actually hold on the only path production uses. The gate
wrappers re-raise GovernanceBackendError but nothing else, so
SourceNotConfiguredError fell into the following `except Exception` and
came back out as a governance_backend_down failure with retry_after_s=5.0.
A misconfigured source name would still retry forever and never surface.

The existing tests missed it because both of them call the private _cfg()
directly, one layer below the wrapper the governance loops actually go
through. The regression test goes through QuotaGate.

telemetry.py has to widen its terminal catch in the same commit: once the
wrapper stops relabeling the error, it is no longer a GovernanceBackendError,
and it is raised before any attempt exists, so the path would have recorded
no telemetry at all.

Also corrects the leak path count from three to five. QuotaGate.stats and
BreakerGate.retry_after_s are not wrapped by _record_quietly either.
2026-08-06 05:57:50 -04:00
47 changed files with 2800 additions and 133 deletions
+1 -1
View File
@@ -38,7 +38,7 @@ LLM_CIRCUIT_BREAKER_COOLDOWN=60 # 或 LLM__BREAKER__COOLDOWN_S
# LLM_TTFT_TIMEOUT=30 # 平铺看门狗缺省(成对生效) # LLM_TTFT_TIMEOUT=30 # 平铺看门狗缺省(成对生效)
# LLM_INTER_TOKEN_TIMEOUT=15 # LLM_INTER_TOKEN_TIMEOUT=15
# LLM__BREAKER__PROBE_TTL_S=240 # 缺省派生: max(2×最大源超时, cooldown, 最大源超时+5);显式值须 ≥ 最大源超时+5 # LLM__BREAKER__PROBE_TTL_S=240 # 缺省派生: max(2×最大源超时, cooldown, 最大源超时+5);显式值须 ≥ 最大源超时+5
# LLM__BACKPRESSURE__STALL_WINDOW_S=300 # stall 双条件判死窗口;须 ≥ 最大源 TTFT # LLM__BACKPRESSURE__STALL_WINDOW_S=300 # stall 双条件判死窗口;只计非生产性等待(429 退避/配额轮询/熔断冷却),与 TIMEOUT_S 无耦合,无需按 timeout×retries 放大
# LLM__BACKPRESSURE__POLL_INTERVAL_S=0.05 # LLM__BACKPRESSURE__POLL_INTERVAL_S=0.05
# LLM__SELECTOR=health_aware # health_aware(默认,M2.5) | round_robin | least_inflight # LLM__SELECTOR=health_aware # health_aware(默认,M2.5) | round_robin | least_inflight
# ── M2.5 失败率熔断通道(可选,缺省即生产推荐值)── # ── M2.5 失败率熔断通道(可选,缺省即生产推荐值)──
+72 -2
View File
@@ -1,5 +1,75 @@
# Changelog # Changelog
## 1.2.0(2026-08-16)
网关拒绝一次调用时,**它说的话不再丢失**(issue #10)。下游一轮 1050 张医学影像的批处理里,1 张在读表格这一步收到 400、被判确定性失败而放弃;事后想知道"这张图到底哪里不合规",无从查起——响应体在 transport 翻译层之后就不存在于进程任何位置了。
根因是三条留存通道同时为空: `_status_to_error` 手上握着 `body_text` 却只用于 429 的类型细分,该模块没有任何 logger 调用,异常类也没有承载响应体的字段。而库的逐次遥测写的是 `str(exc)`,即 message——所以**只给异常加字段并不能让它进遥测表**,必须两者都做。
### 新增
- **四分类错误新增 `body_text` 字段**(加在 `PolyGatewayError` 基类): 非 2xx 响应体的摘要。与 `ResultInvalidError.raw_text` 分工明确——前者是"对方拒绝的理由"(非 2xx),后者是"2xx 但内容不可解析时的模型输出"。scope 级错误(`GatewayUnavailableError` 一族)恒为空串: 它们没有单一响应体可言。
- **同一份摘要同时进入异常 message**,故 SQLite/Postgres 遥测的 `error` 列里直接可查,下游不必为此单独埋点。
### 行为变更
- **非 2xx 的 message 末尾追加 ` | {响应体摘要}`**,覆盖两个 transport 的**全部**分支: chat 的 400 / 401·403 / 4xx 兜底 / 5xx / 429 两支(含 `insufficient_quota`),以及 OCR 的全部分支。issue 只报告了 chat 的 400,但 401 会 `force_open` 整个源、OCR 侧 message 原本只有一个状态码,是同一个缺陷的其余分支。
- 摘要口径: 先折叠空白(错误体常是缩进 JSON,原样拼进 message 会把一行日志炸成多行),再限长 **2048 字符**(对齐 Kubernetes client-go 同场景的 `maxUnstructuredResponseTextBytes`)。超长时**保留头 1400 + 尾 600**并记下省略字数——JSON 错误体的 `code` / `request_id` 收在尾部,头部硬切正好会切掉向网关方追查时唯一有用的那部分。
- 遥测 `error` 列因此变长: 纯 ASCII 约 2KB/条,最坏(5xx 重试 3 次)一次调用约 6KB。
### 不变
- **状态码 → 错误分类的映射逐条未动**(ARCHITECTURE §6.2 表),`retry_after_s` 解析、429 免重试预算、`insufficient_quota` 细分全部保持——429 的类型判定仍解析**未截断的原文**,若改用摘要,超长 body 的配额耗尽会退化成普通限速、该源不再 `force_open`
- 异常类型树、`str(exc)` 之外的字段、遥测 22 字段与列序、DDL 全部未变。**错误面零变更**,下游 `except` 写法不受影响。
- 400 仍按确定性失败处理(不重试不换源)。**但请注意**: 经第三方中转部署时,中转自身抖动也会回 400,从状态码上与"你的输入有问题"分不开(下游实测: 同一份字节 sha256 一致、重发 15 次全部成功,失败那次 `prompt_tokens=0` 且耗时远低于任何成功调用)。库不改默认语义——直连供应商时重试只会白烧配额——但 `body_text` 现在给了下游自行区分的判据。
### 升级提示
README 的安装 pin 由 `==1.1.*` 改为 `>=1.2,<2`。**仍按 `==1.1.*` 安装的下游会静默停在 1.1.2**,拿不到本次修复且没有任何报错,请同步改自己的依赖约束。
- 打包元数据补齐: `readme``[project.urls]`。1.1.2 及之前的包在 registry 页面上**没有任何说明正文**(缺 `readme` 时 twine 只警告不阻塞),也没有仓库链接。代码零变更,自本版生效。
## 1.1.2(2026-08-07)
Postgres 遥测撞上建表权限就整体判死的问题(issue #9)。**最小权限部署会静默丢掉全部遥测**: 应用账号有表级 `INSERT`、表也已存在,但没有 schema 的 `CREATE` 权限时,初始化的 `CREATE TABLE IF NOT EXISTS` 被拒 → recorder 永久 no-op,业务调用一切正常,只留一行 warning。下游 CHSAnalyzer3 首次端到端跑的 150+ 次调用耗时/token/成本因此全部丢失,且事后无法补回。
根因是 **PostgreSQL 对 schema 的 CREATE 权限检查早于 `IF NOT EXISTS` 的存在性判断**(PG 16.14 实测: 同一连接 `INSERT` 通过、`to_regclass` 看得见表,该 DDL 照样被拒)——与 issue #3 修过的 `ALTER TABLE` 是同一类问题,当时只修了补列那一半。
### 行为变更
- **PG 侧建表前先 `to_regclass` 探测,表已存在就一条 DDL 都不发**。探测不需要任何权限,且与 `INSERT` 走同一套 search_path 解析(比裸 DDL 更准: 裸 `CREATE TABLE` 落在首个**可建**的 schema,可能与写入命中的不是同一张表)。表不存在时才建,新建表列已齐全,顺带跳过补列。
- **"结构性失能"的判据收窄为「确定写不进去」**,不再是「初始化时出过异常」。仅两种情形仍永久降级为 no-op: 建池失败(重试要在业务路径上内联吞掉连接超时)、表确定不存在且建不出来(后续 INSERT 必然全败)。探测失败、取连接失败改为**只跳过本条并 warning,下次调用重新准备**——初始化瞬间的一次抖动不再让整个进程失遥测。
- 日志措辞随之细分: `建池失败` / `建表探测失败(跳过本条,下次重试)` / `建表失败(表不存在,记录无处可落)`,原先一律是 `初始化失败`
### 不变
- SQLite 侧**一行未改**。实测其对已存在的表在解析期就把 `CREATE TABLE IF NOT EXISTS` 短路掉(另一连接持 `BEGIN EXCLUSIVE`、文件 `chmod 444` 时该语句均通过,而同条件的 `INSERT` 分别报 database is locked / readonly database),没有同款风险;加探测零收益,故有意不对称,只在 docstring 钉死实测结论。
- 遥测端口签名、22 字段、列序、`ON CONFLICT DO NOTHING` 幂等、单条写失败逐行丢弃的降级方向全部未动。**错误面零变更**。
### 升级提示
若你的部署此前为了绕开本问题给应用账号授了 `CREATE ON SCHEMA`,现在可以收回——表存在时库不再需要该权限。
## 1.1.1(2026-08-06)
stall 判定改为非生产性等待口径(issue #8)。`timeout_s ≥ stall_window_s` 时,**一次耗满超时的请求就会让整个 scope 被判死,配置的重试次数一次都用不上**——而且没有任何报错或 warning,配置方以为自己配了 3 次重试。`stall_window_s` 默认 300 恰是个很容易被 `TIMEOUT_S` 追平的值,"只配 timeout、不配 stall"这种最常见的写法正好踩中。
根因是**两个预算重叠计费**: 真实尝试的耗时同时向重试预算(`max_attempts`)与 stall 预算(`stall_window_s`)计费,而后者更小,必然先耗尽。
### 行为变更(**请先读这一条**)
- **stall 判定的"本地超窗"条件现在只累计非生产性等待**——429 退避、配额 wait 轮询、熔断冷却;消耗重试预算的真实尝试不再计入。两个预算自此正交,划分依据是**谁消耗重试预算**: 烧 `max_attempts` 的时间不烧 `stall_window_s`,不烧 `max_attempts` 的时间(含 429 尝试本身)归 `stall_window_s` 治理。
- **`stall_window_s``timeout_s` 不再有任何耦合**,无需按 `timeout × retries` 放大。若你此前为绕开本 bug 把 `STALL_WINDOW_S` 调大过,现在可以回到默认值。
- **单次调用的最坏耗时由 `stall_window_s` 抬升到约 `max_attempts × timeout_s`**(默认配置下 3 × `TIMEOUT_S`,再加各次退避)。这是重试预算恢复生效的正确表现,但如果你的上游有调用超时,请据此复核。429 路径同样不突破这个量级——429 虽免重试预算,但其尝试耗时计入 stall 账。
**上述量级的前提是 stall 判死能够触发**,即整个 scope 无进展(`progress_age_s() > stall_window_s`)。判死是**双条件合取**,这一条未变: 若同 scope 里其他调用仍在正常出餐,本调用会继续等待换源而不判死——这正是双条件的设计意图("别人还活着,不该因我一路不顺就宣告整个 scope 死亡")。**代价是这种情形下调用级没有硬上限**,持续遭遇慢 429 的调用可以等很久。该性质由条件 B 单独门控,**早于本次修复即如此**(旧口径实测同样无界),不是本次引入;但若你需要调用级硬上限,请在调用方用 `asyncio.wait_for` 自行设置。
- 三条治理循环(chat / embedding / ocr)口径一致。**embedding 与 ocr 此前有同一缺陷**(经"先超时一次、再遇到无可用源"触发),issue 只记录了 chat 路径。
- 遥测收尾属"真实尝试"边界之内,**遥测抖动不会把一次调用推进 stalled 判决**。
### 不变
- 双条件判死的结构、`progress_age_s()``inf` 语义(从未出餐 = 全局超窗)、429 免预算、退避与 jitter 公式、`fail_fast` 分支、`AllSourcesExhausted` 的字段与 `reason` 取值(仍是 `stalled`)全部未动。**错误面零变更**,下游 `except` 写法不受影响。
- 装配期校验 `stall_window_s ≥ 最大源 ttft_timeout_s` 保留。新口径下它已是保守冗余(TTFT 等待属生产性时间),但无害且不误拒合理配置。
## 1.1.0(2026-08-06) ## 1.1.0(2026-08-06)
治理后端故障归位为 scope 级不可用(issue #7)。限流/熔断的状态后端(Redis 等)自身故障时,库按降级方向铁律 fail-closed——**整个 scope 一个请求都发不出去**,语义上就是"scope 级暂时不可用"。但 `GovernanceBackendError` 此前是 `PolyGatewayError` 的直接子类,只写 `except GatewayUnavailableError` 的调用方接不住,后果很具体: Redis 抖一下,积压任务一批批消耗业务失败预算,够到上限就进死信——**而那是运维重启一下就好的故障**。 治理后端故障归位为 scope 级不可用(issue #7)。限流/熔断的状态后端(Redis 等)自身故障时,库按降级方向铁律 fail-closed——**整个 scope 一个请求都发不出去**,语义上就是"scope 级暂时不可用"。但 `GovernanceBackendError` 此前是 `PolyGatewayError` 的直接子类,只写 `except GatewayUnavailableError` 的调用方接不住,后果很具体: Redis 抖一下,积压任务一批批消耗业务失败预算,够到上限就进死信——**而那是运维重启一下就好的故障**。
@@ -19,8 +89,8 @@
### 下游请读 ### 下游请读
- **`GovernanceBackendError` 现携带 `scope` / `reason` / `retry_after_s` / `per_source_reasons`**,与 `AllSourcesExhausted` 同款;`str(exc)` 仍是原来的诊断串(如 `限流后端 try_acquire 失败: ...`),结构化字段与诊断信息并存,排障不受影响。 - **`GovernanceBackendError` 现携带 `scope` / `reason` / `retry_after_s`**,与 `AllSourcesExhausted` 同款(`per_source_reasons` 属性存在但恒为 `{}`——后端故障不针对具体某个源);`str(exc)` 仍是原来的诊断串(如 `限流后端 try_acquire 失败: ...`),结构化字段与诊断信息并存,排障不受影响。
- **条闸门路径**(`try_acquire` / `try_enter` / `progress_age_s`)的后端故障会到达调用方;记账路径(`record_success`)仍被 `_record_quietly` 降级为 warning,这个分工不变。 - **条闸门路径**的后端故障会到达调用方: `QuotaGate``try_acquire` / `stats` / `progress_age_s`,`BreakerGate``try_enter` / `retry_after_s`。记账路径(`record_success` / `record_failure` / `release_probe` / `mark_progress`)仍被 `_record_quietly` 降级为 warning,这个分工不变。
- **CHSAnalyzer 迁移**: `tracking.py` 一条 `except GatewayUnavailableError` 即覆盖完整,无需为后端故障单列分支(`migrations/chsanalyzer.md` G1 已补注)。 - **CHSAnalyzer 迁移**: `tracking.py` 一条 `except GatewayUnavailableError` 即覆盖完整,无需为后端故障单列分支(`migrations/chsanalyzer.md` G1 已补注)。
## 1.0.6(2026-08-02) ## 1.0.6(2026-08-02)
+25
View File
@@ -78,6 +78,31 @@ make ci # 只读验证(check + test)
### 4.4 Git 工作流 ### 4.4 Git 工作流
- 一切开发在 feature 分支,严禁直改 main;频繁语义化提交;提交**必须**调用 `commit` skill;大改动前先提交回滚点。 - 一切开发在 feature 分支,严禁直改 main;频繁语义化提交;提交**必须**调用 `commit` skill;大改动前先提交回滚点。
### 4.4.1 发布流程(每步都是历史欠账换来的,不得跳步)
> [!CRITICAL]
> **发布 = 合并 + push + tag + 构建 + 上传 registry。只 bump 版本号不叫发布。**
> 教训: 1.0.6 与 1.1.0 都完成了版本号 bump 与 CHANGELOG,却从未上传,registry 长期停在 1.0.5——下游 `pip install` 拿不到任何修复,且无人发现。
按顺序执行,**构建之前**必须先改完所有文档:
| # | 动作 | 要点 |
|---|---|---|
| 1 | **更新 README** | 打包会把当时的 README 固化进 sdist,**发布后再改就来不及了**(包里那份永远是旧的)。逐项核对: 安装命令的版本约束(`==1.1.*` 这类**极易漏改**,漏了下游就被锁在旧版)、能力表是否覆盖新行为、数字型断言是否仍成立(如遥测字段数,须用 `inspect.signature` 实测而非凭记忆) |
| 2 | CHANGELOG 定版 | "未发布" → `## X.Y.Z(日期)` |
| 3 | 版本号 | `pyproject.toml` + `src/polygateway/__init__.py` 两处必须一致 |
| 4 | 合并 main + push | `--no-ff`;合并后在 main 上重跑 `make lint` 与全套件 |
| 5 | **打 tag 并 push** | `git tag -a vX.Y.Z -m "..."` + `git push origin vX.Y.Z`。历史上多个版本漏打 |
| 6 | 构建 | `rm -rf dist && python -m build && python -m twine check dist/*` |
| 7 | **上传 registry** | 凭据在 `~/.config/tea/config.yml`(tea CLI 的 Gitea token,**不在** `~/.pypirc`);token 走 `TWINE_PASSWORD` 环境变量,不进命令行<br>`TWINE_USERNAME=iomgaa TWINE_PASSWORD=$TOKEN python -m twine upload --repository-url https://gitea.iomgaa.online/api/packages/iomgaa/pypi dist/*` |
| 8 | **验证已发布** | `pip download --no-deps --index-url .../pypi/simple/ "polygateway==X.Y.Z"`,并解包确认新代码在内。**不验证不算发布完成** |
| 9 | **建 Release + 核对包页面** | `POST /api/v1/repos/iomgaa/PolyGateway/releases`(body 取 CHANGELOG 本版段;历史上只打 tag 不建 release,Releases 页长期为空);随后打开包页面确认有正文与仓库链接,**Link to a repository 只能在网页手动做**(该实例的 link API 返 404) |
> [!CRITICAL]
> **发布完成的判据是外部可见结果,不是本地步骤跑通**: 收尾必须以下游视角逐一打开产物页面(registry 包页面正文与仓库链接、仓库 Releases 页、`pip install` 后包内文件),看到什么算什么,缺的当场补进本清单——1.1.2 三步全绿却出现包页面空白(`pyproject` 缺 `readme`)、Releases 页 0 条、包未挂仓库。
Gitea 包 registry 是 **owner 级**(`/iomgaa/-/packages/`)不是仓库级;PyPI 元数据不含仓库字段,故不会自动挂到 `PolyGateway/packages`,需在包页面手动 Link to a repository。
### 4.5 配置管理 ### 4.5 配置管理
- 工程配置走 `pydantic-settings` + `.env`(模板 `.env.example`,敏感项不提交);严禁硬编码默认值;缺失关键配置直接报错。 - 工程配置走 `pydantic-settings` + `.env`(模板 `.env.example`,敏感项不提交);严禁硬编码默认值;缺失关键配置直接报错。
- 多源命名约定 `{SCOPE}__{PROVIDER}__{N}__{FIELD}`;韧性参数键名沿用三项目习惯(`LLM_TIMEOUT` 等),降低迁移成本。 - 多源命名约定 `{SCOPE}__{PROVIDER}__{N}__{FIELD}`;韧性参数键名沿用三项目习惯(`LLM_TIMEOUT` 等),降低迁移成本。
+5 -2
View File
@@ -15,9 +15,10 @@
| 错误分类重试 | 一切失败落入四分类(见下),由分类决定重试/换源/熔断;429 属 pushback 不消耗重试预算;退避含 jitter 且尊重 Retry-After | | 错误分类重试 | 一切失败落入四分类(见下),由分类决定重试/换源/熔断;429 属 pushback 不消耗重试预算;退避含 jitter 且尊重 Retry-After |
| 熔断 | 双通道(连续失败 + 失败率窗口,健康证据抑制误熔);半开单探针带租约(持有者死亡自动回收);epoch fencing 拒绝迟到写回;开路时长指数递增 | | 熔断 | 双通道(连续失败 + 失败率窗口,健康证据抑制误熔);半开单探针带租约(持有者死亡自动回收);epoch fencing 拒绝迟到写回;开路时长指数递增 |
| 自适应并发 | AIMD:429 削减、成功缓升,防止打爆上游 | | 自适应并发 | AIMD:429 削减、成功缓升,防止打爆上游 |
| 背压与判死 | 配额满可选等待或快速失败;等待期按双条件判死(本地非生产性等待与全局无进展**同时**超窗)。stall 窗口只计**非生产性**等待(429 退避/配额轮询/熔断冷却),与 `TIMEOUT_S` 无耦合 |
| 响应缓存 | Redis/内存;key 含 model + messages 摘要 + namespace/租户 + salt,多模态 content 先摘要再 hash(防毒化);可 per-call 绕过(科研重采样) | | 响应缓存 | Redis/内存;key 含 model + messages 摘要 + namespace/租户 + salt,多模态 content 先摘要再 hash(防毒化);可 per-call 绕过(科研重采样) |
| 流式看门狗 | TTFT / inter-token / 总超时三层活性;thinking token 刷活性不计结果;截断流(缺 `[DONE]`)判瞬时不入缓存 | | 流式看门狗 | TTFT / inter-token / 总超时三层活性;thinking token 刷活性不计结果;截断流(缺 `[DONE]`)判瞬时不入缓存 |
| 遥测与成本 | 每次调用(含缓存命中与失败)必录 18 字段;SQLite / Postgres 后端;按价格表折算成本;多模态内容摘要落库不存原图 | | 遥测与成本 | 每次调用(含缓存命中与失败)必录 22 字段;SQLite / Postgres 后端(表已存在时**不需要** schema 建表权限,最小权限账号可直接用);按价格表折算成本落库(注意 `LLMResponse.cost` 本身恒为 `None`,成本只进遥测);多模态内容摘要落库不存原图 |
| 结构化输出 | json_repair 修复 / 原生 schema 双策略 + 校验失败有界带反馈重问 | | 结构化输出 | json_repair 修复 / 原生 schema 双策略 + 校验失败有界带反馈重问 |
| OCR | MonkeyOCR 双端点(文本转录 + 版面解析),bbox 数值防御下沉,逐源健康预检 `check_health()` | | OCR | MonkeyOCR 双端点(文本转录 + 版面解析),bbox 数值防御下沉,逐源健康预检 `check_health()` |
| Embedding | 分批、维度校验、与 chat 同一治理栈 | | Embedding | 分批、维度校验、与 chat 同一治理栈 |
@@ -30,7 +31,7 @@
```bash ```bash
pip install --extra-index-url https://gitea.iomgaa.online/api/packages/iomgaa/pypi/simple/ \ pip install --extra-index-url https://gitea.iomgaa.online/api/packages/iomgaa/pypi/simple/ \
"polygateway[redis,postgres,structured]==1.0.*" "polygateway[redis,postgres,structured]>=1.2,<2"
``` ```
核心仅依赖 `httpx` + `pydantic`;按需选 extras: 核心仅依赖 `httpx` + `pydantic`;按需选 extras:
@@ -124,6 +125,8 @@ except RequestRejectedError:
预算耗尽/全源熔断时抛 `GatewayUnavailableError` 族(`CircuitOpenError` / `AllSourcesExhausted`),携带 `scope` / `reason` / `retry_after_s` / `per_source_reasons`,供任务队列做延期重投。 预算耗尽/全源熔断时抛 `GatewayUnavailableError` 族(`CircuitOpenError` / `AllSourcesExhausted`),携带 `scope` / `reason` / `retry_after_s` / `per_source_reasons`,供任务队列做延期重投。
**网关拒绝的理由不会丢失**(1.2.0 起):非 2xx 的响应体经折叠与截断后同时进入异常 message 与 `exc.body_text`,故遥测表的 `error` 列里就能看到网关的原话——不必再为查一次 400 单独埋点。截断保头保尾(总长 2048 字符),JSON 错误体尾部的 `code` / `request_id` 不会被切掉。**经中转部署时请注意**:第三方中转服务自身抖动也会回 400,从状态码上与"你的输入有问题"无法区分;库仍按确定性失败处理(直连供应商时重试只会白烧配额),批处理下游宜据 `body_text` 自备兜底分类。
### 哪些异常会到达调用方 ### 哪些异常会到达调用方
上表的"库内行为"一列描述的是**治理动作**,不是调用方要处理的东西。四类里有两类**根本到不了调用方**——它们被重试循环接住,预算耗尽时统一包成 `AllSourcesExhausted`。这个区分只看类型树和 docstring 是读不出来的,曾让下游据此写错整段设计文档,故在此列明: 上表的"库内行为"一列描述的是**治理动作**,不是调用方要处理的东西。四类里有两类**根本到不了调用方**——它们被重试循环接住,预算耗尽时统一包成 `AllSourcesExhausted`。这个区分只看类型树和 docstring 是读不出来的,曾让下游据此写错整段设计文档,故在此列明:
+9 -1
View File
@@ -4,8 +4,11 @@ build-backend = "setuptools.build_meta"
[project] [project]
name = "polygateway" name = "polygateway"
version = "1.1.0" version = "1.2.0"
description = "PolyGateway:实验室统一的大语言模型(LLM/VLM/OCR)调度与中转库——多源、限流、重试、熔断、缓存、遥测" description = "PolyGateway:实验室统一的大语言模型(LLM/VLM/OCR)调度与中转库——多源、限流、重试、熔断、缓存、遥测"
# registry 包页面的正文只认这一项:缺了页面就是一片空白(1.1.2 的教训,twine 会警告
# long_description missing 但不阻塞上传)。README 在打包时被固化进产物,发布后再改无效。
readme = "README.md"
requires-python = ">=3.11" requires-python = ">=3.11"
dependencies = [ dependencies = [
"httpx>=0.27", "httpx>=0.27",
@@ -31,6 +34,11 @@ dev = [
"import-linter>=2.0", "import-linter>=2.0",
] ]
[project.urls]
Homepage = "https://gitea.iomgaa.online/iomgaa/PolyGateway"
Changelog = "https://gitea.iomgaa.online/iomgaa/PolyGateway/src/branch/main/CHANGELOG.md"
Issues = "https://gitea.iomgaa.online/iomgaa/PolyGateway/issues"
[tool.setuptools.packages.find] [tool.setuptools.packages.find]
where = ["src"] where = ["src"]
+7 -2
View File
@@ -396,6 +396,10 @@ flowchart TB
| **空补全**: 200 且流程完整([DONE]/usage 正常)但 content 为空(2026-07-20 M1 验证发现,人类裁决) | `TransientError`(服务抖动,重试/换源;绝不缓存空响应) | | **空补全**: 200 且流程完整([DONE]/usage 正常)但 content 为空(2026-07-20 M1 验证发现,人类裁决) | `TransientError`(服务抖动,重试/换源;绝不缓存空响应) |
| 解析层失败(结构化输出/OCR ZIP) | `ResultInvalidError` | | 解析层失败(结构化输出/OCR ZIP) | `ResultInvalidError` |
**响应体留存(2026-08-16,Gitea issue #10;设计 `designs/2026-08-16-issue10-error-body-retention-design.md`)**: 上表每一条 HTTP 翻译**都必须携带响应体摘要**——摘要同时进入异常 message 与 `PolyGatewayError.body_text`(1.2.0 新增基类字段)。两者都要,因为逐次遥测写的是 `str(exc)`,只加字段进不了遥测表,而"事后可查"正是这条要求的目的。摘要口径由 `transports/_http_errors.summarize_body` 单点实现(折叠空白 → 限长 2048 字符 → 超长保留头 1400 + 尾 600 并记省略字数),两个 transport 共用,**不得各写一份**——issue #10 的成因正是"只有 429 那一支用了响应体"。`body_text` 是旁路数据,不参与任何治理判定;`_translate_429` 的类型细分仍解析未截断原文(摘要会破坏 JSON,改用它会让超长 body 的 `insufficient_quota` 退化成普通限速)。
**400 在中转拓扑下的语义提醒**(同上): 第三方 API 中转服务自身抖动时也会回 400,从状态码上与供应商的"输入非法"无法区分(下游实测: 同一份字节重发 15 次全成功,失败那次 `prompt_tokens=0`、耗时远低于任何成功调用,即请求在推理开始前被挡)。本表**不改** 400 → `RequestRejectedError` 的映射——直连供应商时重试只会白烧配额,且改默认语义等于让所有直连用户为一种部署形态买单;库改为把判据(`body_text`)交给下游自行区分。
### 6.3 "坏结果 ≠ 坏服务"(ResultInvalidError 语义,继承 CHSAnalyzer) ### 6.3 "坏结果 ≠ 坏服务"(ResultInvalidError 语义,继承 CHSAnalyzer)
由输入内容决定的**确定性失败**(这张图就是解析不出表格、这段输出就是修不成 JSON):服务是健康的,换源重试只会白烧配额。因此熔断器记成功、不换源、异常上抛消耗业务侧的失败预算。出处:`CHSAnalyzer governance.py:237-239` 由输入内容决定的**确定性失败**(这张图就是解析不出表格、这段输出就是修不成 JSON):服务是健康的,换源重试只会白烧配额。因此熔断器记成功、不换源、异常上抛消耗业务侧的失败预算。出处:`CHSAnalyzer governance.py:237-239`
@@ -427,8 +431,9 @@ flowchart TB
- `RedisLimiter`: 移植 CHSAnalyzer 六道闸——单条 Lua 原子检查全局并发/单源并发(ZSET 租约)/全局 RPM/单源 RPM/全局 TPM/单源 TPM;窗口 id 用 **Redis 服务器时钟**(TIME 命令)统一多进程口径。随实现移植契约测试。 - `RedisLimiter`: 移植 CHSAnalyzer 六道闸——单条 Lua 原子检查全局并发/单源并发(ZSET 租约)/全局 RPM/单源 RPM/全局 TPM/单源 TPM;窗口 id 用 **Redis 服务器时钟**(TIME 命令)统一多进程口径。随实现移植契约测试。
- `InMemoryLimiter`: 同一契约的进程内实现(semaphore + 滑动窗口计数);单进程场景下语义等价。 - `InMemoryLimiter`: 同一契约的进程内实现(semaphore + 滑动窗口计数);单进程场景下语义等价。
- **配额满行为可配**: `wait`(等待,配 stall 判定——本地等待超窗 + 全局无进展超窗双条件才判卡死)或 `fail-fast`(立即抛)。 - **配额满行为可配**: `wait`(等待,配 stall 判定——本地等待超窗 + 全局无进展超窗双条件才判卡死)或 `fail-fast`(立即抛)。
- **stall 计时口径(2026-08-06 修正,issue #8,设计 `designs/2026-08-06-issue8-stall-budget-design.md`)**: 双条件的**条件 A 只累计非生产性等待**(429 退避、配额 wait 轮询、熔断冷却),真实尝试的耗时由 `StallClock.attempting()` 从 stall 账中扣除。原实现用墙钟总耗时,使真实尝试同时向重试预算与 stall 预算计费;而 stall 预算(默认 300s)小于重试预算(`max_attempts × timeout_s`),必然先耗尽——`timeout_s ≥ stall_window_s` 时一次超时即判 scope 死,`max_attempts` **静默失效**。修正后两个预算正交,**划分依据是"谁消耗重试预算"而非"是否发出请求"**: 烧 `max_attempts` 的时间不烧 `stall_window_s`,不烧 `max_attempts` 的时间归 `stall_window_s`。**429 尝试因此也计入 stall 账**——它免重试预算,若其耗时又算生产性就两个预算都不烧,排队型网关(持满 timeout 才回 429)下调用可挂 25 小时(实施期独立验证实测,见设计 §3.6)。生产性边界即 `_attempt` 边界(含该次记账与遥测收尾),故遥测抖动不参与判死。`stall_window_s``timeout_s` 自此**无耦合**,无需按 `timeout × retries` 放大。三条治理循环(chat/embedding/ocr)共用 `middleware/retry.py``StallClock`。条件 B 的 `inf` 语义未动——新口径下"非生产性排队耗满窗口且 scope 从未出餐"判死本就正当。**残余性质(非本次引入,由条件 B 单独门控)**: 判死是双条件合取,故当同 scope 其他调用仍在正常出餐时本调用不判死(设计意图: 别人还活着就不该宣告 scope 死亡),代价是**该情形下调用级无硬上限**——持续遭遇慢 429 的调用可以等很久;需要硬上限的调用方应自行 `asyncio.wait_for`
- 全局活性信号: `mark_progress()`/`progress_age_s()`("最近一次出餐"时刻)供背压 stall 判定,移植 `CHSAnalyzer limiter.py:193` - 全局活性信号: `mark_progress()`/`progress_age_s()`("最近一次出餐"时刻)供背压 stall 判定,移植 `CHSAnalyzer limiter.py:193`
- **契约补强(2026-07-20,CHS 迁移缺口 G6)**: `settle()`/`release()` 幂等(重复调用无副作用);装配期守卫——`timeout_s ≤ permit 租约 TTL`(防租约先于请求过期)、`stall_window ≥ 最慢源 TTFT 上限`(防误判卡死),违反直接报错拒绝装配。降级方向细化(2026-07-20 M1): "报错不放行"适用于**准入侧**(try_acquire/try_enter 及选源路径消费的 source_stats/retry_after_s);已成功调用后的 settle/release 释放侧失败降级 warning——释放失败不构成放行,且不得掩盖主异常与取消。**勘误(2026-07-20 M2 设计,人类批准)**: 记账侧的 `record_success`/`record_failure`/`mark_progress` 同归此类——调用已真实完成,后端失败若冒泡会丢弃真实成功响应或掩盖原始尝试异常,故降级 warning(CHS 原版一律报错,此为有意反转;丢一次熔断记账最多延迟状态迁移且方向偏保守,epoch fencing 防污染)。 - **契约补强(2026-07-20,CHS 迁移缺口 G6)**: `settle()`/`release()` 幂等(重复调用无副作用);装配期守卫——`timeout_s ≤ permit 租约 TTL`(防租约先于请求过期)、`stall_window ≥ 最慢源 TTFT 上限`(防误判卡死;**issue #8 后为保守冗余**——TTFT 等待属生产性时间已不计入 stall,该误判在机制上不再可能,校验保留因其无害且不误拒合理配置),违反直接报错拒绝装配。降级方向细化(2026-07-20 M1): "报错不放行"适用于**准入侧**(try_acquire/try_enter 及选源路径消费的 source_stats/retry_after_s);已成功调用后的 settle/release 释放侧失败降级 warning——释放失败不构成放行,且不得掩盖主异常与取消。**勘误(2026-07-20 M2 设计,人类批准)**: 记账侧的 `record_success`/`record_failure`/`mark_progress` 同归此类——调用已真实完成,后端失败若冒泡会丢弃真实成功响应或掩盖原始尝试异常,故降级 warning(CHS 原版一律报错,此为有意反转;丢一次熔断记账最多延迟状态迁移且方向偏保守,epoch fencing 防污染)。
### 7.4 熔断 ### 7.4 熔断
@@ -474,7 +479,7 @@ flowchart TB
**`sampling` 列(2026-07-31,issue #4,端口 20 → 21)**: 列语义 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。三个 emit 入口口径必须各自定死,否则同一列在不同行含义不同: `emit_attempt`(RetryMW 调用,**唯一**有生效源者)并上 `source.extra_body`;`emit_cache_hit` / `emit_terminal_failure`(TelemetryMW 最外层调用)无 source 可言,只记调用级——与 `model`/`source_name` 在终态行置空是同一先例,且缓存命中行无损(`sampling` 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同)。三者统一读 `request.sampling` 而非 `request.overlay`(后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处未被污染,直接用必然三行分叉)。OCR/embedding 路径因决策 G 剥离 `extra_body`,该列恒 NULL。 **`sampling` 列(2026-07-31,issue #4,端口 20 → 21)**: 列语义 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。三个 emit 入口口径必须各自定死,否则同一列在不同行含义不同: `emit_attempt`(RetryMW 调用,**唯一**有生效源者)并上 `source.extra_body`;`emit_cache_hit` / `emit_terminal_failure`(TelemetryMW 最外层调用)无 source 可言,只记调用级——与 `model`/`source_name` 在终态行置空是同一先例,且缓存命中行无损(`sampling` 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同)。三者统一读 `request.sampling` 而非 `request.overlay`(后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处未被污染,直接用必然三行分叉)。OCR/embedding 路径因决策 G 剥离 `extra_body`,该列恒 NULL。
(`cached_prompt_tokens`/`model_reported` 为 2026-07-31 issue #3 新增,端口由 18 字段扩为 20;两个后端在初始化期对已存在的旧表幂等补列——`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入都被逐行 warning 丢弃。补列一律**先探测缺列再 ALTER**(`ADD COLUMN IF NOT EXISTS` 即使列已存在也先取 ACCESS EXCLUSIVE 锁,而遥测内联 await,锁共享审计表会拖垮业务调用),且**失败只逐行降级、绝不置结构性失能标志**。新列在 DDL 里必须排在 `created_at` **之后**,与 `ALTER TABLE ADD COLUMN` 的追加位置一致,否则新建库与升级库的物理列序分叉)。链路: `session_id`/`parent_call_id` 由调用方传入贯穿(agent step → LLM call)。`messages` 落库前对多模态 part 先摘要(与缓存 key 共用同一摘要函数,§7.5)——Video-Tree 现状 base64 整段进 SQLite 导致 db 膨胀(`llm.py:330`),库内修复(2026-07-20,VT 迁移缺口 R12)。 (`cached_prompt_tokens`/`model_reported` 为 2026-07-31 issue #3 新增,端口由 18 字段扩为 20;两个后端在初始化期对已存在的旧表幂等补列——`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入都被逐行 warning 丢弃。补列一律**先探测缺列再 ALTER**(`ADD COLUMN IF NOT EXISTS` 即使列已存在也先取 ACCESS EXCLUSIVE 锁,而遥测内联 await,锁共享审计表会拖垮业务调用),且**失败只逐行降级、绝不置结构性失能标志**。**建表同理(2026-08-07,issue #9)**: PG 对 schema 的 CREATE 权限检查早于 `IF NOT EXISTS` 的存在性判断(16.14 实测,只授表级 `SELECT, INSERT` 的角色写得进去却建不了表),故 PG 侧必须**先 `to_regclass` 探测、表在就不发 DDL**;SQLite 侧实测在解析期即短路(持排他锁/只读文件下该语句均通过),无同款风险,**有意不加探测**。由此把"结构性失能"的判据从「初始化时出过异常」收窄为「确定写不进去」——仅建池失败与"表确定不存在且建不出来"判死,探测/取连接失败只跳过本次并留待下次重试。新列在 DDL 里必须排在 `created_at` **之后**,与 `ALTER TABLE ADD COLUMN` 的追加位置一致,否则新建库与升级库的物理列序分叉)。链路: `session_id`/`parent_call_id` 由调用方传入贯穿(agent step → LLM call)。`messages` 落库前对多模态 part 先摘要(与缓存 key 共用同一摘要函数,§7.5)——Video-Tree 现状 base64 整段进 SQLite 导致 db 膨胀(`llm.py:330`),库内修复(2026-07-20,VT 迁移缺口 R12)。
- 后端: `SQLiteRecorder`(默认;WAL + busy_timeout、`INSERT OR IGNORE` 幂等、`asyncio.to_thread` 桥接、初始化/写入失败全降级不冒泡)与 `PostgresRecorder` - 后端: `SQLiteRecorder`(默认;WAL + busy_timeout、`INSERT OR IGNORE` 幂等、`asyncio.to_thread` 桥接、初始化/写入失败全降级不冒泡)与 `PostgresRecorder`
- **单一 helper 铁律**: 遥测调用点收敛为一个内部函数/上下文管理器;Video-Tree 与 GovDoc 各有 4-5 处逐字复制的 `record_llm_call(15 个参数)` 是本条的直接教训。 - **单一 helper 铁律**: 遥测调用点收敛为一个内部函数/上下文管理器;Video-Tree 与 GovDoc 各有 4-5 处逐字复制的 `record_llm_call(15 个参数)` 是本条的直接教训。
@@ -20,7 +20,7 @@
| Issue 原文 | 实际情况 | | Issue 原文 | 实际情况 |
|---|---| |---|---|
| 泄漏路径为 `try_enter` / `try_acquire` 两条 | ****`middleware/retry.py:216` 每轮循环开头的 `progress_age_s()` 同样在 catch 之外,直达调用方 | | 泄漏路径为 `try_enter` / `try_acquire` 两条 | ****(设计初稿写"三条",2026-08-06 独立验证时核出遗漏两条并订正): `QuotaGate``try_acquire` / `stats`(`retry.py:249`)/ `progress_age_s`(`retry.py:216``:305`),`BreakerGate``try_enter` / `retry_after_s`(`retry.py:292``:310`)。判据是该调用点是否被 `_record_quietly` 包裹——未包裹即直达调用方;OCR 与 Embedding 两个治理循环有同构的对应点 |
| (未提及构造点数量) | 全库 **22 处** `raise GovernanceBackendError`,分布于 4 个文件 | | (未提及构造点数量) | 全库 **22 处** `raise GovernanceBackendError`,分布于 4 个文件 |
| 方向 A 只需改类型树 | 其中 **2 处语义完全不同**(见 §3.4),整类归入"可重投"会制造镜像 bug | | 方向 A 只需改类型树 | 其中 **2 处语义完全不同**(见 §3.4),整类归入"可重投"会制造镜像 bug |
| `retry_after_s` 取 0,「docstring 已写 0 = 可立即重试,语义上是通的」 | 语义通,**工程上不通**。见 §3.2 | | `retry_after_s` 取 0,「docstring 已写 0 = 可立即重试,语义上是通的」 | 语义通,**工程上不通**。见 §3.2 |
@@ -131,7 +131,7 @@ Issue 建议取 0。**否决**:下游 `schedule_retry(after_s=0)` 会立刻重
| 测试 | 位置 | 先失败后通过的证据 | | 测试 | 位置 | 先失败后通过的证据 |
|---|---|---| |---|---|---|
| `GovernanceBackendError` 可被 `except GatewayUnavailableError` 接住 | `tests/unit/test_errors.py` | 改前 `pytest.raises(GatewayUnavailableError)` 必失败 | | `GovernanceBackendError` 可被 `except GatewayUnavailableError` 接住 | `tests/unit/test_errors.py` | 改前 `pytest.raises(GatewayUnavailableError)` 必失败 |
| 三条泄漏路径(`try_acquire`/`try_enter`/`progress_age_s`)抛出的异常携带正确 `scope` 与非零 `retry_after_s` | `tests/unit/test_backpressure.py`**三条桩都需新增**(Codex 审计划时核出: `:176-186` 是记账侧 `record_success`/`record_failure`/`mark_progress` 的降级桩,不是闸门路径;`progress_age_s``:243-257` 覆盖包装行为、不验 scope) | 改前无 `scope` 属性,`AttributeError` | | 闸门泄漏路径(五条,§1.1)抛出的异常携带正确 `scope` 与非零 `retry_after_s`;钉住 `try_acquire`/`try_enter`/`progress_age_s` 三条代表路径,余两条由同一注入机制覆盖 | `tests/unit/test_backpressure.py`**三条桩都需新增**(Codex 审计划时核出: `:176-186` 是记账侧 `record_success`/`record_failure`/`mark_progress` 的降级桩,不是闸门路径;`progress_age_s``:243-257` 覆盖包装行为、不验 scope) | 改前无 `scope` 属性,`AttributeError` |
| `str(exc)` 仍为原诊断串 | `tests/unit/test_errors.py` | 防 §3.5 回归 | | `str(exc)` 仍为原诊断串 | `tests/unit/test_errors.py` | 防 §3.5 回归 |
| 未知源抛 `SourceNotConfiguredError` 且**不是** `GatewayUnavailableError` | 改 `tests/unit/test_redis_key_layout.py:70-74`;内存版**当前无覆盖,需新增** | 改前抛 `GovernanceBackendError`,断言"不是 scope 级"必失败 | | 未知源抛 `SourceNotConfiguredError` 且**不是** `GatewayUnavailableError` | 改 `tests/unit/test_redis_key_layout.py:70-74`;内存版**当前无覆盖,需新增** | 改前抛 `GovernanceBackendError`,断言"不是 scope 级"必失败 |
| Redis 真实掉线时准入侧行为 | `tests/integration/test_redis_cross_connection.py:228-245`(真实 Redis,不 mock) | 断言由 `GovernanceBackendError` 收紧为"是 `GatewayUnavailableError``reason == governance_backend_down`" | | Redis 真实掉线时准入侧行为 | `tests/integration/test_redis_cross_connection.py:228-245`(真实 Redis,不 mock) | 断言由 `GovernanceBackendError` 收紧为"是 `GatewayUnavailableError``reason == governance_backend_down`" |
@@ -0,0 +1,254 @@
# stall 判定改为非生产性等待口径设计(Issue #8)
- **日期**: 2026-08-06
- **来源**: Gitea Issue #8(本机全套件跑 391.67s,1 failed;失败源于单次 300s 超时耗尽 stall 窗口,基于 1.1.0 源码核查)
- **状态**: **已批准(2026-08-06)**,待 `writing-plans`
- **触发档位**: 强制(变更治理行为——判死条件的度量口径,是库对下游的承诺)
- **方案范围**: 人类明确要求单一方案,故本文不列平行备选,仅在 §5 记录被否决路线及否决理由(体例沿用 Issue #7 设计)
## 1. 目标与非目标
| | 内容 |
|---|---|
| **G1** | 消除"单次超时即判 scope 级死亡"——`timeout_s``stall_window_s` 的隐式耦合彻底解除,重试预算在超时场景下真实可用 |
| **G2** | 使 stall 判定的度量对象与它的职责一致:**它治理的是无人治理的非生产性循环,不是已被重试预算治理的真实尝试** |
| **G3** | 三条治理循环(chat / embedding / ocr)口径一致,计时逻辑收敛为单一共享单元,杜绝第四次复制 |
| **G4** | 配置方不再需要心算 `stall_window > timeout × max_attempts`;`.env.example` 注释与实际语义对齐 |
| **非目标** | 不改 `progress_age_s()``inf` 语义(见 §3.4);不新增装配期校验(见 §5.2);不新增配置项;不改 `AllSourcesExhausted` 的字段与 `reason` 取值;不给 embedding/ocr 新增主循环判死路径(见 §5.4);不改 429 免预算、AIMD、选源、熔断任何既有行为 |
### 1.1 Issue 前提的三处修正(按 1.1.0 源码核实)
| Issue 原文 | 实际情况 |
|---|---|
| 失效点为 `retry.py:216` 一处 | **三处同构**:`retry.py:216`(主循环)、`retry.py:305` / `embedding.py:247` / `ocr.py:272`(`_on_no_runnable`)。四个判定点共用同一个墙钟 `entered_at`,故 embedding/ocr 在"先超时一次、再遇到无可用源"时同样误判——issue 只覆盖了 chat |
| 建议方向 1:装配期校验 `stall_window_s > max(timeout_s)` | **不采纳**。它把耦合固化成契约而非消除耦合,且约束值须为 `timeout × max_attempts`(本机即 900s),会让 stall 兜底迟钝到近乎失效。详见 §5.1 |
| 建议方向 2:`inf` 不参与判死 | **不采纳**。在新口径下 `inf` 从"有害恒真"变回"正确的保守默认";且它会反转已被测试钉住的既有行为。详见 §3.4 与 §5.3 |
## 2. 根因:两个预算重叠计费
`retry.py:214-215` 的注释自述这处判定是「429 免预算后的兜底,防饱和期无限循环」——它治理的对象是**非生产性循环**。但条件 A `now - entered_at > stall` 度量的是**墙钟总耗时**,无法区分两类性质相反的时间:
| 时间性质 | 构成 | 应由谁治理 | 耗尽后 |
|---|---|---|---|
| **生产性** | 一次尝试的完整生命周期(发请求、等响应含耗满 `timeout_s` 的超时/TTFT/流式读取,以及该次尝试的记账与遥测收尾) | `max_attempts`(重试预算) | `retry_exhausted` |
| **非生产性** | 429 退避、配额 wait 轮询、熔断冷却轮询、AIMD 排队 | **无人治理**(429 不计 `fails`)→ 正是 stall 的职责 | `stalled` |
**缺陷即:生产性时间同时向两个预算计费。** 而 stall 预算(默认 300s)远小于重试预算(`3 × 300s`),必然先耗尽,于是重试预算在超时场景下**永远用不上**——issue 观察到的"静默失效"就是这个重叠计费的直接后果。
`.env``TIMEOUT_S=300``_DEFAULT_STALL_WINDOW_S=300.0`(`config.py:60`)相等只是把它暴露得最快;只要 `timeout_s ≥ stall_window_s / 1`,一次超时就够。
### 2.1 两条佐证:`inf` 恒真是遗漏而非设计
| 证据 | 出处 | 含义 |
|---|---|---|
| `_PROGRESS_TTL_S = 3600 # 远大于任何 stall_window,防进度键过期造成假停滞` | `backends/redis/limiter.py:32` | 「无 progress 记录 ≠ 停滞」早已是设计共识,作者用超长 TTL 规避了"键过期"这一路径,但 TTL 再长也救不了"**从来没写过**"——冷启动是同类情形的漏网之鱼 |
| `test_global_stale_but_local_fresh_keeps_waiting` docstring 写「仅全局超窗(从未出餐 age=inf)」 | `tests/unit/test_backpressure.py:120-121` | 现有测试把 `inf` 当作"全局超窗成立"钉住了;`test_both_windows_exceeded_raises_stalled`(:89)更是**全靠 `inf` 恒真**才能触发判死 |
## 3. 选定方案:双预算正交模型
### 3.1 一句话
**stall 计时器只累计非生产性等待时间**:`stalled_s = (now entered_at) 真实尝试累计耗时`
两个预算自此正交,各管一段,无缝覆盖调用的全部时间:
| 花在哪 | 烧哪个预算 |
|---|---|
| 真实尝试(`_attempt` 内),**429 除外** | 重试预算 `max_attempts` |
| 其余一切等待,**含 429 尝试本身** | stall 预算 `stall_window_s` |
> **划分依据是"谁消耗重试预算",不是"是否发出了请求"**(2026-08-06 实施期订正,见 §3.6)。初稿按后者划分,使 429 尝试两个预算都不烧。
**"生产性"的边界即 `_attempt` 的边界**——包含该次尝试的记账(`record_success`/`mark_progress`)与遥测收尾,而不止于"等响应"。这是有意的:这些收尾是"尝试已有结论"之后的动作,不是"在等待重试机会"的停滞;把它们计入 stall 会让遥测抖动参与判死,与「遥测写失败降级不冒泡」所守的"遥测不得影响主路径判决"同精神。其耗时本也在毫秒量级。
这与库内既有原则**同构**:429 不烧重试预算,所以 429 等待烧 stall 预算;真实尝试烧重试预算,所以它不烧 stall 预算。
### 3.2 为什么取补集,而不是逐处标记 sleep
两种实现都能达到 §3.1 的语义,选**取补集**(总时间减去 `_attempt` 耗时):
| 维度 | 取补集(选定) | 逐处标记 sleep(否决) |
|---|---|---|
| 埋点数量 | 每条循环 **1 处**(`_attempt` 调用点) | chat 3 处、embedding/ocr 各 2 处,共 7 处 |
| 演进安全性 | **默认安全**:将来新增任何等待路径自动计入 stall,兜底不会漏 | 默认危险:新增等待路径若忘记标记,即成新的 stall 盲区 |
| 语义可读性 | 「stall 时间 = 总时间 − 花在真实尝试上的时间」,一句话说清 | 需读者遍历全部标记点才能确认覆盖完整 |
`_attempt` 是纯生产性的:permit 获取、熔断准入、AIMD 判定全部在 `_pick_runnable` 内完成,`_attempt` 进入时已持 permit,内部只做"发请求 + 记账"。故补集口径不会把非生产性时间误算为生产性。
### 3.3 共享单元:`StallClock`
计时逻辑提取为 `middleware/retry.py` 的模块级小类,embedding/ocr 复用——沿用 `backoff_delay` 已被两者复用的既有手法(`tests/unit/test_backpressure.py:258` 记录该先例),不新建模块、不动依赖层次。
```python
class StallClock:
"""调用级 stall 计时器: 只累计非生产性等待(设计 §3.1)。
实例per调用创建, 严禁提升为实例属性——并发调用共享会互相污染。
"""
def __init__(self, now: Callable[[], float]) -> None:
self._now = now
self._entered_at = now()
self._productive_s = 0.0
def stalled_s(self) -> float:
return self._now() - self._entered_at - self._productive_s
@contextlib.asynccontextmanager
async def attempting(self):
started = self._now()
try:
yield
finally:
# 只做算术, 不吞任何异常——CancelledError 照常穿透(库铁律)
self._productive_s += self._now() - started
```
调用点改动(三处循环同款):
```python
clock = StallClock(self._now) # 替换 entered_at = self._now()
...
if clock.stalled_s() > stall and await self._quota.progress_age_s() > stall:
raise AllSourcesExhausted(..., reason="stalled", ...)
...
async with clock.attempting(): # 包裹真实尝试
outcome = await self._attempt(request, *picked, reasons, attempt_fails)
```
`_on_no_runnable` 的形参由 `entered_at: float` 改为 `clock: StallClock`(三处同改)。
### 3.4 `inf` 语义为何不动(本设计的核心权衡)
新口径下第一象限的含义变为:「**非生产性排队已耗满 `stall_window_s`,且整个 scope 从未出餐**」。此时判死是正当的——真的没有任何证据表明这个 scope 还活着,而调用方已经白等了一整个窗口。`inf` 由此从"有害的恒真"回归为"正确的保守默认"。
反过来,若同时改 `inf` 语义:
- 冷启动窗口内 stall 判定**完全失效**,429 饱和场景下 chat 主循环重新暴露无限循环风险(429 不计 `fails`,无其他兜底);
- 会反转 `test_both_windows_exceeded_raises_stalled` 钉住的行为,并与 CHS 保真蓝本分叉。
**一次改动解决问题,优于两次改动互相牵制。** 这是本设计只动条件 A 的理由。
### 3.5 429 饱和场景下兜底仍然有效(正确性验证)
修改后必须确认 stall 兜底没有被削弱。429 免预算使 `fails` 恒为 0,`retry.py``max(fails, 1)` 令退避恒定在 `backoff_base_s` 档(或取 `Retry-After` 提示的较大值),不随轮次增长。每轮构成为「一次 429 往返」+「一段恒定退避 sleep」,后者非生产性且每轮累加,`stalled_s` 单调逼近 `stall_window_s`,兜底有效。
**但这个论证在初稿里依赖一个未加保护的假设**:「429 往返是快速失败,毫秒至秒级」。§3.6 处理它不成立的情形。
### 3.6 订正:429 尝试必须退还给 stall 账(2026-08-06 实施期,独立验证发现)
**缺陷**:初稿按"是否发出请求"划分两个预算,于是 429 尝试的耗时算生产性。但 429 **不消耗重试预算**——它于是**两个预算都不烧**,掉进缝隙。§3.1 初稿声称的"无缝覆盖调用的全部时间"因此不成立。
**后果实测**(排队型网关:持满 `timeout_s` 才回 429,`timeout=300 / stall=300 / backoff_base=2 / rng=0`):
| | 尝试次数 | 墙钟 |
|---|---|---|
| 修复前(main) | 1 | 301s |
| 初稿口径 | **301** | **90,601s ≈ 25.2 小时** |
| 订正后 | 1 | 301s |
即初稿把一个 bug 换成了一个更严重的 bug——25 小时的挂起。
**订正**:划分依据改为**"谁消耗重试预算"**。429 免重试预算 → 429 尝试的耗时归 stall 治理,由 `StallClock.attempting()` yield 的句柄 `refund()` 退还。缝隙就此闭合,且这条规则比初稿更本质:两个预算按"由谁治理"划分,而非按"是否发出请求"这个表象。
**影响范围仅 chat**:embedding/ocr 无 429 免预算(无条件 `fails += 1`),429 照常烧重试预算,不存在缝隙,无需改动(与 §5.4 的分析一致)。
## 4. 旧版行为审计(stall 子系统逐条)
| 既有行为 | 处置 | 说明 |
|---|---|---|
| 双条件判死(本地超窗 ∧ 全局无进展超窗) | **保留** | 结构不变,只改条件 A 的度量口径 |
| 条件 A = 调用级累计、循环内不重置(CHS `governance.py:207`) | **保留** | `StallClock` 同样每调用一个实例、循环内不重置 |
| 条件 A 计入真实尝试耗时 | **替换** | 本设计的唯一行为变更 |
| 条件 B `progress_age_s()`,`inf` = 从未进展 | **保留** | 见 §3.4 |
| 本地 monotonic 与后端时钟刻意不混用 | **保留** | `StallClock` 只用注入的 `self._now`,不读后端时钟 |
| poll jitter ∈ [0.5p, 1.0p] 防惊群 | **保留** | 不触碰 |
| `fail_fast` 不进入 stall 判定 | **保留** | 不触碰 |
| 429 免预算(chat 独有) | **保留** | 不触碰;embedding/ocr 无此逻辑,故无对应缺口(§5.4) |
| `AllSourcesExhausted(reason="stalled")` 及其 `retry_after_s` 取值 | **保留** | 错误面零变更,下游 `except` 写法不受影响 |
**有意放弃**: 无。本设计不删除任何既有行为。
## 5. 被否决的路线
### 5.1 装配期校验 `stall_window_s > max(timeout_s)`(Issue 建议方向 1)
否决理由三条:
1. **治标**。它把"两个预算重叠计费"这个缺陷固化成一条配置契约,要求配置方绕开它,而不是消除它。
2. **约束值不可接受**。要让重试预算真正可用,须 `stall_window > timeout × max_attempts`(本机 900s)。stall 兜底随之迟钝到 900s 才触发,饱和期无限循环的防护近乎失效——**修好一个洞,挖开另一个**。
3. **挡不住残余情形**。即便配到 1200s,一次调用若在 429 轮询与超时上累计超过 1200s,条件 B 的 `inf` 仍恒真,双条件仍退化为单条件。坑只是被推远。
新口径下 `stall_window_s``timeout_s` 不再有任何耦合,**这条校验没有存在的理由**——不加校验、而是消除掉需要校验的耦合。
### 5.2 既有校验 `stall_window_s ≥ max(ttft_timeout_s)` 的处置
`config.py:240-247``_validate_stall``ARCHITECTURE.md` §7.3 记为契约补强 G6。新口径下 TTFT 等待属生产性时间,其 docstring 的理由「防把正常慢首包误判为卡死」**已不成立**。
**人类已定夺:保留校验,改写 docstring 说明新口径**。校验本身无害(不会误拒任何合理配置),保留可避免改动 ARCHITECTURE.md 既有契约、把本次改动的影响面控制在最小。docstring 改为说明"该校验在新口径下为保守冗余,TTFT 已不计入 stall"。
### 5.3 `inf` 不参与判死(Issue 建议方向 2)
见 §3.4:新口径下 `inf` 已无害,单独改它会制造冷启动兜底真空并反转既有测试。
### 5.4 给 embedding/ocr 补主循环 stall 判定
设计过程中一度提出(前提是"429 饱和时它们没有防无限循环兜底"),**核实后前提不成立,故否决**:
| 循环路径 | embedding/ocr 的兜底 |
|---|---|
| `picked is None``_on_no_runnable` 轮询(不烧 `fails`) | `_on_no_runnable` 内已有 stall 判定(`embedding.py:247` / `ocr.py:272`)✓ |
| 尝试失败 → `fails += 1` | `max_attempts` ✓ |
`retry.py:233-234` 的 429 免预算分支是 chat **独有**的(`embedding.py:191``ocr.py:216` 均为无条件 `fails += 1`,两文件亦无 `pacer`),主循环判定正是为它打的补丁。embedding/ocr 两条路径均已封闭,补齐等于凭空新增一条判死路径,使其比 chat 更易判死——纯 gold-plating。
## 6. 非功能维度
| 维度 | 回答 |
|---|---|
| **并发** | `StallClock` **每次调用创建一个实例**,是调用级局部状态,与被替换的 `entered_at` 局部变量同性质。严禁提升为实例属性(并发调用会互相污染计时)——docstring 已写明,单测钉住并发两路调用互不干扰 |
| **取消** | `attempting()``finally` 只做浮点加法,不含 `await`、不捕获任何异常,`CancelledError` 逐字穿透。既有 `test_cancellation_pierces_wait_loop` 继续有效,并新增一条"取消发生在 `_attempt` 内"的用例 |
| **降级方向** | 不变。stall 判定读取的 `progress_age_s()` 属准入侧,后端故障仍 fail-closed 抛 `GovernanceBackendError`(scope 级),不放行 |
| **幂等与重复** | `stalled_s()` 是纯读,可任意次调用;`attempting()` 可重入多次(每次尝试一次),累加语义天然幂等于"总生产性时间" |
| **持久化与原子性** | 不适用。纯进程内计时,无落盘、无后端写入,不新增任何 Redis 往返 |
| **性能** | 每次尝试新增两次 `self._now()` 调用与一次浮点加法,可忽略 |
## 7. 错误处理与测试策略
**错误分类**: 无变更。判死仍抛 `AllSourcesExhausted(reason="stalled")`,属 scope 级不可用(`GatewayUnavailableError` 家族),下游延期重投语义不变。
### 7.1 回归证据(先失败后通过)
核心用例 `test_single_timeout_does_not_exhaust_stall_budget`:`stall_window_s == timeout_s == 300`,第一次尝试推进 `FakeClock` 超过 300s 后抛 `TransientError`,第二次返回成功。
- **改前**:第二次尝试发出前即被判死,抛 `AllSourcesExhausted(reason="stalled")`**失败**
- **改后**:重试预算正常生效,返回成功响应 → **通过**
embedding / ocr 各一条同构用例(经"先超时一次、再遇到无可用源"触发 `_on_no_runnable`)。
### 7.2 其余用例
| 用例 | 钉住什么 |
|---|---|
| 四象限现有四条(`TestStallQuadrants`) | 非生产性路径行为逐字不变;`test_both_windows_exceeded` 全程无真实尝试,`stalled_s` 等价于旧墙钟,**应原样通过** |
| `test_productive_time_excluded_from_stall` | 直接断言:仅靠真实尝试耗时无论多久都不触发判死 |
| `test_nonproductive_wait_still_triggers_stall` | 反向:纯轮询等待累满窗口仍正常判死(兜底未被削弱) |
| `test_saturation_429_still_stalls` | §3.5 的正确性验证:429 连续拒绝 + 退避,最终仍判死而非无限循环 |
| `test_cancel_inside_attempt_pierces` | 取消穿透 `attempting()``finally` |
| `test_concurrent_calls_do_not_share_clock` | 两路并发调用,一路长尝试不影响另一路的 stall 账 |
Redis 后端无需新增用例:本设计不改后端接口与 `progress_age_s()` 语义。
## 8. 交付清单(供 `writing-plans` 展开)
| # | 内容 |
|---|---|
| T1 | `middleware/retry.py` 新增 `StallClock`;主循环与 `_on_no_runnable` 改用之 |
| T2 | `embedding.py` / `ocr.py` 复用 `StallClock`,`_on_no_runnable` 形参改签名 |
| T3 | `config.py:240-247` `_validate_stall` docstring 改写(§5.2) |
| T4 | 测试:§7.1 回归三条 + §7.2 五条;`test_backpressure.py:121` docstring 订正 |
| T5 | `.env.example:41` 注释改写(删除误导性的"须 ≥ 最大源 TTFT",说明新口径);本机 `.env:37` 的临时缓解 `STALL_WINDOW_S=1200` 可回退默认(不入库,仅记录) |
| T6 | `ARCHITECTURE.md` §7.3 背压条目补记新口径与本设计指针;`CHANGELOG.md` 记治理行为变更 |
| T7 | Wiki 同步(`docs-convention.md` §2「治理行为变更」行):`解释-治理行为` + `指南-限流与熔断` |
**副作用提醒**: 修复后单次调用最坏耗时由 `stall_window_s` 抬升至 `max_attempts × timeout_s`(本机 900s)——这是重试预算恢复生效的**正确表现**,但 e2e 冒烟测试的最坏耗时随之变长,`tests/e2e` 的源 `timeout_s` 配置可能需要相应调小。
@@ -0,0 +1,238 @@
# HTTP 错误响应体留存设计(Issue #10)
- **日期**: 2026-08-16
- **来源**: Gitea Issue #10(下游 1050 张医学影像批处理,1 张收到 400 被判确定性失败;事后无从查证原因。基于 1.1.2 源码核查)
- **状态**: **已批准(2026-08-16)**,待 `writing-plans`
- **触发档位**: 强制(`errors.py` 属最内层内核,新增公共字段即变更库对下游的承诺)
- **方案范围**: 人类明确要求单一方案(2026-08-16),故本文不列平行备选,仅在 §6 记录被否决路线及否决理由(体例沿用 Issue #7/#8 设计)
## 1. 目标与非目标
| | 内容 |
|---|---|
| **G1** | 网关拒绝一次调用时,**它说了什么必须可事后查证**——库自己的遥测表里就能查到,不依赖下游额外埋点 |
| **G2** | 留存口径覆盖 transport 层**全部**非 2xx 分支与**全部** transport(chat / embedding / stream / OCR),杜绝"只修 400 → 下次 401 复发" |
| **G3** | 摘要文本单点规范化(折叠空白 + 截断 + 截断标记),message 与结构化字段**取同一份串**,两处永不打架 |
| **G4** | 不改变任何状态码 → 错误分类的映射(ARCHITECTURE §6.2 表原封不动),下游 `except` 写法零影响 |
| **非目标** | 不改 400 的治理语义(不重试不换源,见 §5.2);不新增遥测列(见 §6.1);不新增配置项;不做错误分类可插拔(见 §6.4);不顺手修 `_status_to_error``operation` 硬编码缺陷(见 §5.4) |
### 1.1 Issue 前提的两处修正(按 1.1.2 源码核实)
| Issue 原文 | 实际情况 |
|---|---|
| 建议方向一「让异常带上截断后的响应体……就能让下游把它记进日志和遥测」 | **只做这一半解决不了 Issue 自己陈述的痛点**。库的逐次遥测写的是 `error=str(exc)`(`middleware/retry.py:558``middleware/telemetry.py:89``telemetry/sqlite.py:43``error TEXT` 列),即**异常 message**。新增字段不会进库的遥测表;下游说的"写进遥测表"是他们自己的埋点。故本设计**两件都做,且以 message 为主**(§3.3) |
| 缺陷范围 = 400 分支 + 4xx 兜底 | 实为 **6 处同构**:`openai_compat._status_to_error` 的 400 / 4xx 兜底 / 401·403 / 5xx 四支,`_translate_429` 的两支(读了 body 判 `insufficient_quota`,但 message 仍不带),以及 `monkey_ocr._classify_status:74-88` 的**全部**分支(message 只有 `HTTP {status}`)。Issue 场景是"读表格",极可能正落在 OCR 路径 |
## 2. 根因:诊断信息在翻译层被丢弃,而遥测只看 message
`_status_to_error`(`transports/openai_compat.py:131-143`)手上握着 `body_text`,却只把它用于 429 的类型细分,翻出的异常与 message 都不携带它。响应体在这一层之后**不再存在于进程任何位置**:该模块无 logger(grep `logger|loguru` 零命中),异常类无字段,遥测只写 message。
三条留存通道同时为空,是"永久查不到"的完整解释:
| 通道 | 现状 | 本设计后 |
|---|---|---|
| 日志 | 模块无 logger | 仍无(§6.2:不加日志) |
| 异常字段 | 无承载处 | `body_text`(§3.1) |
| 库遥测 `error` 列 | 只有 `"{源名} 请求被拒: 400"` | message 携带摘要(§3.3) |
## 3. 选定方案
### 3.1 内核:`PolyGatewayError` 基类新增 `body_text`
```python
class PolyGatewayError(Exception):
def __init__(self, message, *, source_name=None, status_code=None,
operation=None, body_text: str = "") -> None:
```
**加在基类而非 `RequestRejectedError`**:这些错误全部由同一个 HTTP 响应翻译而来,"对方说了什么"与"它属于哪一类"正交。只给一个子类加,下次给 `SourceDeadError` 加又是一次公共 API 变更 + 一次人类门。
与既有 `ResultInvalidError.raw_text`(`errors.py:77`)的界限必须在 docstring 钉死,否则两个"原文字段"必然被混用:
| 字段 | 语义 | 来源 |
|---|---|---|
| `body_text` | **非 2xx** 的 HTTP 错误响应体摘要——对方**拒绝**你的理由 | transport 翻译层 |
| `raw_text` | **2xx** 但内容不可解析时的模型输出原文 | 结构化解析层 |
`GatewayUnavailableError` 一族继承到一个恒空的 `body_text` 不是噪音:scope 级失败本就"没有单一响应体可言",空串是对这件事的如实表达。
### 3.2 共享单元:`transports/_http_errors.py`(新建,~40 行)
两个 transport 各有自己的状态码分类逻辑(OCR 无 429 细分,有意保留,见 `monkey_ocr.py:53-54`),但**摘要口径必须同一份**,否则就是下一个"只修一半"。两函数:
| 函数 | 职责 | 关键防御 |
|---|---|---|
| `summarize_body(text) -> str` | 折叠空白 → 按 §3.4 的机械规则截断 | 空/空白入参返回 `""` |
| `response_body(response) -> str` | 从 `httpx.Response` 取已缓冲文本 | `ResponseNotRead` → 返回 `""`,**绝不触发网络读** |
- **折叠空白不是洁癖**:错误体常是缩进 JSON,直接拼进 message 会让一行日志炸成多行、遥测列不可读。
- **截断必须留标记**:不标记,读的人分不清"网关只说了这么多"和"库切的"。
- **`response_body` 的防御是硬要求**:`monkey_ocr._classify_status` 只拿得到 `httpx.HTTPStatusError`,若某天 OCR 走 stream 请求,`.text` 会抛 `ResponseNotRead`,把一次可分类的 4xx 变成泄漏的 httpx 异常——**违反"一切失败必须落入四分类"铁律**。诊断信息缺失绝不能升级为崩溃(降级方向,§4.2)。
放在 `transports/` 私有模块而非 `errors.py`:职责是"HTTP 响应 → 领域错误"的工具,放内核会稀释 `errors.py` 的单一职责(P3)。两个 transport 同 import 一个私有模块,不构成 transport 之间的互相依赖,import-linter 的 layers 契约(同层 `|` 独立性)不受影响。
### 3.3 翻译层:表驱动收口,message 与字段共用一份摘要
`_status_to_error` 现在是五个分支各拼各的 message,新增摘要意味着五处重复。改为**分类表 + 单点拼装**,代码反而变短:
```
summary = summarize_body(body_text) # 全函数只算一次
ctx = {..., "body_text": summary} # 字段
429 → _translate_429(source, body_text, headers, ctx) # 需原文判 type,单列
其余 → cls, label = _STATUS_MAP 查表 → cls(_compose(source, label, status, summary), **ctx)
```
message 形态:`"{源名} {标签}: {状态码} | {摘要}"`;**摘要为空时不拼后缀**,避免出现悬空的 ` | `。分隔符取 ` | ` 而非既有的 `: `,让"库的话"与"网关的话"一眼可分。
**429 也拼,不设例外**:例外就是下一个复发点。`insufficient_quota` 那支尤其需要(配额细节全在 body 里);普通限速 body 通常很短。代价是高频限速场景遥测 `error` 列变长,由 `_ERROR_BODY_CAP` 兜住。
`monkey_ocr._classify_status` 同款处理:`summary = summarize_body(response_body(exc.response))`,message 追加同一后缀,`ctx` 带上字段。
### 3.4 常量取值
**机械规则(实现与测试逐字照此)**:
```
_ERROR_BODY_CAP = 2048 # 字符(非字节),含省略标记在内的最终总长上限
_HEAD_CHARS = 1400
_TAIL_CHARS = 600
折叠空白后 len ≤ 2048 → 原样返回
否则 → s[:1400] + f"…(略 {len(s) - 2000} 字)…" + s[-600:]
```
> 规则必须写成算术而非叙述:"截断至 cap 并补标记"能同时被读成总长 2048 与 2049,两者会让测试断言与遥测长度承诺对不上(Codex 审查 2026-08-16 提出)。
**头尾保留而非头部硬切**(2026-08-16 调研决策)。截断的对象是**结构化 JSON 错误体**,信息分布头重尾也重:人话(`message`)在前,机器可判的 `type` / `code` / `param` / `request_id` 在后。Issue 给出的真实样本即 `"code":"invalid_parameter_error"` 收尾——头部硬切正好切掉向网关方追查时唯一有用的那部分。省略标记记下**被省略的字符数**,读的人才知道自己丢了多少,不会误以为网关只说了这么多。
按**字符**而非字节切:多字节字符不会被切成半个(Sentry 曾为按字节切开 issue #1691),且 `error TEXT` 列无定长约束,无需字节口径。
### 3.4.1 取值依据:同场景开源实践
| 项目 | 场景 | 上限 | 保留策略 |
|---|---|---|---|
| **Kubernetes client-go** `rest/request.go` | **读 HTTP 错误体生成错误信息**(与本设计同构) | `maxUnstructuredResponseTextBytes = 2048` | 头部硬切 |
| OpenAI Python SDK `_exceptions.py` | 异常对象持有 body | **不截断**(内存对象,不落库) | — |
| Sentry Python `strip_string` | 事件写入前 trim | `max_value_length`,2.34.0 前默认 1024 | 头部 + `...`,另用 metadata 记原长 |
| Elastic APM | 长字段 | keyword 1024 / long field 10000 | 截断带省略号 |
| Python 标准库 `reprlib` | 给人读的长字符串 | `maxstring` | **头 + 尾,中间省略** |
**2048 对齐 k8s client-go**——它是唯一与本设计同场景(读 HTTP 错误体做诊断)的成熟先例。初稿的 500 仅以 issue 的单个样本(约 160 字符)为据,是拿一个样本定上限,已废弃。头部硬切在 k8s/Sentry 成立是因为它们截的是任意文本;本设计截的是结构化 JSON,故取 `reprlib` 的头尾策略。
遥测代价:纯 ASCII 约 2KB/条,纯中文最多约 6KB/条;5xx 重试 3 次即一次调用最多约 18KB。批处理场景(1050 次调用、5% 失败)约 300KB,`TEXT` 列可忽略。
message 与 `body_text` **共用同一变量**,不设两个长度:两份不同长度会让"遥测里看到的"与"下游 catch 到的"对不上,排查时反而多一层困惑。
### 3.5 改动清单
| 文件 | 改动 |
|---|---|
| `errors.py` | 基类新增 `body_text` 字段 + 与 `raw_text` 的界限 docstring |
| `transports/_http_errors.py` | **新建**:`summarize_body` / `response_body` / `_ERROR_BODY_CAP` |
| `transports/openai_compat.py` | `_status_to_error` 表驱动重写;`_translate_429``ctx` |
| `transports/monkey_ocr.py` | `_classify_status` 带摘要 |
| `errors.py` docstring + `ARCHITECTURE.md` §6.2 | 中转拓扑下 400 的提醒(§5.2) |
| `README.md:34` | 安装 pin `==1.1.*``>=1.2,<2`(§5.3,发布前置,漏改则下游拿不到本修复) |
三个调用点(`openai_compat.py:402` embed、`:417` stream、`:509` 非流式)**签名不变**,无需改动。
## 4. 非功能维度
### 4.1 并发与取消
新增全部是纯函数与数据字段,无状态、无锁、无 IO、不引入 `await``response_body` 只读已缓冲字节,`ResponseNotRead` 时直接返回空串而**不发起网络读**——否则会在错误路径上凭空插入一次可能挂住的 IO。`CancelledError` 路径逐字不变。
### 4.2 降级方向
响应体不可得(未读缓冲 / 解码失败 / 空体)→ `body_text=""`,**静默降级,绝不报错**。诊断信息属可观测性,按库铁律与缓存/遥测同档:缺了降级,不得把一次本可正确分类的失败变成不可分类的崩溃。流式路径的 `(await resp.aread()).decode("utf-8", errors="replace")`(`:416`)已是这个口径,保持。
### 4.3 幂等与重复
纯函数,同输入同输出。`summarize_body` 对自身输出再调用一次是幂等的:输出总长恒为 `2000 + len(标记) ≤ 2048`(标记形如 `…(略 N 字)…`,8 + N 的位数,现实中远不足 48),且不含需折叠的空白,故第二次调用走"原样返回"分支,不会出现标记被反复嵌套。
### 4.4 持久化与原子性
不新增表、不改 DDL、不动遥测端口的 22 字段与列序。摘要经既有 `error TEXT` 列落盘,原子性由既有单行写入保证。
### 4.5 安全与体积
- **响应体可能回显请求内容**(部分网关的 `error.param` 会带违规字段值)。截断 + 空白折叠是主要止血手段;字段 docstring 须写明"可能包含请求回显,已截断"。库不做内容脱敏——库不知道下游哪些字段敏感,猜测式脱敏只会同时丢掉诊断价值与安全性。
- **本设计不放大既有的读取风险**:`_complete_stream:416``aread()` 对错误响应体无大小上限(超大错误体可打爆内存),该风险今天已经存在(读完即丢),留存后只是更显眼。**不夹带修复**,见 §5.4。
## 5. 错误处理、语义与边界
### 5.1 错误分类
不改任何映射。`body_text` 是**旁路数据**,不参与任何治理判定——不影响重试、换源、熔断计数、AIMD、限流结算。这是本设计能与 ARCHITECTURE §6.1/§6.2 零冲突的根本原因。
### 5.2 400 语义:不改行为,补文档
Issue 报告了一个有说服力的观察:同字节 15 次重发全部成功、`prompt_tokens=0`、耗时 2996ms 远低于同批 631 次成功调用的最快值 7366ms——说明那次 400 来自中转服务自身抖动,而非"你的输入有问题"。
**仍不改分类**:400 重试对直连供应商是纯浪费(确定性坏输入,重试只烧配额并拖延失败);"中转也回 400"是**部署拓扑**引入的信息损失,库从状态码无从分辨。默认改为可重试 = 让所有直连用户为一种部署形态买单,且推翻已冻结的公共契约。
**但本设计本身就是对这个观察最好的答复**:body 留存后,下游能自己区分——中转抖动的 400 体与供应商 `invalid_request_error` 体形态不同。库不替下游做判断,而是把判断所需的信息交出去。配套文档动作:`RequestRejectedError` docstring 与 ARCHITECTURE §6.2 各加一句"经中转部署时 400 可能源于中转自身抖动,批处理场景下游宜自备兜底分类"。
### 5.3 兼容性
`LLMResponse` 一族的"字段只增不删不改名"约束(ARCHITECTURE §5.1)同样适用于异常。本次是**纯新增关键字参数且带默认值**:既有构造点、既有 `except` 写法、既有 `str(exc)` 消费方全部不受影响。message 文本变化不构成破坏——现有测试对这些 message 无格式依赖(仅 `test_openai_compat.py:558` match 源名)。
版本 **1.2.0**(公共类型新增字段属 minor;2026-08-16 人类定夺)。
**发布时必须同步改 README 的安装 pin**:`README.md:34` 现为 `"polygateway[redis,postgres,structured]==1.1.*"`,发 1.2.0 后照此命令安装的下游会**静默停在 1.1.2**——无报错、无警告,与 CLAUDE.md §4.4.1 点名的"极易漏改"完全同款(registry 长期停在 1.0.5 即此类事故)。本次改为 **`>=1.2,<2`**,把"每发一个 minor 就要通知三个下游改 pin"这一反复出现的麻烦一次性消除。此项列入实现计划的发布前置步骤,不是发布日的临时动作。
### 5.4 有意不夹带的两项(建议单开 issue)
| 项 | 说明 |
|---|---|
| `_status_to_error``operation` 硬编码 `"chat"`(`:134`),而 `embed()` 也调它(`:402`) | embedding 的 HTTP 错误在遥测里被标成 `operation="chat"`,是既有数据正确性缺陷,与本 issue 无关 |
| `_complete_stream:416``aread()` 无大小上限 | 恶意/故障网关的超大错误体可打爆内存,属独立的健壮性问题 |
两项都在本次重构触及的函数附近,但修它们既不服务 G1-G4,也各自需要独立的行为讨论——按反 gold-plating 铁律留给独立 issue。
## 6. 被否决的路线
### 6.1 给遥测端口加一列(22 → 23 字段)
最"正统"的结构化留存,但成本极不相称:端口 Protocol 签名变更 + SQLite/Postgres 双后端 DDL 迁移 + 下游已有表的 ALTER + 列序契约测试全线改动——为一个诊断串付出一次跨三项目的迁移。而复用既有 `error TEXT` 列可达成同样的可查证性。
### 6.2 只在 `_status_to_error` 打一条 WARNING 日志(Issue 方向二)
不采纳为**主**手段:日志与遥测是两套留存,日志轮转后仍然查不到,而 Issue 的痛点恰是"事后"。且库铁律要求库不擅自向下游日志流写入高频内容(4xx/5xx 在批处理下可能极高频)。message 携带摘要已让 loguru 侧的下游在捕获点自然拿到同一份信息,再加一条独立日志属重复留存。
### 6.3 截断放在异常构造器内
构造器自动规范化更"防遗漏",但会让下游自建异常时传入的文本被悄悄改写,违反 P4;且 message 里的摘要仍需在翻译层单独算一次,反而出现两条规范化路径。选定方案在翻译层算一次、两处共用,更简且更显式。
### 6.4 错误分类映射可插拔(provider profile 注入 classifier)
Issue 的中转 400 场景确实指向这个方向,但当前只有一个使用方且他们已用自己的兜底分类解决。`ProviderProfile`(`providers.py:17-45`)目前也没有这个扩展点,加它是新子系统级的设计。YAGNI:等第二个使用方提出。
## 7. 测试策略(先失败后通过)
**验收主张**:一次 400 调用后,注入的 recorder 收到的 `error` 串含网关响应体摘要。这条端到端断言直接对应 Issue 的痛点,是本设计成立与否的唯一硬判据;其余为覆盖性用例。
| # | 用例 | 覆盖 |
|---|---|---|
| 1 | **端到端遥测**:mock transport 返回 400 + 真实样本体 → 断言 recorder 收到的 `error` 含摘要 | G1 |
| 2 | 参数化状态码(400 / 401 / 404 兜底 / 429 普通 / 429 `insufficient_quota` / 500)→ 断言 message 含摘要且 `exc.body_text` 非空,**分类与既有断言逐一不变** | G2, G4 |
| 3 | 超长体 → 前 1400 字符与原文头部逐字相同、**末 600 字符与原文尾部逐字相同**、中段为 `…(略 N 字)…` 且 N 等于实际省略数;`body_text` 与 message 中的摘要逐字相同 | G3 |
| 3b | 长度恰为 2048 / 2049 的体 → 前者原样无标记,后者走头尾保留(边界) | §3.4 |
| 3c | **尾部关键字段可见**:以 issue 的真实样本尾部 `"code":"invalid_parameter_error"}}` 构造超长体 → 断言该串出现在摘要中 | §3.4 头尾决策的验收 |
| 3d | 摘要对自身幂等(再摘要一次不嵌套标记) | §4.3 |
| 4 | 多行缩进 JSON → 折叠为单行 | G3 |
| 5 | 空体 / 纯空白体 → 不拼悬空分隔符,`body_text == ""` | §3.3 |
| 6 | 非 JSON 体、非 UTF-8 字节 → 不抛异常,分类不变 | §4.2 |
| 7 | 流式错误路径(`_complete_stream` 415-417)同样带摘要 | G2 |
| 8 | `monkey_ocr._classify_status` 同款(含 `ResponseNotRead` 时降级为空串而非抛出) | G2, §4.2 |
| 9 | embedding 路径(`:402`)HTTP 错误带摘要 | G2 |
`tests/unit/test_errors.py:29` 现有的"四类构造形态"参数化用例需扩展 `body_text` 默认值断言(默认 `""`、可传入、`GatewayUnavailableError` 一族恒空)。
## 8. 人类定夺记录(2026-08-16)
| 议题 | 定夺 |
|---|---|
| 摘要上限与保留策略 | 初稿 500 + 头部硬切被否:上限提至 **2048**(对齐 k8s client-go 同场景先例),策略改为**头 1400 + 尾 600 + 省略字数标记**——人类指出"有用的信息可能只在后半部分",经调研证实 JSON 错误体的 `code`/`request_id` 确实收尾(§3.4、§3.4.1) |
| 429 是否设例外 | **不设**,一律拼摘要(§3.3) |
| 版本 | **1.2.0**,并同步把 README pin 由 `==1.1.*` 改为 `>=1.2,<2`(§5.3) |
@@ -33,7 +33,7 @@ date: 2026-08-06
## 对 issue 前提的四处修正 ## 对 issue 前提的四处修正
泄漏路径是**条**不是两条(`retry.py:216``progress_age_s()` 同样在 catch 之外);构造点 **22 处**;其中 2 处语义完全不同(未知源);`retry_after_s=0` 语义通但工程不通。 泄漏路径是**条**不是两条(判据: 该 gate 调用点是否被 `_record_quietly` 包裹——`QuotaGate` 的 try_acquire / stats / progress_age_s 与 `BreakerGate` 的 try_enter / retry_after_s 均未包裹,直达调用方);构造点 **22 处**;其中 2 处语义完全不同(未知源);`retry_after_s=0` 语义通但工程不通。
根因记录: `ARCHITECTURE.md` §6.1 错误分类表里 `GovernanceBackendError` **一次都没出现**——它是 M2 引入分布式后端时新增的,当时未回补架构表,于是它在"调用方视角的分类学"中从来没有位置,README 的遗漏是这个遗漏的下游后果。 根因记录: `ARCHITECTURE.md` §6.1 错误分类表里 `GovernanceBackendError` **一次都没出现**——它是 M2 引入分布式后端时新增的,当时未回补架构表,于是它在"调用方视角的分类学"中从来没有位置,README 的遗漏是这个遗漏的下游后果。
@@ -0,0 +1,62 @@
---
type: design
node_id: design:issue10-error-body-retention
title: "HTTP 错误响应体留存(Issue #10)"
date: 2026-08-16
---
# HTTP 错误响应体留存(Issue #10)
**来源**: Gitea issue #10(CHSAnalyzer3 现场,1050 张影像批处理中 1 张 400 被判确定性失败、事后无从查证)|**范围**: `errors.py` + 两个 transport|**全文**: `designs/2026-08-16-issue10-error-body-retention-design.md`|**相关**: [[design:issue8-stall-budget]](同为下游实测反馈驱动的治理修正)
## 问题
网关拒绝一次调用时,它说的话在 transport 翻译层被丢弃,进程中不再有任何副本:该模块无 logger、异常类无承载字段、库遥测只写 message。三条留存通道同时为空,故"永久查不到"。
## 根因(Issue 前提的关键修正)
Issue 建议"给异常加 `body_text` 字段,下游就能记进遥测"——**只做这一半解决不了它自己陈述的痛点**。库的逐次遥测写的是 `error=str(exc)`(`retry.py:552``telemetry.py``sqlite.py``error TEXT` 列),即**异常 message**;新增字段不进库的遥测表。下游说的"写进遥测表"是他们自己的埋点。
缺陷范围也大于 issue 所述:实为 6 处同构——`_status_to_error` 的 400 / 4xx 兜底 / 401·403 / 5xx 四支,`_translate_429` 两支(读了 body 判类型却不带),以及 `monkey_ocr._classify_status` 全部分支(message 只有 `HTTP {status}`)。Issue 场景"读表格"极可能正落在 OCR 路径。
## 选定方案
**摘要在翻译层算一次,同一份串同时进 message 与新增的基类字段**——前者解决"事后可查"(走既有遥测列,零 DDL),后者解决下游结构化留存。
| 决策 | 理由 |
|---|---|
| 字段加在 `PolyGatewayError` 基类,非 `RequestRejectedError` | 这些错误全由同一个 HTTP 响应翻译而来,"对方说了什么"与"属于哪一类"正交;只加子类,下次给 `SourceDeadError` 加又是一次公共 API 变更 + 人类门 |
| 与 `ResultInvalidError.raw_text` 的界限写进 docstring | `body_text` = 非 2xx 的拒绝理由;`raw_text` = 2xx 但不可解析的模型输出。两个"原文字段"不钉死必被混用 |
| 新建 `transports/_http_errors.py` 共用摘要口径 | 两 transport 各有分类逻辑(OCR 无 429 细分,有意保留),但摘要必须同一份,否则就是下一个"只修一半" |
| `_status_to_error` 改表驱动 | 五分支各拼各的 message,加摘要即五处重复;查表 + 单点拼装后代码更短 |
| 摘要 = 折叠空白 + 总长 ≤ 2048,超出则**保留头 1400 + 尾 600**,中段记省略字数 | 折叠是因错误体常是缩进 JSON,拼进 message 会炸成多行。2048 对齐 k8s client-go 的 `maxUnstructuredResponseTextBytes`(唯一同场景先例);**头尾保留取自 `reprlib`**——JSON 错误体的 `code`/`request_id` 收尾,头部硬切正好切掉向网关追查唯一有用的部分(人类质疑 + 2026-08-16 调研,初稿的 500 + 头部硬切已废) |
| 429 不设例外 | 例外就是下一个复发点;`insufficient_quota` 那支的配额细节全在 body 里 |
| 400 治理语义不动,只补文档 | 见下 |
## 400 语义:不改行为,本次修复本身就是答复
Issue 给出有力证据(同字节 15 次重发全成功、`prompt_tokens=0`、2996ms 远低于同批成功最快的 7366ms),说明那次 400 来自中转服务抖动而非坏输入。**仍不改分类**:400 重试对直连供应商是纯浪费,而"中转也回 400"是部署拓扑引入的信息损失,库从状态码无从分辨;默认改可重试 = 让所有直连用户为一种部署形态买单。
但 body 留存后**下游能自己区分**——中转抖动体与供应商 `invalid_request_error` 体形态不同。库不替下游判断,把判断所需的信息交出去。配套在 docstring 与 ARCHITECTURE §6.2 加一句中转拓扑提醒。
## 被否决的备选
| 备选 | 否决理由 |
|---|---|
| 遥测端口加一列(22 → 23 字段) | 端口签名变更 + 双后端 DDL + 下游 ALTER + 列序契约全线改动,为一个诊断串付出跨三项目迁移;复用既有 `error` 列可达成同样可查证性 |
| 只打一条 WARNING 日志(issue 方向二) | 日志轮转后仍查不到,而痛点恰是"事后";且 4xx/5xx 在批处理下可能极高频 |
| 截断放进异常构造器 | 下游自建异常的文本被悄悄改写(违反 P4),且 message 侧仍需单独算一次,反出现两条规范化路径 |
| 错误分类映射可插拔 | Issue 场景确实指向它,但当前只有一个使用方且已用自己的兜底分类解决;`ProviderProfile` 无此扩展点,加它是子系统级设计。YAGNI |
## 有意不夹带(留独立 issue)
- `_status_to_error``operation` 硬编码 `"chat"`,而 `embed()` 也调它 → embedding 的 HTTP 错误在遥测里被标成 chat。
- `_complete_stream``aread()` 对错误响应体无大小上限,超大错误体可打爆内存(既有风险,留存后更显眼)。
两项都在本次触及的函数附近,但均不服务本 issue 目标,且各需独立行为讨论。
## 验收主张
一次 400 调用后,注入的 recorder 收到的 `error` 串含网关响应体摘要——这条端到端断言是本设计成立与否的唯一硬判据,其余用例为覆盖性(状态码参数化、截断边界 2048/2049、**尾部关键字段可见**、空白折叠、空体不拼悬空分隔符、非 UTF-8 不炸、流式路径、OCR 路径含 `ResponseNotRead` 降级)。
**发布约束**:版本 1.2.0,且 README 安装 pin 必须由 `==1.1.*` 改为 `>=1.2,<2`——否则照 README 安装的下游静默停在 1.1.2,拿不到本修复。
@@ -0,0 +1,46 @@
---
type: design
node_id: design:issue8-stall-budget
title: "stall 判定改为非生产性等待口径"
date: 2026-08-06
---
# stall 判定改为非生产性等待口径
**全文**: `designs/2026-08-06-issue8-stall-budget-design.md`(已批准 2026-08-06)|**来源**: Gitea issue #8 |**实施**: [[plan:issue8-stall-budget-plan]]
## 问题
`timeout_s ≥ stall_window_s` 时,一次耗满超时的请求即判 scope 死,`max_attempts` **静默失效**(无报错无 warning)。`stall_window_s` 默认 300 极易被 `TIMEOUT_S` 追平,"只配 timeout 不配 stall"这种最常见写法正好踩中。
## 根因
**两个预算重叠计费**:真实尝试的耗时同时向重试预算(`max_attempts`)与 stall 预算(`stall_window_s`)计费,而后者更小,必然先耗尽。
## 选定方案
`StallClock` 让 stall 只累计非生产性等待。**划分依据是"谁消耗重试预算"**,不是"是否发出请求"——烧 `max_attempts` 的时间不烧 `stall_window_s`,不烧 `max_attempts` 的时间(含 429 尝试本身)归 stall 治理。
关键理由:
- **消除耦合而非守护耦合**。`stall_window_s``timeout_s` 自此无关系,配置方不必心算 `stall > timeout × retries`
- **`inf` 语义因此不必改**。新口径下"非生产性排队耗满窗口且 scope 从未出餐"判死本就正当,`inf` 从"有害恒真"回归为"正确的保守默认"。一次改动解决问题,优于两次改动互相牵制。
- **取补集实现**(总时间减 `_attempt` 耗时)而非逐处标记 sleep:埋点 7 处降到 3 处,且将来新增等待路径自动计入 stall,默认安全。
## 被否决的备选
| 备选 | 否决理由 |
|---|---|
| **装配期校验 `stall_window_s > max(timeout_s)`**(issue 建议方向 1) | 治标:把缺陷固化成配置契约。且约束值须为 `timeout × max_attempts`(本机 900s),使 stall 兜底迟钝到近乎失效——修好一个洞挖开另一个。仍挡不住残余情形 |
| **`inf` 不参与判死**(issue 建议方向 2) | 新口径下 `inf` 已无害。单独改它会制造冷启动兜底真空(429 免预算无其他兜底),并反转 `test_both_windows_exceeded_raises_stalled` 钉住的行为、与 CHS 蓝本分叉 |
| **逐处标记 sleep** | 埋点 7 处且默认危险:新增等待路径忘记标记即成 stall 盲区 |
| **给 embedding/ocr 补主循环 stall 判定** | 前提不成立。429 免预算是 chat 独有,embedding/ocr 无条件 `fails += 1`,两条循环路径均已封闭,补齐等于凭空新增判死路径 |
| **删除既有 ttft 装配校验** | 其理由虽已消失(TTFT 属生产性时间),但校验无害且不误拒合理配置;删除需动 ARCHITECTURE §7.3 契约 G6,超出本 issue 范围(人类定夺:保留并改注释) |
## 实施期订正(§3.6)
初稿按"是否发出请求"划分,使 429 尝试**两个预算都不烧**(429 免重试预算,其耗时又算生产性)。排队型网关持满 timeout 才回 429 时实测挂 **25.2 小时**(301 次尝试),而改前只有 301s——**把一个 bug 换成了更严重的 bug**。由独立验证发现。订正为按"谁消耗重试预算"划分,429 尝试耗时退还 stall 账,实测回到 301s。
## 不变量
双条件结构、`progress_age_s()``inf` 语义、429 免预算、退避与 jitter 公式、`fail_fast` 分支、`AllSourcesExhausted` 字段与 `reason` 取值全部未动——**错误面零变更**。
@@ -0,0 +1,62 @@
---
type: design
node_id: design:issue9-telemetry-ddl-probe
title: "建表前先探测,判死只认「确定写不进去」"
date: 2026-08-07
---
# 建表前先探测,判死只认「确定写不进去」
**来源**: Gitea issue #9(CHSAnalyzer3 现场)|**范围**: `telemetry/postgres.py` 单模块,无独立 plan(小改动自判)|**相关**: [[design:response-observability-fields]](issue #3 修的是同一个坑的另一半)
## 问题
应用账号有表级 `INSERT`、表也已存在,但没有 schema 的 `CREATE` 权限时,`_ensure_ready()``CREATE TABLE IF NOT EXISTS` 被拒 → `_failed = True`**整个进程遥测永久 no-op**。业务调用一切正常,只留一行 warning,从外部完全看不出异常;下游 CHSAnalyzer3 首次端到端跑的 150+ 次调用数据因此全丢且无法补回。
## 根因
**PostgreSQL 对 schema 的 CREATE 权限检查早于 `IF NOT EXISTS` 的存在性判断**(`RangeVarGetAndCheckCreationNamespace()` 先 aclcheck 后查 relid)。这与 issue #3`ALTER TABLE` 的 ownership 检查早于 `IF NOT EXISTS` 是同一类问题——当时只修了补列那一半,建表这一半原样留着,于是同一账号形态下"补列失败只丢一行日志接着干活,建表失败却把整个 recorder 判死"。
**实测(PostgreSQL 16.14,临时角色只授 `SELECT, INSERT ON llm_calls`)**:
| 语句 | 结果 |
|---|---|
| `SELECT to_regclass('llm_calls')` | 非 NULL(表就在那儿) |
| `CREATE TABLE IF NOT EXISTS llm_calls (...)` | **被拒 InsufficientPrivilegeError: permission denied for schema** |
| `INSERT INTO llm_calls ...` | 通过 |
| `ALTER TABLE ... ADD COLUMN IF NOT EXISTS` | 被拒 must be owner(即 issue #3 那条) |
## 选定方案
两条,第二条才是治本的那条:
1. **表存在就绝不发 DDL**。探测走 `to_regclass`(不需要任何权限,且与 `INSERT` 走同一套 search_path 解析——裸 `CREATE TABLE` 落在首个**可建**的 schema,可能与写入命中的不是同一张表,故探测优先反而更准)。表不存在才建;新建表列已齐全,顺带跳过补列。
2. **"结构性失能"的判据从「初始化时出过异常」收窄为「确定写不进去」**:
| 情形 | 处置 | 理由 |
|---|---|---|
| 建池失败 | 永久 no-op | 重试要在业务调用路径上内联吞掉 connect 超时 |
| 表存在 | 不发 DDL,只补列(失败仅 warning) | 本 issue 的直接修复 |
| 表不存在 → 建表成功 | 就绪,跳过补列 | 新建表列已齐 |
| 表不存在 → 建表失败 | 永久 no-op | 后续 INSERT 必然全败,重试无意义、日志纯噪音 |
| 探测/取连接失败 | 只跳过本条,下次调用重试 | 瞬时抖动,判死代价远大于多一次往返 |
**SQLite 侧有意不对称**:实测其对已存在的表在**解析期**就把 `CREATE TABLE IF NOT EXISTS` 短路掉——另一连接持 `BEGIN EXCLUSIVE`、或文件 `chmod 444` 时该语句均通过(同条件下 `INSERT` 与新表名建表分别报 database is locked / readonly database),既不抢写锁也不检查可写性。故 PG 侧的坑在此不存在,加探测零收益。**需要对称的是保证(表存在就不该因建表失败而失能),不是代码**;结论已钉进 `sqlite.py` 模块 docstring,防止后人为"对称"加回来。
## 被否决的备选
| 备选 | 否决理由 |
|---|---|
| 只加探测,`_failed` 语义不动(issue 原方案) | 治标。初始化瞬间的 DB 抖动、一次 `pool.acquire` 失败、search_path 配错仍会让整个进程永久失遥测——同一个开关,换个触发口 |
| 除建池外一律不判死 | 方向最统一,但表真的不存在时每次调用都发一条注定失败的 INSERT + 一条 warning(150 次调用 = 150 行噪音),而这种情形是**可确定判定**的,没必要留活路 |
| 捕获 `InsufficientPrivilegeError` 特判放行 | 按异常类型打补丁,漏一种错误码就复发;探测是把"该不该发这条 DDL"判断在前,与错误面无关 |
| SQLite 侧同步加探测 | 实测证明零收益,属为对称而对称的 gold-plating |
## 遗留
**SQLite 的窄缝**:表不存在 + 构造瞬间库被排他锁(多进程共库)→ `__init__` 里的建表失败 → recorder 永久失能。修它要把 SQLite 也改成 lazy 重试结构,超出本 issue 范围,记此备查。
## 测试证据
- 单测 `TestPostgresTableProbe`(5 例,fake conn):表存在不发 DDL / DDL 被拒仍照常 INSERT 且 `_failed` 不置位 / 表缺失则建表且不补列 / 表缺失且建不出来才判死 / 探测失败下次重试。
- 集成 `TestLeastPrivilegeDeployment`(真实 PG,临时 schema + 临时角色,teardown 删净):先钉死"该角色确实建不了表"这条库外事实,再验两行记录照常落库。**修复前该用例复现 issue 原文那行 warning 并失败**。
+34
View File
@@ -140,6 +140,26 @@
"id": "plan:governance-backend-error", "id": "plan:governance-backend-error",
"label": "实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)", "label": "实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)",
"type": "plan" "type": "plan"
},
{
"id": "design:issue8-stall-budget",
"label": "stall 判定改为非生产性等待口径",
"type": "design"
},
{
"id": "plan:issue8-stall-budget-plan",
"label": "issue #8 实施计划: stall 非生产性等待口径",
"type": "plan"
},
{
"id": "design:issue10-error-body-retention",
"label": "HTTP 错误响应体留存(Issue #10)",
"type": "design"
},
{
"id": "plan:issue10-error-body-retention-plan",
"label": "实现计划: HTTP 错误响应体留存(Issue #10)",
"type": "plan"
} }
], ],
"links": [ "links": [
@@ -254,6 +274,20 @@
"relation": "implements", "relation": "implements",
"evidence": "T1-T5 逐任务实现设计 §3 的五项决策与 §8 影响面清单", "evidence": "T1-T5 逐任务实现设计 §3 的五项决策与 §8 影响面清单",
"added": "2026-08-06T08:08:51.865565+00:00" "added": "2026-08-06T08:08:51.865565+00:00"
},
{
"source": "plan:issue8-stall-budget-plan",
"target": "design:issue8-stall-budget",
"relation": "implements",
"evidence": "T1-T6 实施该设计,含 §3.6 订正",
"added": "2026-08-06T14:58:01.673693+00:00"
},
{
"source": "plan:issue10-error-body-retention-plan",
"target": "design:issue10-error-body-retention",
"relation": "implements",
"evidence": "7 任务覆盖设计 G1-G4 与 §7 全部验收用例",
"added": "2026-08-16T09:50:57.830855+00:00"
} }
] ]
} }
+12 -3
View File
@@ -1,8 +1,8 @@
# Research Wiki 索引 # Research Wiki 索引
> 自动生成,更新时间:2026-08-06 08:11 UTC > 自动生成,更新时间:2026-08-16 09:50 UTC
## design (23) ## design (28)
- [2026-07-20-m1-core-design](designs/2026-07-20-m1-core-design.md) `design:2026-07-20-m1-core-design` - [2026-07-20-m1-core-design](designs/2026-07-20-m1-core-design.md) `design:2026-07-20-m1-core-design`
- [2026-07-20-m2-distributed-design](designs/2026-07-20-m2-distributed-design.md) `design:2026-07-20-m2-distributed-design` - [2026-07-20-m2-distributed-design](designs/2026-07-20-m2-distributed-design.md) `design:2026-07-20-m2-distributed-design`
- [2026-07-21-m25-resilience-design](designs/2026-07-21-m25-resilience-design.md) `design:2026-07-21-m25-resilience-design` - [2026-07-21-m25-resilience-design](designs/2026-07-21-m25-resilience-design.md) `design:2026-07-21-m25-resilience-design`
@@ -14,15 +14,20 @@
- [2026-07-31-response-observability-fields-design](designs/2026-07-31-response-observability-fields-design.md) `design:2026-07-31-response-observability-fields-design` - [2026-07-31-response-observability-fields-design](designs/2026-07-31-response-observability-fields-design.md) `design:2026-07-31-response-observability-fields-design`
- [2026-07-31-sampling-params-design](designs/2026-07-31-sampling-params-design.md) `design:2026-07-31-sampling-params-design` - [2026-07-31-sampling-params-design](designs/2026-07-31-sampling-params-design.md) `design:2026-07-31-sampling-params-design`
- [2026-08-06-governance-backend-error-design](designs/2026-08-06-governance-backend-error-design.md) `design:2026-08-06-governance-backend-error-design` - [2026-08-06-governance-backend-error-design](designs/2026-08-06-governance-backend-error-design.md) `design:2026-08-06-governance-backend-error-design`
- [2026-08-06-issue8-stall-budget-design](designs/2026-08-06-issue8-stall-budget-design.md) `design:2026-08-06-issue8-stall-budget-design`
- [2026-08-16-issue10-error-body-retention-design](designs/2026-08-16-issue10-error-body-retention-design.md) `design:2026-08-16-issue10-error-body-retention-design`
- [est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)](designs/est-tokens-decoupling.md) `design:est-tokens-decoupling` - [est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)](designs/est-tokens-decoupling.md) `design:est-tokens-decoupling`
- [GatewaySettings 装配校验补齐(第二轮)](designs/settings-invariants-round-2.md) `design:settings-invariants-round-2` - [GatewaySettings 装配校验补齐(第二轮)](designs/settings-invariants-round-2.md) `design:settings-invariants-round-2`
- [GatewaySettings 跨字段不变量守卫的生效范围](designs/settings-invariant-guards.md) `design:settings-invariant-guards` - [GatewaySettings 跨字段不变量守卫的生效范围](designs/settings-invariant-guards.md) `design:settings-invariant-guards`
- [HTTP 错误响应体留存(Issue #10)](designs/issue10-error-body-retention.md) `design:issue10-error-body-retention`
- [M1 核心里程碑设计:公共签名冻结与治理栈落地](designs/m1-core-design.md) `design:m1-core-design` - [M1 核心里程碑设计:公共签名冻结与治理栈落地](designs/m1-core-design.md) `design:m1-core-design`
- [M2 分布式:Redis 治理后端+背压+Postgres 遥测+pricing+Embedding+压测 harness](designs/m2-distributed.md) `design:m2-distributed` - [M2 分布式:Redis 治理后端+背压+Postgres 遥测+pricing+Embedding+压测 harness](designs/m2-distributed.md) `design:m2-distributed`
- [M2.5 治理韧性: 半死源隔离与健康感知调度](designs/m25-resilience.md) `design:m25-resilience` - [M2.5 治理韧性: 半死源隔离与健康感知调度](designs/m25-resilience.md) `design:m25-resilience`
- [M3 OCR 端口族设计](designs/m3-ocr.md) `design:m3-ocr` - [M3 OCR 端口族设计](designs/m3-ocr.md) `design:m3-ocr`
- [M4 迁移验证设计(GovDoc→CHS,发 v1.0)](designs/m4-migration.md) `design:m4-migration` - [M4 迁移验证设计(GovDoc→CHS,发 v1.0)](designs/m4-migration.md) `design:m4-migration`
- [stall 判定改为非生产性等待口径](designs/issue8-stall-budget.md) `design:issue8-stall-budget`
- [响应可观测字段扩展(Issue #3)](designs/response-observability-fields.md) `design:response-observability-fields` - [响应可观测字段扩展(Issue #3)](designs/response-observability-fields.md) `design:response-observability-fields`
- [建表前先探测,判死只认「确定写不进去」](designs/issue9-telemetry-ddl-probe.md) `design:issue9-telemetry-ddl-probe`
- [推理开关能力建模与 reasoning_tokens 采集(issue #5 + #6)](designs/2026-08-02-thinking-capability-design.md) `design:2026-08-02-thinking-capability-design` - [推理开关能力建模与 reasoning_tokens 采集(issue #5 + #6)](designs/2026-08-02-thinking-capability-design.md) `design:2026-08-02-thinking-capability-design`
- [治理后端故障归位为 scope 级不可用(Issue #7)](designs/governance-backend-error.md) `design:governance-backend-error` - [治理后端故障归位为 scope 级不可用(Issue #7)](designs/governance-backend-error.md) `design:governance-backend-error`
- [采样参数透传设计(issue #4)](designs/sampling-params.md) `design:sampling-params` - [采样参数透传设计(issue #4)](designs/sampling-params.md) `design:sampling-params`
@@ -41,7 +46,7 @@
- [P7 OCR soak 验收: 99.73% 与 13 不变量全 PASS](findings/p7-ocr-soak.md) `finding:p7-ocr-soak` - [P7 OCR soak 验收: 99.73% 与 13 不变量全 PASS](findings/p7-ocr-soak.md) `finding:p7-ocr-soak`
- [推理开关与 reasoning_tokens: 供应商实测与业界做法](findings/2026-08-02-thinking-switch-and-reasoning-tokens.md) `finding:2026-08-02-thinking-switch-and-reasoning-tokens` - [推理开关与 reasoning_tokens: 供应商实测与业界做法](findings/2026-08-02-thinking-switch-and-reasoning-tokens.md) `finding:2026-08-02-thinking-switch-and-reasoning-tokens`
## plan (19) ## plan (23)
- [2026-07-20-m1-core-plan](plans/2026-07-20-m1-core-plan.md) `plan:2026-07-20-m1-core-plan` - [2026-07-20-m1-core-plan](plans/2026-07-20-m1-core-plan.md) `plan:2026-07-20-m1-core-plan`
- [2026-07-20-m2-distributed-plan](plans/2026-07-20-m2-distributed-plan.md) `plan:2026-07-20-m2-distributed-plan` - [2026-07-20-m2-distributed-plan](plans/2026-07-20-m2-distributed-plan.md) `plan:2026-07-20-m2-distributed-plan`
- [2026-07-21-m25-resilience-plan](plans/2026-07-21-m25-resilience-plan.md) `plan:2026-07-21-m25-resilience-plan` - [2026-07-21-m25-resilience-plan](plans/2026-07-21-m25-resilience-plan.md) `plan:2026-07-21-m25-resilience-plan`
@@ -51,13 +56,17 @@
- [2026-07-31-response-observability-fields](plans/2026-07-31-response-observability-fields.md) `plan:2026-07-31-response-observability-fields` - [2026-07-31-response-observability-fields](plans/2026-07-31-response-observability-fields.md) `plan:2026-07-31-response-observability-fields`
- [2026-07-31-sampling-params](plans/2026-07-31-sampling-params.md) `plan:2026-07-31-sampling-params` - [2026-07-31-sampling-params](plans/2026-07-31-sampling-params.md) `plan:2026-07-31-sampling-params`
- [2026-08-06-governance-backend-error-plan](plans/2026-08-06-governance-backend-error-plan.md) `plan:2026-08-06-governance-backend-error-plan` - [2026-08-06-governance-backend-error-plan](plans/2026-08-06-governance-backend-error-plan.md) `plan:2026-08-06-governance-backend-error-plan`
- [2026-08-06-issue8-stall-budget](plans/2026-08-06-issue8-stall-budget.md) `plan:2026-08-06-issue8-stall-budget`
- [2026-08-16-issue10-error-body-retention](plans/2026-08-16-issue10-error-body-retention.md) `plan:2026-08-16-issue10-error-body-retention`
- [est_tokens 解耦实施计划](plans/est-tokens-decoupling.md) `plan:est-tokens-decoupling` - [est_tokens 解耦实施计划](plans/est-tokens-decoupling.md) `plan:est-tokens-decoupling`
- [issue #8 实施计划: stall 非生产性等待口径](plans/issue8-stall-budget-plan.md) `plan:issue8-stall-budget-plan`
- [M1 核心里程碑实现计划](plans/m1-core-plan.md) `plan:m1-core-plan` - [M1 核心里程碑实现计划](plans/m1-core-plan.md) `plan:m1-core-plan`
- [M2 分布式实现计划](plans/m2-distributed.md) `plan:m2-distributed` - [M2 分布式实现计划](plans/m2-distributed.md) `plan:m2-distributed`
- [M2.5 治理韧性实现计划](plans/m25-resilience.md) `plan:m25-resilience` - [M2.5 治理韧性实现计划](plans/m25-resilience.md) `plan:m25-resilience`
- [M3 OCR 实现计划](plans/m3-ocr.md) `plan:m3-ocr` - [M3 OCR 实现计划](plans/m3-ocr.md) `plan:m3-ocr`
- [M4 迁移实现计划(T0-T14)](plans/m4-migration.md) `plan:m4-migration` - [M4 迁移实现计划(T0-T14)](plans/m4-migration.md) `plan:m4-migration`
- [响应可观测字段扩展实现计划](plans/response-observability-fields.md) `plan:response-observability-fields` - [响应可观测字段扩展实现计划](plans/response-observability-fields.md) `plan:response-observability-fields`
- [实现计划: HTTP 错误响应体留存(Issue #10)](plans/issue10-error-body-retention-plan.md) `plan:issue10-error-body-retention-plan`
- [实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)](plans/governance-backend-error.md) `plan:governance-backend-error` - [实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)](plans/governance-backend-error.md) `plan:governance-backend-error`
- [推理开关能力建模与 reasoning_tokens 采集实施计划(issue #5 + #6)](plans/2026-08-02-thinking-capability.md) `plan:2026-08-02-thinking-capability` - [推理开关能力建模与 reasoning_tokens 采集实施计划(issue #5 + #6)](plans/2026-08-02-thinking-capability.md) `plan:2026-08-02-thinking-capability`
- [采样参数透传实现计划(issue #4)](plans/sampling-params-plan.md) `plan:sampling-params-plan` - [采样参数透传实现计划(issue #4)](plans/sampling-params-plan.md) `plan:sampling-params-plan`
+12
View File
@@ -90,3 +90,15 @@
- [2026-08-06 08:08 UTC] 新增边: plan:governance-backend-error --implements--> design:governance-backend-error - [2026-08-06 08:08 UTC] 新增边: plan:governance-backend-error --implements--> design:governance-backend-error
- [2026-08-06 08:08 UTC] 重建索引: 57 篇页面 - [2026-08-06 08:08 UTC] 重建索引: 57 篇页面
- [2026-08-06 08:11 UTC] 重建索引: 57 篇页面 - [2026-08-06 08:11 UTC] 重建索引: 57 篇页面
- [2026-08-06 14:57 UTC] 新增 design: stall 判定改为非生产性等待口径 (design:issue8-stall-budget)
- [2026-08-06 14:58 UTC] 新增 plan: issue #8 实施计划: stall 非生产性等待口径 (plan:issue8-stall-budget-plan)
- [2026-08-06 14:58 UTC] 新增边: plan:issue8-stall-budget-plan --implements--> design:issue8-stall-budget
- [2026-08-06 14:58 UTC] 重建索引: 61 篇页面
- [2026-08-07 15:20 UTC] 新增 design: 建表前先探测,判死只认「确定写不进去」 (design:issue9-telemetry-ddl-probe)
- [2026-08-07 15:20 UTC] 重建索引: 62 篇页面
- [2026-08-07 15:11 UTC] 重建索引: 62 篇页面
- [2026-08-16 09:09 UTC] 新增 design: HTTP 错误响应体留存(Issue #10) (design:issue10-error-body-retention)
- [2026-08-16 09:12 UTC] 重建索引: 64 篇页面
- [2026-08-16 09:50 UTC] 新增 plan: 实现计划: HTTP 错误响应体留存(Issue #10) (plan:issue10-error-body-retention-plan)
- [2026-08-16 09:50 UTC] 新增边: plan:issue10-error-body-retention-plan --implements--> design:issue10-error-body-retention
- [2026-08-16 09:50 UTC] 重建索引: 66 篇页面
@@ -107,7 +107,7 @@ class BreakerGate:
## 任务清单 ## 任务清单
### - [ ] T1: ARCHITECTURE §6.1 回补(必须先行) ### - [x] T1: ARCHITECTURE §6.1 回补(必须先行)
**文件**: `research-wiki/ARCHITECTURE.md`(§6.1,约 372-380 行) **文件**: `research-wiki/ARCHITECTURE.md`(§6.1,约 372-380 行)
@@ -125,7 +125,7 @@ class BreakerGate:
--- ---
### - [ ] T2: errors.py 纯增量(新常量、新 reason、新类)+ 导出 ### - [x] T2: errors.py 纯增量(新常量、新 reason、新类)+ 导出
**文件**: 改 `src/polygateway/errors.py``src/polygateway/__init__.py`;改 `tests/unit/test_errors.py` **文件**: 改 `src/polygateway/errors.py``src/polygateway/__init__.py`;改 `tests/unit/test_errors.py`
@@ -148,7 +148,7 @@ class BreakerGate:
--- ---
### - [ ] T3: `GovernanceBackendError` 归位 + 22 处构造点 + scope 注入(原子) ### - [x] T3: `GovernanceBackendError` 归位 + 22 处构造点 + scope 注入(原子)
**文件**: 改 `src/polygateway/errors.py``backends/redis/limiter.py``backends/redis/breaker.py``backends/memory/limiter.py``middleware/ratelimit.py``middleware/breaker.py``middleware/retry.py``ocr.py``embedding.py`;改 `tests/unit/test_errors.py``tests/unit/test_backpressure.py``tests/unit/test_redis_key_layout.py``tests/integration/test_redis_cross_connection.py` **文件**: 改 `src/polygateway/errors.py``backends/redis/limiter.py``backends/redis/breaker.py``backends/memory/limiter.py``middleware/ratelimit.py``middleware/breaker.py``middleware/retry.py``ocr.py``embedding.py`;改 `tests/unit/test_errors.py``tests/unit/test_backpressure.py``tests/unit/test_redis_key_layout.py``tests/integration/test_redis_cross_connection.py`
@@ -170,7 +170,7 @@ class BreakerGate:
- `backends/redis/breaker.py``:370 / :388 / :410 / :422 / :432`(5 处) - `backends/redis/breaker.py``:370 / :388 / :410 / :422 / :432`(5 处)
4. **两个 gate 包装器**: 构造函数改为上文"关键接口"的签名;`QuotaGate` 4 处(`ratelimit.py:30/38/46/54`)与 `BreakerGate` 5 处(`breaker.py:26/36/46/54/62`)的 `raise``scope=self._scope` 4. **两个 gate 包装器**: 构造函数改为上文"关键接口"的签名;`QuotaGate` 4 处(`ratelimit.py:30/38/46/54`)与 `BreakerGate` 5 处(`breaker.py:26/36/46/54/62`)的 `raise``scope=self._scope`
- 各方法开头的 `except GovernanceBackendError: raise` **保持不变**(后端层已填好 scope,重建实例只会重复构造,设计 §3.3) - ~~各方法开头的 `except GovernanceBackendError: raise` **保持不变**~~ **← 这条是错的,2026-08-06 独立验证时炸出(见 §T6)**。正确做法: 该放行必须扩为 `except (GovernanceBackendError, SourceNotConfiguredError): raise`,否则新增的兄弟类型会落进下一行的 `except Exception` 被**重新包成** `GovernanceBackendError`,使 Q1 的拆分在唯一的生产路径上完全失效
5. **三处装配各传 scope**(三处的 `self._scope` 均已在装配前赋值,无需调整顺序): 5. **三处装配各传 scope**(三处的 `self._scope` 均已在装配前赋值,无需调整顺序):
@@ -186,7 +186,7 @@ class BreakerGate:
|---|---|---| |---|---|---|
| `GovernanceBackendError` 可被 `except GatewayUnavailableError` 接住,且 `reason == "governance_backend_down"``retry_after_s == 5.0` | `tests/unit/test_errors.py` | 改前非其子类,`pytest.raises(GatewayUnavailableError)` 不匹配 | | `GovernanceBackendError` 可被 `except GatewayUnavailableError` 接住,且 `reason == "governance_backend_down"``retry_after_s == 5.0` | `tests/unit/test_errors.py` | 改前非其子类,`pytest.raises(GatewayUnavailableError)` 不匹配 |
| `str(exc)` 仍为构造时的诊断串(防 §3.5 回归) | `tests/unit/test_errors.py` | 改前无该风险但改后若漏写 `self.args` 即失败,是回归护栏 | | `str(exc)` 仍为构造时的诊断串(防 §3.5 回归) | `tests/unit/test_errors.py` | 改前无该风险但改后若漏写 `self.args` 即失败,是回归护栏 |
| 三条泄漏路径(`try_acquire` / `try_enter` / `progress_age_s`)抛出的异常带正确 `scope`、且可被 `except GatewayUnavailableError` 接住 | `tests/unit/test_backpressure.py`**三条都要新增桩**。现状: `progress_age_s` 只有 `TestQuotaGateProgressAge`(`:243-257`)覆盖包装行为、不验 scope;`try_acquire`(`QuotaGate`)与 `try_enter`(`BreakerGate`)**完全无桩** | 改前异常无 `scope` 属性 → `AttributeError`;两条新路径改前无覆盖 | | 闸门泄漏路径(共五条,见设计 §1.1)抛出的异常带正确 `scope`、且可被 `except GatewayUnavailableError` 接住;钉住 `try_acquire` / `try_enter` / `progress_age_s` 三条代表路径 | `tests/unit/test_backpressure.py`**三条都要新增桩**。现状: `progress_age_s` 只有 `TestQuotaGateProgressAge`(`:243-257`)覆盖包装行为、不验 scope;`try_acquire`(`QuotaGate`)与 `try_enter`(`BreakerGate`)**完全无桩** | 改前异常无 `scope` 属性 → `AttributeError`;两条新路径改前无覆盖 |
| 未知源抛 `SourceNotConfiguredError`,且断言它**不是** `GatewayUnavailableError` | 改 `tests/unit/test_redis_key_layout.py:70-74`(`test_unknown_source_rejected`,现断言 `GovernanceBackendError`);内存版**当前无对应用例,需新增**一条同款(`backends/memory/limiter.py:92``_cfg("nope")`) | 改前 redis 版类型断言失败;内存版改前无覆盖(该分支从未被测过) | | 未知源抛 `SourceNotConfiguredError`,且断言它**不是** `GatewayUnavailableError` | 改 `tests/unit/test_redis_key_layout.py:70-74`(`test_unknown_source_rejected`,现断言 `GovernanceBackendError`);内存版**当前无对应用例,需新增**一条同款(`backends/memory/limiter.py:92``_cfg("nope")`) | 改前 redis 版类型断言失败;内存版改前无覆盖(该分支从未被测过) |
| Redis 真实掉线时准入侧抛 scope 级异常且 `reason == "governance_backend_down"` | `tests/integration/test_redis_cross_connection.py:228-245`(真实 Redis,不 mock) | 改前无 `reason` 属性 | | Redis 真实掉线时准入侧抛 scope 级异常且 `reason == "governance_backend_down"` | `tests/integration/test_redis_cross_connection.py:228-245`(真实 Redis,不 mock) | 改前无 `reason` 属性 |
@@ -211,7 +211,7 @@ class BreakerGate:
--- ---
### - [ ] T4: 公开错误面文档(issue #7 第二诉求) ### - [x] T4: 公开错误面文档(issue #7 第二诉求)
**文件**: 改 `README.md`(§"错误模型(四分类)",约 114-125 行)、`research-wiki/migrations/chsanalyzer.md` **文件**: 改 `README.md`(§"错误模型(四分类)",约 114-125 行)、`research-wiki/migrations/chsanalyzer.md`
@@ -238,7 +238,7 @@ class BreakerGate:
--- ---
### - [ ] T5: 版本 1.1.0 + CHANGELOG + Wiki 同步 ### - [x] T5: 版本 1.1.0 + CHANGELOG + Wiki 同步
**文件**: 改 `pyproject.toml`(version)、`src/polygateway/__init__.py`(`__version__`)、`CHANGELOG.md`;按 `research-wiki/docs-convention.md` §2 同步 Gitea Wiki **文件**: 改 `pyproject.toml`(version)、`src/polygateway/__init__.py`(`__version__`)、`CHANGELOG.md`;按 `research-wiki/docs-convention.md` §2 同步 Gitea Wiki
@@ -258,6 +258,33 @@ class BreakerGate:
--- ---
### - [x] T6: 修复独立验证炸出的阻塞缺陷(计划外,2026-08-06)
T1–T5 全绿、全部门禁通过之后,全新上下文的 verifier 用一个**走 `QuotaGate` 的**端到端用例炸出:装配缺陷在唯一的生产路径上根本没有拆出去。
**缺陷**: `QuotaGate`/`BreakerGate``except GovernanceBackendError: raise` 只放行了旧类型,新增的 `SourceNotConfiguredError` 落进下一行 `except Exception` 被重新包成 `GovernanceBackendError`(`reason=governance_backend_down``retry_after_s=5.0`)。实证:
```
RAISED: GovernanceBackendError | isGatewayUnavailable=True | isSourceNotConfigured=False
| 限流后端故障(source_stats): 未知源 's1'(scope=llm)
```
即配置写错的任务照样落进"可延期重投"家族,**永远重投、永不进死信、无人告警**——正是 Q1 要防的镜像 bug,G2 等于没做。
**为什么原有测试测不出来**: T3 写的两条用例(`test_backpressure.py``test_redis_key_layout.py`)都直接打私有 `_cfg()`,绕过了包装器;而治理循环只经包装器访问后端。**盲区在于测试打的层次比生产路径低一层。**
**修复**(三处):
| 文件 | 改动 |
|---|---|
| `middleware/ratelimit.py` | 4 个方法的放行扩为 `except (GovernanceBackendError, SourceNotConfiguredError): raise` |
| `middleware/breaker.py` | 同上,5 个方法 |
| `middleware/telemetry.py:254` | 终态捕获元组加 `SourceNotConfiguredError`。**连带坑**: 放行生效后该异常不再是 `GovernanceBackendError`,而它在任何 attempt 之前抛出,若不显式捕获则 `emit_terminal_failure` 不触发、该路径**遥测归零**,违反"遥测必录"铁律 |
**回归测试**: `test_backpressure.py::TestUnknownSourceIsAssemblyDefect::test_survives_the_quota_gate_wrapper`(参数化覆盖 `try_acquire` / `stats`),**走包装器而非私有方法**。修前 2 failed,修后 PASS。
**同批文档订正**: 泄漏路径由"三条"改为**五条**(遗漏了 `QuotaGate.stats``BreakerGate.retry_after_s`,判据是该调用点是否被 `_record_quietly` 包裹);CHANGELOG 的 `per_source_reasons` 表述改为"属性存在但恒为 `{}`"。
## 完成后 ## 完成后
按 CLAUDE.md §3 Phase 2,合并前须派**全新上下文**的 verifier subagent 做独立验证(`verification-before-completion`),并按新规则**前台运行**。随后走 `finishing-a-development-branch` 决定合并方式,并在 Gitea 关闭 issue #7 按 CLAUDE.md §3 Phase 2,合并前须派**全新上下文**的 verifier subagent 做独立验证(`verification-before-completion`),并按新规则**前台运行**。随后走 `finishing-a-development-branch` 决定合并方式,并在 Gitea 关闭 issue #7
@@ -0,0 +1,276 @@
# 实施计划: stall 判定改为非生产性等待口径(Issue #8)
- **依据设计**: `research-wiki/designs/2026-08-06-issue8-stall-budget-design.md`(**已批准 2026-08-06**)
- **分支**: `feat/issue-8-stall-budget`(已建,已含设计提交 `bfe423d` + `ce2dda7`)
- **目标**: 让 stall 计时器只累计非生产性等待,解除 `timeout_s``stall_window_s` 的隐式耦合,使重试预算在超时场景下真实可用。
- **方案概述**: 新增调用级 `StallClock`(总时间减去 `_attempt` 耗时),替换三条治理循环里的墙钟 `entered_at`。判死双条件的结构、`inf` 语义、错误面、429 免预算全部不动。
- **涉及技术**: Python 3.11 asyncio、`contextlib.asynccontextmanager`、pytest + `FakeClock`
## 保真校验适用性
**适用**。三条治理循环均为 `reference/CHSAnalyzer app/providers/governance.py:200-285` 的移植物(ARCHITECTURE.md §1.4 关键资产)。但 **`reference/` 当前不在工作区**,无法逐段比对源码,故保真基准改为两处已入库的等价证据:
1. 设计文档 §4「旧版行为审计」表——9 条既有行为逐条标注保留/替换,实施时逐条核对;
2. 代码内既有的 CHS 行号注释(`retry.py:212`「调用级累计计时,循环内不重置(CHS governance.py:207)」、`:303-304`「双条件 stall 判死(CHS governance.py:270-281)」、`:315`「jitter 防惊群(CHS governance.py:283-285)」)与 `tests/unit/test_backpressure.py:1-6` 的蓝本 docstring。
**唯一允许的语义变更是条件 A 的度量口径**(设计 §4 中标"替换"的那一行)。其余任何条件分支、退避公式、jitter 区间、状态迁移若发生行为改变,即为违规,必须回退。
## 文件结构
| 文件 | 动作 | 职责 |
|---|---|---|
| `src/polygateway/middleware/retry.py` | 修改 | 新增模块级 `StallClock`(共享单元);主循环与 `_on_no_runnable` 改用之 |
| `src/polygateway/embedding.py` | 修改 | 复用 `StallClock`;`_on_no_runnable` 改签名 |
| `src/polygateway/ocr.py` | 修改 | 同上 |
| `src/polygateway/config.py` | 修改 | `_validate_stall` docstring 改写(仅注释,不改逻辑) |
| `tests/unit/test_backpressure.py` | 修改 | 新增 `TestStallBudget` 类;订正 `:121` docstring |
| `tests/unit/test_embedding.py` | 修改 | 新增 embedding 回归用例 |
| `tests/unit/test_ocr_client.py` | 修改 | 新增 ocr 回归用例 |
| `.env.example` | 修改 | 第 41 行注释改写 |
| `research-wiki/ARCHITECTURE.md` | 修改 | §7.3 背压条目补记新口径 |
| `CHANGELOG.md` | 修改 | 记治理行为变更 |
**不创建任何新模块**`StallClock` 放在 `retry.py`,沿用 `backoff_delay` 已被 embedding/ocr 复用的既有手法(依赖方向不变:`embedding.py:42``ocr.py:38` 已在 import 该模块)。
## 关键接口(跨任务消费,此处给出实际代码)
`StallClock` 由 T1 落地,T2/T3 直接消费,签名以此为准:
```python
class StallClock:
"""调用级 stall 计时器: 只累计非生产性等待(设计 §3.1)。
stall 预算治理的是"无人治理的等待"(429 退避、配额轮询、熔断冷却),
真实尝试已由重试预算 max_attempts 治理,故须从 stall 账里扣除——
两者重叠计费正是 issue #8 的根因。
每次调用创建一个实例。严禁提升为实例属性: 并发调用共享会互相污染计时。
"""
__slots__ = ("_now", "_entered_at", "_productive_s")
def __init__(self, now: Callable[[], float]) -> None:
self._now = now
self._entered_at = now()
self._productive_s = 0.0
def stalled_s(self) -> float:
"""非生产性等待累计秒数 = 总耗时 - 真实尝试耗时。"""
return self._now() - self._entered_at - self._productive_s
@contextlib.asynccontextmanager
async def attempting(self) -> AsyncIterator[None]:
"""包裹一次真实尝试, 其耗时记为生产性(边界即 _attempt 的边界)。"""
started = self._now()
try:
yield
finally:
# 只做算术, 不吞任何异常——CancelledError 逐字穿透(库铁律)
self._productive_s += self._now() - started
```
需在 `retry.py` 新增 `import contextlib`;`AsyncIterator``collections.abc` 引入(该文件已有 `from __future__ import annotations`,类型注解延迟求值,若 `TYPE_CHECKING` 块中已有 `Callable` 则复用)。
三条循环的改造模式一致:
```python
clock = StallClock(self._now) # 替换 entered_at = self._now()
...
await self._on_no_runnable(gate_rejections, reasons, clock) # 形参改类型
...
async with clock.attempting():
outcome = await self._attempt(...) # 原调用不变, 仅被包裹
```
判定式由 `self._now() - entered_at > stall` 改为 `clock.stalled_s() > stall`,**条件 B 与 `and` 结构逐字不动**。
---
## T1 — `StallClock` 落地与 chat 路径改造
- [ ] **文件**: `src/polygateway/middleware/retry.py`(修改)、`tests/unit/test_backpressure.py`(修改)
### 行为与验收标准
1. 按上文「关键接口」实现 `StallClock`,置于模块级(建议紧邻既有 `backoff_delay` 纯函数,便于 embedding/ocr 一并 import)。
2. `RetryMW.__call__`:`entered_at = self._now()`(`retry.py:212`)改为 `clock = StallClock(self._now)`;主循环判定(`:217`)改为 `clock.stalled_s() > stall`;`_attempt` 调用(`:228`)用 `async with clock.attempting():` 包裹。
3. `_on_no_runnable`(`:286-288`)形参 `entered_at: float` 改为 `clock: StallClock`,其内判定(`:306`)同步改为 `clock.stalled_s() > stall`
4. **不得改动**:429 免预算分支(`:233-234`)、`max(fails, 1)` 退避(`:243`)、jitter 公式(`:315`)、`fail_fast` 分支、条件 B `await self._quota.progress_age_s() > stall``AllSourcesExhausted` 的任何字段。
5. 保留 `retry.py:212` 的 CHS 行号注释并补记新口径(说明"调用级累计、循环内不重置"仍然成立,变的只是不再计入真实尝试)。
### 测试要求(先失败后通过)
`tests/unit/test_backpressure.py` 新增 `class TestStallBudget`:
| 用例 | 构造 | 断言 |
|---|---|---|
| `test_single_timeout_does_not_exhaust_stall_budget` | `_STALL=300`,源 `timeout_s` 等价;脚本 `[TransientError(耗时 350s), _ok()]`——用 `FakeTransport` 配合在尝试中推进 `FakeClock` 350s | 返回成功响应。**改前**:抛 `AllSourcesExhausted(reason="stalled")` |
| `test_productive_time_excluded_from_stall` | 连续两次尝试各推进时钟 `_STALL+100`,第三次成功 | 返回成功;全程不触发 `stalled` |
| `test_nonproductive_wait_still_triggers_stall` | 沿用 `_blocked_limiter`,轮询中推进时钟超窗且不 `mark_progress` | 抛 `stalled`(兜底未被削弱) |
| `test_saturation_429_still_stalls` | 源持续抛 429(`TransientError(status_code=429)`),退避 sleep 中推进时钟 | 抛 `stalled` 而非无限循环(`BoundedSleep` 上限内)。钉住设计 §3.5 |
| `test_cancel_inside_attempt_pierces` | 在 `_attempt` 内挂起后 `task.cancel()` | 抛 `CancelledError`(`attempting()` 的 finally 不吞) |
| `test_concurrent_calls_do_not_share_clock` | 两路并发调用,一路长尝试、一路正常 | 两路互不影响;钉住 `StallClock` 不得为实例属性 |
| `test_telemetry_time_counts_as_productive` | 注入一个在 `emit_attempt` 中推进 `FakeClock` 超过 `_STALL` 的慢 emitter,transport 正常成功 | 返回成功响应,不触发 `stalled`。**钉住设计 §3.1 的边界声明**:遥测收尾属生产性,遥测抖动不得参与判死。若将来有人把 `attempting()` 的包裹范围收窄到只包 transport 调用,该不变式会被悄悄破坏而其余用例抓不到 |
**既有四象限用例(`TestStallQuadrants` 四条)必须原样通过,不得修改断言**——它们全程无真实尝试或真实尝试耗时为 0,`stalled_s()` 与旧墙钟等价。若其中任何一条需要改断言才能通过,说明实现越界,停下来复核。
同时订正 `tests/unit/test_backpressure.py:121` 的 docstring:「仅全局超窗(从未出餐 age=inf)」保持不变(该语义确实不变),但补一句说明本地口径已是非生产性等待。
### 验证命令
```bash
conda run -n PolyGateway pytest tests/unit/test_backpressure.py -v
conda run -n PolyGateway pytest tests/unit/test_retry.py -v
```
预期:全部 PASS。先在实现前跑新增用例,记录 `test_single_timeout_does_not_exhaust_stall_budget` 的 FAILED 输出作为红证据。
- [ ] **提交点**: `fix: bill only non-productive waiting against the chat stall budget`
---
## T2 — embedding 路径改造
- [ ] **文件**: `src/polygateway/embedding.py`(修改)、`tests/unit/test_embedding.py`(修改)
### 行为与验收标准
1. `embedding.py:42` 的 import 增加 `StallClock`(该行已 import `_failure_reason, backoff_delay`)。
2. `_embed_batch`(`:182`):`entered_at = self._now()` 改为 `clock = StallClock(self._now)`;`_attempt` 调用(`:188`)用 `async with clock.attempting():` 包裹;`_on_no_runnable` 传参(`:186`)改为 `clock`
3. `_on_no_runnable`(`:230-232`)形参改 `clock: StallClock`,判定(`:248`)改 `clock.stalled_s() > stall`
4. **不得新增主循环 stall 判定**(设计 §5.4:embedding 无 429 免预算,`fails += 1` 无条件,缺口不存在;新增等于凭空多一条判死路径)。
5. **不得改动**:`fails += 1` 的无条件性(`:191`)、`max_attempts` 判定、退避调用(`:200`)。
### 测试要求(先失败后通过)
`tests/unit/test_embedding.py` 新增 `test_single_timeout_does_not_exhaust_stall_budget`
**构造方式(已核实可行,不必绕过既有 helper)**:`_embed_client`(`:219`)的 `**overrides` 直通 `EmbeddingClient.__init__`,而后者接受 `now`/`sleep`/`rng`(`embedding.py:109-111`),故可写 `_embed_client([src], script, now=clock, sleep=<推进时钟的 fake>)`。制造一轮 `_on_no_runnable` 沿用 `test_backpressure.py:75-85` `_blocked_limiter` 的手法:源 `max_concurrency=1`,测试先 `try_acquire` 占满 permit,在 fake sleep 回调里释放。helper 内的 `InMemoryLimiter` 未注入 `now` 不影响本用例——判定要的是 `progress_age_s()` 返回 `inf`(从未 `mark_progress`),与 limiter 时钟无关。
**断言**:第一次尝试推进 `FakeClock` 超过 `stall_window_s` 后抛 `TransientError`,随后经一轮 `_on_no_runnable` 再恢复,最终返回成功的 `EmbeddingResponse`。改前应抛 `AllSourcesExhausted(reason="stalled")`
**取消穿透验收点**:既有 `test_cancel_releases_permit`(`test_embedding.py:331-339`)的取消路径**将被新的 `async with clock.attempting()` 包住**,故它是本任务的必过回归项,不得因改动而修改其断言。若它转红,说明 `attempting()``finally` 吞了 `CancelledError` 或泄漏了 permit,停下来复核而非改测试。
### 验证命令
```bash
conda run -n PolyGateway pytest tests/unit/test_embedding.py -v
```
预期:全部 PASS(含既有取消与遥测用例)。
- [ ] **提交点**: `fix: apply the non-productive stall budget to the embedding loop`
---
## T3 — ocr 路径改造
- [ ] **文件**: `src/polygateway/ocr.py`(修改)、`tests/unit/test_ocr_client.py`(修改)
### 行为与验收标准
与 T2 同构,对应行号:import(`:38`)、`_call``entered_at`(`:207`)、`_on_no_runnable` 传参(`:211`)、`_attempt` 调用(`:213`)、`_on_no_runnable` 签名(`:255-257`)与判定(`:273`)。同样**不得新增主循环 stall 判定**,不得改动 `fails += 1`(`:216`)与退避(`:225`)。
### 测试要求(先失败后通过)
`tests/unit/test_ocr_client.py` 新增与 T2 同构的 `test_single_timeout_does_not_exhaust_stall_budget`,覆盖 `recognize_text``parse_layout` 任一端点即可(两者共用 `_call`)。
**构造方式**:同 T2——经该文件既有的 `_client(...)` helper 传 `now=clock` 与推进时钟的 fake `sleep`;`_on_no_runnable` 一轮用"源 `max_concurrency=1` + 测试预先占满 permit + 在 fake sleep 回调里释放"制造。
**取消穿透验收点**:既有 `test_cancel_during_transport_releases_permit`(`test_ocr_client.py:339-348`)与其上方的退避期取消用例同样会被新包裹覆盖,均为必过回归项,不得修改断言。
### 验证命令
```bash
conda run -n PolyGateway pytest tests/unit/test_ocr_client.py tests/unit/test_monkey_ocr.py -v
```
预期:全部 PASS。
- [ ] **提交点**: `fix: apply the non-productive stall budget to the ocr loop`
---
## T4 — 配置侧注释对齐(无逻辑变更)
- [ ] **文件**: `src/polygateway/config.py`(修改)、`.env.example`(修改)
### 行为与验收标准
1. `config.py:240-241` `_validate_stall` 的 docstring 由「stall 窗口须 ≥ 最慢源 TTFT 上限,防把正常慢首包误判为卡死」改写为:说明该校验在新口径下**属保守冗余**——TTFT 等待是生产性时间,已不计入 stall;保留校验是为不改动 ARCHITECTURE.md §7.3 契约 G6(人类 2026-08-06 定夺)。**校验逻辑本身一字不改**。
2. `.env.example:41` 注释由「stall 双条件判死窗口;须 ≥ 最大源 TTFT」改写为说明它度量的是**非生产性等待**(429 退避/配额轮询/熔断冷却)累计,与 `TIMEOUT_S` 无耦合、无需按 `timeout × retries` 放大。
3. 不新增、不改名任何配置键(`_DEFAULT_STALL_WINDOW_S = 300.0` 保持不变)。
### 验证命令
```bash
conda run -n PolyGateway pytest tests/unit/test_config.py -v
conda run -n PolyGateway make lint
```
预期:全部 PASS(本任务不改逻辑,`test_config.py` 应零变化通过)。
- [ ] **提交点**: `docs: align the stall window comments with the new metering`
---
## T5 — 全套件回归与文档同步
- [ ] **文件**: `research-wiki/ARCHITECTURE.md`(修改)、`CHANGELOG.md`(修改)
### 行为与验收标准
1. 跑全套件确认零回归。**命令末尾不得接管道**(CLAUDE.md 执行模式:管道会掩盖真实退出码),需要后台跑时用 `wait`/轮询 PID 判完成。
2. `ARCHITECTURE.md` §7.3 背压条目(第 429-431 行区域)补记:stall 双条件的条件 A 现为**非生产性等待累计**,并给出本设计文档指针。既有 G6 契约行保留,补注其在新口径下为保守冗余。
3. `CHANGELOG.md` 记治理行为变更(属公共行为变更,须显式列出:单次调用最坏耗时由 `stall_window_s` 抬升至 `max_attempts × timeout_s`)。
4. 核对设计 §4 行为审计表 9 条,逐条确认实现与标注一致(保真校验检查点)。
### 副作用处置
修复后单次调用最坏耗时变为 `max_attempts × timeout_s`(本机 900s)。跑 e2e 前先评估 `tests/e2e` 的源 `timeout_s` 是否需调小,以免冒烟耗时失控。本机 `.env:37` 的临时缓解 `STALL_WINDOW_S=1200` 可回退默认值(`.env` 不入库,仅在本任务记录该动作)。
### 验证命令
```bash
conda run -n PolyGateway make lint
conda run -n PolyGateway pytest tests/unit tests/integration -v
conda run -n PolyGateway make test
```
预期:lint 通过(含 import-linter 依赖契约);单元与集成全绿;覆盖率不低于既有水平。
- [ ] **提交点**: `docs: record the stall metering change in architecture and changelog`
---
## T6 — 独立验证与 Wiki 同步
- [ ] **文件**: Gitea Wiki(独立仓库)
### 行为与验收标准
1. **派全新上下文 verifier subagent**(`verification-before-completion`,MANDATORY:跨 3 模块属里程碑级),**前台运行**(`run_in_background: false`,CLAUDE.md 执行模式)。核验对象:设计 §1 的 G1-G4 是否逐条兑现、§4 行为审计表 9 条是否与实现一致、是否出现设计未声明的语义变更、测试是否真的覆盖"先失败后通过"。
2.`docs-convention.md` §2「治理行为变更」行同步 Wiki:`解释-治理行为`(stall 判定口径)、`指南-限流与熔断`(配置说明中删除"须按 timeout×retries 放大 stall"一类误导)。
3. 在 Gitea issue #8 下回帖:根因、方案、被否决的两个原建议方向及理由、影响面。
4. Wiki 注册:
```bash
.claude/tools/research_wiki.py add_entity research-wiki/ --type plan --id issue8-stall-budget --title "stall 判定改为非生产性等待口径"
.claude/tools/research_wiki.py add_edge research-wiki/ --from "plan:issue8-stall-budget" --to "design:issue8-stall-budget" --type implements --evidence "本计划实施该设计的 T1-T6"
.claude/tools/research_wiki.py rebuild_index research-wiki/
```
### 验证命令
```bash
conda run -n PolyGateway make ci
```
预期:只读验证全绿。verifier 报告须逐条对应本会话内的工具输出(证据化声明,禁止虚报)。
- [ ] **提交点**: `chore: register the issue #8 plan and sync the wiki`
---
## 任务依赖
T1 → (T2 ‖ T3) → T4 → T5 → T6。T2 与 T3 相互独立,但都依赖 T1 落地的 `StallClock`。
@@ -0,0 +1,259 @@
# 实现计划: HTTP 错误响应体留存(Issue #10)
- **设计**: `research-wiki/designs/2026-08-16-issue10-error-body-retention-design.md`(**已批准 2026-08-16**)
- **分支**: `feat/issue-10-error-body-retention`
- **目标**: 网关拒绝一次调用时,它说的话必须能在库自己的遥测表里被事后查到。
- **方案概述**: transport 翻译层把 HTTP 错误响应体折叠空白并按头尾策略摘要,**同一份串**同时拼进异常 message(经既有 `error` 列落遥测)与新增的基类字段 `body_text`(供下游结构化留存)。覆盖两个 transport 的全部非 2xx 分支。不改任何状态码→分类的映射。
- **涉及技术**: Python 3.11 / httpx / pytest。无新增依赖。
- **保真校验**: 本计划**不涉及** `reference/` 参考实现迁移——错误分类映射逐条不变,保真体现为"既有分类断言全部保留、无一条被改写"(Task 3/4 验收项)。
## 文件结构
| 文件 | 动作 | 职责 |
|---|---|---|
| `src/polygateway/errors.py` | 改 | 基类 `PolyGatewayError` 新增 `body_text` 字段;与 `raw_text` 的界限 docstring;`RequestRejectedError` 补中转拓扑提醒 |
| `src/polygateway/transports/_http_errors.py` | **新建** | 摘要口径单一实现:`summarize_body` / `compose_message` / `response_body` + 三个常量 |
| `src/polygateway/transports/openai_compat.py` | 改 | `_status_to_error` 表驱动重写;`_translate_429``ctx` |
| `src/polygateway/transports/monkey_ocr.py` | 改 | `_classify_status` 带摘要 |
| `tests/unit/test_http_error_body.py` | **新建** | 摘要单元的纯函数用例(截断边界、头尾保留、折叠、幂等) |
| `tests/unit/test_openai_compat.py` | 改 | 状态码参数化断言 message + 字段;超长 `insufficient_quota` 回归 |
| `tests/unit/test_monkey_ocr.py` | 改 | OCR 分支同款 + `ResponseNotRead` 降级 |
| `tests/unit/test_errors.py` | 改 | `body_text` 默认值与可传性 |
| `tests/integration/test_governance_stack.py` | 改 | **端到端验收**:400 调用后 SQLite `error` 列含摘要 |
| `README.md` / `CHANGELOG.md` / `pyproject.toml` / `src/polygateway/__init__.py` / `research-wiki/ARCHITECTURE.md` | 改 | 文档与 1.2.0 版本号 |
## 关键接口(跨任务消费,必须逐字一致)
```python
# src/polygateway/transports/_http_errors.py
from __future__ import annotations
import httpx # response_body 的类型与 ResponseNotRead 都来自它
_ERROR_BODY_CAP = 2048 # 字符(非字节),含省略标记在内的最终总长上限
_HEAD_CHARS = 1400
_TAIL_CHARS = 600
def summarize_body(text: str) -> str:
"""折叠空白后按头尾策略摘要;空/空白入参返回空串。"""
collapsed = " ".join(text.split())
if len(collapsed) <= _ERROR_BODY_CAP:
return collapsed
omitted = len(collapsed) - _HEAD_CHARS - _TAIL_CHARS
return f"{collapsed[:_HEAD_CHARS]}…(略 {omitted} 字)…{collapsed[-_TAIL_CHARS:]}"
def compose_message(message: str, summary: str) -> str:
"""摘要非空才拼后缀,避免悬空分隔符。"""
return f"{message} | {summary}" if summary else message
def response_body(response: httpx.Response) -> str:
"""取已缓冲的响应文本;未读缓冲一律降级空串,绝不触发网络读。"""
try:
return response.text
except httpx.ResponseNotRead:
return ""
```
```python
# src/polygateway/errors.py
class PolyGatewayError(Exception):
def __init__(
self,
message: str,
*,
source_name: str | None = None,
status_code: int | None = None,
operation: str | None = None,
body_text: str = "",
) -> None:
```
```python
# src/polygateway/transports/openai_compat.py
# 既有 errors 导入(:18-23)须补入 PolyGatewayError —— 当前只导了四个子类,
# 直接写 _classify 的返回注解会让 ruff 报 F821 未定义名。
from polygateway.errors import (
PolyGatewayError, # ← 新增
RequestRejectedError,
ResultInvalidError,
SourceDeadError,
TransientError,
)
from polygateway.transports._http_errors import compose_message, summarize_body
def _classify(status: int) -> tuple[type[PolyGatewayError], str]:
"""状态码 → (错误类, message 标签);映射与 1.1.2 逐条相同。"""
if status in (401, 403):
return SourceDeadError, "凭据失效/欠费"
if status == 400:
return RequestRejectedError, "请求被拒"
if status >= 500:
return TransientError, "瞬时错误"
return RequestRejectedError, "客户端错误"
def _status_to_error(
source: SourceConfig, status: int, body_text: str, headers: Mapping[str, str]
) -> Exception:
summary = summarize_body(body_text) # 全函数只算一次
ctx: dict[str, Any] = {
"source_name": source.name,
"status_code": status,
"operation": "chat",
"body_text": summary,
}
if status == 429:
return _translate_429(source, body_text, headers, ctx) # 传**原文**,见下
cls, label = _classify(status)
return cls(compose_message(f"{source.name} {label}: {status}", summary), **ctx)
```
> **实现红线**:`_translate_429` 判 `insufficient_quota` 必须解析**未截断的原文** `body_text`,不得改用 `summary`。摘要会破坏 JSON 结构,超长体一旦改用摘要解析,配额耗尽的源将不再 `force_open`——那是把一个诊断改进变成治理 bug。Task 3 有专门的回归用例钉死这条。
## 任务清单
### - [ ] Task 1: 内核新增 `body_text` 字段
**改**: `src/polygateway/errors.py`
- `PolyGatewayError.__init__` 按上文签名新增 `body_text: str = ""`,存为实例属性。
- 类 docstring 增补与 `ResultInvalidError.raw_text` 的界限:`body_text` = 非 2xx 的 HTTP 错误响应体摘要(对方拒绝的理由);`raw_text` = 2xx 但内容不可解析时的模型输出。并写明"可能包含请求回显,已截断"。
- **不动** `TransientError` / `SourceDeadError` / `RequestRejectedError` / `ResultInvalidError` / `GatewayUnavailableError` 的任何既有签名与行为。
**测试**(`tests/unit/test_errors.py`,扩展 `:29` 的四类构造形态参数化):
- 四个 transport 错误类默认 `body_text == ""`;显式传入后可读回。
- `GatewayUnavailableError` / `CircuitOpenError` / `AllSourcesExhausted` / `GovernanceBackendError``body_text` 恒为 `""`(它们不经 HTTP 响应翻译)。
**验收**: 新增字段不改变任何既有异常的 `str()` 输出。
**验证**: `conda run -n PolyGateway pytest tests/unit/test_errors.py -v` → 全 PASS。
### - [ ] Task 2: 共享摘要单元
**新建**: `src/polygateway/transports/_http_errors.py`(按上文"关键接口"逐字实现,**含其中的 `import httpx`**,加中文模块/函数 docstring 解释**为什么**折叠空白、为什么头尾保留、为什么 `response_body` 必须降级)
**新建测试**: `tests/unit/test_http_error_body.py`
| 用例 | 断言 |
|---|---|
| 短体原样 | `summarize_body('{"a":1}') == '{"a":1}'` |
| 空白折叠 | 多行缩进 JSON → 单行,无连续空格 |
| 空 / 纯空白入参 | 返回 `""` |
| 长度恰 2048 | 原样返回,无标记 |
| 长度 2049 | 走头尾策略 |
| 超长体头尾 | 前 1400 字符 == 原文前 1400;**末 600 字符 == 原文末 600**;中段标记内 N == `len(原文) - 2000` |
| **尾部关键字段可见**(设计 §7 用例 3c) | 以 issue 真实样本尾部 `"code":"invalid_parameter_error"}}` 收尾构造超长体 → 断言该串出现在摘要中 |
| 幂等 | `summarize_body(summarize_body(x)) == summarize_body(x)`(标记不嵌套) |
| `compose_message` | 摘要为空时返回原 message 不变;非空时以竖线分隔符拼接 |
| `response_body` 降级 | `httpx.Response(400, stream=<未读 SyncByteStream>)` → 返回 `""` 且不抛(构造法见下) |
未读响应的构造(已实测可用):
```python
class _Unread(httpx.SyncByteStream):
def __iter__(self):
yield b"body"
resp = httpx.Response(400, stream=_Unread()) # 未 read → .text 抛 ResponseNotRead
```
**验收**: 摘要总长恒 ≤ `2000 + len(标记)`;头尾各自与原文逐字对应。
**验证**: `conda run -n PolyGateway pytest tests/unit/test_http_error_body.py -v` → 全 PASS。
### - [ ] Task 3: openai_compat 翻译层收口
**改**: `src/polygateway/transports/openai_compat.py`
- 新增 `_classify`,`_status_to_error` 按上文骨架重写(五分支各拼各的 message → 查表 + 单点拼装)。
- `_translate_429` 签名改为 `(source, body_text, headers, ctx)`,两支 message 各自追加 `compose_message` 后缀,构造改用 `**ctx`;**`json.loads` 仍读原文 `body_text`**。
- message 主体逐字保持 1.1.2 原样(`凭据失效/欠费: {status}` / `请求被拒: 400` / `瞬时错误: {status}` / `客户端错误: {status}` / `配额耗尽(insufficient_quota)` / `限速: 429`),只在末尾追加 ` | {摘要}`
- 三个调用点(`:402` embed、`:417` stream、`:509` 非流式)签名不变,**不改动**。
- **补 import**:`PolyGatewayError`(errors)与 `compose_message` / `summarize_body`(`._http_errors`),见上文关键接口——漏补则 `make lint` 报 F821(Codex 审查 2026-08-16 提出)。
- **不改** `operation` 硬编码 `"chat"`(设计 §5.4 有意留给独立 issue)。
**测试**(`tests/unit/test_openai_compat.py`,沿用既有 `_transport_for(handler)` + `httpx.MockTransport`):
| # | 用例 | 断言 |
|---|---|---|
| 3.1 | 状态码参数化 400 / 401 / 403 / 404 / 500 / 503,handler 返回带真实样本体 | 异常类型与 1.1.2 **逐条相同**;message 含摘要;`exc.body_text` == 摘要 |
| 3.2 | 429 普通限速(body 无 `insufficient_quota`) | `TransientError`,message 含摘要,`retry_after_s` 解析不受影响 |
| 3.3 | 429 + `insufficient_quota` | `SourceDeadError`,message 含摘要 |
| 3.4 | **回归红线**:429 + `insufficient_quota` 且 body 长度 > 2048(前置大量填充字段) | 仍判 `SourceDeadError`——证明类型判定读的是原文而非摘要 |
| 3.5 | 空 body 的 400 | message 无悬空分隔符,`body_text == ""` |
| 3.6 | 非 JSON body、非 UTF-8 字节 body | 不抛额外异常,分类不变 |
| 3.7 | 流式路径(handler 对 stream 请求返回 400 + body) | 经 `_complete_stream:415-417` 抛出的异常同样带摘要 |
| 3.8 | embedding 路径(`transport.embed(...)` 遇 400) | 同样带摘要 |
**验收**: 既有测试零修改通过(除 3.x 新增外);`test_openai_compat.py:558`(match 源名)仍 PASS。
**验证**: `conda run -n PolyGateway pytest tests/unit/test_openai_compat.py -v` → 全 PASS。
### - [ ] Task 4: monkey_ocr 同款收口
**改**: `src/polygateway/transports/monkey_ocr.py`
- `_classify_status`:`summary = summarize_body(response_body(exc.response))`,三支 message 统一经 `compose_message` 追加后缀,`ctx``body_text=summary`
- message 主体保持 `f"{source_name} OCR {operation} HTTP {status}"` 不变。
- 429/5xx → `TransientError`、401/403 → `SourceDeadError`、其余 → `RequestRejectedError` 的映射**逐条不变**(OCR 无 429 细分是设计有意保留,见模块 docstring `:53-54`)。
**测试**(`tests/unit/test_monkey_ocr.py`):
- 扩展 `:300` 的状态码参数化:各分支 message 含摘要且 `body_text` 非空,分类不变。
- `ResponseNotRead` 降级:`exc.response` 为未读流 → `body_text == ""`,message 无悬空分隔符,**分类仍正确**(不得因取 body 失败而改变错误类型或抛出 httpx 异常)。
**验收**: `:195``:300` 既有断言不被改写。
**验证**: `conda run -n PolyGateway pytest tests/unit/test_monkey_ocr.py -v` → 全 PASS。
### - [ ] Task 5: 端到端遥测验收(**本计划的硬判据**)
**改**: `tests/integration/test_governance_stack.py`
新增用例,沿用既有 `_full_client(handler, telemetry=SQLiteRecorder(...))``:135``SELECT error FROM llm_calls` 断言模式:
- handler 对 chat 请求返回 `httpx.Response(400, content=<issue #10 真实样本体>)`
- `client.chat(...)``RequestRejectedError`(400 不重试不换源,行为不变)。
- `recorder.close()` 后查 `SELECT error FROM llm_calls`:该行 `error` 串**含样本体里的 `InvalidParameter` 与结尾的 `invalid_parameter_error`**。
真实样本体(取自 issue #10 原文,一字不改):
```json
{"error":{"message":"<400> ***.***.InvalidParameter: The image format is illegal and cannot be opened","type":"invalid_request_error","param":"","code":"invalid_parameter_error"}}
```
**验收**: 这条断言在 Task 1-4 之前**必然失败**(1.1.2 的 `error` 列只有 `"qwen_1 请求被拒: 400"`),之后通过——这就是本 issue 的"先失败后通过"证据主体,执行时须保留失败输出截图/文本进提交说明。
**验证**: `conda run -n PolyGateway pytest tests/integration/test_governance_stack.py -v` → 全 PASS。
### - [ ] Task 6: 文档与版本
**改**:
| 文件 | 内容 |
|---|---|
| `src/polygateway/errors.py` | `RequestRejectedError` docstring 加一句:经中转部署时 400 可能源于中转自身抖动,批处理场景下游宜自备兜底分类(设计 §5.2) |
| `research-wiki/ARCHITECTURE.md` §6.2 | 同一提醒 + 注明四分类错误自 1.2.0 起携带 `body_text` |
| `CHANGELOG.md` | 新增 `## 1.2.0(2026-08-16)` 段:行为变更(message 追加摘要 → 遥测 `error` 列变长)、新增字段、不变项(分类映射零变更、错误面零变更) |
| `README.md:34` | 安装 pin `==1.1.*`**`>=1.2,<2`**(2026-08-16 人类定夺;漏改则下游静默停在 1.1.2) |
| `pyproject.toml` + `src/polygateway/__init__.py` | 版本号 `1.1.2``1.2.0`,**两处必须一致** |
**验收**: `grep -rn "1\.1\.\*" README.md` 零命中;两处版本号一致。
**验证**: `conda run -n PolyGateway python -c "import polygateway; print(polygateway.__version__)"``1.2.0`
### - [ ] Task 7: 合并前全量门
1. `make lint`(ruff + import-linter)→ 零违规,**重点确认新建 `transports/_http_errors.py` 未触发洋葱分层契约**。
2. `make test` 全套件 → 全 PASS,覆盖率不低于既有水平。
3. 派**全新上下文** verifier subagent 独立验证(CLAUDE.md §3 Phase 2 硬门):逐条核对 Task 1-6 验收项与本会话工具输出。
4. `finishing-a-development-branch` 合并回 main(`--no-ff`),合并后在 main 上重跑 `make lint` 与全套件。
**发布**(合并后)严格按 CLAUDE.md §4.4.1 九步执行,不在本计划展开;其中步骤 1(更新 README)已在 Task 6 前置完成,**构建前须再次确认 pin 已是 `>=1.2,<2`**。
## 执行顺序与提交点
```
Task 1 ──┐
├── Task 3 ──┐
Task 2 ──┴── Task 4 ──┴── Task 5 ── Task 6 ── Task 7
```
Task 1 与 2 可并行(互不依赖);Task 3、4 都依赖 1+2;Task 5 依赖 3;Task 6 独立于代码但须在 Task 7 之前。每个 Task 一次语义化提交(`commit` skill),Task 5 的提交说明须附"修复前失败、修复后通过"的实际输出。
@@ -34,4 +34,12 @@ date: 2026-08-06
Codex 同时独立核实了计划的可执行性锚点: 22 处构造点、三处 gate 装配、后端层 `self._scope` 位置、README/ARCH 章节行号,均与 `src/` 现状相符。 Codex 同时独立核实了计划的可执行性锚点: 22 处构造点、三处 gate 装配、后端层 `self._scope` 位置、README/ARCH 章节行号,均与 `src/` 现状相符。
## 独立验证炸出的阻塞缺陷(2026-08-06,全新上下文 verifier)
T1–T5 全绿、四道门禁全过之后,verifier 用一个**走 `QuotaGate` 的**端到端用例证明: 装配缺陷在唯一的生产路径上根本没拆出去——包装器的 `except GovernanceBackendError: raise` 只放行旧类型,`SourceNotConfiguredError` 落进下一行 `except Exception` 被重新包回去,配置写错照样永远重投。**盲区在于 T3 写的两条用例都直接打私有 `_cfg()`,比生产路径低一层。**
修复见正文 §T6(9 处放行 + 遥测终态捕获 + 走包装器的回归测试)。复核时 verifier 又指出一颗雷: 新放行让该异常能穿透 `_record_quietly`,而那层降级的存在理由是"调用已真实完成,写回失败不该丢弃成功响应"——同批把三处 `_record_quietly` 一并放宽并加了回归断言。
两轮都订正了同一处事实错误: 闸门泄漏路径是**五条**不是三条(`QuotaGate.stats``BreakerGate.retry_after_s` 同样未被 `_record_quietly` 包裹)。
相关: [[governance-backend-error]](design)、[[m2-distributed]] 相关: [[governance-backend-error]](design)、[[m2-distributed]]
@@ -0,0 +1,37 @@
---
type: plan
node_id: plan:issue10-error-body-retention-plan
title: "实现计划: HTTP 错误响应体留存(Issue #10)"
date: 2026-08-16
---
# 实现计划: HTTP 错误响应体留存(Issue #10)
**全文**: `plans/2026-08-16-issue10-error-body-retention.md`|**实现**: [[design:issue10-error-body-retention]]|**分支**: `feat/issue-10-error-body-retention`
## 任务分解
| # | 任务 | 产出 |
|---|---|---|
| 1 | 内核字段 | `PolyGatewayError.body_text`(默认空串)+ 与 `raw_text` 的界限 docstring |
| 2 | 共享摘要单元 | 新建 `transports/_http_errors.py`:`summarize_body` / `compose_message` / `response_body` |
| 3 | openai_compat 收口 | `_status_to_error` 表驱动;429 判类型仍读原文 |
| 4 | monkey_ocr 收口 | `_classify_status` 同款,含 `ResponseNotRead` 降级 |
| 5 | **端到端验收** | 400 调用后 SQLite `error` 列含摘要——本计划的硬判据 |
| 6 | 文档与版本 | 1.2.0、README pin `>=1.2,<2`、CHANGELOG、ARCHITECTURE §6.2 |
| 7 | 合并前门 | lint + 全套件 + 全新上下文 verifier |
依赖:1‖2 → 3‖4 → 5 → 6 → 7。
## 计划期钉死的两条实现红线
1. **429 判 `insufficient_quota` 必须解析未截断原文**,不得改用摘要——摘要会破坏 JSON,超长体一旦改用摘要解析,配额耗尽的源将不再 `force_open`,把诊断改进变成治理 bug。Task 3.4 有专门回归用例。
2. **message 主体逐字保持 1.1.2 原样**,只在末尾追加摘要后缀;状态码→分类映射逐条不变,验收要求既有分类断言零改写。
## 保真校验
不涉及 `reference/` 迁移。错误分类映射不变,保真体现为"既有分类断言全部保留"。
## 执行期观察
Task 计划提交时 pre-commit 钩子的全套件跑出现一次 `tests/e2e/test_compat_projects.py::TestGovDocOnboarding::test_call_site_shape_runs_governed` 失败,单独跑与 e2e 全目录跑(7 passed / 23.30s,每例 2-3.5s)均通过,重跑全套件亦通过 → 判为真实网关抖动,非回归。该用例正是 [[design:issue8-stall-budget]] 当年记录的三个漂移用例之一,e2e 打真实网关的固有 flaky 面仍在。
@@ -0,0 +1,35 @@
---
type: plan
node_id: plan:issue8-stall-budget-plan
title: "issue #8 实施计划: stall 非生产性等待口径"
date: 2026-08-06
---
# issue #8 实施计划: stall 非生产性等待口径
**全文**: `plans/2026-08-06-issue8-stall-budget.md` |**实现**: [[design:issue8-stall-budget]] |**分支**: `feat/issue-8-stall-budget`
## 交付
| 任务 | 内容 | 提交 |
|---|---|---|
| T1 | `StallClock` 落地 + chat 路径改造 + 8 条测试 | `02c3d06` |
| T2 | embedding 路径复用 | `6d0f3c9` |
| T3 | ocr 路径复用 | `0477d95` |
| T4 | `config.py` docstring 与 `.env.example` 注释对齐(无逻辑变更) | `d05114e` |
| T5 | 全套件回归 + `ARCHITECTURE.md` §7.3 与 `CHANGELOG` 同步 | `3645e57` |
| T6 | 独立验证(全新上下文 verifier)+ 三个问题的修复 | `bc4683d``a0a5cf7` |
## 测试证据(先失败后通过)
三条路径的失效链条各有一条回归用例,改前均转红于 `reason="stalled"`:`retry.py:218``embedding.py:250``ocr.py:275`。issue 只记录了 chat 路径,embedding/ocr 两条为本次核出。
## 独立验证发现的三个问题(均已修)
1. **429 缝隙(中)**: 初稿使 429 尝试两个预算都不烧,慢 429 场景实测挂 25.2 小时——**修复引入的回归**。见设计 §3.6。
2. **测试假证据(中)**: 并发用例用了两个 `RetryMW` 实例,实例级共享被对象隔离掩盖,clock 提升为实例属性时 7 条用例全部逃逸。改为复用同一 `mw` 并补"两次调用间空转超窗"用例,变异测试确认可抓。
3. **文档遗漏(轻)**: 计划要求的 `test_backpressure.py` docstring 订正漏做。
## 保真校验
治理主循环为 CHS `governance.py:200-285` 移植物,但 `reference/` 不在工作区,故以设计 §4 行为审计表 9 条 + 代码内 CHS 行号注释为基准。核对结果:标"保留"的 8 条在 `git diff` 中零出现,唯一"替换"项为条件 A 度量口径。
+1 -1
View File
@@ -32,7 +32,7 @@ from polygateway.types import (
SourceConfig, SourceConfig,
) )
__version__ = "1.1.0" __version__ = "1.2.0"
__all__ = [ __all__ = [
"DEFAULT_PROFILES", "DEFAULT_PROFILES",
+8 -1
View File
@@ -238,7 +238,14 @@ class GatewaySettings:
) )
def _validate_stall(self) -> None: def _validate_stall(self) -> None:
"""stall 窗口须 ≥ 最慢源 TTFT 上限,防把正常慢首包误判为卡死。""" """stall 窗口须 ≥ 最慢源 TTFT 上限(保守冗余,见下)。
原理由是"防把正常慢首包误判为卡死"。issue #8 起 stall 只累计**非
生产性等待**(429 退避、配额轮询、熔断冷却),TTFT 等待属生产性时间、
已不计入 stall 账,该误判在机制上不再可能。校验本身无害且不会误拒
任何合理配置,故保留——删除它需同步改动 ARCHITECTURE.md §7.3 的契约
补强 G6,超出 issue #8 的范围(2026-08-06 人类定夺)。
"""
ttfts = [s.ttft_timeout_s for s in self.sources if s.ttft_timeout_s is not None] ttfts = [s.ttft_timeout_s for s in self.sources if s.ttft_timeout_s is not None]
if ttfts and self.backpressure.stall_window_s < max(ttfts): if ttfts and self.backpressure.stall_window_s < max(ttfts):
raise ValueError( raise ValueError(
+9 -6
View File
@@ -34,11 +34,12 @@ from polygateway.errors import (
RequestRejectedError, RequestRejectedError,
ResultInvalidError, ResultInvalidError,
SourceDeadError, SourceDeadError,
SourceNotConfiguredError,
TransientError, TransientError,
) )
from polygateway.middleware.breaker import BreakerGate from polygateway.middleware.breaker import BreakerGate
from polygateway.middleware.ratelimit import QuotaGate from polygateway.middleware.ratelimit import QuotaGate
from polygateway.middleware.retry import _failure_reason, backoff_delay from polygateway.middleware.retry import StallClock, _failure_reason, backoff_delay
from polygateway.middleware.telemetry import TelemetryEmitter from polygateway.middleware.telemetry import TelemetryEmitter
from polygateway.sources import SourceCooldownMemo from polygateway.sources import SourceCooldownMemo
from polygateway.types import ( from polygateway.types import (
@@ -178,12 +179,14 @@ class EmbeddingClient:
) -> _BatchOutcome: ) -> _BatchOutcome:
fails = 0 fails = 0
reasons: dict[str, str] = {} reasons: dict[str, str] = {}
entered_at = self._now() # 只计非生产性等待(issue #8): 真实尝试由重试预算治理,不重复烧 stall 预算
clock = StallClock(self._now)
while True: while True:
picked, gate_rejections = await self._pick_runnable(reasons) picked, gate_rejections = await self._pick_runnable(reasons)
if picked is None: if picked is None:
await self._on_no_runnable(gate_rejections, reasons, entered_at) await self._on_no_runnable(gate_rejections, reasons, clock)
continue continue
async with clock.attempting():
outcome = await self._attempt(batch, *picked, reasons, session_id, parent_call_id) outcome = await self._attempt(batch, *picked, reasons, session_id, parent_call_id)
if isinstance(outcome, _BatchOutcome): if isinstance(outcome, _BatchOutcome):
return outcome return outcome
@@ -227,7 +230,7 @@ class EmbeddingClient:
return None, gate_rejections return None, gate_rejections
async def _on_no_runnable( async def _on_no_runnable(
self, gate_rejections: int, reasons: dict[str, str], entered_at: float self, gate_rejections: int, reasons: dict[str, str], clock: StallClock
) -> None: ) -> None:
if gate_rejections == len(self._sources): if gate_rejections == len(self._sources):
names = tuple(s.name for s in self._sources) names = tuple(s.name for s in self._sources)
@@ -244,7 +247,7 @@ class EmbeddingClient:
per_source_reasons=reasons, per_source_reasons=reasons,
) )
stall = self._bp.stall_window_s stall = self._bp.stall_window_s
if self._now() - entered_at > stall and await self._quota.progress_age_s() > stall: if clock.stalled_s() > stall and await self._quota.progress_age_s() > stall:
names = tuple(s.name for s in self._sources) names = tuple(s.name for s in self._sources)
raise AllSourcesExhausted( raise AllSourcesExhausted(
scope=self._scope, scope=self._scope,
@@ -326,7 +329,7 @@ class EmbeddingClient:
await write_back await write_back
except asyncio.CancelledError: except asyncio.CancelledError:
raise raise
except GovernanceBackendError as exc: except (GovernanceBackendError, SourceNotConfiguredError) as exc:
logger.warning("embedding 治理记账写回降级(不冒泡): {}", exc) logger.warning("embedding 治理记账写回降级(不冒泡): {}", exc)
async def _settle_and_release(self, permit: Permit, actual: int) -> None: async def _settle_and_release(self, permit: Permit, actual: int) -> None:
+32 -2
View File
@@ -35,7 +35,26 @@ SOURCE_REASONS = frozenset(
class PolyGatewayError(Exception): class PolyGatewayError(Exception):
"""库内一切领域错误的基类,携带来源上下文便于遥测与日志定位。""" """库内一切领域错误的基类,携带来源上下文便于遥测与日志定位。
`body_text` 是**非 2xx 响应体的摘要**——网关拒绝这次调用时说的话(issue #10)。
它与 `ResultInvalidError.raw_text` 是两回事,严禁混用:
============== ==================================================
``body_text`` **非 2xx** 的 HTTP 错误响应体: 对方**拒绝**的理由
``raw_text`` **2xx** 但内容不可解析时的模型输出原文
============== ==================================================
加在基类而非某个子类,是因为这些错误全部由同一个 HTTP 响应翻译而来——
"对方说了什么""它属于哪一类"正交。scope 级错误(`GatewayUnavailableError`
一族)继承到的恒空值不是噪音,而是"没有单一响应体可言"的如实表达。
**内容已由 transport 层截断**(`transports/_http_errors.summarize_body`),
且可能包含网关对请求的回显——库不做脱敏: 它不知道下游哪些字段敏感,
猜测式脱敏只会同时丢掉诊断价值与安全性。
本字段是**旁路数据**,不参与任何治理判定(重试/换源/熔断计数/限流结算)。
"""
def __init__( def __init__(
self, self,
@@ -44,11 +63,13 @@ class PolyGatewayError(Exception):
source_name: str | None = None, source_name: str | None = None,
status_code: int | None = None, status_code: int | None = None,
operation: str | None = None, operation: str | None = None,
body_text: str = "",
) -> None: ) -> None:
super().__init__(message) super().__init__(message)
self.source_name = source_name self.source_name = source_name
self.status_code = status_code self.status_code = status_code
self.operation = operation self.operation = operation
self.body_text = body_text
class TransientError(PolyGatewayError): class TransientError(PolyGatewayError):
@@ -64,7 +85,16 @@ class SourceDeadError(PolyGatewayError):
class RequestRejectedError(PolyGatewayError): class RequestRejectedError(PolyGatewayError):
"""请求被拒(400/坏输入): 不重试不换源,直接上抛。""" """请求被拒(400/坏输入): 不重试不换源,直接上抛。
**经中转部署时请注意**(issue #10 下游实测): 第三方 API 中转服务自身抖动
时也会回 400,从状态码上与供应商说"你的输入有问题"无法区分。下游曾观测到
同一份字节(sha256 一致)重发 15 次全部成功,且失败那次 `prompt_tokens=0`、
耗时远低于任何成功调用——请求在推理开始前就被挡了。本库仍按确定性失败处理
(对直连供应商而言重试只会白烧配额),批处理场景的下游宜自备兜底分类;
`body_text` 即为此提供判据: 中转抖动的响应体与供应商的 `invalid_request_error`
形态不同。
"""
class ResultInvalidError(PolyGatewayError): class ResultInvalidError(PolyGatewayError):
+21 -11
View File
@@ -4,7 +4,7 @@ from __future__ import annotations
from typing import TYPE_CHECKING from typing import TYPE_CHECKING
from polygateway.errors import GovernanceBackendError from polygateway.errors import GovernanceBackendError, SourceNotConfiguredError
if TYPE_CHECKING: if TYPE_CHECKING:
from polygateway.ports import GateDecision, GateUpdate, ProviderGate from polygateway.ports import GateDecision, GateUpdate, ProviderGate
@@ -22,43 +22,53 @@ class BreakerGate:
async def try_enter(self, source: SourceConfig, owner: str) -> GateDecision: async def try_enter(self, source: SourceConfig, owner: str) -> GateDecision:
try: try:
return await self._gate.try_enter(source.name, owner) return await self._gate.try_enter(source.name, owner)
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"熔断后端故障(try_enter): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"熔断后端故障(try_enter): {exc}", scope=self._scope
) from exc
async def record_success( async def record_success(
self, entry: GateDecision, *, count_attempt: bool = True self, entry: GateDecision, *, count_attempt: bool = True
) -> GateUpdate: ) -> GateUpdate:
try: try:
return await self._gate.record_success(entry, count_attempt=count_attempt) return await self._gate.record_success(entry, count_attempt=count_attempt)
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"熔断后端故障(record_success): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"熔断后端故障(record_success): {exc}", scope=self._scope
) from exc
async def record_failure( async def record_failure(
self, entry: GateDecision, reason: str, force_open: bool self, entry: GateDecision, reason: str, force_open: bool
) -> GateUpdate: ) -> GateUpdate:
try: try:
return await self._gate.record_failure(entry, reason, force_open) return await self._gate.record_failure(entry, reason, force_open)
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"熔断后端故障(record_failure): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"熔断后端故障(record_failure): {exc}", scope=self._scope
) from exc
async def release_probe(self, entry: GateDecision) -> GateUpdate: async def release_probe(self, entry: GateDecision) -> GateUpdate:
try: try:
return await self._gate.release_probe(entry) return await self._gate.release_probe(entry)
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"熔断后端故障(release_probe): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"熔断后端故障(release_probe): {exc}", scope=self._scope
) from exc
async def retry_after_s(self, sources: tuple[str, ...]) -> float: async def retry_after_s(self, sources: tuple[str, ...]) -> float:
try: try:
return await self._gate.retry_after_s(sources) return await self._gate.retry_after_s(sources)
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"熔断后端故障(retry_after_s): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"熔断后端故障(retry_after_s): {exc}", scope=self._scope
) from exc
+17 -9
View File
@@ -8,7 +8,7 @@ from __future__ import annotations
from typing import TYPE_CHECKING from typing import TYPE_CHECKING
from polygateway.errors import GovernanceBackendError from polygateway.errors import GovernanceBackendError, SourceNotConfiguredError
if TYPE_CHECKING: if TYPE_CHECKING:
from polygateway.ports import Permit, RateLimiter from polygateway.ports import Permit, RateLimiter
@@ -26,31 +26,39 @@ class QuotaGate:
async def try_acquire(self, source: SourceConfig) -> Permit | None: async def try_acquire(self, source: SourceConfig) -> Permit | None:
try: try:
return await self._limiter.try_acquire(source.name, source.effective_est_tokens()) return await self._limiter.try_acquire(source.name, source.effective_est_tokens())
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"限流后端故障(try_acquire): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"限流后端故障(try_acquire): {exc}", scope=self._scope
) from exc
async def stats(self, source: SourceConfig) -> SourceStats: async def stats(self, source: SourceConfig) -> SourceStats:
try: try:
return await self._limiter.source_stats(source.name) return await self._limiter.source_stats(source.name)
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"限流后端故障(source_stats): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"限流后端故障(source_stats): {exc}", scope=self._scope
) from exc
async def mark_progress(self) -> None: async def mark_progress(self) -> None:
try: try:
await self._limiter.mark_progress() await self._limiter.mark_progress()
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"限流后端故障(mark_progress): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"限流后端故障(mark_progress): {exc}", scope=self._scope
) from exc
async def progress_age_s(self) -> float: async def progress_age_s(self) -> float:
try: try:
return await self._limiter.progress_age_s() return await self._limiter.progress_age_s()
except GovernanceBackendError: except (GovernanceBackendError, SourceNotConfiguredError):
raise raise
except Exception as exc: except Exception as exc:
raise GovernanceBackendError(f"限流后端故障(progress_age_s): {exc}", scope=self._scope) from exc raise GovernanceBackendError(
f"限流后端故障(progress_age_s): {exc}", scope=self._scope
) from exc
+98 -16
View File
@@ -11,6 +11,7 @@ httpx 是库的核心依赖而非实现层内部件,不违反"middleware 只依
from __future__ import annotations from __future__ import annotations
import asyncio import asyncio
import contextlib
import random import random
import time import time
import uuid import uuid
@@ -28,6 +29,7 @@ from polygateway.errors import (
RequestRejectedError, RequestRejectedError,
ResultInvalidError, ResultInvalidError,
SourceDeadError, SourceDeadError,
SourceNotConfiguredError,
TransientError, TransientError,
) )
from polygateway.middleware.breaker import BreakerGate from polygateway.middleware.breaker import BreakerGate
@@ -38,7 +40,7 @@ from polygateway.streaming import StreamLivenessTimeout
from polygateway.types import LLMResponse from polygateway.types import LLMResponse
if TYPE_CHECKING: if TYPE_CHECKING:
from collections.abc import Awaitable, Callable from collections.abc import AsyncIterator, Awaitable, Callable
from polygateway.ports import ( from polygateway.ports import (
GateDecision, GateDecision,
@@ -73,6 +75,65 @@ def backoff_delay(
return max(delay, retry_after) return max(delay, retry_after)
class _Attempt:
"""一次尝试的计时句柄;`refund()` 把它退还给 stall 账(见 `StallClock`)。"""
__slots__ = ("productive",)
def __init__(self) -> None:
self.productive = True
def refund(self) -> None:
"""该次尝试不消耗重试预算(429),故其耗时归 stall 治理而非重试治理。"""
self.productive = False
class StallClock:
"""调用级 stall 计时器: 只累计非生产性等待(issue #8 设计 §3.1)。
**划分依据是"谁消耗重试预算"**,不是"是否发出了请求"。消耗 `max_attempts`
的时间已被重试预算治理,从 stall 账扣除;不消耗它的时间无人治理,归 stall。
两者重叠计费正是 issue #8 的根因: stall 预算(默认 300s)小于重试预算
(3 × timeout_s),必然先耗尽,于是重试预算在超时场景下永远用不上。
"生产性"的边界即 `_attempt` 的边界,含该次尝试的记账与遥测收尾——它们是
"尝试已有结论"之后的动作,不是在等待重试机会;把它们计入 stall 会让遥测
抖动参与判死。
**例外: 429 尝试须 `refund()`**。429 免重试预算(饱和期等待而非死亡),若其
耗时又算生产性,就掉进两个预算的缝隙——排队型网关持满 timeout 才回 429 时,
每轮只有退避那一两秒进 stall 账,调用可挂满 `stall_window/backoff_base` 轮
(实测 timeout=300/base=2 时达 25 小时)。退还后缝隙闭合。
每次调用创建一个实例。严禁提升为实例属性: `_entered_at` 会固定在进程启动
时刻,使 `stalled_s()` 随进程运行时长单调增长,最终所有调用被误判 stalled。
模块级共享单元, EmbeddingClient 与 OcrClient 复用(同 `backoff_delay`)。
"""
__slots__ = ("_now", "_entered_at", "_productive_s")
def __init__(self, now: Callable[[], float]) -> None:
self._now = now
self._entered_at = now()
self._productive_s = 0.0
def stalled_s(self) -> float:
"""非生产性等待累计秒数 = 调用总耗时 - 消耗重试预算的时间。"""
return self._now() - self._entered_at - self._productive_s
@contextlib.asynccontextmanager
async def attempting(self) -> AsyncIterator[_Attempt]:
"""包裹一次真实尝试,其耗时默认记为生产性(除非被 `refund()`)。"""
handle = _Attempt()
started = self._now()
try:
yield handle
finally:
# 只做算术与取值, 不吞任何异常——CancelledError 逐字穿透(库铁律)
if handle.productive:
self._productive_s += self._now() - started
def _demote_call_failures( def _demote_call_failures(
ordered: list[SourceConfig], ordered: list[SourceConfig],
attempt_fails: dict[str, int], attempt_fails: dict[str, int],
@@ -156,6 +217,15 @@ class _Failed:
immediate: bool immediate: bool
def _is_rate_limited(outcome: LLMResponse | _Failed) -> bool:
"""429 = 服务端调度指令(gRPC pushback 语义,迭代 5): 按 Retry-After 退避但
**不消耗重试预算**——饱和窗口里等待而非死亡;其余失败照常计数。
因其免重试预算,该次尝试的耗时必须归 stall 治理(`StallClock` 的 refund)。
"""
return isinstance(outcome, _Failed) and _failure_reason(outcome.exc) == "rate_limited"
class RetryMW: class RetryMW:
"""尝试编排器;时钟/睡眠/随机全部注入,纯确定性可测(P6)。""" """尝试编排器;时钟/睡眠/随机全部注入,纯确定性可测(P6)。"""
@@ -208,12 +278,12 @@ class RetryMW:
reasons: dict[str, str] = {} reasons: dict[str, str] = {}
# 调用内失败计数(设计 §3.3): 局部状态,调用结束即弃;严禁实例属性(并发共享) # 调用内失败计数(设计 §3.3): 局部状态,调用结束即弃;严禁实例属性(并发共享)
attempt_fails: dict[str, int] = {} attempt_fails: dict[str, int] = {}
entered_at = self._now() # 调用级累计计时,循环内不重置(CHS governance.py:207) # 调用级累计计时,循环内不重置(CHS governance.py:207);issue #8 起只计
# 非生产性等待——真实尝试由重试预算治理,不再重复烧 stall 预算
clock = StallClock(self._now)
while True: while True:
# 调用级时间上限(迭代 5): 429 免预算后的兜底,防饱和期无限循环 # 调用级时间上限(迭代 5): 429 免预算后的兜底,防饱和期无限循环
# 与 _on_no_runnable 同款双条件(CHS 口径): 本地超窗且全局无进展才判死 if await self._stalled(clock):
stall = self._bp.stall_window_s
if self._now() - entered_at > stall and await self._quota.progress_age_s() > stall:
raise AllSourcesExhausted( raise AllSourcesExhausted(
scope=self._scope, scope=self._scope,
reason="stalled", reason="stalled",
@@ -222,14 +292,17 @@ class RetryMW:
) )
picked, gate_rejections = await self._pick_runnable(reasons, attempt_fails) picked, gate_rejections = await self._pick_runnable(reasons, attempt_fails)
if picked is None: if picked is None:
await self._on_no_runnable(gate_rejections, reasons, entered_at) await self._on_no_runnable(gate_rejections, reasons, clock)
continue continue
async with clock.attempting() as attempt:
outcome = await self._attempt(request, *picked, reasons, attempt_fails) outcome = await self._attempt(request, *picked, reasons, attempt_fails)
rate_limited = _is_rate_limited(outcome)
if rate_limited:
# 免了重试预算就得进 stall 账,否则这段耗时无人治理(见 StallClock)
attempt.refund()
if isinstance(outcome, LLMResponse): if isinstance(outcome, LLMResponse):
return outcome return outcome
# 429 = 服务端调度指令(gRPC pushback 语义,迭代 5): 按 Retry-After if not rate_limited:
# 退避但不消耗重试预算——饱和窗口里等待而非死亡;其余失败照常计数
if _failure_reason(outcome.exc) != "rate_limited":
fails += 1 fails += 1
if fails >= self._retry.max_attempts: if fails >= self._retry.max_attempts:
raise AllSourcesExhausted( raise AllSourcesExhausted(
@@ -282,8 +355,20 @@ class RetryMW:
await self._settle_and_release(permit, 0) await self._settle_and_release(permit, 0)
return None, gate_rejections return None, gate_rejections
# —— 背压与 stall 判死(CHS governance.py:270-285)——
async def _stalled(self, clock: StallClock) -> bool:
"""双条件 stall 判死(CHS governance.py:270-281): 本地累计等待与全局
无进展**同时**超窗才判死——本地 monotonic 与后端时钟刻意不混用。
本地一侧只计非生产性等待(issue #8,见 `StallClock`)。短路顺序有意为之:
本地未超窗就不问后端,省一次 Redis 往返。
"""
stall = self._bp.stall_window_s
return clock.stalled_s() > stall and await self._quota.progress_age_s() > stall
async def _on_no_runnable( async def _on_no_runnable(
self, gate_rejections: int, reasons: dict[str, str], entered_at: float self, gate_rejections: int, reasons: dict[str, str], clock: StallClock
) -> None: ) -> None:
if gate_rejections == len(self._sources): if gate_rejections == len(self._sources):
names = tuple(s.name for s in self._sources) names = tuple(s.name for s in self._sources)
@@ -299,10 +384,7 @@ class RetryMW:
retry_after_s=self._bp.poll_interval_s, retry_after_s=self._bp.poll_interval_s,
per_source_reasons=reasons, per_source_reasons=reasons,
) )
# 双条件 stall 判死(CHS governance.py:270-281): 本地累计等待与全局 if await self._stalled(clock):
# 无进展**同时**超窗才判死——本地 monotonic 与后端时钟刻意不混用。
stall = self._bp.stall_window_s
if self._now() - entered_at > stall and await self._quota.progress_age_s() > stall:
names = tuple(s.name for s in self._sources) names = tuple(s.name for s in self._sources)
raise AllSourcesExhausted( raise AllSourcesExhausted(
scope=self._scope, scope=self._scope,
@@ -401,7 +483,7 @@ class RetryMW:
await write_back await write_back
except asyncio.CancelledError: except asyncio.CancelledError:
raise raise
except GovernanceBackendError as exc: except (GovernanceBackendError, SourceNotConfiguredError) as exc:
logger.warning("治理记账写回降级(不冒泡): {}", exc) logger.warning("治理记账写回降级(不冒泡): {}", exc)
def _feed_outcome(self, source_name: str, ok: bool) -> None: def _feed_outcome(self, source_name: str, ok: bool) -> None:
+6 -2
View File
@@ -17,7 +17,11 @@ from typing import TYPE_CHECKING
from loguru import logger from loguru import logger
from polygateway.errors import GatewayUnavailableError, GovernanceBackendError from polygateway.errors import (
GatewayUnavailableError,
GovernanceBackendError,
SourceNotConfiguredError,
)
from polygateway.middleware.cache import digest_messages from polygateway.middleware.cache import digest_messages
from polygateway.types import canonical_sampling_json, merge_sampling from polygateway.types import canonical_sampling_json, merge_sampling
@@ -247,7 +251,7 @@ class TelemetryMW:
started = self._now() started = self._now()
try: try:
response = await call_next(request) response = await call_next(request)
except (GatewayUnavailableError, GovernanceBackendError) as exc: except (GatewayUnavailableError, GovernanceBackendError, SourceNotConfiguredError) as exc:
await self._emitter.emit_terminal_failure( await self._emitter.emit_terminal_failure(
request=request, request=request,
call_id=str(uuid.uuid4()), call_id=str(uuid.uuid4()),
+12 -7
View File
@@ -30,11 +30,12 @@ from polygateway.errors import (
RequestRejectedError, RequestRejectedError,
ResultInvalidError, ResultInvalidError,
SourceDeadError, SourceDeadError,
SourceNotConfiguredError,
TransientError, TransientError,
) )
from polygateway.middleware.breaker import BreakerGate from polygateway.middleware.breaker import BreakerGate
from polygateway.middleware.ratelimit import QuotaGate from polygateway.middleware.ratelimit import QuotaGate
from polygateway.middleware.retry import _failure_reason, backoff_delay from polygateway.middleware.retry import StallClock, _failure_reason, backoff_delay
from polygateway.middleware.telemetry import TelemetryEmitter from polygateway.middleware.telemetry import TelemetryEmitter
from polygateway.ports import OutcomeAwareSelector from polygateway.ports import OutcomeAwareSelector
from polygateway.sources import SourceCooldownMemo from polygateway.sources import SourceCooldownMemo
@@ -203,13 +204,17 @@ class OcrClient:
raise AllSourcesExhausted(scope=self._scope, reason="no_sources", retry_after_s=0.0) raise AllSourcesExhausted(scope=self._scope, reason="no_sources", retry_after_s=0.0)
fails = 0 fails = 0
reasons: dict[str, str] = {} reasons: dict[str, str] = {}
entered_at = self._now() # 只计非生产性等待(issue #8): 真实尝试由重试预算治理,不重复烧 stall 预算
clock = StallClock(self._now)
while True: while True:
picked, gate_rejections = await self._pick_runnable(reasons) picked, gate_rejections = await self._pick_runnable(reasons)
if picked is None: if picked is None:
await self._on_no_runnable(gate_rejections, reasons, entered_at) await self._on_no_runnable(gate_rejections, reasons, clock)
continue continue
outcome = await self._attempt(kind, image, *picked, reasons, session_id, parent_call_id) async with clock.attempting():
outcome = await self._attempt(
kind, image, *picked, reasons, session_id, parent_call_id
)
if isinstance(outcome, _AttemptOutcome): if isinstance(outcome, _AttemptOutcome):
return outcome return outcome
fails += 1 fails += 1
@@ -252,7 +257,7 @@ class OcrClient:
return None, gate_rejections return None, gate_rejections
async def _on_no_runnable( async def _on_no_runnable(
self, gate_rejections: int, reasons: dict[str, str], entered_at: float self, gate_rejections: int, reasons: dict[str, str], clock: StallClock
) -> None: ) -> None:
if gate_rejections == len(self._sources): if gate_rejections == len(self._sources):
names = tuple(s.name for s in self._sources) names = tuple(s.name for s in self._sources)
@@ -269,7 +274,7 @@ class OcrClient:
per_source_reasons=reasons, per_source_reasons=reasons,
) )
stall = self._bp.stall_window_s stall = self._bp.stall_window_s
if self._now() - entered_at > stall and await self._quota.progress_age_s() > stall: if clock.stalled_s() > stall and await self._quota.progress_age_s() > stall:
names = tuple(s.name for s in self._sources) names = tuple(s.name for s in self._sources)
raise AllSourcesExhausted( raise AllSourcesExhausted(
scope=self._scope, scope=self._scope,
@@ -360,7 +365,7 @@ class OcrClient:
await write_back await write_back
except asyncio.CancelledError: except asyncio.CancelledError:
raise raise
except GovernanceBackendError as exc: except (GovernanceBackendError, SourceNotConfiguredError) as exc:
logger.warning("OCR 治理记账写回降级(不冒泡): {}", exc) logger.warning("OCR 治理记账写回降级(不冒泡): {}", exc)
async def _settle_and_release(self, permit: Permit) -> None: async def _settle_and_release(self, permit: Permit) -> None:
+1 -1
View File
@@ -245,7 +245,7 @@ class StructuredOutputStrategy(Protocol):
@runtime_checkable @runtime_checkable
class TelemetryRecorder(Protocol): class TelemetryRecorder(Protocol):
"""遥测后端;20 字段冻结(M1 设计 §4.4 + issue #3),唯一调用点是 TelemetryEmitter。 """遥测后端;22 字段冻结(M1 设计 §4.4 + issue #3/#4),唯一调用点是 TelemetryEmitter。
新增参数不设默认值: 库外无第三方实现者(三项目迁移时删除了各自的同名 新增参数不设默认值: 库外无第三方实现者(三项目迁移时删除了各自的同名
Protocol),完整签名的成本为零,而少写一列会被 emitter 的降级吞成 warning。 Protocol),完整签名的成本为零,而少写一列会被 emitter 的降级吞成 warning。
+73 -14
View File
@@ -1,12 +1,17 @@
"""Postgres 遥测后端(M2 设计 §5): asyncpg lazy 池 + 两级降级。 """Postgres 遥测后端(M2 设计 §5): asyncpg lazy 池 + 两级降级。
参考仓无先例(三项目遥测全 SQLite);asyncpg 工程写法取 GovDoc 参考仓无先例(三项目遥测全 SQLite);asyncpg 工程写法取 GovDoc
`taskrun/postgres_store.py`($n 占位、`CREATE TABLE IF NOT EXISTS`、 `taskrun/postgres_store.py`($n 占位、`ON CONFLICT DO NOTHING`),但其
`ON CONFLICT DO NOTHING`),但其"失败冒泡"方向按遥测铁律**有意反转**: "失败冒泡"方向按遥测铁律**有意反转**:
① 结构性失败(建池/建表)→ warning 一次后永久降级(池置 None 短路); ① 结构性失败 → warning 一次后永久降级(所有写入短路);
② 运行时单条写失败 → 逐条 warning 丢弃,不降级不重试(连接抖动由 ② 运行时单条写失败 → 逐条 warning 丢弃,不降级不重试(连接抖动由
asyncpg 池自恢复;避免浸泡开头一次抖动导致后续全程失遥测)。 asyncpg 池自恢复;避免浸泡开头一次抖动导致后续全程失遥测)。
构造不连库(lazy),20 列 schema 与 SQLite 版同名同序。 构造不连库(lazy),22 列 schema 与 SQLite 版同名同序。
**"结构性"的判据是「确定写不进去」,不是「初始化时出过错」**(issue #9):
只有建池失败(重试要在业务路径上内联吞掉 connect 超时)与"表确定不存在
且建不出来"(后续 INSERT 必然全败)才判死;探测失败、补列失败、取连接
失败一律只 warning,让写入照常尝试或下次调用重试。
""" """
from __future__ import annotations from __future__ import annotations
@@ -55,6 +60,9 @@ _BACKFILL = (
("reasoning_tokens", "ALTER TABLE llm_calls ADD COLUMN reasoning_tokens INTEGER"), ("reasoning_tokens", "ALTER TABLE llm_calls ADD COLUMN reasoning_tokens INTEGER"),
) )
# 探测表是否存在;不需要任何权限,且与 INSERT 走同一套 search_path 解析
_TABLE_EXISTS = "SELECT to_regclass('llm_calls')"
# 探测现有列;尊重 search_path(to_regclass 按当前 search_path 解析) # 探测现有列;尊重 search_path(to_regclass 按当前 search_path 解析)
_EXISTING_COLUMNS = ( _EXISTING_COLUMNS = (
"SELECT attname FROM pg_attribute " "SELECT attname FROM pg_attribute "
@@ -111,30 +119,81 @@ class PostgresRecorder:
self._init_lock = asyncio.Lock() self._init_lock = asyncio.Lock()
async def _ensure_ready(self) -> asyncpg.Pool | None: async def _ensure_ready(self) -> asyncpg.Pool | None:
"""lazy 建池+表;结构性失败 warning 一次后永久降级(设计 §5 两级之一)""" """lazy 建池+表;判死只认「确定写不进去」(issue #9),其余失败都留活路"""
if self._failed: if self._failed:
return None return None
if self._schema_ready: if self._schema_ready:
return self._pool return self._pool
async with self._init_lock: async with self._init_lock:
if self._failed or self._schema_ready: if self._failed:
return None if self._failed else self._pool return None
if self._schema_ready:
return self._pool
pool = await self._open_pool()
if pool is None:
return None
return await self._prepare_schema(pool)
async def _open_pool(self) -> asyncpg.Pool | None:
"""建池;失败即永久降级(唯一一处「无条件判死」)。"""
if self._pool is not None:
return self._pool
try: try:
if self._pool is None:
import asyncpg import asyncpg
self._pool = await asyncpg.create_pool(self._dsn, timeout=10) self._pool = await asyncpg.create_pool(self._dsn, timeout=10)
async with self._pool.acquire() as conn:
await conn.execute(_DDL)
await self._backfill_columns(conn)
self._schema_ready = True
return self._pool
except asyncio.CancelledError: except asyncio.CancelledError:
raise raise
except Exception as exc: except Exception as exc:
# 池建不出来 = 确定写不进去;且每次调用重试都要内联吞掉 connect
# 超时,而遥测是业务路径上的 await —— 此处必须永久降级
self._failed = True self._failed = True
logger.warning("Postgres 遥测初始化失败,后续记录降级为 no-op: {}", exc) logger.warning("Postgres 遥测建池失败,后续记录降级为 no-op: {}", exc)
return None return None
return self._pool
async def _prepare_schema(self, pool: asyncpg.Pool) -> asyncpg.Pool | None:
"""备好表并交回可用的池;瞬时失败只跳过本次,确定写不进去才判死。"""
try:
async with pool.acquire() as conn:
writable = await self._prepare_table(conn)
except asyncio.CancelledError:
raise
except Exception as exc:
# 池已在手,取连接/探测失败多为瞬时抖动: 不判死也不标就绪,
# 只跳过本次记录,下次调用重新准备
logger.warning("Postgres 遥测建表探测失败(跳过本条,下次重试): {}", exc)
return None
if not writable:
self._failed = True
return None
self._schema_ready = True
return pool
async def _prepare_table(self, conn: object) -> bool:
"""备好 `llm_calls`;**表存在就绝不发 DDL**。返回 False 仅表示表确定不存在。
`CREATE TABLE IF NOT EXISTS` 不能无条件发: PostgreSQL 对 schema 的
CREATE 权限检查**早于** `IF NOT EXISTS` 的存在性判断(PG 16.14 实测:
只授 `SELECT, INSERT ON llm_calls` 的角色,表明明在、也写得进去,这一句
照样被拒 `permission denied for schema`)。这与 `_backfill_columns` 撞的
是同一类问题(issue #3/#9),故守卫也必须同款: 先探测,后 DDL。
探测走 `to_regclass`,不需要任何权限,且与 INSERT 的 search_path 解析
口径一致——比裸 DDL 更准(裸 `CREATE TABLE` 落在首个**可建**的 schema,
可能与 INSERT 命中的不是同一张表)。
"""
exists = await conn.fetchval(_TABLE_EXISTS) is not None # type: ignore[attr-defined]
if exists:
await self._backfill_columns(conn) # 旧表可能缺列;失败只逐行降级
return True
try:
await conn.execute(_DDL) # type: ignore[attr-defined]
except asyncio.CancelledError:
raise
except Exception as exc:
logger.warning("Postgres 遥测建表失败(表不存在,记录无处可落): {}", exc)
return False
return True # 新建表列已齐全,无需再走补列
async def _backfill_columns(self, conn: object) -> None: async def _backfill_columns(self, conn: object) -> None:
"""给已存在的旧表补新列(issue #3);**先探测再 ALTER,失败绝不置 `_failed`**。 """给已存在的旧表补新列(issue #3);**先探测再 ALTER,失败绝不置 `_failed`**。
+8
View File
@@ -3,6 +3,14 @@
蓝本 VT `adapters/telemetry.py`: 构造期建连接与表,失败降级为 no-op 蓝本 VT `adapters/telemetry.py`: 构造期建连接与表,失败降级为 no-op
(记录基础设施不得拖垮业务调用);`INSERT OR IGNORE` 幂等(call_id 主键); (记录基础设施不得拖垮业务调用);`INSERT OR IGNORE` 幂等(call_id 主键);
写入经 threading.Lock 串行化后由 `asyncio.to_thread` 执行,不阻塞事件循环。 写入经 threading.Lock 串行化后由 `asyncio.to_thread` 执行,不阻塞事件循环。
**这里不做 postgres.py 那样的建表前探测,是实测后的有意不对称**(issue #9):
SQLite 对已存在的表在**解析期**就把 `CREATE TABLE IF NOT EXISTS` 短路掉,
既不抢写锁也不检查可写性——实测同一时刻另一连接持 `BEGIN EXCLUSIVE`、或
文件 `chmod 444`,该语句均通过,而同条件下的 `INSERT` 与新表名建表分别报
database is locked / readonly database。故 PG 侧"权限检查早于存在性判断"
的坑在此不存在,加探测零收益。**别为了代码对称把它加回来**;需要对称的是
保证(表存在就不该因建表失败而失能),这一条两侧都已满足。
""" """
from __future__ import annotations from __future__ import annotations
@@ -0,0 +1,63 @@
"""HTTP 错误响应体的取用与摘要(issue #10 设计 §3.2)。
两个 transport 各有自己的状态码分类逻辑(OCR 有意不做 429 细分),但**摘要口径
必须是同一份**——issue #10 的教训正是"只有一个分支用了响应体",一处例外就是
下一次事后查不到原因。故本模块是全库唯一的摘要实现,不得在别处复制。
"""
from __future__ import annotations
import httpx
_ERROR_BODY_CAP = 2048
"""摘要总长上限(**字符**,含省略标记在内)。
取值对齐 Kubernetes client-go `rest/request.go` 的 `maxUnstructuredResponseTextBytes
= 2048`——它是唯一与本设计同场景(读 HTTP 错误体做诊断)的成熟先例。按字符而非
字节切,多字节字符不会被切成半个;`error` 列是 TEXT,无定长约束,不需要字节口径。
"""
_HEAD_CHARS = 1400
_TAIL_CHARS = 600
def summarize_body(text: str) -> str:
"""折叠空白后按头尾策略摘要;空/空白入参返回空串。
**折叠空白**不是洁癖: 错误体常是缩进 JSON,原样拼进 message 会把一行日志
炸成多行、把遥测列变得不可读。
**保头保尾**而非头部硬切: 截断的对象是结构化 JSON,信息分布头重尾也重——
人话(`message`)在前,机器可判的 `type`/`code`/`param`/`request_id` 在后。
k8s/Sentry 用头部硬切是因为它们截的是任意文本;本函数截的是错误 JSON,
头部硬切正好切掉向网关方追查时唯一有用的那部分。策略取自标准库 `reprlib`
"给人读的长字符串"的处置。
**标记记下省略字数**,读的人才知道自己丢了多少,不会误以为网关只说了这么多。
"""
collapsed = " ".join(text.split())
if len(collapsed) <= _ERROR_BODY_CAP:
return collapsed
omitted = len(collapsed) - _HEAD_CHARS - _TAIL_CHARS
return f"{collapsed[:_HEAD_CHARS]}…(略 {omitted} 字)…{collapsed[-_TAIL_CHARS:]}"
def compose_message(message: str, summary: str) -> str:
"""摘要非空才拼后缀,避免留下悬空的分隔符。
分隔符取 ` | ` 而非既有的 `: `,让"库说的话""网关说的话"一眼可分。
"""
return f"{message} | {summary}" if summary else message
def response_body(response: httpx.Response) -> str:
"""取**已缓冲**的响应文本;未读缓冲一律降级空串。
绝不在此触发网络读: 那会在错误路径上凭空插入一次可能挂住的 IO。降级方向
与缓存/遥测同档(库铁律)——诊断信息缺失不得把一次本可正确分类的失败变成
不可分类的崩溃,那正是 `ResponseNotRead` 泄漏出四分类之外的后果。
"""
try:
return response.text
except httpx.ResponseNotRead:
return ""
+13 -1
View File
@@ -24,6 +24,11 @@ from polygateway.errors import (
SourceDeadError, SourceDeadError,
TransientError, TransientError,
) )
from polygateway.transports._http_errors import (
compose_message,
response_body,
summarize_body,
)
from polygateway.types import ( from polygateway.types import (
OcrLayoutElement, OcrLayoutElement,
OcrLayoutTransportResult, OcrLayoutTransportResult,
@@ -74,13 +79,20 @@ def _translate_http_errors(source_name: str, operation: str) -> Iterator[None]:
def _classify_status( def _classify_status(
exc: httpx.HTTPStatusError, source_name: str, operation: str exc: httpx.HTTPStatusError, source_name: str, operation: str
) -> TransientError | SourceDeadError | RequestRejectedError: ) -> TransientError | SourceDeadError | RequestRejectedError:
"""HTTP 状态码 → 错误四分类,**全部分支**携带响应体摘要(issue #10)。
分类映射本身零变更;摘要口径与 chat 侧共用同一实现,不得在此另起一份——
"只有一个分支用了响应体"正是 issue #10 的成因。
"""
status = exc.response.status_code status = exc.response.status_code
summary = summarize_body(response_body(exc.response))
ctx: dict[str, Any] = { ctx: dict[str, Any] = {
"source_name": source_name, "source_name": source_name,
"status_code": status, "status_code": status,
"operation": operation, "operation": operation,
"body_text": summary,
} }
message = f"{source_name} OCR {operation} HTTP {status}" message = compose_message(f"{source_name} OCR {operation} HTTP {status}", summary)
if status >= 500 or status == 429: if status >= 500 or status == 429:
return TransientError(message, **ctx) return TransientError(message, **ctx)
if status in (401, 403): if status in (401, 403):
+39 -18
View File
@@ -16,6 +16,7 @@ from typing import TYPE_CHECKING, Any
import httpx import httpx
from polygateway.errors import ( from polygateway.errors import (
PolyGatewayError,
RequestRejectedError, RequestRejectedError,
ResultInvalidError, ResultInvalidError,
SourceDeadError, SourceDeadError,
@@ -30,6 +31,7 @@ from polygateway.providers import (
resolve_thinking, resolve_thinking,
) )
from polygateway.streaming import StreamLivenessTimeout, stream_with_liveness_timeouts from polygateway.streaming import StreamLivenessTimeout, stream_with_liveness_timeouts
from polygateway.transports._http_errors import compose_message, summarize_body
from polygateway.types import EmbeddingTransportResult, SourceConfig, TransportResult from polygateway.types import EmbeddingTransportResult, SourceConfig, TransportResult
if TYPE_CHECKING: if TYPE_CHECKING:
@@ -107,40 +109,59 @@ def _parse_retry_after(raw: str | None) -> float | None:
return seconds if seconds > 0 else None return seconds if seconds > 0 else None
def _translate_429(source: SourceConfig, body_text: str, headers: Mapping[str, str]) -> Exception: def _translate_429(
source: SourceConfig, body_text: str, headers: Mapping[str, str], ctx: dict[str, Any]
) -> Exception:
"""429 细分。`body_text` 必须是**未截断的原文**——`ctx["body_text"]` 是摘要,
头尾保留会破坏 JSON 结构,拿它解析会让超长 body 的配额耗尽退化成普通限速
(该源不再 force_open),把一个诊断改进变成治理 bug(issue #10 实现红线)。
"""
try: try:
err_type = json.loads(body_text).get("error", {}).get("type", "") err_type = json.loads(body_text).get("error", {}).get("type", "")
except (json.JSONDecodeError, AttributeError): except (json.JSONDecodeError, AttributeError):
err_type = "" err_type = ""
summary = ctx["body_text"]
if err_type == "insufficient_quota": if err_type == "insufficient_quota":
return SourceDeadError( return SourceDeadError(
f"{source.name} 配额耗尽(insufficient_quota)", compose_message(f"{source.name} 配额耗尽(insufficient_quota)", summary), **ctx
source_name=source.name,
status_code=429,
operation="chat",
) )
return TransientError( return TransientError(
f"{source.name} 限速: 429", compose_message(f"{source.name} 限速: 429", summary),
retry_after_s=_parse_retry_after(headers.get("retry-after")), retry_after_s=_parse_retry_after(headers.get("retry-after")),
source_name=source.name, **ctx,
status_code=429,
operation="chat",
) )
def _classify(status: int) -> tuple[type[PolyGatewayError], str]:
"""状态码 → (错误类, message 标签);映射与 ARCH §6.2 逐条相同,本次零变更。"""
if status in (401, 403):
return SourceDeadError, "凭据失效/欠费"
if status == 400:
return RequestRejectedError, "请求被拒"
if status >= 500:
return TransientError, "瞬时错误"
return RequestRejectedError, "客户端错误"
def _status_to_error( def _status_to_error(
source: SourceConfig, status: int, body_text: str, headers: Mapping[str, str] source: SourceConfig, status: int, body_text: str, headers: Mapping[str, str]
) -> Exception: ) -> Exception:
ctx: dict[str, Any] = {"source_name": source.name, "status_code": status, "operation": "chat"} """非 2xx → 领域错误,**全部分支**携带响应体摘要(issue #10)。
if status in (401, 403):
return SourceDeadError(f"{source.name} 凭据失效/欠费: {status}", **ctx) 摘要只算一次,message 与 `body_text` 共用同一份串: 两份不同长度会让"遥测里
if status == 400: 看到的""下游 catch 到的"对不上,排查时反而多一层困惑。
return RequestRejectedError(f"{source.name} 请求被拒: 400", **ctx) """
summary = summarize_body(body_text)
ctx: dict[str, Any] = {
"source_name": source.name,
"status_code": status,
"operation": "chat",
"body_text": summary,
}
if status == 429: if status == 429:
return _translate_429(source, body_text, headers) return _translate_429(source, body_text, headers, ctx)
if status >= 500: cls, label = _classify(status)
return TransientError(f"{source.name} 瞬时错误: {status}", **ctx) return cls(compose_message(f"{source.name} {label}: {status}", summary), **ctx)
return RequestRejectedError(f"{source.name} 客户端错误: {status}", **ctx)
def _strip_think(content: str) -> tuple[str, str]: def _strip_think(content: str) -> tuple[str, str]:
+40 -1
View File
@@ -11,7 +11,12 @@ import sqlite3
import httpx import httpx
import pytest import pytest
from polygateway import CircuitOpenError, GatewayClient, TransientError from polygateway import (
CircuitOpenError,
GatewayClient,
RequestRejectedError,
TransientError,
)
from polygateway.backends.memory.breaker import InMemoryGate from polygateway.backends.memory.breaker import InMemoryGate
from polygateway.backends.memory.cache import InMemoryCache from polygateway.backends.memory.cache import InMemoryCache
from polygateway.backends.memory.limiter import InMemoryLimiter from polygateway.backends.memory.limiter import InMemoryLimiter
@@ -173,6 +178,40 @@ class TestTelemetryAcrossPaths:
assert len({cid for _, cid in rows}) == 2 # call_id 逐次独立 assert len({cid for _, cid in rows}) == 2 # call_id 逐次独立
class TestRejectionReasonIsQueryable:
"""issue #10 的验收主张: 400 之后,网关说的话必须能在遥测表里查到。
下游一轮 1050 张影像的批处理里,1 张在读表格时收到 400 被判确定性失败,
事后"这张图到底哪里不合规"无从查起——响应体在 transport 翻译层就没了。
"""
# issue #10 原文给出的真实响应体(一字不改)
_BODY = (
'{"error":{"message":"<400> ***.***.InvalidParameter: The image format is illegal '
'and cannot be opened","type":"invalid_request_error","param":"",'
'"code":"invalid_parameter_error"}}'
)
async def test_rejected_call_leaves_the_reason_in_telemetry(self, tmp_path):
recorder = SQLiteRecorder(tmp_path / "t.db")
client = _full_client(
lambda req: httpx.Response(400, content=self._BODY.encode()), telemetry=recorder
)
with pytest.raises(RequestRejectedError):
await client.chat([{"role": "user", "content": "hi"}])
recorder.close()
rows = sqlite3.connect(tmp_path / "t.db").execute("SELECT error FROM llm_calls").fetchall()
assert rows, "400 必须留下遥测行(遥测必录)"
errors = " ".join(r[0] or "" for r in rows)
# 修复前这里只有 "qwen_1 请求被拒: 400"——诊断信息一个字都不在
assert "InvalidParameter" in errors
assert "The image format is illegal" in errors
# 尾部的 code 才是向网关方追查的凭据,头部硬切正好会丢掉它
assert "invalid_parameter_error" in errors
class TestStructuredThroughStack: class TestStructuredThroughStack:
async def test_feedback_reask_passes_through_governance(self): async def test_feedback_reask_passes_through_governance(self):
"""重问经过内层治理: 第二次真实请求同样被限流/熔断记账。""" """重问经过内层治理: 第二次真实请求同样被限流/熔断记账。"""
@@ -13,6 +13,7 @@ from __future__ import annotations
import asyncio import asyncio
import json import json
import os import os
import re
from uuid import uuid4 from uuid import uuid4
import pytest import pytest
@@ -299,3 +300,88 @@ class TestDegradation:
await _record_minimal(recorder) await _record_minimal(recorder)
await recorder.aclose() await recorder.aclose()
await recorder.aclose() await recorder.aclose()
_PROBE_PASSWORD = "pgw_issue9_probe" # 临时角色,teardown 删除;非任何真实凭据
@pytest.fixture
async def least_privilege_dsn(dsn):
"""临时 schema + 临时角色: 只授表级 SELECT/INSERT,**不授 schema CREATE**。
这是 issue #9 的现场——最小权限部署的标准形态。fixture 建的一切
(schema、表、角色)都在 teardown 里删净,共享的 public.llm_calls 不受影响;
连不上或无权建角色(非超级用户)时 skip,不让 CI 假绿。
"""
import asyncpg
from polygateway.telemetry.postgres import _DDL
name = f"pgwtest_lp_{uuid4().hex[:8]}"
admin = await asyncpg.connect(dsn, timeout=10)
try:
if not await admin.fetchval(
"SELECT rolcreaterole OR rolsuper FROM pg_roles WHERE rolname = current_user"
):
pytest.skip("当前账号无权建临时角色,跳过最小权限用例")
await admin.execute(f"CREATE ROLE {name} LOGIN PASSWORD '{_PROBE_PASSWORD}'")
await admin.execute(f"CREATE SCHEMA {name}")
await admin.execute(f"SET search_path = {name}")
await admin.execute(_DDL) # 表由**别的账号**建好,与现场一致
await admin.execute(f"GRANT USAGE ON SCHEMA {name} TO {name}")
await admin.execute(f"GRANT SELECT, INSERT ON {name}.llm_calls TO {name}")
# 关键: 绝不 GRANT CREATE ON SCHEMA —— 缺的正是这一项
finally:
await admin.close()
low = re.sub(r"//[^@/]+@", f"//{name}:{_PROBE_PASSWORD}@", dsn, count=1)
sep = "&" if "?" in low else "?"
yield f"{low}{sep}options=-csearch_path%3D{name}", name
admin = await asyncpg.connect(dsn, timeout=10)
try:
await admin.execute(f"DROP SCHEMA IF EXISTS {name} CASCADE")
await admin.execute(f"DROP OWNED BY {name}")
await admin.execute(f"DROP ROLE IF EXISTS {name}")
finally:
await admin.close()
class TestLeastPrivilegeDeployment:
"""issue #9: 只有表级写权限的账号,遥测必须照常落库而不是整体判死。"""
async def test_create_table_if_not_exists_is_denied_for_this_role(self, least_privilege_dsn):
"""库外事实先钉死: 表存在、写得进去,DDL 仍被拒——PG 的权限检查早于 IF NOT EXISTS。
修复依赖的是这条 PG 语义;若某天它变了,这里先红,而不是让下面那条
用例悄悄变成"永远通过"的空断言。
"""
import asyncpg
low_dsn, _ = least_privilege_dsn
conn = await asyncpg.connect(low_dsn, timeout=10)
try:
assert await conn.fetchval("SELECT to_regclass('llm_calls')") is not None
with pytest.raises(asyncpg.exceptions.InsufficientPrivilegeError):
await conn.execute("CREATE TABLE IF NOT EXISTS llm_calls (call_id TEXT)")
finally:
await conn.close()
async def test_records_land_without_schema_create_privilege(self, least_privilege_dsn):
"""修复前: 建表被拒 → _failed → 整个进程一条不落(下游 150 次调用全丢)。"""
low_dsn, schema = least_privilege_dsn
recorder = PostgresRecorder(low_dsn)
try:
await _record_minimal(recorder, call_id=_cid("lp1"))
await _record_minimal(recorder, call_id=_cid("lp2"), cost=1.5)
assert recorder._failed is False # 判死开关不得被建表权限触发
rows = await _fetch(
low_dsn,
"SELECT call_id, cost FROM llm_calls WHERE call_id LIKE $1 ORDER BY call_id",
f"{_RUN_PREFIX}-lp%",
)
assert [(r["call_id"], r["cost"]) for r in rows] == [
(_cid("lp1"), None),
(_cid("lp2"), 1.5),
]
assert schema # teardown 会连表带角色删净
finally:
await recorder.aclose()
+277 -6
View File
@@ -18,6 +18,7 @@ from polygateway.errors import (
SourceNotConfiguredError, SourceNotConfiguredError,
TransientError, TransientError,
) )
from polygateway.middleware.ratelimit import QuotaGate
from polygateway.middleware.retry import RetryMW, backoff_delay from polygateway.middleware.retry import RetryMW, backoff_delay
from polygateway.sources import RoundRobinSelector, SourceCooldownMemo from polygateway.sources import RoundRobinSelector, SourceCooldownMemo
from polygateway.types import ( from polygateway.types import (
@@ -52,19 +53,31 @@ class BoundedSleep:
await self._side_effect(len(self.delays)) await self._side_effect(len(self.delays))
def _mw(sources, limiter, script, *, clock, sleep, rng=lambda: 0.0, quota_full="wait", gate=None): def _mw(
sources,
limiter,
script,
*,
clock,
sleep,
rng=lambda: 0.0,
quota_full="wait",
gate=None,
transport=None,
emitter=None,
):
return RetryMW( return RetryMW(
scope="llm", scope="llm",
sources=sources, sources=sources,
selector=RoundRobinSelector(), selector=RoundRobinSelector(),
limiter=limiter, limiter=limiter,
gate=gate or InMemoryGate(config=_BREAKER, now=clock), gate=gate or InMemoryGate(config=_BREAKER, now=clock),
transport=FakeTransport(script), transport=transport or FakeTransport(script),
retry=RetryPolicy(max_attempts=3, backoff_base_s=2.0, backoff_max_s=30.0), retry=RetryPolicy(max_attempts=3, backoff_base_s=2.0, backoff_max_s=30.0),
backpressure=BackpressurePolicy(stall_window_s=_STALL, poll_interval_s=0.01), backpressure=BackpressurePolicy(stall_window_s=_STALL, poll_interval_s=0.01),
quota_full=quota_full, quota_full=quota_full,
cooldown_memo=SourceCooldownMemo(now=clock), cooldown_memo=SourceCooldownMemo(now=clock),
emitter=None, emitter=emitter,
now=clock, now=clock,
sleep=sleep, sleep=sleep,
rng=rng, rng=rng,
@@ -117,7 +130,11 @@ class TestStallQuadrants:
assert resp.content == "ok" assert resp.content == "ok"
async def test_global_stale_but_local_fresh_keeps_waiting(self): async def test_global_stale_but_local_fresh_keeps_waiting(self):
"""仅全局超窗(从未出餐 age=inf): 本地才刚开始等 → 不判死。""" """仅全局超窗(从未出餐 age=inf): 本地才刚开始等 → 不判死。
`inf` 语义在 issue #8 后未变;变的是"本地"的口径——它现在度量的是
非生产性等待累计,不再是墙钟总耗时(见 TestStallBudget)。
"""
clock = FakeClock() clock = FakeClock()
src, limiter = _blocked_limiter(clock) src, limiter = _blocked_limiter(clock)
held = await limiter.try_acquire("s1", 0) held = await limiter.try_acquire("s1", 0)
@@ -177,6 +194,216 @@ class TestStallQuadrants:
await task await task
class ClockAdvancingTransport:
"""按脚本 [(推进秒数, 动作), ...] 执行: 在一次尝试内部推进时钟, 模拟真实耗时。
动作语义同 `FakeTransport`(异常即抛、"hang" 即挂起、其余为返回值)。
stall 口径的关键区分在于"时间花在哪", 故必须能让时钟只在 transport 内前进。
"""
def __init__(self, script, clock):
self.script = list(script)
self.clock = clock
self.calls = []
async def complete(self, *, messages, source, stream, overlay, call_id):
self.calls.append((source.name, call_id))
advance, action = self.script.pop(0)
self.clock.advance(advance)
if isinstance(action, Exception):
raise action
if action == "hang":
await asyncio.Event().wait()
return action
class _SlowEmitter:
"""遥测收尾中推进时钟: 钉住"遥测耗时属生产性"(设计 §3.1 边界声明)。"""
def __init__(self, clock, advance):
self._clock = clock
self._advance = advance
async def emit_attempt(self, *args, **kwargs):
self._clock.advance(self._advance)
class TestStallBudget:
"""stall 预算只计非生产性等待(issue #8 设计 §3.1)。
根因是两个预算重叠计费: 真实尝试的耗时同时烧重试预算与 stall 预算,
而 stall 预算更小必然先耗尽, 于是 max_attempts 在超时场景下永不生效。
"""
def _free_limiter(self, clock):
src = make_source()
limiter = InMemoryLimiter(
scope="llm", sources={"s1": src}, global_limits=_NO_GLOBAL, now=clock
)
return src, limiter
async def test_single_timeout_does_not_exhaust_stall_budget(self):
"""timeout_s == stall_window_s 时, 一次超时不得判死——重试预算须真实可用。"""
clock = FakeClock()
src, limiter = self._free_limiter(clock)
# 第一次尝试耗满 300s 超时后失败, 第二次立即成功
transport = ClockAdvancingTransport(
[(_STALL + 1, TransientError("timeout", status_code=504)), (0.0, _ok())], clock
)
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
resp = await mw(_REQ)
assert resp.content == "ok"
assert len(transport.calls) == 2 # 第二次尝试确实发出了
async def test_productive_time_excluded_from_stall(self):
"""连续多次长尝试也不烧 stall 预算: 它们烧的是重试预算。"""
clock = FakeClock()
src, limiter = self._free_limiter(clock)
transport = ClockAdvancingTransport(
[
(_STALL + 100, TransientError("slow", status_code=500)),
(_STALL + 100, TransientError("slow", status_code=500)),
(0.0, _ok()),
],
clock,
)
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
resp = await mw(_REQ)
assert resp.content == "ok"
async def test_telemetry_time_counts_as_productive(self):
"""遥测收尾属 `_attempt` 边界内: 遥测抖动不得参与判死(设计 §3.1)。
必须走**失败**路径才有判别力: 成功后直接 return, 循环开头的 stall
判定根本不会再执行。此处让首次尝试快速失败、而遥测收尾慢得超窗,
下一轮循环开头即检验遥测耗时有没有被算进 stall 账。
"""
clock = FakeClock()
src, limiter = self._free_limiter(clock)
transport = ClockAdvancingTransport(
[(0.1, TransientError("boom", status_code=500)), (0.0, _ok())], clock
)
mw = _mw(
[src],
limiter,
[],
clock=clock,
sleep=BoundedSleep(),
transport=transport,
emitter=_SlowEmitter(clock, _STALL + 100),
)
resp = await mw(_REQ)
assert resp.content == "ok"
async def test_nonproductive_wait_still_triggers_stall(self):
"""兜底未被削弱: 纯轮询等待累满窗口仍判死。"""
clock = FakeClock()
src, limiter = _blocked_limiter(clock)
_held = await limiter.try_acquire("s1", 0)
async def advance(_n):
clock.advance(_STALL + 100)
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(advance))
with pytest.raises(AllSourcesExhausted) as ei:
await mw(_REQ)
assert ei.value.reason == "stalled"
async def test_saturation_429_still_stalls(self):
"""429 免预算不烧 fails, 主循环兜底须仍能判死而非无限循环(设计 §3.5)。"""
clock = FakeClock()
src, limiter = self._free_limiter(clock)
# 429 往返本身极快(生产性可忽略), 退避 sleep 才是非生产性的大头
transport = ClockAdvancingTransport(
[(0.1, TransientError("429", status_code=429)) for _ in range(10)], clock
)
async def advance(_n):
clock.advance(_STALL)
mw = _mw(
[src], limiter, [], clock=clock, sleep=BoundedSleep(advance), transport=transport
)
with pytest.raises(AllSourcesExhausted) as ei:
await mw(_REQ)
assert ei.value.reason == "stalled" # 不是 retry_exhausted: 429 确实没烧重试预算
async def test_slow_429_does_not_escape_both_budgets(self):
"""排队型网关: 持满 timeout 才回 429。该耗时必须落进 stall 账。
429 免重试预算, 所以它的耗时若又算生产性就**两个预算都不烧**——调用
会挂满 stall_window/backoff_base 轮。修复前实测 301 次尝试、25.2 小时;
此处钉住"一轮 429 就把 stall 账推满"这个上界。
"""
clock = FakeClock()
src, limiter = self._free_limiter(clock)
transport = ClockAdvancingTransport(
[(_STALL + 1, TransientError("429", status_code=429))] * 20, clock
)
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
with pytest.raises(AllSourcesExhausted) as ei:
await mw(_REQ)
assert ei.value.reason == "stalled"
# 一次持满超时的 429 即耗尽 stall 窗口, 不再无限排队
assert len(transport.calls) <= 2
async def test_cancel_inside_attempt_pierces(self):
"""取消发生在 `attempting()` 包裹内仍逐字穿透(库铁律)。"""
clock = FakeClock()
src, limiter = self._free_limiter(clock)
transport = ClockAdvancingTransport([(0.0, "hang")], clock)
mw = _mw([src], limiter, [], clock=clock, sleep=asyncio.sleep, transport=transport)
task = asyncio.create_task(mw(_REQ))
while not transport.calls:
await asyncio.sleep(0.01)
task.cancel()
with pytest.raises(asyncio.CancelledError):
await task
assert (await limiter.source_stats("s1")).inflight == 0 # permit 在 finally 释放
async def test_clock_is_per_call_not_per_instance(self):
"""StallClock 必须是**调用级**局部状态,不得提升为 RetryMW 实例属性。
生产形态是一个长寿命 RetryMW 跑成千上万次调用。若 clock 成了实例属性,
`_entered_at` 会固定在进程启动时刻, 每次调用的 stalled_s() 随进程运行
时长单调增长, 最终所有调用被误判 stalled——这是本用例要拦的灾难。
判别力的关键是**复用同一个 mw**: 两个 mw 实例天然隔离, 抓不到实例共享。
"""
clock = FakeClock()
src = make_source()
limiter = InMemoryLimiter(
scope="llm", sources={"s1": src}, global_limits=_NO_GLOBAL, now=clock
)
transport = ClockAdvancingTransport([(0.0, _ok("first")), (0.0, _ok("second"))], clock)
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
first = await mw(_REQ)
clock.advance(_STALL + 100) # 两次调用之间进程空转远超窗
second = await mw(_REQ)
assert (first.content, second.content) == ("first", "second")
async def test_concurrent_calls_do_not_share_clock(self):
"""并发两路共用同一个 mw: 一快一慢都能正常完成(形态冒烟)。
**这条不是回归防线**: 实测它在"clock 提为实例属性""去掉 refund""去掉
生产性扣减"三种变异下均保持绿色——共享 clock 时慢调用的耗时是作为
credit 记进共享账的,污染方向是让 stall 账**变小**(更宽松),而本用例
断言两路都成功。真正钉住调用级隔离的是上面那条
`test_clock_is_per_call_not_per_instance`。保留此条只为覆盖并发形态。
"""
clock = FakeClock()
src = make_source(max_concurrency=2)
limiter = InMemoryLimiter(
scope="llm", sources={"s1": src}, global_limits=_NO_GLOBAL, now=clock
)
transport = ClockAdvancingTransport(
[(_STALL + 100, _ok("slow")), (0.0, _ok("fast"))], clock
)
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
results = await asyncio.gather(mw(_REQ), mw(_REQ))
assert {r.content for r in results} == {"slow", "fast"}
class _GateSuccessBroken(InMemoryGate): class _GateSuccessBroken(InMemoryGate):
async def record_success(self, entry): async def record_success(self, entry):
raise GovernanceBackendError("redis 抖动", scope="llm") raise GovernanceBackendError("redis 抖动", scope="llm")
@@ -192,6 +419,12 @@ class _LimiterProgressBroken(InMemoryLimiter):
raise GovernanceBackendError("redis 抖动", scope="llm") raise GovernanceBackendError("redis 抖动", scope="llm")
class _GateSuccessMisconfigured(InMemoryGate):
# 签名须与端口一致(含 count_attempt),否则抛的是 TypeError 而非本类要测的异常
async def record_success(self, entry, *, count_attempt: bool = True):
raise SourceNotConfiguredError("未知源 's1'(scope=llm)")
class TestAccountingDegradation: class TestAccountingDegradation:
"""记账侧降级(设计 §10,ARCH §7.3 勘误): 调用已完成,写回失败不冒泡。""" """记账侧降级(设计 §10,ARCH §7.3 勘误): 调用已完成,写回失败不冒泡。"""
@@ -206,6 +439,24 @@ class TestAccountingDegradation:
resp = await mw(_REQ) resp = await mw(_REQ)
assert resp.content == "ok" # 真实成功响应不因记账失败被丢弃 assert resp.content == "ok" # 真实成功响应不因记账失败被丢弃
async def test_assembly_defect_on_accounting_path_also_degrades(self):
"""记账侧降级按"路径性质"而非异常类型: 装配缺陷同样不得毁掉已完成的调用。
`SourceNotConfiguredError` 被放行穿透闸门包装器(issue #7 §T6)后,若
`_record_quietly` 只降级 `GovernanceBackendError`,它就会从记账侧冒泡、
销毁一个真实成功的响应——反转本类钉住的既有行为。当前无后端会从记账
方法抛它,此用例是为将来加了源名校验的后端守住这条不变式。
"""
clock = FakeClock()
src = make_source()
limiter = InMemoryLimiter(
scope="llm", sources={"s1": src}, global_limits=_NO_GLOBAL, now=clock
)
gate = _GateSuccessMisconfigured(config=_BREAKER, now=clock)
mw = _mw([src], limiter, [_ok()], clock=clock, sleep=BoundedSleep(), gate=gate)
resp = await mw(_REQ)
assert resp.content == "ok"
async def test_mark_progress_failure_does_not_lose_response(self): async def test_mark_progress_failure_does_not_lose_response(self):
clock = FakeClock() clock = FakeClock()
src = make_source() src = make_source()
@@ -280,14 +531,34 @@ class TestUnknownSourceIsAssemblyDefect:
# 关键: 若归入 scope 级家族,配置写错的任务会永远延期重投、永不进死信 # 关键: 若归入 scope 级家族,配置写错的任务会永远延期重投、永不进死信
assert not isinstance(ei.value, GatewayUnavailableError) assert not isinstance(ei.value, GatewayUnavailableError)
@pytest.mark.parametrize("method", ["try_acquire", "stats"])
async def test_survives_the_quota_gate_wrapper(self, method):
"""必须穿透 QuotaGate,否则整个拆分在生产路径上等于没做。
上面两条(以及 redis 版)打的都是私有 `_cfg`,绕过了包装器。而治理循环
只经 QuotaGate 访问后端,包装器的 `except Exception` 会把装配缺陷重新
包成 `GovernanceBackendError`——下游又拿到可重投异常,永远重投不告警。
"""
src = make_source("s1")
# 限流后端的源名单与治理循环拿到的源对不上 = 装配缺陷
limiter = InMemoryLimiter(
scope="llm", sources={"other": src}, global_limits=_NO_GLOBAL
)
gate = QuotaGate(limiter, scope="llm")
with pytest.raises(SourceNotConfiguredError) as ei:
await getattr(gate, method)(src)
assert not isinstance(ei.value, GatewayUnavailableError)
class TestGateFailuresReachCallersAsScopeLevel: class TestGateFailuresReachCallersAsScopeLevel:
"""三条闸门泄漏路径必须以 scope 级不可用的形态到达调用方(issue #7)。 """闸门泄漏路径必须以 scope 级不可用的形态到达调用方(issue #7)。
记账路径由 `_record_quietly` 降级为 warning,但闸门路径没有那层包裹,会一路 记账路径由 `_record_quietly` 降级为 warning,但闸门路径没有那层包裹,会一路
抛给调用方。只写 `except GatewayUnavailableError` 的调用方此前接不住,后果 抛给调用方。只写 `except GatewayUnavailableError` 的调用方此前接不住,后果
是 Redis 抖一下就让积压任务烧掉业务失败预算进死信——而那是运维重启即可恢复 是 Redis 抖一下就让积压任务烧掉业务失败预算进死信——而那是运维重启即可恢复
的故障。三条路径逐一钉住,防止将来任何一条被漏掉。 的故障。全部五条为: `QuotaGate` 的 try_acquire / stats / progress_age_s,
`BreakerGate` 的 try_enter / retry_after_s(判据是该调用点未被 `_record_quietly`
包裹)。此处钉住其中三条代表路径,余两条由同一注入机制覆盖。
""" """
async def test_try_acquire_failure_is_scope_level(self): async def test_try_acquire_failure_is_scope_level(self):
+60
View File
@@ -173,6 +173,8 @@ from polygateway.types import ( # noqa: E402
GlobalLimits, GlobalLimits,
RetryPolicy, RetryPolicy,
) )
from tests.contracts.conftest import FakeClock # noqa: E402
from tests.unit.test_backpressure import BoundedSleep # noqa: E402
_BREAKER = BreakerConfig(fail_threshold=3, cooldown_s=60.0, probe_ttl_s=120.0) _BREAKER = BreakerConfig(fail_threshold=3, cooldown_s=60.0, probe_ttl_s=120.0)
_NO_GLOBAL = GlobalLimits(max_concurrency=0, rpm=0, tpm=0) _NO_GLOBAL = GlobalLimits(max_concurrency=0, rpm=0, tpm=0)
@@ -208,6 +210,31 @@ class ScriptedEmbedTransport:
return action return action
class _ClockAdvancingEmbedTransport:
"""按脚本 [(推进秒数, 动作), ...] 执行: 在一次尝试内部推进时钟(issue #8)。
动作语义同 `ScriptedEmbedTransport`。stall 口径要区分"时间花在哪",
故必须能让时钟只在 transport 内前进。
"""
def __init__(self, script, clock):
self.script = list(script)
self.clock = clock
self.calls = []
async def embed(self, *, texts, source, call_id):
self.calls.append((source.name, list(texts), call_id))
advance, action = self.script.pop(0)
self.clock.advance(advance)
if isinstance(action, Exception):
raise action
if action == "hang":
await asyncio.Event().wait()
if action == "ok":
return _vec_for(texts)
return action
class _MemoryRecorder: class _MemoryRecorder:
def __init__(self): def __init__(self):
self.rows = [] self.rows = []
@@ -338,6 +365,39 @@ class TestEmbedGovernance:
await task await task
assert (await limiter.source_stats("e1")).inflight == 0 assert (await limiter.source_stats("e1")).inflight == 0
async def test_single_timeout_does_not_exhaust_stall_budget(self):
"""issue #8: 一次耗满 timeout 的尝试不得吃掉 stall 预算而使重试失效。
embedding 只有 `_on_no_runnable` 一处 stall 判定, 故失效链条是
"先超时一次(墙钟耗尽) → 再遇到无可用源 → 判死"。此处正是这条路径。
"""
clock = FakeClock()
limiter = held = None # 闭包延迟求值: client 建好后才有 limiter
async def toggle_permit(n):
"""首次退避占满 permit, 迫使下一轮走 _on_no_runnable; 之后放行。"""
nonlocal held
if n == 1:
held = await limiter.try_acquire("e1", 0)
else:
await held.release()
# 第一次尝试耗满 300s 超时失败, 随后被迫走一轮 _on_no_runnable——
# stall 判定就在那里, 检验它有没有把这 300s 生产性时间算进 stall 账
transport = _ClockAdvancingEmbedTransport(
[(300.1, TransientError("timeout", status_code=504)), (0.0, "ok")], clock
)
client, limiter = _embed_client(
[_src(max_concurrency=1)],
[],
now=clock,
transport=transport,
sleep=BoundedSleep(toggle_permit),
)
resp = await client.embed(["a"])
assert resp.vectors == [[1.0]]
assert len(transport.calls) == 2 # 第二次尝试确实发出了
class TestEmbedTelemetry: class TestEmbedTelemetry:
async def test_per_batch_rows_with_digest(self): async def test_per_batch_rows_with_digest(self):
+38 -3
View File
@@ -34,6 +34,43 @@ class TestBaseShape:
assert TransientError("429", retry_after_s=2.5).retry_after_s == 2.5 assert TransientError("429", retry_after_s=2.5).retry_after_s == 2.5
class TestBodyText:
"""issue #10: 非 2xx 的响应体摘要必须有承载处,否则拒绝理由事后不可查。"""
@pytest.mark.parametrize(
"cls", (PolyGatewayError, TransientError, SourceDeadError, RequestRejectedError)
)
def test_defaults_empty_and_accepts_summary(self, cls):
assert cls("boom").body_text == ""
assert cls("boom", body_text='{"error":{"code":"bad"}}').body_text == (
'{"error":{"code":"bad"}}'
)
def test_result_invalid_keeps_both_fields_apart(self):
"""`body_text`(非 2xx 的拒绝理由)与 `raw_text`(2xx 的不可解析输出)不得混用。"""
exc = ResultInvalidError("bad json", raw_text="{oops", body_text="")
assert exc.raw_text == "{oops"
assert exc.body_text == ""
@pytest.mark.parametrize(
"exc",
(
AllSourcesExhausted(scope="llm", reason="stalled", retry_after_s=1.0),
CircuitOpenError(scope="llm", retry_after_s=1.0),
GovernanceBackendError("redis down", scope="llm"),
),
)
def test_scope_level_errors_carry_no_body(self, exc):
"""scope 级失败没有单一响应体可言,空串是如实表达而非噪音。"""
assert exc.body_text == ""
def test_body_text_does_not_leak_into_str(self):
"""字段是旁路数据: 加了它不得改变任何既有异常的 str() 输出。"""
assert str(RequestRejectedError("qwen_1 请求被拒: 400", body_text="whatever")) == (
"qwen_1 请求被拒: 400"
)
class TestResultInvalid: class TestResultInvalid:
def test_carries_diagnosis(self): def test_carries_diagnosis(self):
exc = ResultInvalidError( exc = ResultInvalidError(
@@ -134,9 +171,7 @@ class TestGovernanceBackendReason:
assert "governance_backend_down" in SCOPE_REASONS assert "governance_backend_down" in SCOPE_REASONS
def test_gateway_unavailable_accepts_the_new_reason(self): def test_gateway_unavailable_accepts_the_new_reason(self):
exc = AllSourcesExhausted( exc = AllSourcesExhausted(scope="LLM", reason="governance_backend_down", retry_after_s=0.0)
scope="LLM", reason="governance_backend_down", retry_after_s=0.0
)
assert exc.reason == "governance_backend_down" assert exc.reason == "governance_backend_down"
def test_retry_after_default_is_non_zero(self): def test_retry_after_default_is_non_zero(self):
+89
View File
@@ -0,0 +1,89 @@
"""HTTP 错误响应体摘要口径(issue #10 设计 §3.2/§3.4)。
摘要是 message `body_text` 共用的**同一份串**,故它的边界行为直接决定
遥测里看到的与下游 catch 到的是否一致本组用例把规则钉成算术
"""
import httpx
import pytest
from polygateway.transports._http_errors import (
_ERROR_BODY_CAP,
_HEAD_CHARS,
_TAIL_CHARS,
compose_message,
response_body,
summarize_body,
)
# issue #10 原文给出的真实响应体(一字不改),关键在于 code 收尾
_REAL_SAMPLE = (
'{"error":{"message":"<400> ***.***.InvalidParameter: The image format is illegal '
'and cannot be opened","type":"invalid_request_error","param":"",'
'"code":"invalid_parameter_error"}}'
)
class TestSummarizeBody:
def test_short_body_passes_through(self):
assert summarize_body(_REAL_SAMPLE) == _REAL_SAMPLE
def test_whitespace_collapsed(self):
"""错误体常是缩进 JSON: 不折叠会把一行日志炸成多行、遥测列不可读。"""
assert (
summarize_body('{\n "error": {\n "code": "x"\n }\n}')
== '{ "error": { "code": "x" } }'
)
@pytest.mark.parametrize("raw", ("", " ", "\n\t \n"))
def test_blank_yields_empty(self, raw):
assert summarize_body(raw) == ""
def test_exactly_at_cap_is_untouched(self):
body = "x" * _ERROR_BODY_CAP
assert summarize_body(body) == body
def test_one_over_cap_is_summarized(self):
summary = summarize_body("x" * (_ERROR_BODY_CAP + 1))
assert summary != "x" * (_ERROR_BODY_CAP + 1)
assert "" in summary
def test_head_and_tail_both_survive(self):
"""头部硬切会丢掉尾部,而 JSON 错误体的 code/request_id 正在尾部。"""
body = "H" * 5000 + "T" * 5000
summary = summarize_body(body)
assert summary[:_HEAD_CHARS] == body[:_HEAD_CHARS]
assert summary[-_TAIL_CHARS:] == body[-_TAIL_CHARS:]
assert f"…(略 {10000 - _HEAD_CHARS - _TAIL_CHARS} 字)…" in summary
def test_real_sample_tail_visible_in_oversized_body(self):
"""设计 §7 用例 3c: 超长体里,追查网关方所需的 code 仍须可见。"""
summary = summarize_body("PADDING" * 1000 + _REAL_SAMPLE)
assert '"code":"invalid_parameter_error"}}' in summary
def test_idempotent(self):
"""再摘要一次不得嵌套标记,否则重复经手的串会层层套娃。"""
once = summarize_body("y" * 9999)
assert summarize_body(once) == once
class TestComposeMessage:
def test_empty_summary_leaves_message_intact(self):
assert compose_message("qwen_1 请求被拒: 400", "") == "qwen_1 请求被拒: 400"
def test_non_empty_summary_is_appended(self):
assert compose_message("qwen_1 请求被拒: 400", "{}") == "qwen_1 请求被拒: 400 | {}"
class TestResponseBody:
def test_reads_buffered_text(self):
assert response_body(httpx.Response(400, content=b'{"e":1}')) == '{"e":1}'
def test_unread_stream_degrades_to_empty(self):
"""取不到诊断信息绝不能升级为崩溃: 未读缓冲返回空串,且不触发网络读。"""
class _Unread(httpx.SyncByteStream):
def __iter__(self):
yield b"body"
assert response_body(httpx.Response(400, stream=_Unread())) == ""
+44 -1
View File
@@ -19,7 +19,11 @@ from polygateway.errors import (
SourceDeadError, SourceDeadError,
TransientError, TransientError,
) )
from polygateway.transports.monkey_ocr import MonkeyOcrTransport, _parse_middle_json from polygateway.transports.monkey_ocr import (
MonkeyOcrTransport,
_classify_status,
_parse_middle_json,
)
from polygateway.types import SourceConfig from polygateway.types import SourceConfig
@@ -306,6 +310,45 @@ class TestErrorTranslation:
await t.recognize_text(image=b"jpg", source=_source(), call_id="c1") await t.recognize_text(image=b"jpg", source=_source(), call_id="c1")
assert ei.value.status_code == status assert ei.value.status_code == status
@pytest.mark.parametrize(
("status", "exc_type"),
[
(502, TransientError),
(429, TransientError), # 与 5xx 共用分支,但仍单列: 漏分支正是 issue #10 的成因
(401, SourceDeadError),
(404, RequestRejectedError),
],
)
async def test_body_survives_every_branch(self, status, exc_type):
"""issue #10: OCR 侧 message 原本只有 HTTP 状态码,拒绝理由同样丢失。"""
body = '{"detail":"unsupported image mode CMYK"}'
t = _transport_for(_routes(text_resp=httpx.Response(status, content=body.encode())))
with pytest.raises(exc_type) as ei:
await t.recognize_text(image=b"jpg", source=_source(), call_id="c1")
assert ei.value.body_text == body
assert str(ei.value).endswith(f" | {body}")
def test_unread_body_degrades_without_changing_class(self):
"""取不到 body 时降级空串: 绝不能让 ResponseNotRead 逃出错误四分类。
直接测纯函数而非走 MockTransport真实客户端对非 stream 请求总会读完
响应,未读态只可能在将来给 OCR stream 时出现,而那正是要防的场景
"""
class _Unread(httpx.SyncByteStream):
def __iter__(self):
yield b"body"
exc = httpx.HTTPStatusError(
"404",
request=httpx.Request("POST", "http://ocr.example/ocr/text"),
response=httpx.Response(404, stream=_Unread()),
)
err = _classify_status(exc, "monkey_1", "text")
assert isinstance(err, RequestRejectedError)
assert err.body_text == ""
assert str(err) == "monkey_1 OCR text HTTP 404"
async def test_connect_error_transient(self): async def test_connect_error_transient(self):
def handler(request): def handler(request):
raise httpx.ConnectError("refused", request=request) raise httpx.ConnectError("refused", request=request)
+48
View File
@@ -80,6 +80,22 @@ class ScriptedOcrTransport:
raise NotImplementedError raise NotImplementedError
class ClockAdvancingOcrTransport(ScriptedOcrTransport):
"""按脚本 [(推进秒数, 动作), ...] 在一次尝试内部推进时钟(issue #8)。
stall 口径要区分"时间花在哪",故必须能让时钟只在 transport 内前进
"""
def __init__(self, script, clock):
super().__init__([a for _, a in script])
self._advances = [d for d, _ in script]
self.clock = clock
async def _next(self, method, source, call_id):
self.clock.advance(self._advances.pop(0))
return await super()._next(method, source, call_id)
class StaticSelector: class StaticSelector:
def order(self, sources, stats): def order(self, sources, stats):
return list(sources) return list(sources)
@@ -292,6 +308,38 @@ class TestBackpressure:
await permit.settle(0) await permit.settle(0)
await permit.release() await permit.release()
async def test_single_timeout_does_not_exhaust_stall_budget(self):
"""issue #8: 一次耗满 timeout 的尝试不得吃掉 stall 预算而使重试失效。
OCR 只有 `_on_no_runnable` 一处 stall 判定,故失效链条是"先超时一次
(墙钟耗尽) 再遇到无可用源 判死"。此处正是这条路径。
"""
clock = FakeClock()
limiter = held = None
rounds = []
async def toggle_permit(_seconds):
"""首次退避占满 permit,迫使下一轮走 _on_no_runnable;之后放行。"""
nonlocal held
rounds.append(_seconds)
if len(rounds) > 10:
raise RuntimeError("超过 10 次轮询仍未判死/未获 permit")
if len(rounds) == 1:
held = await limiter.acquire("m1", 0)
else:
await held.settle(0)
await held.release()
transport = ClockAdvancingOcrTransport(
[(300.1, TransientError("timeout", status_code=504)), (0.0, "text")], clock
)
client, limiter, _ = _client(
[_src(max_concurrency=1)], [], now=clock, sleep=toggle_permit, transport=transport
)
r = await client.recognize_text(b"jpg")
assert r.text == "LINE-1"
assert len(transport.calls) == 2 # 第二次尝试确实发出了
class FakeClock: class FakeClock:
def __init__(self, start=1000.0): def __init__(self, start=1000.0):
+120
View File
@@ -16,6 +16,7 @@ from polygateway.errors import (
) )
from polygateway.middleware.telemetry import TelemetryEmitter from polygateway.middleware.telemetry import TelemetryEmitter
from polygateway.pricing import ModelPrice, PricingTable from polygateway.pricing import ModelPrice, PricingTable
from polygateway.transports._http_errors import summarize_body
from polygateway.transports.openai_compat import ( from polygateway.transports.openai_compat import (
OpenAICompatTransport, OpenAICompatTransport,
_iter_sse_deltas, _iter_sse_deltas,
@@ -691,6 +692,125 @@ class TestErrorTranslation:
await _complete(_transport_for(handler), _source()) await _complete(_transport_for(handler), _source())
# issue #10 原文给出的真实响应体(一字不改): 关键在于 code 收尾
_REJECT_BODY = (
'{"error":{"message":"<400> ***.***.InvalidParameter: The image format is illegal '
'and cannot be opened","type":"invalid_request_error","param":"",'
'"code":"invalid_parameter_error"}}'
)
class TestErrorBodyRetention:
"""issue #10: 网关说了什么必须活着离开翻译层——message 与字段各留一份。"""
@pytest.mark.parametrize(
("status", "exc"),
[
(400, RequestRejectedError),
(401, SourceDeadError),
(403, SourceDeadError),
(404, RequestRejectedError),
(500, TransientError),
(503, TransientError),
],
)
async def test_every_non_2xx_branch_keeps_the_body(self, status, exc):
def handler(request):
return httpx.Response(status, content=_REJECT_BODY.encode())
with pytest.raises(exc) as ei:
await _complete(_transport_for(handler), _source())
assert ei.value.body_text == _REJECT_BODY
# message 与字段共用同一份串: 遥测里看到的与下游 catch 到的不得打架
assert str(ei.value).endswith(f" | {_REJECT_BODY}")
assert "invalid_parameter_error" in str(ei.value)
async def test_rate_limited_429_keeps_body_and_retry_after(self):
def handler(request):
return httpx.Response(
429,
content=b'{"error":{"message":"per-minute cap 3"}}',
headers={"retry-after": "2.5"},
)
with pytest.raises(TransientError) as ei:
await _complete(_transport_for(handler), _source())
assert "per-minute cap 3" in str(ei.value)
assert ei.value.retry_after_s == 2.5 # 摘要不得干扰既有解析
async def test_insufficient_quota_429_keeps_body(self):
body = json.dumps(
{"error": {"type": "insufficient_quota", "message": "daily budget spent"}}
)
def handler(request):
return httpx.Response(429, content=body.encode())
with pytest.raises(SourceDeadError) as ei:
await _complete(_transport_for(handler), _source())
assert "daily budget spent" in str(ei.value)
async def test_oversized_insufficient_quota_still_classified_dead(self):
"""实现红线: 类型判定必须读**原文**。
摘要会破坏 JSON 结构,若改用摘要解析,超长 body 的配额耗尽将退化成普通
限速配额已耗尽的源不再 force_open,一个诊断改进就变成了治理 bug
**填充必须是多个键**,不能是单个超长字符串值: 后者的截断点落在字符串
*内部*,省略标记成了合法的字符串内容,而头尾保留又让尾部的 error 对象
幸存摘要照样解析得出 `insufficient_quota`,用例即告空转(2026-08-16
verifier 变异测试发现: 按错误写法实现,全套件 824 项依然全绿)多键
填充让截断点落在结构记号之间,摘要才真正不可解析
"""
body = json.dumps(
{**{f"k{i}": "v" * 10 for i in range(300)}, "error": {"type": "insufficient_quota"}}
)
assert len(body) > 2048
with pytest.raises(json.JSONDecodeError):
# 判别力的前提: 摘要确实不再是合法 JSON,读它必然拿不到 type
json.loads(summarize_body(body))
def handler(request):
return httpx.Response(429, content=body.encode())
with pytest.raises(SourceDeadError):
await _complete(_transport_for(handler), _source())
async def test_empty_body_leaves_no_dangling_separator(self):
def handler(request):
return httpx.Response(400, content=b"")
with pytest.raises(RequestRejectedError) as ei:
await _complete(_transport_for(handler), _source())
assert str(ei.value) == "qwen_1 请求被拒: 400"
assert ei.value.body_text == ""
@pytest.mark.parametrize("body", [b"<html>gateway down</html>", b"\xff\xfe not utf-8"])
async def test_non_json_and_non_utf8_bodies_do_not_explode(self, body):
def handler(request):
return httpx.Response(400, content=body)
with pytest.raises(RequestRejectedError) as ei:
await _complete(_transport_for(handler), _source())
assert ei.value.status_code == 400 # 分类不受 body 形态影响
async def test_non_stream_path_keeps_the_body(self):
def handler(request):
return httpx.Response(400, content=_REJECT_BODY.encode())
with pytest.raises(RequestRejectedError) as ei:
await _complete(_transport_for(handler), _source(), stream=False)
assert ei.value.body_text == _REJECT_BODY
async def test_embedding_path_keeps_the_body(self):
def handler(request):
return httpx.Response(400, content=_REJECT_BODY.encode())
with pytest.raises(RequestRejectedError) as ei:
await _transport_for(handler).embed(texts=["hi"], source=_source(), call_id="cid-embed")
assert ei.value.body_text == _REJECT_BODY
class TestLifecycle: class TestLifecycle:
async def test_aclose_idempotent(self): async def test_aclose_idempotent(self):
def handler(request): def handler(request):
+96 -2
View File
@@ -257,17 +257,41 @@ class TestSQLiteColumnBackfill:
class _FakePgConn: class _FakePgConn:
"""记录执行过的语句;可让 ALTER 抛错以模拟权限不足。""" """记录执行过的语句;可让 ALTER/CREATE/探测抛错以模拟权限不足与抖动
def __init__(self, existing: list[str], *, fail_alter: bool = False): `existing` 为空列表即表示**表不存在**(与真实 PG 一致: `to_regclass` NULL
时列探测必然零行), `fetchval` `fetch` 共用同一份事实
"""
def __init__(
self,
existing: list[str],
*,
fail_alter: bool = False,
fail_create: bool = False,
probe_errors: int = 0,
):
self.existing = existing self.existing = existing
self.fail_alter = fail_alter self.fail_alter = fail_alter
self.fail_create = fail_create
self.probe_errors = probe_errors
self.statements: list[str] = [] self.statements: list[str] = []
async def execute(self, sql, *args): async def execute(self, sql, *args):
self.statements.append(sql) self.statements.append(sql)
if sql.startswith("ALTER TABLE") and self.fail_alter: if sql.startswith("ALTER TABLE") and self.fail_alter:
raise RuntimeError("must be owner of table llm_calls") raise RuntimeError("must be owner of table llm_calls")
if sql.lstrip().startswith("CREATE TABLE"):
if self.fail_create:
raise RuntimeError("permission denied for schema public")
self.existing = list(_EXPECTED_COLUMNS)
async def fetchval(self, sql, *args):
self.statements.append(sql)
if self.probe_errors > 0:
self.probe_errors -= 1
raise RuntimeError("connection was closed in the middle of operation")
return "llm_calls" if self.existing else None
async def fetch(self, sql, *args): async def fetch(self, sql, *args):
self.statements.append(sql) self.statements.append(sql)
@@ -338,6 +362,76 @@ class TestPostgresBackfillDiscipline:
assert all("IF NOT EXISTS" not in s for s in altered) # 探测已确认缺列,无需再判 assert all("IF NOT EXISTS" not in s for s in altered) # 探测已确认缺列,无需再判
class TestPostgresTableProbe:
"""建表必须先探测,且"判死"只认"确定写不进去"(issue #9)。
实测(PostgreSQL 16.14,只有表级 SELECT/INSERT 的角色): `CREATE TABLE IF NOT
EXISTS` 被拒 permission denied for schema,而同一连接的 `INSERT` 通过
PG schema CREATE 权限检查早于 `IF NOT EXISTS` 的存在性判断无条件发
DDL 会让这类最小权限部署的整个进程静默失遥测
"""
_CURRENT = [
"call_id",
"cost",
"created_at",
"cached_prompt_tokens",
"model_reported",
"sampling",
"reasoning_tokens",
]
def _recorder(self, conn):
from polygateway.telemetry.postgres import PostgresRecorder
return PostgresRecorder("postgresql://u:p@h:5432/polygateway", pool=_FakePgPool(conn))
def _created(self, conn):
return [s for s in conn.statements if s.lstrip().startswith("CREATE TABLE")]
async def test_existing_table_is_never_recreated(self):
"""表已存在就一条 DDL 都不发——这是权限被拒的唯一根治办法。"""
conn = _FakePgConn(self._CURRENT)
await _record_minimal(self._recorder(conn))
assert not self._created(conn)
async def test_create_denied_on_existing_table_keeps_recording(self):
"""就算 DDL 仍被发出并被拒,表存在时也不得判死整个 recorder。"""
conn = _FakePgConn(self._CURRENT, fail_create=True)
recorder = self._recorder(conn)
await _record_minimal(recorder) # 不得抛
assert recorder._failed is False
assert any(s.startswith("INSERT INTO llm_calls") for s in conn.statements)
async def test_missing_table_is_created_and_not_backfilled(self):
"""表不存在→建表;新建表列已齐全,不得再发补列 ALTER。"""
conn = _FakePgConn([])
recorder = self._recorder(conn)
await _record_minimal(recorder)
assert len(self._created(conn)) == 1
assert not [s for s in conn.statements if s.startswith("ALTER TABLE")]
assert recorder._failed is False
assert any(s.startswith("INSERT INTO llm_calls") for s in conn.statements)
async def test_create_failure_on_missing_table_degrades_to_noop(self):
"""表确定不存在且建不出来 = 确定写不进去: 此时才允许永久 no-op。"""
conn = _FakePgConn([], fail_create=True)
recorder = self._recorder(conn)
await _record_minimal(recorder) # 不得抛
assert recorder._failed is True
assert not [s for s in conn.statements if s.startswith("INSERT INTO llm_calls")]
async def test_probe_failure_is_transient_not_terminal(self):
"""探测失败多为连接抖动: 跳过本次,下次调用必须重试,绝不永久判死。"""
conn = _FakePgConn(self._CURRENT, probe_errors=1)
recorder = self._recorder(conn)
await _record_minimal(recorder, call_id="first") # 不得抛
assert recorder._failed is False
assert not [s for s in conn.statements if s.startswith("INSERT INTO llm_calls")]
await _record_minimal(recorder, call_id="second")
assert [s for s in conn.statements if s.startswith("INSERT INTO llm_calls")]
class _MemoryRecorder: class _MemoryRecorder:
def __init__(self): def __init__(self):
self.rows = [] self.rows = []