docs: admit governance backend failures into the scope-level error model
GovernanceBackendError arrived with the M2 distributed backends but never made it into the section 6.1 table, so it had no place in the taxonomy callers actually read. That omission is why the README missed it too. Records the reparenting, the new governance_backend_down reason, and why SourceNotConfiguredError deliberately stays outside the reparented family: a misconfigured source name must burn its failure budget and surface, not retry forever in silence.
This commit is contained in:
@@ -376,8 +376,14 @@ flowchart TB
|
|||||||
| `RequestRejectedError` | 400/请求格式错/坏输入(如不支持的图像格式) | ❌ | ❌ | ❌ |
|
| `RequestRejectedError` | 400/请求格式错/坏输入(如不支持的图像格式) | ❌ | ❌ | ❌ |
|
||||||
| `ResultInvalidError` | 调用成功但内容不可解析(JSON 修不好、ZIP 缺关键文件) | ❌(仅 D14 结构化阶梯的有界带反馈重问,不入 transport 重试计数) | ❌ | ❌(熔断记**成功**) |
|
| `ResultInvalidError` | 调用成功但内容不可解析(JSON 修不好、ZIP 缺关键文件) | ❌(仅 D14 结构化阶梯的有界带反馈重问,不入 transport 重试计数) | ❌ | ❌(熔断记**成功**) |
|
||||||
| `CircuitOpenError` / `AllSourcesExhausted` | 开路 / 全源耗尽 | 调用方决定: wait / fail-fast 可配 | — | — |
|
| `CircuitOpenError` / `AllSourcesExhausted` | 开路 / 全源耗尽 | 调用方决定: wait / fail-fast 可配 | — | — |
|
||||||
|
| `GovernanceBackendError` | 限流/熔断**状态后端自身**故障(Redis 挂等);降级方向 fail-closed,故一个请求都发不出去 | 调用方决定(同 scope 级: 延期重投) | — | — |
|
||||||
|
| `SourceNotConfiguredError` | 源名不在限流后端配置字典中——**装配缺陷**,非调用失败,正常不可达 | ❌ | ❌ | ❌ |
|
||||||
|
|
||||||
**scope 级不可用的结构化语义(2026-07-20,CHS 迁移缺口 G1;2026-07-20 M1 设计勘误修订)**: `AllSourcesExhausted`/`CircuitOpenError` 必须携带结构化字段——`retry_after_s: float`(**非可选**,承 CHS `ProviderUnavailableError` 同款,0 表示可立即重试;取各源冷却与 Retry-After 的最小值)、`reason` 枚举、`per_source_reasons: dict[str, str]`。reason 两层值域(M1 设计 §3 勘误: 本节初版所列 7 值与 CHS `errors.py:143-153` 实际值域不符,重组如下)——scope 级 `reason`: circuit_open / retry_exhausted / stalled / quota_exhausted / no_sources;`per_source_reasons` 值: network_error / timeout / rate_limited / source_dead / circuit_open / cooldown。CHS 的"scope 级不可用 → arq 延期重投、不消耗业务失败预算"(`workers/tracking.py:406-428`)依赖 `retry_after_s` 复现。
|
**scope 级不可用的结构化语义(2026-07-20,CHS 迁移缺口 G1;2026-07-20 M1 设计勘误修订)**: `AllSourcesExhausted`/`CircuitOpenError` 必须携带结构化字段——`retry_after_s: float`(**非可选**,承 CHS `ProviderUnavailableError` 同款,0 表示可立即重试;取各源冷却与 Retry-After 的最小值)、`reason` 枚举、`per_source_reasons: dict[str, str]`。reason 两层值域(M1 设计 §3 勘误: 本节初版所列 7 值与 CHS `errors.py:143-153` 实际值域不符,重组如下)——scope 级 `reason`: circuit_open / retry_exhausted / stalled / quota_exhausted / no_sources / **governance_backend_down**(2026-08-06 增,见下);`per_source_reasons` 值: network_error / timeout / rate_limited / source_dead / circuit_open / cooldown。CHS 的"scope 级不可用 → arq 延期重投、不消耗业务失败预算"(`workers/tracking.py:406-428`)依赖 `retry_after_s` 复现。
|
||||||
|
|
||||||
|
**治理后端故障归位(2026-08-06,Gitea issue #7;设计 `designs/2026-08-06-governance-backend-error-design.md`)**: `GovernanceBackendError` 自 M2 引入分布式后端时新增,但**当时未回补本表**,于是它在"调用方视角的分类学"里一直没有位置——本次归位同时补上这个遗漏。它此前是 `PolyGatewayError` 的直接子类,而语义上 fail-closed 意味着整个 scope 发不出任何请求,正是 scope 级不可用;下游只写 `except GatewayUnavailableError` 会把它落进兜底分支,导致"Redis 抖一下 → 积压任务消耗业务失败预算 → 进死信",而那是运维重启即可恢复的故障。现改为继承 `GatewayUnavailableError`,`reason` 恒为 `governance_backend_down`,`retry_after_s` 默认取常量 `GOVERNANCE_BACKEND_RETRY_AFTER_S = 5.0`——**不取 0**,因为后端恢复时间物理上不可知(不同于熔断冷却有确定到期时刻),而 0 会让积压任务零延迟同时冲击已挂掉的后端。
|
||||||
|
|
||||||
|
同批拆出 `SourceNotConfiguredError`: 限流后端 `_cfg()` 遇到源名不在配置字典中时原先也抛 `GovernanceBackendError`,但那是装配缺陷而非后端故障。若随整类归入"可延期重投",配置写错的任务会**永远重投、永不进死信、无人告警**——恰是本次要修的 bug 的镜像。故它有意留在 `GatewayUnavailableError` 之外,让缺陷消耗失败预算并浮出水面。它与四分类的关系见 §6.3 之外的第三论域说明: 四分类的论域是 transport 层翻译的**调用失败**(§6.2),scope 级不可用回答"整个 scope 还能不能用",而装配缺陷根本不该进入治理循环被"决定"。
|
||||||
|
|
||||||
### 6.2 翻译规则(transport 层职责)
|
### 6.2 翻译规则(transport 层职责)
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user