Compare commits
60 Commits
v1.0.5
...
17dcff41c3
| Author | SHA1 | Date | |
|---|---|---|---|
| 17dcff41c3 | |||
| fa4a7e220b | |||
| 658086e2c0 | |||
| 9dada0be9d | |||
| a3f4cc323f | |||
| 0edb9d397a | |||
| 484900d300 | |||
| e302247022 | |||
| 1489aab95d | |||
| c2dd4a1cf4 | |||
| 1801289277 | |||
| 7462cad166 | |||
| 3cbe8aab91 | |||
| 707f8f7317 | |||
| 10fbc5441e | |||
| 114fc8b1b3 | |||
| 4f1ab21562 | |||
| 2be89c47d8 | |||
| 7c60199680 | |||
| 2e028d38f2 | |||
| c2e9f5396c | |||
| d2cb8770df | |||
| 80bc94c42d | |||
| bb69ecf0da | |||
| 014fc2bfa7 | |||
| f3e06eac89 | |||
| a0a5cf7ecc | |||
| bc4683d1f5 | |||
| 3645e574d3 | |||
| d05114e895 | |||
| 0477d9534b | |||
| 6d0f3c9044 | |||
| 02c3d06ec6 | |||
| 573e505a4b | |||
| ce2dda7d45 | |||
| bfe423ddf8 | |||
| 9c2824ce8a | |||
| 5853c3f8ff | |||
| a57a5cea72 | |||
| 8ced49a515 | |||
| 77f9260189 | |||
| 45073486a7 | |||
| dd540496a1 | |||
| c634cab35e | |||
| 1fa91cf73d | |||
| 3a104fcce4 | |||
| a8cca51164 | |||
| b1109e9fe9 | |||
| 2f5abb6a55 | |||
| a1a9212ba1 | |||
| 7adcfff0fa | |||
| 5eb01a0096 | |||
| 48805cb9fb | |||
| 4c135075b3 | |||
| 82f4ec4910 | |||
| 89ff916bc8 | |||
| e5871cccd2 | |||
| 781579bf36 | |||
| acc419a29b | |||
| de7273598e |
+1
-1
@@ -38,7 +38,7 @@ LLM_CIRCUIT_BREAKER_COOLDOWN=60 # 或 LLM__BREAKER__COOLDOWN_S
|
|||||||
# LLM_TTFT_TIMEOUT=30 # 平铺看门狗缺省(成对生效)
|
# LLM_TTFT_TIMEOUT=30 # 平铺看门狗缺省(成对生效)
|
||||||
# LLM_INTER_TOKEN_TIMEOUT=15
|
# LLM_INTER_TOKEN_TIMEOUT=15
|
||||||
# LLM__BREAKER__PROBE_TTL_S=240 # 缺省派生: max(2×最大源超时, cooldown, 最大源超时+5);显式值须 ≥ 最大源超时+5
|
# LLM__BREAKER__PROBE_TTL_S=240 # 缺省派生: max(2×最大源超时, cooldown, 最大源超时+5);显式值须 ≥ 最大源超时+5
|
||||||
# LLM__BACKPRESSURE__STALL_WINDOW_S=300 # stall 双条件判死窗口;须 ≥ 最大源 TTFT
|
# LLM__BACKPRESSURE__STALL_WINDOW_S=300 # stall 双条件判死窗口;只计非生产性等待(429 退避/配额轮询/熔断冷却),与 TIMEOUT_S 无耦合,无需按 timeout×retries 放大
|
||||||
# LLM__BACKPRESSURE__POLL_INTERVAL_S=0.05
|
# LLM__BACKPRESSURE__POLL_INTERVAL_S=0.05
|
||||||
# LLM__SELECTOR=health_aware # health_aware(默认,M2.5) | round_robin | least_inflight
|
# LLM__SELECTOR=health_aware # health_aware(默认,M2.5) | round_robin | least_inflight
|
||||||
# ── M2.5 失败率熔断通道(可选,缺省即生产推荐值)──
|
# ── M2.5 失败率熔断通道(可选,缺省即生产推荐值)──
|
||||||
|
|||||||
+119
@@ -1,5 +1,124 @@
|
|||||||
# Changelog
|
# Changelog
|
||||||
|
|
||||||
|
## 1.2.0(2026-08-16)
|
||||||
|
|
||||||
|
网关拒绝一次调用时,**它说的话不再丢失**(issue #10)。下游一轮 1050 张医学影像的批处理里,1 张在读表格这一步收到 400、被判确定性失败而放弃;事后想知道"这张图到底哪里不合规",无从查起——响应体在 transport 翻译层之后就不存在于进程任何位置了。
|
||||||
|
|
||||||
|
根因是三条留存通道同时为空: `_status_to_error` 手上握着 `body_text` 却只用于 429 的类型细分,该模块没有任何 logger 调用,异常类也没有承载响应体的字段。而库的逐次遥测写的是 `str(exc)`,即 message——所以**只给异常加字段并不能让它进遥测表**,必须两者都做。
|
||||||
|
|
||||||
|
### 新增
|
||||||
|
|
||||||
|
- **四分类错误新增 `body_text` 字段**(加在 `PolyGatewayError` 基类): 非 2xx 响应体的摘要。与 `ResultInvalidError.raw_text` 分工明确——前者是"对方拒绝的理由"(非 2xx),后者是"2xx 但内容不可解析时的模型输出"。scope 级错误(`GatewayUnavailableError` 一族)恒为空串: 它们没有单一响应体可言。
|
||||||
|
- **同一份摘要同时进入异常 message**,故 SQLite/Postgres 遥测的 `error` 列里直接可查,下游不必为此单独埋点。
|
||||||
|
|
||||||
|
### 行为变更
|
||||||
|
|
||||||
|
- **非 2xx 的 message 末尾追加 ` | {响应体摘要}`**,覆盖两个 transport 的**全部**分支: chat 的 400 / 401·403 / 4xx 兜底 / 5xx / 429 两支(含 `insufficient_quota`),以及 OCR 的全部分支。issue 只报告了 chat 的 400,但 401 会 `force_open` 整个源、OCR 侧 message 原本只有一个状态码,是同一个缺陷的其余分支。
|
||||||
|
- 摘要口径: 先折叠空白(错误体常是缩进 JSON,原样拼进 message 会把一行日志炸成多行),再限长 **2048 字符**(对齐 Kubernetes client-go 同场景的 `maxUnstructuredResponseTextBytes`)。超长时**保留头 1400 + 尾 600**并记下省略字数——JSON 错误体的 `code` / `request_id` 收在尾部,头部硬切正好会切掉向网关方追查时唯一有用的那部分。
|
||||||
|
- 遥测 `error` 列因此变长: 纯 ASCII 约 2KB/条,最坏(5xx 重试 3 次)一次调用约 6KB。
|
||||||
|
|
||||||
|
### 不变
|
||||||
|
|
||||||
|
- **状态码 → 错误分类的映射逐条未动**(ARCHITECTURE §6.2 表),`retry_after_s` 解析、429 免重试预算、`insufficient_quota` 细分全部保持——429 的类型判定仍解析**未截断的原文**,若改用摘要,超长 body 的配额耗尽会退化成普通限速、该源不再 `force_open`。
|
||||||
|
- 异常类型树、`str(exc)` 之外的字段、遥测 22 字段与列序、DDL 全部未变。**错误面零变更**,下游 `except` 写法不受影响。
|
||||||
|
- 400 仍按确定性失败处理(不重试不换源)。**但请注意**: 经第三方中转部署时,中转自身抖动也会回 400,从状态码上与"你的输入有问题"分不开(下游实测: 同一份字节 sha256 一致、重发 15 次全部成功,失败那次 `prompt_tokens=0` 且耗时远低于任何成功调用)。库不改默认语义——直连供应商时重试只会白烧配额——但 `body_text` 现在给了下游自行区分的判据。
|
||||||
|
|
||||||
|
### 升级提示
|
||||||
|
|
||||||
|
README 的安装 pin 由 `==1.1.*` 改为 `>=1.2,<2`。**仍按 `==1.1.*` 安装的下游会静默停在 1.1.2**,拿不到本次修复且没有任何报错,请同步改自己的依赖约束。
|
||||||
|
|
||||||
|
- 打包元数据补齐: `readme` 与 `[project.urls]`。1.1.2 及之前的包在 registry 页面上**没有任何说明正文**(缺 `readme` 时 twine 只警告不阻塞),也没有仓库链接。代码零变更,自本版生效。
|
||||||
|
|
||||||
|
## 1.1.2(2026-08-07)
|
||||||
|
|
||||||
|
Postgres 遥测撞上建表权限就整体判死的问题(issue #9)。**最小权限部署会静默丢掉全部遥测**: 应用账号有表级 `INSERT`、表也已存在,但没有 schema 的 `CREATE` 权限时,初始化的 `CREATE TABLE IF NOT EXISTS` 被拒 → recorder 永久 no-op,业务调用一切正常,只留一行 warning。下游 CHSAnalyzer3 首次端到端跑的 150+ 次调用耗时/token/成本因此全部丢失,且事后无法补回。
|
||||||
|
|
||||||
|
根因是 **PostgreSQL 对 schema 的 CREATE 权限检查早于 `IF NOT EXISTS` 的存在性判断**(PG 16.14 实测: 同一连接 `INSERT` 通过、`to_regclass` 看得见表,该 DDL 照样被拒)——与 issue #3 修过的 `ALTER TABLE` 是同一类问题,当时只修了补列那一半。
|
||||||
|
|
||||||
|
### 行为变更
|
||||||
|
|
||||||
|
- **PG 侧建表前先 `to_regclass` 探测,表已存在就一条 DDL 都不发**。探测不需要任何权限,且与 `INSERT` 走同一套 search_path 解析(比裸 DDL 更准: 裸 `CREATE TABLE` 落在首个**可建**的 schema,可能与写入命中的不是同一张表)。表不存在时才建,新建表列已齐全,顺带跳过补列。
|
||||||
|
- **"结构性失能"的判据收窄为「确定写不进去」**,不再是「初始化时出过异常」。仅两种情形仍永久降级为 no-op: 建池失败(重试要在业务路径上内联吞掉连接超时)、表确定不存在且建不出来(后续 INSERT 必然全败)。探测失败、取连接失败改为**只跳过本条并 warning,下次调用重新准备**——初始化瞬间的一次抖动不再让整个进程失遥测。
|
||||||
|
- 日志措辞随之细分: `建池失败` / `建表探测失败(跳过本条,下次重试)` / `建表失败(表不存在,记录无处可落)`,原先一律是 `初始化失败`。
|
||||||
|
|
||||||
|
### 不变
|
||||||
|
|
||||||
|
- SQLite 侧**一行未改**。实测其对已存在的表在解析期就把 `CREATE TABLE IF NOT EXISTS` 短路掉(另一连接持 `BEGIN EXCLUSIVE`、文件 `chmod 444` 时该语句均通过,而同条件的 `INSERT` 分别报 database is locked / readonly database),没有同款风险;加探测零收益,故有意不对称,只在 docstring 钉死实测结论。
|
||||||
|
- 遥测端口签名、22 字段、列序、`ON CONFLICT DO NOTHING` 幂等、单条写失败逐行丢弃的降级方向全部未动。**错误面零变更**。
|
||||||
|
|
||||||
|
### 升级提示
|
||||||
|
|
||||||
|
若你的部署此前为了绕开本问题给应用账号授了 `CREATE ON SCHEMA`,现在可以收回——表存在时库不再需要该权限。
|
||||||
|
|
||||||
|
## 1.1.1(2026-08-06)
|
||||||
|
|
||||||
|
stall 判定改为非生产性等待口径(issue #8)。`timeout_s ≥ stall_window_s` 时,**一次耗满超时的请求就会让整个 scope 被判死,配置的重试次数一次都用不上**——而且没有任何报错或 warning,配置方以为自己配了 3 次重试。`stall_window_s` 默认 300 恰是个很容易被 `TIMEOUT_S` 追平的值,"只配 timeout、不配 stall"这种最常见的写法正好踩中。
|
||||||
|
|
||||||
|
根因是**两个预算重叠计费**: 真实尝试的耗时同时向重试预算(`max_attempts`)与 stall 预算(`stall_window_s`)计费,而后者更小,必然先耗尽。
|
||||||
|
|
||||||
|
### 行为变更(**请先读这一条**)
|
||||||
|
|
||||||
|
- **stall 判定的"本地超窗"条件现在只累计非生产性等待**——429 退避、配额 wait 轮询、熔断冷却;消耗重试预算的真实尝试不再计入。两个预算自此正交,划分依据是**谁消耗重试预算**: 烧 `max_attempts` 的时间不烧 `stall_window_s`,不烧 `max_attempts` 的时间(含 429 尝试本身)归 `stall_window_s` 治理。
|
||||||
|
- **`stall_window_s` 与 `timeout_s` 不再有任何耦合**,无需按 `timeout × retries` 放大。若你此前为绕开本 bug 把 `STALL_WINDOW_S` 调大过,现在可以回到默认值。
|
||||||
|
- **单次调用的最坏耗时由 `stall_window_s` 抬升到约 `max_attempts × timeout_s`**(默认配置下 3 × `TIMEOUT_S`,再加各次退避)。这是重试预算恢复生效的正确表现,但如果你的上游有调用超时,请据此复核。429 路径同样不突破这个量级——429 虽免重试预算,但其尝试耗时计入 stall 账。
|
||||||
|
**上述量级的前提是 stall 判死能够触发**,即整个 scope 无进展(`progress_age_s() > stall_window_s`)。判死是**双条件合取**,这一条未变: 若同 scope 里其他调用仍在正常出餐,本调用会继续等待换源而不判死——这正是双条件的设计意图("别人还活着,不该因我一路不顺就宣告整个 scope 死亡")。**代价是这种情形下调用级没有硬上限**,持续遭遇慢 429 的调用可以等很久。该性质由条件 B 单独门控,**早于本次修复即如此**(旧口径实测同样无界),不是本次引入;但若你需要调用级硬上限,请在调用方用 `asyncio.wait_for` 自行设置。
|
||||||
|
- 三条治理循环(chat / embedding / ocr)口径一致。**embedding 与 ocr 此前有同一缺陷**(经"先超时一次、再遇到无可用源"触发),issue 只记录了 chat 路径。
|
||||||
|
- 遥测收尾属"真实尝试"边界之内,**遥测抖动不会把一次调用推进 stalled 判决**。
|
||||||
|
|
||||||
|
### 不变
|
||||||
|
|
||||||
|
- 双条件判死的结构、`progress_age_s()` 的 `inf` 语义(从未出餐 = 全局超窗)、429 免预算、退避与 jitter 公式、`fail_fast` 分支、`AllSourcesExhausted` 的字段与 `reason` 取值(仍是 `stalled`)全部未动。**错误面零变更**,下游 `except` 写法不受影响。
|
||||||
|
- 装配期校验 `stall_window_s ≥ 最大源 ttft_timeout_s` 保留。新口径下它已是保守冗余(TTFT 等待属生产性时间),但无害且不误拒合理配置。
|
||||||
|
|
||||||
|
## 1.1.0(2026-08-06)
|
||||||
|
|
||||||
|
治理后端故障归位为 scope 级不可用(issue #7)。限流/熔断的状态后端(Redis 等)自身故障时,库按降级方向铁律 fail-closed——**整个 scope 一个请求都发不出去**,语义上就是"scope 级暂时不可用"。但 `GovernanceBackendError` 此前是 `PolyGatewayError` 的直接子类,只写 `except GatewayUnavailableError` 的调用方接不住,后果很具体: Redis 抖一下,积压任务一批批消耗业务失败预算,够到上限就进死信——**而那是运维重启一下就好的故障**。
|
||||||
|
|
||||||
|
### 行为变更(**请先读这一条**)
|
||||||
|
|
||||||
|
- **`GovernanceBackendError` 现在能被 `except GatewayUnavailableError` 捕获。** 它改为继承该类,`reason` 恒为新增的 `governance_backend_down`。**下游对后端故障的处置路线因此改变**: 从"落进兜底分支、按业务失败处置"变为"按 scope 级不可用延期重投、不消耗失败预算"。这正是本次修复的目标,但升级前请确认下游的兜底分支没有依赖旧行为(例如靠它触发告警)。既有的 `except GovernanceBackendError` **继续有效**——加父类是扩大捕获面,不是破坏。
|
||||||
|
- **配置写错(源名与限流后端配置不匹配)现在抛 `SourceNotConfiguredError` 而非 `GovernanceBackendError`。** 该类**有意不在** `GatewayUnavailableError` 之下: 那是装配缺陷不是暂时故障,必须消耗失败预算、进死信、让人看见。若随整类归入可重投家族,配置写错的任务会永远重投且无人告警——恰是本次要修的 bug 的镜像。
|
||||||
|
- **`GovernanceBackendError` 的构造签名增加必填 keyword `scope`。** 库内 20 处构造点已全部更新;若下游有自行构造该异常的代码(罕见)需同步补 `scope`。
|
||||||
|
|
||||||
|
### 新增
|
||||||
|
|
||||||
|
- **`SourceNotConfiguredError`**(公共导出)。源名不在限流后端配置字典中时抛出,正常不可达,属装配缺陷。
|
||||||
|
- **`GOVERNANCE_BACKEND_RETRY_AFTER_S = 5.0`**,`GovernanceBackendError.retry_after_s` 的默认值。**不是环境配置项**——后端恢复时间物理上不可知(不同于熔断冷却有确定到期时刻),故取保守固定值。**不取 0**: 那会让积压任务零延迟同时冲击已挂掉的后端,把一次故障放大成一场风暴。
|
||||||
|
- **scope 级 `reason` 值域增 `governance_backend_down`**(由 5 值扩为 6 值)。
|
||||||
|
- **README 新增"哪些异常会到达调用方"两列表**。四分类里 `TransientError` / `SourceDeadError` 被重试循环接住、耗尽时包成 `AllSourcesExhausted`,**根本到不了调用方**,而这只看类型树与 docstring 读不出来——曾让下游据此写错整段设计文档。
|
||||||
|
|
||||||
|
### 下游请读
|
||||||
|
|
||||||
|
- **`GovernanceBackendError` 现携带 `scope` / `reason` / `retry_after_s`**,与 `AllSourcesExhausted` 同款(`per_source_reasons` 属性存在但恒为 `{}`——后端故障不针对具体某个源);`str(exc)` 仍是原来的诊断串(如 `限流后端 try_acquire 失败: ...`),结构化字段与诊断信息并存,排障不受影响。
|
||||||
|
- **五条闸门路径**的后端故障会到达调用方: `QuotaGate` 的 `try_acquire` / `stats` / `progress_age_s`,`BreakerGate` 的 `try_enter` / `retry_after_s`。记账路径(`record_success` / `record_failure` / `release_probe` / `mark_progress`)仍被 `_record_quietly` 降级为 warning,这个分工不变。
|
||||||
|
- **CHSAnalyzer 迁移**: `tracking.py` 一条 `except GatewayUnavailableError` 即覆盖完整,无需为后端故障单列分支(`migrations/chsanalyzer.md` G1 已补注)。
|
||||||
|
|
||||||
|
## 1.0.6(2026-08-02)
|
||||||
|
|
||||||
|
推理开关能力建模与 `reasoning_tokens` 采集。`enable_thinking=False` 此前对 `minimax` / `openai` 两类源**完全不产生效果**——两个 profile 的 thinking 两档皆为空字典,`payload.update({})` 是空操作,而配置方以为关掉了推理。这比"不提供这个开关"更危险:不提供的话调用方会去找别的办法,提供了但静默失效,调用方就带着一个错误的前提往下走。一个下游项目正卡在这上面。
|
||||||
|
|
||||||
|
### 行为变更(**请先读这一条**)
|
||||||
|
|
||||||
|
- **MiniMax 源的 `ENABLE_THINKING` 从"无效"变为"生效"。** 经实测,MiniMax 认的开关是 `reasoning_effort` 而非 `enable_thinking` / `thinking`(后两者被静默丢弃);现在 `False` 注入 `reasoning_effort: none`、`True` 注入 `medium`。此前依赖"设了 false 但其实没关"这一实际行为的调用方,行为会变。
|
||||||
|
- **`MiniMax-M2.7` / `MiniMax-M2.5` 配 `ENABLE_THINKING=false` 会在装配期报错。** 这两个模型的推理**关不掉**,是模型固有属性(三种参数形态各 15 轮实测全部无效,OpenRouter 与 models.dev 两个外部注册表独立登记为强制推理)。调用方要的是"不推理"的语义保证,给不了就必须说,而不是装出一个骗人的 client。
|
||||||
|
- **`provider=openai` 的源配任何非 `None` 的 `ENABLE_THINKING` 会在装配期报错。** 该段名实践中被复用为任意 OpenAI 兼容厂商的兜底,向未知厂商下发厂商方言参数会 400。要控制推理请 `register_provider` 注册形态,或用 `SourceConfig.extra_body` 直接下发。
|
||||||
|
- **`enable_thinking` 进入缓存指纹。** 它现在真的改变请求体,不进指纹就会出现"关掉推理后重启读到开着推理时的旧响应"。**配了该项的 scope 会有一次性冷启动**;未配的 scope 指纹字面量逐字不变,不受影响。
|
||||||
|
|
||||||
|
### 新增
|
||||||
|
|
||||||
|
- **`LLMResponse` / `TransportResult` 新增 `reasoning_tokens: int | None`**(issue #6)。推理 token 已计入 `completion_tokens`,故**成本总额一直是对的**——这不是计费缺口,是归因缺口:缺了它,"这次调用花的钱里有多少花在推理上"无法区分。
|
||||||
|
- **遥测表 `llm_calls` 新增 `reasoning_tokens` 列**,`TelemetryRecorder` 端口由 21 字段扩为 22;补列纪律与 issue #3/#4 逐字相同(排末尾、先探测再 ALTER、失败只逐行降级)。
|
||||||
|
- **`ProviderProfile` 的 thinking 两档类型放宽为 `Mapping | None`**,三值语义互不重叠:`{...}` 已知注入片段 / `{}` 已知无需注入 / `None` **未知**。空字典曾同时承载后两种含义,那正是本次 bug 的根因。
|
||||||
|
- **新增 model 级能力表** `ThinkingCapability` / `DEFAULT_CAPABILITIES` / `get_capability` / `register_capability`,以及单一判定函数 `resolve_thinking`。形态(参数长什么样)按 provider 变、数年不变一次;能力(能否关闭)按 model 变、每代都变——provider 级的表在物理上表达不了同厂代际差异。每条登记都附实测证据与日期。
|
||||||
|
|
||||||
|
### 下游请读
|
||||||
|
|
||||||
|
- **`reasoning_tokens` 的 `None` 是"本次调用未上报",不是"该源不上报"**,与 `cached_prompt_tokens` 的 NULL 语义**不同**。中转网关在上游不返回 usage 时会用本地 tokenizer 补算并整体替换 usage 对象,把 `completion_tokens_details` 一并吃掉(实测同一请求 10 轮呈 6:4 双峰)。故判据须写 `in (None, 0)`;**写 `== 0` 的条件永远不成立**——实测三家供应商在未推理时都是整个 details 缺失,无人上报字面 `0`。
|
||||||
|
- **不要用输出长度反推是否发生了推理。** 两档的 `completion_tokens` 分布是重叠的(实测关闭档最高 46、开启档最低 13),按阈值判两个方向都会误判。唯一可靠的判别量是 `reasoning_tokens`。
|
||||||
|
- **`enable_thinking=True` 对 MiniMax 映射到 `medium` 档。** 它是五档旋钮而库给的是布尔开关,这个映射是库做的选择:`medium` 对应"厂商正常强度",与 qwen 的 `enable_thinking:true`、deepseek 的 `thinking:{enabled}` 同为"不指定预算、由模型自定"的语义。要精确控制档位用 `extra_body={"reasoning_effort": "..."}`,它的优先级高于 profile 注入。
|
||||||
|
- **未登记的模型不会被挡住**,按 provider 形态尽力注入并发一条 warning。新模型上线不该被库拦下,但也不该假装成功;实测后请用 `register_capability` 登记。
|
||||||
|
- **`pricing.py` 一行未改。** 推理 token 已含在 `completion_tokens` 内,单列计价即重复计费。
|
||||||
|
|
||||||
## 1.0.5(2026-07-31)
|
## 1.0.5(2026-07-31)
|
||||||
|
|
||||||
采样参数透传(issue #4)。`chat()` 此前没有任何途径设置 `temperature` / `seed` / `max_tokens`——全库检索 `temperature` 零命中,`ChatRequest.overlay` 虽会被并进请求体却只由结构化中间件填充,调用方够不着。对受控实验而言这是阻塞性的:解码温度未知且可能随供应商默认值变化,每格配置跑 5 个 seed 报出的标准差无从解释。
|
采样参数透传(issue #4)。`chat()` 此前没有任何途径设置 `temperature` / `seed` / `max_tokens`——全库检索 `temperature` 零命中,`ChatRequest.overlay` 虽会被并进请求体却只由结构化中间件填充,调用方够不着。对受控实验而言这是阻塞性的:解码温度未知且可能随供应商默认值变化,每格配置跑 5 个 seed 报出的标准差无从解释。
|
||||||
|
|||||||
@@ -28,6 +28,12 @@ make ci # 只读验证(check + test)
|
|||||||
|
|
||||||
> **档位原则(Fable 5 适配,2026-07 调研决策)**: 约束"边界与验收",不规定思考步骤。强制档(MANDATORY)是硬门;其余由模型按 skill description 自判,自判标准是任务实质(规模/风险/是否触及公共承诺),不是省事。硬边界(reference/ 只读、危险命令、提交质量门)由 `.claude/settings.json` 注册的 hooks **确定性执行**,不依赖提示词自觉。
|
> **档位原则(Fable 5 适配,2026-07 调研决策)**: 约束"边界与验收",不规定思考步骤。强制档(MANDATORY)是硬门;其余由模型按 skill description 自判,自判标准是任务实质(规模/风险/是否触及公共承诺),不是省事。硬边界(reference/ 只读、危险命令、提交质量门)由 `.claude/settings.json` 注册的 hooks **确定性执行**,不依赖提示词自觉。
|
||||||
|
|
||||||
|
> [!CRITICAL]
|
||||||
|
> **执行模式: subagent 与 Codex 一律前台(2026-08-06 人类指令)**
|
||||||
|
> 一切 subagent(verifier、`subagent-driven-development` 执行器、Explore 等)与 Codex 调用**必须前台运行**——`Agent` 工具传 `run_in_background: false`,`/codex:rescue` 带 `--wait`,**禁止**后台派发后继续做别的事。
|
||||||
|
> **理由(实测教训)**: 后台完成通知不可靠——管道会掩盖真实退出码(`pytest ... | tail` 让失败跑报成 exit 0),等待脚本的 `pgrep -f` 会自匹配成死循环,于是出现"任务早完成却没人知道"和"任务挂了也没人知道"两种失败,且两种都以"看起来还在跑"的形态呈现,无法从外部区分。前台运行牺牲并行度换取状态确定性,这个交换在本项目是划算的。
|
||||||
|
> **同一理由适用于长跑命令**: 需要后台跑时(如全套件测试),命令末尾**不得接管道**,否则退出码失真;要判完成用 `wait`/轮询 PID,不要用会匹配到自身的 `pgrep -f "<完整命令串>"`。
|
||||||
|
|
||||||
### Phase 1: 规划与设计
|
### Phase 1: 规划与设计
|
||||||
1. 涉及**公共 API、端口签名、架构边界、新子系统**的变更**必须**调用 `brainstorming`(产出 2-3 备选方案+权衡)并经**人类确认**后实施;其余任务自判(判据: 是否改变库对下游的承诺)。动手前查阅 `research-wiki/`(单一事实源)。
|
1. 涉及**公共 API、端口签名、架构边界、新子系统**的变更**必须**调用 `brainstorming`(产出 2-3 备选方案+权衡)并经**人类确认**后实施;其余任务自判(判据: 是否改变库对下游的承诺)。动手前查阅 `research-wiki/`(单一事实源)。
|
||||||
2. 功能产生运行时数据时**必须**调用 `structured-logging`。
|
2. 功能产生运行时数据时**必须**调用 `structured-logging`。
|
||||||
@@ -72,6 +78,31 @@ make ci # 只读验证(check + test)
|
|||||||
### 4.4 Git 工作流
|
### 4.4 Git 工作流
|
||||||
- 一切开发在 feature 分支,严禁直改 main;频繁语义化提交;提交**必须**调用 `commit` skill;大改动前先提交回滚点。
|
- 一切开发在 feature 分支,严禁直改 main;频繁语义化提交;提交**必须**调用 `commit` skill;大改动前先提交回滚点。
|
||||||
|
|
||||||
|
### 4.4.1 发布流程(每步都是历史欠账换来的,不得跳步)
|
||||||
|
|
||||||
|
> [!CRITICAL]
|
||||||
|
> **发布 = 合并 + push + tag + 构建 + 上传 registry。只 bump 版本号不叫发布。**
|
||||||
|
> 教训: 1.0.6 与 1.1.0 都完成了版本号 bump 与 CHANGELOG,却从未上传,registry 长期停在 1.0.5——下游 `pip install` 拿不到任何修复,且无人发现。
|
||||||
|
|
||||||
|
按顺序执行,**构建之前**必须先改完所有文档:
|
||||||
|
|
||||||
|
| # | 动作 | 要点 |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | **更新 README** | 打包会把当时的 README 固化进 sdist,**发布后再改就来不及了**(包里那份永远是旧的)。逐项核对: 安装命令的版本约束(`==1.1.*` 这类**极易漏改**,漏了下游就被锁在旧版)、能力表是否覆盖新行为、数字型断言是否仍成立(如遥测字段数,须用 `inspect.signature` 实测而非凭记忆) |
|
||||||
|
| 2 | CHANGELOG 定版 | "未发布" → `## X.Y.Z(日期)` |
|
||||||
|
| 3 | 版本号 | `pyproject.toml` + `src/polygateway/__init__.py` 两处必须一致 |
|
||||||
|
| 4 | 合并 main + push | `--no-ff`;合并后在 main 上重跑 `make lint` 与全套件 |
|
||||||
|
| 5 | **打 tag 并 push** | `git tag -a vX.Y.Z -m "..."` + `git push origin vX.Y.Z`。历史上多个版本漏打 |
|
||||||
|
| 6 | 构建 | `rm -rf dist && python -m build && python -m twine check dist/*` |
|
||||||
|
| 7 | **上传 registry** | 凭据在 `~/.config/tea/config.yml`(tea CLI 的 Gitea token,**不在** `~/.pypirc`);token 走 `TWINE_PASSWORD` 环境变量,不进命令行<br>`TWINE_USERNAME=iomgaa TWINE_PASSWORD=$TOKEN python -m twine upload --repository-url https://gitea.iomgaa.online/api/packages/iomgaa/pypi dist/*` |
|
||||||
|
| 8 | **验证已发布** | `pip download --no-deps --index-url .../pypi/simple/ "polygateway==X.Y.Z"`,并解包确认新代码在内。**不验证不算发布完成** |
|
||||||
|
| 9 | **建 Release + 核对包页面** | `POST /api/v1/repos/iomgaa/PolyGateway/releases`(body 取 CHANGELOG 本版段;历史上只打 tag 不建 release,Releases 页长期为空);随后打开包页面确认有正文与仓库链接,**Link to a repository 只能在网页手动做**(该实例的 link API 返 404) |
|
||||||
|
|
||||||
|
> [!CRITICAL]
|
||||||
|
> **发布完成的判据是外部可见结果,不是本地步骤跑通**: 收尾必须以下游视角逐一打开产物页面(registry 包页面正文与仓库链接、仓库 Releases 页、`pip install` 后包内文件),看到什么算什么,缺的当场补进本清单——1.1.2 三步全绿却出现包页面空白(`pyproject` 缺 `readme`)、Releases 页 0 条、包未挂仓库。
|
||||||
|
|
||||||
|
Gitea 包 registry 是 **owner 级**(`/iomgaa/-/packages/`)不是仓库级;PyPI 元数据不含仓库字段,故不会自动挂到 `PolyGateway/packages`,需在包页面手动 Link to a repository。
|
||||||
|
|
||||||
### 4.5 配置管理
|
### 4.5 配置管理
|
||||||
- 工程配置走 `pydantic-settings` + `.env`(模板 `.env.example`,敏感项不提交);严禁硬编码默认值;缺失关键配置直接报错。
|
- 工程配置走 `pydantic-settings` + `.env`(模板 `.env.example`,敏感项不提交);严禁硬编码默认值;缺失关键配置直接报错。
|
||||||
- 多源命名约定 `{SCOPE}__{PROVIDER}__{N}__{FIELD}`;韧性参数键名沿用三项目习惯(`LLM_TIMEOUT` 等),降低迁移成本。
|
- 多源命名约定 `{SCOPE}__{PROVIDER}__{N}__{FIELD}`;韧性参数键名沿用三项目习惯(`LLM_TIMEOUT` 等),降低迁移成本。
|
||||||
|
|||||||
@@ -1,4 +1,4 @@
|
|||||||
.PHONY: install test lint format check ci wiki
|
.PHONY: install test lint format check ci wiki wiki-check
|
||||||
|
|
||||||
ENV := PolyGateway
|
ENV := PolyGateway
|
||||||
|
|
||||||
@@ -24,3 +24,10 @@ ci: check test
|
|||||||
|
|
||||||
wiki:
|
wiki:
|
||||||
conda run -n $(ENV) python3 .claude/tools/research_wiki.py rebuild_index research-wiki/
|
conda run -n $(ENV) python3 .claude/tools/research_wiki.py rebuild_index research-wiki/
|
||||||
|
|
||||||
|
# 用户文档站(Gitea Wiki)与源码的机械对齐校验。wiki 是独立仓库,须显式给路径:
|
||||||
|
# make wiki-check WIKI=~/PolyGateway.wiki
|
||||||
|
# 不并入 ci: 仓库里没有 wiki,自动跳过等于静默降级(违 P5),宁可让人显式跑。
|
||||||
|
wiki-check:
|
||||||
|
@test -n "$(WIKI)" || (echo "用法: make wiki-check WIKI=<PolyGateway.wiki 克隆路径>" && exit 1)
|
||||||
|
conda run -n $(ENV) python3 tools/check_wiki_alignment.py --wiki $(WIKI)
|
||||||
|
|||||||
@@ -15,9 +15,10 @@
|
|||||||
| 错误分类重试 | 一切失败落入四分类(见下),由分类决定重试/换源/熔断;429 属 pushback 不消耗重试预算;退避含 jitter 且尊重 Retry-After |
|
| 错误分类重试 | 一切失败落入四分类(见下),由分类决定重试/换源/熔断;429 属 pushback 不消耗重试预算;退避含 jitter 且尊重 Retry-After |
|
||||||
| 熔断 | 双通道(连续失败 + 失败率窗口,健康证据抑制误熔);半开单探针带租约(持有者死亡自动回收);epoch fencing 拒绝迟到写回;开路时长指数递增 |
|
| 熔断 | 双通道(连续失败 + 失败率窗口,健康证据抑制误熔);半开单探针带租约(持有者死亡自动回收);epoch fencing 拒绝迟到写回;开路时长指数递增 |
|
||||||
| 自适应并发 | AIMD:429 削减、成功缓升,防止打爆上游 |
|
| 自适应并发 | AIMD:429 削减、成功缓升,防止打爆上游 |
|
||||||
|
| 背压与判死 | 配额满可选等待或快速失败;等待期按双条件判死(本地非生产性等待与全局无进展**同时**超窗)。stall 窗口只计**非生产性**等待(429 退避/配额轮询/熔断冷却),与 `TIMEOUT_S` 无耦合 |
|
||||||
| 响应缓存 | Redis/内存;key 含 model + messages 摘要 + namespace/租户 + salt,多模态 content 先摘要再 hash(防毒化);可 per-call 绕过(科研重采样) |
|
| 响应缓存 | Redis/内存;key 含 model + messages 摘要 + namespace/租户 + salt,多模态 content 先摘要再 hash(防毒化);可 per-call 绕过(科研重采样) |
|
||||||
| 流式看门狗 | TTFT / inter-token / 总超时三层活性;thinking token 刷活性不计结果;截断流(缺 `[DONE]`)判瞬时不入缓存 |
|
| 流式看门狗 | TTFT / inter-token / 总超时三层活性;thinking token 刷活性不计结果;截断流(缺 `[DONE]`)判瞬时不入缓存 |
|
||||||
| 遥测与成本 | 每次调用(含缓存命中与失败)必录 18 字段;SQLite / Postgres 后端;按价格表折算成本;多模态内容摘要落库不存原图 |
|
| 遥测与成本 | 每次调用(含缓存命中与失败)必录 22 字段;SQLite / Postgres 后端(表已存在时**不需要** schema 建表权限,最小权限账号可直接用);按价格表折算成本落库(注意 `LLMResponse.cost` 本身恒为 `None`,成本只进遥测);多模态内容摘要落库不存原图 |
|
||||||
| 结构化输出 | json_repair 修复 / 原生 schema 双策略 + 校验失败有界带反馈重问 |
|
| 结构化输出 | json_repair 修复 / 原生 schema 双策略 + 校验失败有界带反馈重问 |
|
||||||
| OCR | MonkeyOCR 双端点(文本转录 + 版面解析),bbox 数值防御下沉,逐源健康预检 `check_health()` |
|
| OCR | MonkeyOCR 双端点(文本转录 + 版面解析),bbox 数值防御下沉,逐源健康预检 `check_health()` |
|
||||||
| Embedding | 分批、维度校验、与 chat 同一治理栈 |
|
| Embedding | 分批、维度校验、与 chat 同一治理栈 |
|
||||||
@@ -30,7 +31,7 @@
|
|||||||
|
|
||||||
```bash
|
```bash
|
||||||
pip install --extra-index-url https://gitea.iomgaa.online/api/packages/iomgaa/pypi/simple/ \
|
pip install --extra-index-url https://gitea.iomgaa.online/api/packages/iomgaa/pypi/simple/ \
|
||||||
"polygateway[redis,postgres,structured]==1.0.*"
|
"polygateway[redis,postgres,structured]>=1.2,<2"
|
||||||
```
|
```
|
||||||
|
|
||||||
核心仅依赖 `httpx` + `pydantic`;按需选 extras:
|
核心仅依赖 `httpx` + `pydantic`;按需选 extras:
|
||||||
@@ -124,6 +125,23 @@ except RequestRejectedError:
|
|||||||
|
|
||||||
预算耗尽/全源熔断时抛 `GatewayUnavailableError` 族(`CircuitOpenError` / `AllSourcesExhausted`),携带 `scope` / `reason` / `retry_after_s` / `per_source_reasons`,供任务队列做延期重投。
|
预算耗尽/全源熔断时抛 `GatewayUnavailableError` 族(`CircuitOpenError` / `AllSourcesExhausted`),携带 `scope` / `reason` / `retry_after_s` / `per_source_reasons`,供任务队列做延期重投。
|
||||||
|
|
||||||
|
**网关拒绝的理由不会丢失**(1.2.0 起):非 2xx 的响应体经折叠与截断后同时进入异常 message 与 `exc.body_text`,故遥测表的 `error` 列里就能看到网关的原话——不必再为查一次 400 单独埋点。截断保头保尾(总长 2048 字符),JSON 错误体尾部的 `code` / `request_id` 不会被切掉。**经中转部署时请注意**:第三方中转服务自身抖动也会回 400,从状态码上与"你的输入有问题"无法区分;库仍按确定性失败处理(直连供应商时重试只会白烧配额),批处理下游宜据 `body_text` 自备兜底分类。
|
||||||
|
|
||||||
|
### 哪些异常会到达调用方
|
||||||
|
|
||||||
|
上表的"库内行为"一列描述的是**治理动作**,不是调用方要处理的东西。四类里有两类**根本到不了调用方**——它们被重试循环接住,预算耗尽时统一包成 `AllSourcesExhausted`。这个区分只看类型树和 docstring 是读不出来的,曾让下游据此写错整段设计文档,故在此列明:
|
||||||
|
|
||||||
|
| 会到达调用方 | 库内吸收(不必 catch) |
|
||||||
|
|---|---|
|
||||||
|
| `GatewayUnavailableError` 族——`CircuitOpenError` / `AllSourcesExhausted` / `GovernanceBackendError` | `TransientError`(退避后换源重试,耗尽即转为 `AllSourcesExhausted`) |
|
||||||
|
| `RequestRejectedError` | `SourceDeadError`(立即熔断该源并换源,同上) |
|
||||||
|
| `ResultInvalidError` | |
|
||||||
|
| `SourceNotConfiguredError` | |
|
||||||
|
|
||||||
|
**`GovernanceBackendError` 属于第一列**: 限流/熔断的状态后端(如 Redis)自身故障时库 fail-closed——一个请求都发不出去,这就是"整个 scope 暂时不可用"。它继承 `GatewayUnavailableError`,所以 §4 那段 `except GatewayUnavailableError` 一条即覆盖完整,无需为它单列分支。`retry_after_s` 默认 5 秒(后端恢复时间不可知,取 0 会让积压任务零延迟冲击已挂掉的后端)。
|
||||||
|
|
||||||
|
**`SourceNotConfiguredError` 有意不在第一列的族内**: 源名不在限流后端的配置字典中是**装配缺陷**而非暂时故障,它应当消耗失败预算、进死信、让人看见——归入可重投家族只会让配置写错的任务永远重投且无人告警。
|
||||||
|
|
||||||
## 配置参考
|
## 配置参考
|
||||||
|
|
||||||
配置只有两条装配路径:`from_env()`(读 `.env`/环境变量)或构造函数全量注入(测试/高级);库内部任何组件不自读环境变量。键名全集见 [.env.example](.env.example),约定速览:
|
配置只有两条装配路径:`from_env()`(读 `.env`/环境变量)或构造函数全量注入(测试/高级);库内部任何组件不自读环境变量。键名全集见 [.env.example](.env.example),约定速览:
|
||||||
|
|||||||
+9
-1
@@ -4,8 +4,11 @@ build-backend = "setuptools.build_meta"
|
|||||||
|
|
||||||
[project]
|
[project]
|
||||||
name = "polygateway"
|
name = "polygateway"
|
||||||
version = "1.0.5"
|
version = "1.2.0"
|
||||||
description = "PolyGateway:实验室统一的大语言模型(LLM/VLM/OCR)调度与中转库——多源、限流、重试、熔断、缓存、遥测"
|
description = "PolyGateway:实验室统一的大语言模型(LLM/VLM/OCR)调度与中转库——多源、限流、重试、熔断、缓存、遥测"
|
||||||
|
# registry 包页面的正文只认这一项:缺了页面就是一片空白(1.1.2 的教训,twine 会警告
|
||||||
|
# long_description missing 但不阻塞上传)。README 在打包时被固化进产物,发布后再改无效。
|
||||||
|
readme = "README.md"
|
||||||
requires-python = ">=3.11"
|
requires-python = ">=3.11"
|
||||||
dependencies = [
|
dependencies = [
|
||||||
"httpx>=0.27",
|
"httpx>=0.27",
|
||||||
@@ -31,6 +34,11 @@ dev = [
|
|||||||
"import-linter>=2.0",
|
"import-linter>=2.0",
|
||||||
]
|
]
|
||||||
|
|
||||||
|
[project.urls]
|
||||||
|
Homepage = "https://gitea.iomgaa.online/iomgaa/PolyGateway"
|
||||||
|
Changelog = "https://gitea.iomgaa.online/iomgaa/PolyGateway/src/branch/main/CHANGELOG.md"
|
||||||
|
Issues = "https://gitea.iomgaa.online/iomgaa/PolyGateway/issues"
|
||||||
|
|
||||||
[tool.setuptools.packages.find]
|
[tool.setuptools.packages.find]
|
||||||
where = ["src"]
|
where = ["src"]
|
||||||
|
|
||||||
|
|||||||
@@ -376,8 +376,14 @@ flowchart TB
|
|||||||
| `RequestRejectedError` | 400/请求格式错/坏输入(如不支持的图像格式) | ❌ | ❌ | ❌ |
|
| `RequestRejectedError` | 400/请求格式错/坏输入(如不支持的图像格式) | ❌ | ❌ | ❌ |
|
||||||
| `ResultInvalidError` | 调用成功但内容不可解析(JSON 修不好、ZIP 缺关键文件) | ❌(仅 D14 结构化阶梯的有界带反馈重问,不入 transport 重试计数) | ❌ | ❌(熔断记**成功**) |
|
| `ResultInvalidError` | 调用成功但内容不可解析(JSON 修不好、ZIP 缺关键文件) | ❌(仅 D14 结构化阶梯的有界带反馈重问,不入 transport 重试计数) | ❌ | ❌(熔断记**成功**) |
|
||||||
| `CircuitOpenError` / `AllSourcesExhausted` | 开路 / 全源耗尽 | 调用方决定: wait / fail-fast 可配 | — | — |
|
| `CircuitOpenError` / `AllSourcesExhausted` | 开路 / 全源耗尽 | 调用方决定: wait / fail-fast 可配 | — | — |
|
||||||
|
| `GovernanceBackendError` | 限流/熔断**状态后端自身**故障(Redis 挂等);降级方向 fail-closed,故一个请求都发不出去 | 调用方决定(同 scope 级: 延期重投) | — | — |
|
||||||
|
| `SourceNotConfiguredError` | 源名不在限流后端配置字典中——**装配缺陷**,非调用失败,正常不可达 | ❌ | ❌ | ❌ |
|
||||||
|
|
||||||
**scope 级不可用的结构化语义(2026-07-20,CHS 迁移缺口 G1;2026-07-20 M1 设计勘误修订)**: `AllSourcesExhausted`/`CircuitOpenError` 必须携带结构化字段——`retry_after_s: float`(**非可选**,承 CHS `ProviderUnavailableError` 同款,0 表示可立即重试;取各源冷却与 Retry-After 的最小值)、`reason` 枚举、`per_source_reasons: dict[str, str]`。reason 两层值域(M1 设计 §3 勘误: 本节初版所列 7 值与 CHS `errors.py:143-153` 实际值域不符,重组如下)——scope 级 `reason`: circuit_open / retry_exhausted / stalled / quota_exhausted / no_sources;`per_source_reasons` 值: network_error / timeout / rate_limited / source_dead / circuit_open / cooldown。CHS 的"scope 级不可用 → arq 延期重投、不消耗业务失败预算"(`workers/tracking.py:406-428`)依赖 `retry_after_s` 复现。
|
**scope 级不可用的结构化语义(2026-07-20,CHS 迁移缺口 G1;2026-07-20 M1 设计勘误修订)**: `AllSourcesExhausted`/`CircuitOpenError` 必须携带结构化字段——`retry_after_s: float`(**非可选**,承 CHS `ProviderUnavailableError` 同款,0 表示可立即重试;取各源冷却与 Retry-After 的最小值)、`reason` 枚举、`per_source_reasons: dict[str, str]`。reason 两层值域(M1 设计 §3 勘误: 本节初版所列 7 值与 CHS `errors.py:143-153` 实际值域不符,重组如下)——scope 级 `reason`: circuit_open / retry_exhausted / stalled / quota_exhausted / no_sources / **governance_backend_down**(2026-08-06 增,见下);`per_source_reasons` 值: network_error / timeout / rate_limited / source_dead / circuit_open / cooldown。CHS 的"scope 级不可用 → arq 延期重投、不消耗业务失败预算"(`workers/tracking.py:406-428`)依赖 `retry_after_s` 复现。
|
||||||
|
|
||||||
|
**治理后端故障归位(2026-08-06,Gitea issue #7;设计 `designs/2026-08-06-governance-backend-error-design.md`)**: `GovernanceBackendError` 自 M2 引入分布式后端时新增,但**当时未回补本表**,于是它在"调用方视角的分类学"里一直没有位置——本次归位同时补上这个遗漏。它此前是 `PolyGatewayError` 的直接子类,而语义上 fail-closed 意味着整个 scope 发不出任何请求,正是 scope 级不可用;下游只写 `except GatewayUnavailableError` 会把它落进兜底分支,导致"Redis 抖一下 → 积压任务消耗业务失败预算 → 进死信",而那是运维重启即可恢复的故障。现改为继承 `GatewayUnavailableError`,`reason` 恒为 `governance_backend_down`,`retry_after_s` 默认取常量 `GOVERNANCE_BACKEND_RETRY_AFTER_S = 5.0`——**不取 0**,因为后端恢复时间物理上不可知(不同于熔断冷却有确定到期时刻),而 0 会让积压任务零延迟同时冲击已挂掉的后端。
|
||||||
|
|
||||||
|
同批拆出 `SourceNotConfiguredError`: 限流后端 `_cfg()` 遇到源名不在配置字典中时原先也抛 `GovernanceBackendError`,但那是装配缺陷而非后端故障。若随整类归入"可延期重投",配置写错的任务会**永远重投、永不进死信、无人告警**——恰是本次要修的 bug 的镜像。故它有意留在 `GatewayUnavailableError` 之外,让缺陷消耗失败预算并浮出水面。它与四分类的关系见 §6.3 之外的第三论域说明: 四分类的论域是 transport 层翻译的**调用失败**(§6.2),scope 级不可用回答"整个 scope 还能不能用",而装配缺陷根本不该进入治理循环被"决定"。
|
||||||
|
|
||||||
### 6.2 翻译规则(transport 层职责)
|
### 6.2 翻译规则(transport 层职责)
|
||||||
|
|
||||||
@@ -390,6 +396,10 @@ flowchart TB
|
|||||||
| **空补全**: 200 且流程完整([DONE]/usage 正常)但 content 为空(2026-07-20 M1 验证发现,人类裁决) | `TransientError`(服务抖动,重试/换源;绝不缓存空响应) |
|
| **空补全**: 200 且流程完整([DONE]/usage 正常)但 content 为空(2026-07-20 M1 验证发现,人类裁决) | `TransientError`(服务抖动,重试/换源;绝不缓存空响应) |
|
||||||
| 解析层失败(结构化输出/OCR ZIP) | `ResultInvalidError` |
|
| 解析层失败(结构化输出/OCR ZIP) | `ResultInvalidError` |
|
||||||
|
|
||||||
|
**响应体留存(2026-08-16,Gitea issue #10;设计 `designs/2026-08-16-issue10-error-body-retention-design.md`)**: 上表每一条 HTTP 翻译**都必须携带响应体摘要**——摘要同时进入异常 message 与 `PolyGatewayError.body_text`(1.2.0 新增基类字段)。两者都要,因为逐次遥测写的是 `str(exc)`,只加字段进不了遥测表,而"事后可查"正是这条要求的目的。摘要口径由 `transports/_http_errors.summarize_body` 单点实现(折叠空白 → 限长 2048 字符 → 超长保留头 1400 + 尾 600 并记省略字数),两个 transport 共用,**不得各写一份**——issue #10 的成因正是"只有 429 那一支用了响应体"。`body_text` 是旁路数据,不参与任何治理判定;`_translate_429` 的类型细分仍解析未截断原文(摘要会破坏 JSON,改用它会让超长 body 的 `insufficient_quota` 退化成普通限速)。
|
||||||
|
|
||||||
|
**400 在中转拓扑下的语义提醒**(同上): 第三方 API 中转服务自身抖动时也会回 400,从状态码上与供应商的"输入非法"无法区分(下游实测: 同一份字节重发 15 次全成功,失败那次 `prompt_tokens=0`、耗时远低于任何成功调用,即请求在推理开始前被挡)。本表**不改** 400 → `RequestRejectedError` 的映射——直连供应商时重试只会白烧配额,且改默认语义等于让所有直连用户为一种部署形态买单;库改为把判据(`body_text`)交给下游自行区分。
|
||||||
|
|
||||||
### 6.3 "坏结果 ≠ 坏服务"(ResultInvalidError 语义,继承 CHSAnalyzer)
|
### 6.3 "坏结果 ≠ 坏服务"(ResultInvalidError 语义,继承 CHSAnalyzer)
|
||||||
|
|
||||||
由输入内容决定的**确定性失败**(这张图就是解析不出表格、这段输出就是修不成 JSON):服务是健康的,换源重试只会白烧配额。因此熔断器记成功、不换源、异常上抛消耗业务侧的失败预算。出处:`CHSAnalyzer governance.py:237-239`。
|
由输入内容决定的**确定性失败**(这张图就是解析不出表格、这段输出就是修不成 JSON):服务是健康的,换源重试只会白烧配额。因此熔断器记成功、不换源、异常上抛消耗业务侧的失败预算。出处:`CHSAnalyzer governance.py:237-239`。
|
||||||
@@ -421,8 +431,9 @@ flowchart TB
|
|||||||
- `RedisLimiter`: 移植 CHSAnalyzer 六道闸——单条 Lua 原子检查全局并发/单源并发(ZSET 租约)/全局 RPM/单源 RPM/全局 TPM/单源 TPM;窗口 id 用 **Redis 服务器时钟**(TIME 命令)统一多进程口径。随实现移植契约测试。
|
- `RedisLimiter`: 移植 CHSAnalyzer 六道闸——单条 Lua 原子检查全局并发/单源并发(ZSET 租约)/全局 RPM/单源 RPM/全局 TPM/单源 TPM;窗口 id 用 **Redis 服务器时钟**(TIME 命令)统一多进程口径。随实现移植契约测试。
|
||||||
- `InMemoryLimiter`: 同一契约的进程内实现(semaphore + 滑动窗口计数);单进程场景下语义等价。
|
- `InMemoryLimiter`: 同一契约的进程内实现(semaphore + 滑动窗口计数);单进程场景下语义等价。
|
||||||
- **配额满行为可配**: `wait`(等待,配 stall 判定——本地等待超窗 + 全局无进展超窗双条件才判卡死)或 `fail-fast`(立即抛)。
|
- **配额满行为可配**: `wait`(等待,配 stall 判定——本地等待超窗 + 全局无进展超窗双条件才判卡死)或 `fail-fast`(立即抛)。
|
||||||
|
- **stall 计时口径(2026-08-06 修正,issue #8,设计 `designs/2026-08-06-issue8-stall-budget-design.md`)**: 双条件的**条件 A 只累计非生产性等待**(429 退避、配额 wait 轮询、熔断冷却),真实尝试的耗时由 `StallClock.attempting()` 从 stall 账中扣除。原实现用墙钟总耗时,使真实尝试同时向重试预算与 stall 预算计费;而 stall 预算(默认 300s)小于重试预算(`max_attempts × timeout_s`),必然先耗尽——`timeout_s ≥ stall_window_s` 时一次超时即判 scope 死,`max_attempts` **静默失效**。修正后两个预算正交,**划分依据是"谁消耗重试预算"而非"是否发出请求"**: 烧 `max_attempts` 的时间不烧 `stall_window_s`,不烧 `max_attempts` 的时间归 `stall_window_s`。**429 尝试因此也计入 stall 账**——它免重试预算,若其耗时又算生产性就两个预算都不烧,排队型网关(持满 timeout 才回 429)下调用可挂 25 小时(实施期独立验证实测,见设计 §3.6)。生产性边界即 `_attempt` 边界(含该次记账与遥测收尾),故遥测抖动不参与判死。`stall_window_s` 与 `timeout_s` 自此**无耦合**,无需按 `timeout × retries` 放大。三条治理循环(chat/embedding/ocr)共用 `middleware/retry.py` 的 `StallClock`。条件 B 的 `inf` 语义未动——新口径下"非生产性排队耗满窗口且 scope 从未出餐"判死本就正当。**残余性质(非本次引入,由条件 B 单独门控)**: 判死是双条件合取,故当同 scope 其他调用仍在正常出餐时本调用不判死(设计意图: 别人还活着就不该宣告 scope 死亡),代价是**该情形下调用级无硬上限**——持续遭遇慢 429 的调用可以等很久;需要硬上限的调用方应自行 `asyncio.wait_for`。
|
||||||
- 全局活性信号: `mark_progress()`/`progress_age_s()`("最近一次出餐"时刻)供背压 stall 判定,移植 `CHSAnalyzer limiter.py:193`。
|
- 全局活性信号: `mark_progress()`/`progress_age_s()`("最近一次出餐"时刻)供背压 stall 判定,移植 `CHSAnalyzer limiter.py:193`。
|
||||||
- **契约补强(2026-07-20,CHS 迁移缺口 G6)**: `settle()`/`release()` 幂等(重复调用无副作用);装配期守卫——`timeout_s ≤ permit 租约 TTL`(防租约先于请求过期)、`stall_window ≥ 最慢源 TTFT 上限`(防误判卡死),违反直接报错拒绝装配。降级方向细化(2026-07-20 M1): "报错不放行"适用于**准入侧**(try_acquire/try_enter 及选源路径消费的 source_stats/retry_after_s);已成功调用后的 settle/release 释放侧失败降级 warning——释放失败不构成放行,且不得掩盖主异常与取消。**勘误(2026-07-20 M2 设计,人类批准)**: 记账侧的 `record_success`/`record_failure`/`mark_progress` 同归此类——调用已真实完成,后端失败若冒泡会丢弃真实成功响应或掩盖原始尝试异常,故降级 warning(CHS 原版一律报错,此为有意反转;丢一次熔断记账最多延迟状态迁移且方向偏保守,epoch fencing 防污染)。
|
- **契约补强(2026-07-20,CHS 迁移缺口 G6)**: `settle()`/`release()` 幂等(重复调用无副作用);装配期守卫——`timeout_s ≤ permit 租约 TTL`(防租约先于请求过期)、`stall_window ≥ 最慢源 TTFT 上限`(防误判卡死;**issue #8 后为保守冗余**——TTFT 等待属生产性时间已不计入 stall,该误判在机制上不再可能,校验保留因其无害且不误拒合理配置),违反直接报错拒绝装配。降级方向细化(2026-07-20 M1): "报错不放行"适用于**准入侧**(try_acquire/try_enter 及选源路径消费的 source_stats/retry_after_s);已成功调用后的 settle/release 释放侧失败降级 warning——释放失败不构成放行,且不得掩盖主异常与取消。**勘误(2026-07-20 M2 设计,人类批准)**: 记账侧的 `record_success`/`record_failure`/`mark_progress` 同归此类——调用已真实完成,后端失败若冒泡会丢弃真实成功响应或掩盖原始尝试异常,故降级 warning(CHS 原版一律报错,此为有意反转;丢一次熔断记账最多延迟状态迁移且方向偏保守,epoch fencing 防污染)。
|
||||||
|
|
||||||
### 7.4 熔断
|
### 7.4 熔断
|
||||||
|
|
||||||
@@ -468,7 +479,7 @@ flowchart TB
|
|||||||
|
|
||||||
**`sampling` 列(2026-07-31,issue #4,端口 20 → 21)**: 列语义 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。三个 emit 入口口径必须各自定死,否则同一列在不同行含义不同: `emit_attempt`(RetryMW 调用,**唯一**有生效源者)并上 `source.extra_body`;`emit_cache_hit` / `emit_terminal_failure`(TelemetryMW 最外层调用)无 source 可言,只记调用级——与 `model`/`source_name` 在终态行置空是同一先例,且缓存命中行无损(`sampling` 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同)。三者统一读 `request.sampling` 而非 `request.overlay`(后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处未被污染,直接用必然三行分叉)。OCR/embedding 路径因决策 G 剥离 `extra_body`,该列恒 NULL。
|
**`sampling` 列(2026-07-31,issue #4,端口 20 → 21)**: 列语义 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。三个 emit 入口口径必须各自定死,否则同一列在不同行含义不同: `emit_attempt`(RetryMW 调用,**唯一**有生效源者)并上 `source.extra_body`;`emit_cache_hit` / `emit_terminal_failure`(TelemetryMW 最外层调用)无 source 可言,只记调用级——与 `model`/`source_name` 在终态行置空是同一先例,且缓存命中行无损(`sampling` 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同)。三者统一读 `request.sampling` 而非 `request.overlay`(后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处未被污染,直接用必然三行分叉)。OCR/embedding 路径因决策 G 剥离 `extra_body`,该列恒 NULL。
|
||||||
|
|
||||||
(`cached_prompt_tokens`/`model_reported` 为 2026-07-31 issue #3 新增,端口由 18 字段扩为 20;两个后端在初始化期对已存在的旧表幂等补列——`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入都被逐行 warning 丢弃。补列一律**先探测缺列再 ALTER**(`ADD COLUMN IF NOT EXISTS` 即使列已存在也先取 ACCESS EXCLUSIVE 锁,而遥测内联 await,锁共享审计表会拖垮业务调用),且**失败只逐行降级、绝不置结构性失能标志**。新列在 DDL 里必须排在 `created_at` **之后**,与 `ALTER TABLE ADD COLUMN` 的追加位置一致,否则新建库与升级库的物理列序分叉)。链路: `session_id`/`parent_call_id` 由调用方传入贯穿(agent step → LLM call)。`messages` 落库前对多模态 part 先摘要(与缓存 key 共用同一摘要函数,§7.5)——Video-Tree 现状 base64 整段进 SQLite 导致 db 膨胀(`llm.py:330`),库内修复(2026-07-20,VT 迁移缺口 R12)。
|
(`cached_prompt_tokens`/`model_reported` 为 2026-07-31 issue #3 新增,端口由 18 字段扩为 20;两个后端在初始化期对已存在的旧表幂等补列——`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入都被逐行 warning 丢弃。补列一律**先探测缺列再 ALTER**(`ADD COLUMN IF NOT EXISTS` 即使列已存在也先取 ACCESS EXCLUSIVE 锁,而遥测内联 await,锁共享审计表会拖垮业务调用),且**失败只逐行降级、绝不置结构性失能标志**。**建表同理(2026-08-07,issue #9)**: PG 对 schema 的 CREATE 权限检查早于 `IF NOT EXISTS` 的存在性判断(16.14 实测,只授表级 `SELECT, INSERT` 的角色写得进去却建不了表),故 PG 侧必须**先 `to_regclass` 探测、表在就不发 DDL**;SQLite 侧实测在解析期即短路(持排他锁/只读文件下该语句均通过),无同款风险,**有意不加探测**。由此把"结构性失能"的判据从「初始化时出过异常」收窄为「确定写不进去」——仅建池失败与"表确定不存在且建不出来"判死,探测/取连接失败只跳过本次并留待下次重试。新列在 DDL 里必须排在 `created_at` **之后**,与 `ALTER TABLE ADD COLUMN` 的追加位置一致,否则新建库与升级库的物理列序分叉)。链路: `session_id`/`parent_call_id` 由调用方传入贯穿(agent step → LLM call)。`messages` 落库前对多模态 part 先摘要(与缓存 key 共用同一摘要函数,§7.5)——Video-Tree 现状 base64 整段进 SQLite 导致 db 膨胀(`llm.py:330`),库内修复(2026-07-20,VT 迁移缺口 R12)。
|
||||||
|
|
||||||
- 后端: `SQLiteRecorder`(默认;WAL + busy_timeout、`INSERT OR IGNORE` 幂等、`asyncio.to_thread` 桥接、初始化/写入失败全降级不冒泡)与 `PostgresRecorder`。
|
- 后端: `SQLiteRecorder`(默认;WAL + busy_timeout、`INSERT OR IGNORE` 幂等、`asyncio.to_thread` 桥接、初始化/写入失败全降级不冒泡)与 `PostgresRecorder`。
|
||||||
- **单一 helper 铁律**: 遥测调用点收敛为一个内部函数/上下文管理器;Video-Tree 与 GovDoc 各有 4-5 处逐字复制的 `record_llm_call(15 个参数)` 是本条的直接教训。
|
- **单一 helper 铁律**: 遥测调用点收敛为一个内部函数/上下文管理器;Video-Tree 与 GovDoc 各有 4-5 处逐字复制的 `record_llm_call(15 个参数)` 是本条的直接教训。
|
||||||
|
|||||||
@@ -0,0 +1,270 @@
|
|||||||
|
---
|
||||||
|
type: design
|
||||||
|
node_id: design:2026-08-02-thinking-capability-design
|
||||||
|
title: "推理开关能力建模与 reasoning_tokens 采集(issue #5 + #6)"
|
||||||
|
date: 2026-08-02
|
||||||
|
---
|
||||||
|
|
||||||
|
# 推理开关能力建模与 reasoning_tokens 采集(issue #5 + #6)
|
||||||
|
|
||||||
|
> 类型:design|日期:2026-08-02|状态:待人类确认
|
||||||
|
> 事实基础见 `findings/2026-08-02-thinking-switch-and-reasoning-tokens.md`(本文所有实测引用均出自该文)。
|
||||||
|
> 本设计经 2026-08-02 充分讨论后直接给出单一方案,不列备选。
|
||||||
|
|
||||||
|
## 1. 问题
|
||||||
|
|
||||||
|
**issue #5——静默失效。** `SourceConfig.enable_thinking` 是给上层的统一推理开关,靠 `providers.py` 的 `ProviderProfile.thinking_on/thinking_off` 落地。`minimax` 与 `openai` 两格皆为空 dict,`_build_payload` 的 `payload.update({})` 是空操作:`enable_thinking=False` 对这两类源**完全不产生效果**,而配置方以为关掉了。
|
||||||
|
|
||||||
|
这不是理论缺陷。`dissect/.env:84,99` 两个 scope 均写 `ENABLE_THINKING=false`,并在 `:67-70` 记为明确阻塞项——Phase-0 要求关闭思维链以隔离变量。
|
||||||
|
|
||||||
|
**issue #6——归因缺口。** `usage.completion_tokens_details.reasoning_tokens` 未被采集。成本总额正确(推理 token 已含在 `completion_tokens` 内),但"本次调用有多少钱花在推理上"无法区分,而这正是 dissect 要测的因子的主要成本通道。
|
||||||
|
|
||||||
|
**两者的耦合。** #6 是 #5 的验收仪器:修完 #5 后判断"这次是否真的没推理",靠正文长度不可靠,靠 `reasoning_content` 也不行(MiniMax 非流式恒为空、正文无 `<think>` 标签)。因此 **#6 先落地,#5 的测试断言它**。
|
||||||
|
|
||||||
|
## 2. 根因
|
||||||
|
|
||||||
|
空 dict 同时承载了两种语义:「本 provider 无需注入任何参数」与「我们不知道本 provider 怎么表达」。二者混同,就只能靠"表里没有 = 不发"兜底,静默失效随之产生。
|
||||||
|
|
||||||
|
更深一层:`ProviderProfile` 的注册单位是 **provider**,而"能否关闭推理"是 **model** 的属性。实测证明同一 provider 内部代际差异是决定性的——MiniMax-M3 可关,M2.7 / M2.5 **固有不可关**(三种参数形态实测全部无效,OpenRouter 与 models.dev 独立登记为 mandatory)。provider 级的表在物理上表达不了这件事。
|
||||||
|
|
||||||
|
业界佐证:注册单位下沉到 model 级的(LiteLLM、models.dev、LangChain、OpenRouter、Helicone)都有显式失败通道;仍停在 provider 级的(Portkey、LlamaIndex)恰是失败语义最差的两家,均静默丢弃。**注册粒度与失败语义是同一个问题的两面。**
|
||||||
|
|
||||||
|
## 3. 决策摘要
|
||||||
|
|
||||||
|
| # | 决策 |
|
||||||
|
|---|---|
|
||||||
|
| D1 | **形态留 provider 级,能力下沉 model 级**。形态 = 参数长什么样(数年不变);能力 = 能否关闭(每代都变) |
|
||||||
|
| D2 | **「未知 / 不支持 / 不干预」必须是三个不同的值**,落在三个不同层次 |
|
||||||
|
| D3 | **遇到"关不掉"的模型报错,不静默放行**;报错在装配期,请求期兜底 |
|
||||||
|
| D4 | **「开」的默认档定 `medium`,允许 per-source 覆盖**(经已有 `extra_body`,不新增字段) |
|
||||||
|
| D5 | `enable_thinking` **纳入缓存指纹**(配套,必做) |
|
||||||
|
| D6 | `reasoning_tokens` 的文档措辞为「**本次调用**未上报」,非「该源未上报」(配套,必做) |
|
||||||
|
|
||||||
|
D4 的依据:业界对「开」映射到哪一档**无语义共识**(LiteLLM 用 2 的幂、OpenRouter 用百分比、Helicone 一律折半),唯一的工程共识是**该映射必须是可覆盖的常量**。选 `medium` 是因为 qwen 的 `enable_thinking:true` 与 deepseek 的 `thinking:{enabled}` 都不指定预算、由模型自定,`medium` 是五档中语义最接近"厂商正常强度"的一档;选 `high` 等于库替所有下游做"加钱换质量"的业务判断,违反零业务假设。
|
||||||
|
|
||||||
|
## 4. 数据模型
|
||||||
|
|
||||||
|
### 4.1 形态层(provider 级)
|
||||||
|
|
||||||
|
`ProviderProfile` 两档由 `dict` 放宽为 `dict | None`:
|
||||||
|
|
||||||
|
| 值 | 含义 | 当前实例 |
|
||||||
|
|---|---|---|
|
||||||
|
| `{...}` | 已知的注入片段 | qwen / deepseek / minimax |
|
||||||
|
| `{}` | 已知**无需注入**即处于该档 | 无(保留为自然零值) |
|
||||||
|
| `None` | **未知**:库不知道该 provider 如何表达 | `openai` 两档 |
|
||||||
|
|
||||||
|
```python
|
||||||
|
"minimax": ProviderProfile(
|
||||||
|
name="minimax",
|
||||||
|
thinking_on={"reasoning_effort": "medium"},
|
||||||
|
thinking_off={"reasoning_effort": "none"},
|
||||||
|
strip_think_tags=False,
|
||||||
|
),
|
||||||
|
"openai": ProviderProfile(
|
||||||
|
name="openai", thinking_on=None, thinking_off=None, strip_think_tags=False,
|
||||||
|
),
|
||||||
|
```
|
||||||
|
|
||||||
|
`openai` 填 `None` 而非补 `reasoning_effort`,理由是该段名在实践中已被复用为**任意 OpenAI 兼容厂商的兜底**(`dissect/.env:116` 把 `kimi-k3` 挂在 `provider=openai` 下)。向未知厂商下发 `reasoning_effort` 会招致 400;标为未知则让误配在装配期显式暴露。真·OpenAI 推理模型的使用者走 `register_provider`——这正是 D11 承诺的"新 provider = 一个条目"。
|
||||||
|
|
||||||
|
qwen / deepseek 两条实测正确,**不动**。
|
||||||
|
|
||||||
|
### 4.2 能力层(model 级,新增)
|
||||||
|
|
||||||
|
```python
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ThinkingCapability:
|
||||||
|
"""某个具体模型的推理能力(model 级);登记必须附实测证据与日期。"""
|
||||||
|
can_disable: bool
|
||||||
|
evidence: str
|
||||||
|
```
|
||||||
|
|
||||||
|
登记表键为模型名精确匹配,**只登记在用的模型**,未登记即"未知"并走退化路径:
|
||||||
|
|
||||||
|
| 模型 | `can_disable` | 证据 |
|
||||||
|
|---|---|---|
|
||||||
|
| `MiniMax-M3` | `True` | 2026-08-02 实测 N=10,`reasoning_effort=none` 稳定关闭 |
|
||||||
|
| `MiniMax-M2.7` | `False` | 三形态各 N=3 全无效;OpenRouter `mandatory:true` |
|
||||||
|
| `MiniMax-M2.5` | `False` | 同上 |
|
||||||
|
| `qwen3.7-plus` | `True` | 实测 `enable_thinking=false` 关闭 |
|
||||||
|
| `deepseek-v4-pro` | `True` | 实测 `thinking:{disabled}` 关闭 |
|
||||||
|
|
||||||
|
注入方式沿用 D11 的纯函数注册纪律:`get_capability(model, *, table=None)` 与 `register_capability(...)` 返回新表,经 `capabilities` 参数注入,与现有 `registry` 参数同形,**不引入模块级可变状态**。
|
||||||
|
|
||||||
|
**不引入 models.dev / LiteLLM 的 JSON 作为运行时依赖**——违反依赖极简与纯 asyncio 中立(import 期发网络请求)。二者仅作为写表时的对照参考;本次三条 MiniMax 实测与它们的登记 100% 吻合,这本身就是表可信的旁证。
|
||||||
|
|
||||||
|
### 4.3 三个值的层次归属(D2)
|
||||||
|
|
||||||
|
| 语义 | 载体 | 层次 |
|
||||||
|
|---|---|---|
|
||||||
|
| **不干预**(调用方不表态) | `SourceConfig.enable_thinking is None` | 调用方意图 |
|
||||||
|
| **未知**(库不知道怎么表达) | `ProviderProfile` 该档为 `None` | 形态层 |
|
||||||
|
| **不支持**(模型做不到) | `ThinkingCapability.can_disable is False` | 能力层 |
|
||||||
|
|
||||||
|
三者不可互相替代:不干预是意图缺失,未知是知识缺失,不支持是能力缺失。当前实现把后两者塌缩成空 dict,是 issue #5 的根因。
|
||||||
|
|
||||||
|
## 5. 判定与失败语义(D3)
|
||||||
|
|
||||||
|
单一判定函数收口,形态层与能力层在此相遇:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def resolve_thinking(profile, capability, enable_thinking) -> Mapping[str, Any]:
|
||||||
|
"""三态 + 两层能力 → 注入片段;不可满足时 ValueError(由调用点翻译为领域错误)。"""
|
||||||
|
```
|
||||||
|
|
||||||
|
真值表:
|
||||||
|
|
||||||
|
| # | 条件 | 行为 |
|
||||||
|
|---|---|---|
|
||||||
|
| R1 | `enable_thinking is None` | 不注入。与 `False` 严格区分 |
|
||||||
|
| R2 | 形态层该档为 `None` | **报错**,文案指路 `register_provider` 或 `extra_body` |
|
||||||
|
| R3 | `enable_thinking is False` 且 `can_disable is False` | **报错**:调用方要的是"不推理"的语义保证,给不了必须说 |
|
||||||
|
| R4 | 模型未登记(能力未知) | 按形态层注入 + `loguru.warning`,不阻断 |
|
||||||
|
| R5 | 其余 | 按形态层注入 |
|
||||||
|
|
||||||
|
R3 与 R4 的极性相反,这是刻意的,借鉴 LiteLLM 的两极性纪律:**"关不掉"用错的后果是下游带着错误前提做实验(opt-in,从严);"未登记"多为新模型上线(opt-out,从宽)**,误拒会让库成为升级路上的绊脚石。
|
||||||
|
|
||||||
|
### 5.1 报错位置:两处,共用同一份判定
|
||||||
|
|
||||||
|
| 位置 | 异常 | 覆盖 |
|
||||||
|
|---|---|---|
|
||||||
|
| `client.py:from_settings`(`:248` 已在此解析 profiles) | `ValueError`(装配期) | `from_env` / `from_settings` 两条工厂路径,即 90% 场景 |
|
||||||
|
| `OpenAICompatTransport` | `RequestRejectedError`(四分类之一,不重试不换源) | 构造函数全量注入路径 |
|
||||||
|
|
||||||
|
这不是重复判定:`get_provider` 现在就是同一形态(`client.py:248` + `openai_compat.py:313`)。双点校验的必要性来自 issue #1 的教训——**装配守卫必须任何构造路径都生效**。
|
||||||
|
|
||||||
|
**绝不在 `_build_payload` 里抛裸 `ValueError`**:该处位于 RetryMW 内侧,裸异常不属错误四分类、`TelemetryMW` 也不捕,会导致一行遥测都没有就逃出 `chat()`。
|
||||||
|
|
||||||
|
## 6. reasoning_tokens 采集(issue #6)
|
||||||
|
|
||||||
|
照搬 issue #3 的 `_coerce_cached_tokens` 形态:只收非负整数,显式排除 `bool`(`isinstance(True, int)` 为真,放行会把 `True` 记成 1)。
|
||||||
|
|
||||||
|
`LLMResponse` / `TransportResult` **尾部**各加 `reasoning_tokens: int | None = None`——字段顺序是公共承诺(`types.py:1-5`),只增不删不改名。
|
||||||
|
|
||||||
|
流式与非流式对称取值:`completion_tokens_details` 在最后的 usage 帧里,`missing_done="salvage"` 打捞路径拿不到时记 `None` 而非 `0`(现有代码天然满足:`sink` 无 usage 时 `_coerce_*` 返回 `None`)。
|
||||||
|
|
||||||
|
**`pricing.py` 一行不改**:推理 token 已含在 `completion_tokens` 内,单列计价即重复计费。这是归因缺口,不是计费缺口。
|
||||||
|
|
||||||
|
**缓存路径无需改动**:`CacheMW._rehydrate` 按 `_RESPONSE_FIELDS` 动态过滤(`cache.py:28,133`),旧条目缺该字段自动落 `None`,语义正确。
|
||||||
|
|
||||||
|
### 6.1 语义澄清(D6)
|
||||||
|
|
||||||
|
实测三家在未推理时都是**整个 `completion_tokens_details` 对象缺失**,无一上报 `0`。且 new-api 在上游不返回 usage 时会用本地 tokenizer 补算并整体替换 usage,把 ctd 一并吃掉(实测同一请求 10 轮呈 6:4 双峰)。因此:
|
||||||
|
|
||||||
|
- docstring 写「**本次调用**未上报」,**不可**写「该源未上报」
|
||||||
|
- 下游判据必须是 `reasoning_tokens in (None, 0)`,写 `== 0` 的条件永远不成立
|
||||||
|
- 这三句要同时进 docstring、CHANGELOG 与 wiki
|
||||||
|
|
||||||
|
## 7. 缓存指纹配套(D5)
|
||||||
|
|
||||||
|
`build_model_fingerprint`(`client.py:63-80`)当前只摘要 `(model, extra_body)`。#5 一旦让 thinking 真正改变请求体,就会出现"关掉推理后重启读到开着推理时的旧缓存"——issue #4 为 `temperature` 写过逐字相同的理由。
|
||||||
|
|
||||||
|
做法:marks 的判据由 `if s.extra_body` 扩为 `if s.extra_body or s.enable_thinking is not None`,摘要对象并入该值。**全源不配 `enable_thinking` 时字面量与现值逐字相同,不触发存量缓存冷启动**;dissect 会有一次性冷启动,这是正确行为(旧缓存来自推理开着的调用)。
|
||||||
|
|
||||||
|
## 8. 落点清单
|
||||||
|
|
||||||
|
| 文件 | 改动 |
|
||||||
|
|---|---|
|
||||||
|
| `providers.py` | 两档放宽为 `dict \| None`;填 minimax、`openai` 改 `None`;新增 `ThinkingCapability` / `DEFAULT_CAPABILITIES` / `get_capability` / `register_capability` / `resolve_thinking` |
|
||||||
|
| `transports/openai_compat.py` | `_build_payload` 两分支收敛为一行 `resolve_thinking(...)`;新增 `_coerce_reasoning_tokens`;流式 `:401` 与非流式 `:485` 填值;构造函数收 `capabilities` |
|
||||||
|
| `client.py` | `from_settings` / `from_env` 加 `capabilities`;`:248` 后加装配守卫;`build_model_fingerprint` 纳入 `enable_thinking` |
|
||||||
|
| `types.py` | `LLMResponse` / `TransportResult` 尾部加 `reasoning_tokens` |
|
||||||
|
| `middleware/retry.py` | `_build_response` 透传 |
|
||||||
|
| `ports.py` | `record_llm_call` 21 → 22 字段 |
|
||||||
|
| `telemetry/{sqlite,postgres}.py` | 建表列 + `_BACKFILL_COLUMNS` 迁移 + `_COLUMNS`,**新列排末尾**(两处注释均有明文要求) |
|
||||||
|
| `middleware/telemetry.py` | `_record` + 三个 `emit_*` 入口 |
|
||||||
|
|
||||||
|
## 9. 测试策略
|
||||||
|
|
||||||
|
本次改动的正确性**与具体模型强相关**,mock 只能验证代码路径、无法验证"这个参数在这个模型上是否真的关掉了推理"。因此核心行为**必须由真实 API 多轮调用验证**。
|
||||||
|
|
||||||
|
### 9.1 三层分工
|
||||||
|
|
||||||
|
| 层 | 内容 | 是否门控合并 |
|
||||||
|
|---|---|---|
|
||||||
|
| unit | `resolve_thinking` 真值表(R1–R5)、`_coerce_reasoning_tokens` 形态防御、注入优先级、装配守卫报错、缓存指纹变化与不变性 | **是**(CI 可跑) |
|
||||||
|
| integration | 遥测两后端新列写入与 ALTER 迁移 | **是** |
|
||||||
|
| **e2e(真实 API)** | 见 9.2 | 打 `slow` 标记被默认排除;**合并前必须 `-m slow` 真跑并存档报告** |
|
||||||
|
|
||||||
|
不让本组阻断 CI 的理由是外部不可用会误伤:实测中 kimi 渠道在 429 后被中转下线并返回 404,另有一次 `network_error` 连续三次耗尽源导致 L2 假红。让外部波动阻断合并,会把测试变成噪声源。
|
||||||
|
|
||||||
|
**实现机制**:给本组打项目既有的 `slow` 标记。`pyproject.toml` 的 `addopts = "-m 'not slow'"` 默认排除它(该配置的注释原文:「慢速测试,CI 按需跑」),合并前用 `pytest -m slow tests/e2e/test_thinking_live.py` 显式真跑。实测效果:`make ci` 由 7 分钟降至 91 秒。
|
||||||
|
|
||||||
|
**一处必须澄清的事实**:`make test` 跑的是 `pytest tests/`,**包含 `tests/e2e/`**——只要 `.env` 有凭据,既有的轻量 e2e 冒烟就会真跑。所以「e2e 不进 CI」这句对本项目**并不成立**,只有打了 `slow` 的才被排除;本节初稿写成前者,是错的。「不自动门控」也不等于「可跳过」——沿用既有口径(`tests/e2e/test_smoke_gateway.py:22` 的 reason 写着「验收前必须真跑」)。
|
||||||
|
|
||||||
|
### 9.2 e2e 覆盖矩阵
|
||||||
|
|
||||||
|
沿用既有 e2e 约定:`dotenv_values(".env")` + `pytestmark = pytest.mark.skipif(not _HAS_SOURCE, ...)`,结构化报告输出至 `tests/outputs/e2e/`。
|
||||||
|
|
||||||
|
| # | 场景 | 源 | 轮数 | 判据 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| L1 | `enable_thinking=False` | MiniMax-M3 | ≥10 | 每轮 `completion_tokens < 30` 且 `reasoning_tokens` 恒 `None` |
|
||||||
|
| L2 | `enable_thinking=True` | MiniMax-M3 | ≥10 | 多数轮 `completion_tokens > 100`;请求体实发 `reasoning_effort=medium` |
|
||||||
|
| L3 | `enable_thinking=None` | MiniMax-M3 | ≥10 | 不注入任何 thinking 参数(基线) |
|
||||||
|
| L4 | `extra_body` 覆盖 profile | MiniMax-M3 | ≥5 | 实发 `high`,profile 的 `medium` 被覆盖 |
|
||||||
|
| L5 | L1 / L2 的**流式**重跑 | MiniMax-M3 | 各 ≥10 | 同 L1 / L2(库默认 `stream=True`,这是主路径) |
|
||||||
|
| L6 | `enable_thinking=False` | qwen | ≥10 | 关闭 |
|
||||||
|
| L7 | `enable_thinking=False` | deepseek | ≥10 | 关闭 |
|
||||||
|
| L8 | **能力表漂移哨兵** | 全部登记模型 | 各 ≥5 | 实测行为与 `can_disable` 声明一致 |
|
||||||
|
| L9 | `enable_thinking=False` + M2.7 → 装配期报错 | — | — | 纯本地,无需真实调用 |
|
||||||
|
|
||||||
|
轮数由环境变量可调高,默认 ≥10。总量约 100–150 次调用。
|
||||||
|
|
||||||
|
### 9.3 三条必须遵守的测试纪律
|
||||||
|
|
||||||
|
**(a)判别量只能是 `reasoning_tokens`。**(2026-08-02 e2e 实测修正:本节初稿写的是"主判据用 `completion_tokens`",被数据推翻。)两档的输出长度分布**重叠**——关闭档实测最高 46(模型偶尔把解题过程写进正文),开启档最低 13(medium 档想得少的轮次),按长度阈值判两个方向都会误判;而 `reasoning_tokens` 在同一批 30 轮里干净分开。`completion_tokens` 仅作 `reasoning_tokens` 被中转吃掉时的退路。另配一个不含魔数的确定性锚点:关闭档 `prompt_tokens` 严格小于开启档(实测 194 < 207)。
|
||||||
|
|
||||||
|
**(b)多轮 + 计数判定,不用单轮判定。** 关闭方向要求**每轮**都满足(关掉后 `completion_tokens` 极稳定,实测 4–10);开启方向只要求**多数轮**满足(推理量方差大)。
|
||||||
|
|
||||||
|
**(c)源不可用必须跳过并显式记录为"未覆盖",不得静默计入通过。** 报告里要能一眼看出哪些矩阵行没跑到。
|
||||||
|
|
||||||
|
### 9.4 漂移哨兵(L8)的定位
|
||||||
|
|
||||||
|
能力表过期是必然事件(LiteLLM 有过 `gpt-5.1-mini` 漏登记导致误拒的真实事故)。L8 用真实调用反向校验每条登记,是这张表的**过期告警**——模型升级后若 `can_disable` 声明失真,这里会先炸。建议纳入发版前清单定期执行。
|
||||||
|
|
||||||
|
## 10. 明确不做
|
||||||
|
|
||||||
|
不为中转的观测漂移在库内加任何机制(多轮取众数、渠道探测、重试到拿到 `reasoning_tokens`)——中转路由不受请求参数影响,探测结果不可迁移,属 YAGNI 违规;该问题在运维侧解决,写入 wiki 前提。
|
||||||
|
|
||||||
|
不改 `SourceConfig` 的公开字段形态:`enable_thinking` 保持 `bool | None`。分档需求走已有的 `extra_body` / `overlay`,两条路径已进缓存 key 与 `sampling` 遥测列,新增字段则要额外接这两处,是隐藏成本。
|
||||||
|
|
||||||
|
不动 qwen / deepseek 的 profile;不碰 `pricing.py`;不引入任何新依赖。
|
||||||
|
|
||||||
|
## 11. 验收标准
|
||||||
|
|
||||||
|
1. `ENABLE_THINKING=false` + MiniMax-M3 → 请求体含 `reasoning_effort: none`,响应 `reasoning_tokens is None`,真实 API 多轮验证
|
||||||
|
2. `ENABLE_THINKING=false` + MiniMax-M2.7 → **装配期报错**,文案说明该模型无法关闭推理
|
||||||
|
3. `ENABLE_THINKING` 任意非 `None` + `provider=openai` → **装配期报错**,指路 `register_provider` / `extra_body`
|
||||||
|
4. 未登记模型 + 任意 `enable_thinking` → 正常注入 + 一条 warning
|
||||||
|
5. `extra_body={"reasoning_effort":"high"}` 仍覆盖 profile 注入
|
||||||
|
6. 流式与非流式均能采到 `reasoning_tokens`;打捞路径记 `None` 而非 `0`
|
||||||
|
7. 改 `enable_thinking` → 缓存 key 变化;不配该项的存量 scope key 逐字不变
|
||||||
|
8. 遥测两后端新列可写、旧库经 ALTER 迁移后可写
|
||||||
|
9. e2e 报告存档于 `tests/outputs/e2e/`,矩阵覆盖情况可核
|
||||||
|
|
||||||
|
每条均需"先失败后通过"的证据(测试结果门)。
|
||||||
|
|
||||||
|
## 12. 影响与风险
|
||||||
|
|
||||||
|
**这是行为变更,不是纯修复。** MiniMax 源的 `ENABLE_THINKING` 从"无效"变为"生效",CHANGELOG 须醒目标注;dissect 会有一次性缓存冷启动。
|
||||||
|
|
||||||
|
**dissect 的 Phase-0 实验设计需调整。** M2.7 上做不了"开思考 vs 关思考"的对照——这是模型固有属性,任何库层改动都无法改变。可行替代是只在 M3 上做该对照,或将因子改为"高档 vs 低档"。此结论须同步给 dissect。
|
||||||
|
|
||||||
|
**能力表的正确性依赖实测,且经中转。** 三条 MiniMax 结论均在自建 new-api 中转下取得,直连官方端点未验证;表中每条 `evidence` 须写明这一点。若下游改为直连,L8 漂移哨兵是发现失真的第一道防线。
|
||||||
|
|
||||||
|
**新增两处失败面。**(本段两次修正:初稿只列了 `openai` 那一处、遗漏 M2.x;二稿又把 M2.x 那处写成「合并即打挂 dissect」,同样不准确——见下。)
|
||||||
|
|
||||||
|
其一是 `provider=openai` + 配了 `ENABLE_THINKING`,经全仓与 dissect 检索当前无此用法(dissect 的 K3 scope 用 `provider=openai` 但未配该项)。
|
||||||
|
|
||||||
|
其二是**关不掉推理的模型 + `ENABLE_THINKING=false`**,而 `dissect/.env:80,85` 正是 `MiniMax-M2.7` + `false`。准确的影响是:**dissect 升到 1.0.6 之后**,该 scope 装配会抛 `ValueError`;它当前跑着的版本不受本次发布影响。但 `dissect/requirements.txt:7` 声明的是 `polygateway>=1.0.1,<1.1` —— 一个**范围**而非精确 pin,`1.0.6` 落在范围内,所以任何一次 `pip install -U`、重建环境或 CI 重装依赖都会**自动**装上它,无需谁刻意升级。换言之不是「突然挂」,而是「下次装依赖时挂」。
|
||||||
|
|
||||||
|
这是本设计的**预期行为**(给不了「不推理」的语义保证就必须说),dissect 侧的处置是改配置:该对照只能在 M3 上做,或把因子改为「高档 vs 低档」。
|
||||||
|
|
||||||
|
**三个参考下游零破坏**:VT / CHS / GovDoc 的 thinking 用法均为二元,本方案不改公开字段形态。
|
||||||
|
|
||||||
|
## 13. 另立 issue(不在本次范围)
|
||||||
|
|
||||||
|
`kimi-k3` 拒绝 `temperature=0`(400),而 400 归 `RequestRejectedError` 不重试不换源,下游统一下发 `temperature=0` 会导致此类源 100% 硬失败。与本次两条 issue 同源(供应商能力差异未被建模),但属采样参数域,独立处理。
|
||||||
|
|
||||||
|
`qwen` 的 `strip_think_tags=True` 已过时(实测走 `reasoning_content`,正文无 `<think>` 标签),无害死代码,可顺带清理或另记。
|
||||||
@@ -0,0 +1,181 @@
|
|||||||
|
# 治理后端故障归位为 scope 级不可用设计(Issue #7)
|
||||||
|
|
||||||
|
- **日期**: 2026-08-06
|
||||||
|
- **来源**: Gitea Issue #7(下游 CHSAnalyzer3 按异常类型分流失败,基于 1.0.1 源码核查)
|
||||||
|
- **状态**: **已批准(2026-08-06)**,待 `writing-plans`
|
||||||
|
- **触发档位**: 强制(变更 `errors.py` 公共错误类型树 = 库对下游的承诺)
|
||||||
|
- **方案范围**: 人类已选定方向 A′ 并明确要求单一方案,故本文不列平行备选,仅在 §4 记录被否决路线及否决理由
|
||||||
|
|
||||||
|
## 1. 目标与非目标
|
||||||
|
|
||||||
|
| | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| **G1** | `GovernanceBackendError` 归入 `GatewayUnavailableError` 之下,使"该延期重投的失败"在类型上闭合——调用方一条 `except GatewayUnavailableError` 覆盖完整,漏接在物理上不可能 |
|
||||||
|
| **G2** | 把混在同一类里的**装配期缺陷**("未知源")拆出去,使其**不**被误判为可重投 |
|
||||||
|
| **G3** | `retry_after_s` 取非零值,避免后端故障期间下游零延迟批量重投形成忙循环 |
|
||||||
|
| **G4** | 公开错误面文档化:README 增"会到达调用方 / 库内吸收"两列表,`ARCHITECTURE.md` §6.1 回补缺失的 `GovernanceBackendError` 行 |
|
||||||
|
| **非目标** | 不改 fail-closed 降级方向(限流/熔断后端不可用 → 报错而非放行,库铁律不动);不改后端重连/健康探测;不新增配置项;不改 `TransientError`/`SourceDeadError` 的库内吸收行为 |
|
||||||
|
|
||||||
|
### 1.1 Issue 前提的四处修正(按 1.0.6 源码核实)
|
||||||
|
|
||||||
|
| Issue 原文 | 实际情况 |
|
||||||
|
|---|---|
|
||||||
|
| 泄漏路径为 `try_enter` / `try_acquire` 两条 | **五条**(设计初稿写"三条",2026-08-06 独立验证时核出遗漏两条并订正): `QuotaGate` 的 `try_acquire` / `stats`(`retry.py:249`)/ `progress_age_s`(`retry.py:216`、`:305`),`BreakerGate` 的 `try_enter` / `retry_after_s`(`retry.py:292`、`:310`)。判据是该调用点是否被 `_record_quietly` 包裹——未包裹即直达调用方;OCR 与 Embedding 两个治理循环有同构的对应点 |
|
||||||
|
| (未提及构造点数量) | 全库 **22 处** `raise GovernanceBackendError`,分布于 4 个文件 |
|
||||||
|
| 方向 A 只需改类型树 | 其中 **2 处语义完全不同**(见 §3.4),整类归入"可重投"会制造镜像 bug |
|
||||||
|
| `retry_after_s` 取 0,「docstring 已写 0 = 可立即重试,语义上是通的」 | 语义通,**工程上不通**。见 §3.2 |
|
||||||
|
|
||||||
|
另需记录一处根因:`ARCHITECTURE.md:372-378` §6.1 的错误分类表里 `GovernanceBackendError` **一次都没出现**。它是 M2 引入分布式后端时新增的,当时未回补架构表,于是它在"调用方视角的分类学"中从来就没有位置——README 的遗漏是这个遗漏的下游后果。
|
||||||
|
|
||||||
|
## 2. 影响面的决定性前提(改动安全性的依据)
|
||||||
|
|
||||||
|
| 事实 | 证据 | 含义 |
|
||||||
|
|---|---|---|
|
||||||
|
| 库内仅一处 `except GatewayUnavailableError` | `middleware/telemetry.py:250`,写法为 `except (GatewayUnavailableError, GovernanceBackendError)` | 变成父子关系后该处由"并列捕获"退化为"父类捕获",**行为逐字不变**,库内零回归 |
|
||||||
|
| 加父类是纯扩大 | 下游既有 `except GovernanceBackendError` 全部照旧命中 | 不违反 CLAUDE.md §4.3「已被下游消费的公共类型只增不删不改名」 |
|
||||||
|
| `QuotaGate`/`BreakerGate` 是后端异常的唯一入口 | 两类 docstring 自述,三处装配 `retry.py:186` / `ocr.py:122` / `embedding.py:123` | scope 注入点收敛为 2 个类、3 处装配 |
|
||||||
|
| 三个装配点都持有 `self._scope` | `retry.py:183`、`ocr.py:116`、`embedding.py:118` | 注入无需新增上游参数传递链 |
|
||||||
|
|
||||||
|
## 3. 选定方案
|
||||||
|
|
||||||
|
### 3.1 类型树变更
|
||||||
|
|
||||||
|
`SCOPE_REASONS` 增枚举值 `governance_backend_down`;`GovernanceBackendError` 改继承 `GatewayUnavailableError`,`reason` 恒为该值(与 `CircuitOpenError` 恒为 `circuit_open` 同构,是本库已有的表达手法)。
|
||||||
|
|
||||||
|
构造签名保持"首参为 message"的位置参数形态,以免 22 处构造点与既有测试全部改写:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class GovernanceBackendError(GatewayUnavailableError):
|
||||||
|
def __init__(self, message, *, scope, retry_after_s=GOVERNANCE_BACKEND_RETRY_AFTER_S,
|
||||||
|
source_name=None):
|
||||||
|
super().__init__(scope=scope, reason="governance_backend_down",
|
||||||
|
retry_after_s=retry_after_s, source_name=source_name)
|
||||||
|
self.args = (message,) # 见 §3.5
|
||||||
|
```
|
||||||
|
|
||||||
|
`scope` 为必填 keyword(P4 显式优于隐式:它在三层调用点全部可得,给默认值只会掩盖装配疏漏)。
|
||||||
|
|
||||||
|
### 3.2 `retry_after_s` 的取值(本设计的核心权衡)
|
||||||
|
|
||||||
|
Issue 建议取 0。**否决**:下游 `schedule_retry(after_s=0)` 会立刻重投,Redis 挂掉期间队列里积压的任务将以零延迟批量重投,对着一个已经挂掉的后端打忙循环——把一次故障放大成一场风暴。这与本 issue 想修的问题同源:都是"分类正确但处置参数错误"。
|
||||||
|
|
||||||
|
已考虑并否决的两个替代取值:
|
||||||
|
|
||||||
|
| 取值 | 否决理由 |
|
||||||
|
|---|---|
|
||||||
|
| 复用 `BackpressureConfig.poll_interval_s`(与 `quota_exhausted` 同源,`retry.py:299` 有先例) | 该值只有三个装配点持有,后端层 11 处构造点拿不到;为此给 `RedisLimiter`/`RedisBreaker` 增构造参数,是让后端层去持有"重投策略"——违反 P7(决策逻辑与状态存储分离),后端只该知道"我坏了",不该知道这在治理上意味着什么 |
|
||||||
|
| 新增配置项 `PGW_GOVERNANCE_BACKEND_RETRY_AFTER_S` | YAGNI。目前无任何下游表达过需要调它;真需要时下游可完全忽略 `exc.retry_after_s` 用自有退避 |
|
||||||
|
|
||||||
|
**选定**:`errors.py` 模块级常量 `GOVERNANCE_BACKEND_RETRY_AFTER_S = 5.0`,作为构造默认值,docstring 写明理由——后端恢复时间物理上不可知(不同于熔断冷却有确定到期时刻),取一个保守固定值;下游若有自己的退避策略可忽略此值。本库对 scope 级异常硬编码语义值已有先例(`retry.py:206` 的 `no_sources` 取 `0.0`)。
|
||||||
|
|
||||||
|
它**不是环境配置项**,故不落 CLAUDE.md §4.5「严禁硬编码默认值」的论域——§4.5 约束的是 `pydantic-settings` + `.env` 管辖的工程配置(超时、并发、限额),而本常量是异常自身携带的语义默认值,与 `no_sources` 取 `0.0` 同性质。docstring 需显式写明这一点,避免后来者误加环境键。
|
||||||
|
|
||||||
|
### 3.3 `scope` 的三层来源
|
||||||
|
|
||||||
|
| 层 | 构造点数 | scope 来源 | 改动 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| `backends/redis/limiter.py` | 6 | `self._scope`(`:170`) | 补 `scope=self._scope` |
|
||||||
|
| `backends/redis/breaker.py` | 5 | `self._scope`(`:291`) | 补 `scope=self._scope` |
|
||||||
|
| `middleware/breaker.py` `BreakerGate` | 5 | **需注入** | 构造函数增 `scope: str`,三处装配传 `self._scope` |
|
||||||
|
| `middleware/ratelimit.py` `QuotaGate` | 4 | **需注入** | 同上 |
|
||||||
|
|
||||||
|
包装器对后端自抛异常的 `except GovernanceBackendError: raise` 原样放行**保持不变**——后端层已填好 scope,重建实例只会制造"同一异常构造两次"的怪味且覆盖值相同。
|
||||||
|
|
||||||
|
### 3.4 "未知源"拆分为独立错误类
|
||||||
|
|
||||||
|
`backends/memory/limiter.py:92` 与 `backends/redis/limiter.py:198` 的 `_cfg()` 在源名不在配置字典中时抛 `GovernanceBackendError`。**这不是后端故障**,是限流后端拿到的源列表与治理循环的对不上——装配期缺陷,正常不可达。
|
||||||
|
|
||||||
|
若随整类归入"延期重投、不扣失败预算",配置写错的任务将**永远重投、永远不进死信**,运维永远收不到告警——正是本 issue 要修的 bug 的镜像。
|
||||||
|
|
||||||
|
新增 `SourceNotConfiguredError(PolyGatewayError)`,**有意不放在** `GatewayUnavailableError` 之下:下游默认按"任务的错"处置 → 扣失败预算 → 进死信 → 人能看见。这是缺陷该有的可见性。该类进 `__init__.py` 公共导出(下游可选择性识别,但不识别也能得到正确处置)。
|
||||||
|
|
||||||
|
### 3.5 message 保全
|
||||||
|
|
||||||
|
`GatewayUnavailableError.__init__` 会把 message 覆盖为 `f"{scope} 网关暂时不可用: {reason}"`,而 22 处构造点携带的诊断串(如 `限流后端 try_acquire 失败: {exc}`)是排障的主要线索,不可丢。方案是 `super().__init__()` 后覆写 `self.args = (message,)`,使 `str(exc)` 仍为原诊断串,而 `scope`/`reason`/`retry_after_s` 作为结构化字段并存。父类不动——它的 message 生成逻辑对 `CircuitOpenError`/`AllSourcesExhausted` 仍然正确。
|
||||||
|
|
||||||
|
## 4. 被否决的路线
|
||||||
|
|
||||||
|
| 路线 | 否决理由 |
|
||||||
|
|---|---|
|
||||||
|
| **B: 只补文档,类型树不动** | 正确性依赖每个下游都读到那句话。本库下游不止一个,且本 issue 本身就是"文档读不出来"引发的——同一个失效模式不能用同一种药治 |
|
||||||
|
| **C: 类型树不动,在 RetryMW 边界包成 `AllSourcesExhausted`** | 比 A′ 更具破坏性:下游现有 `except GovernanceBackendError` 会直接失效。加父类是扩大,换类型是破坏 |
|
||||||
|
| **D: 后端层不再构造该异常,原始异常穿透由包装器统一翻译**(初评时倾向,已否决) | `backends/redis/limiter.py:133,151` 的 `RedisPermit.release/settle` 依赖 `except GovernanceBackendError` 实现**释放侧降级**(失败只 warning 不冒泡)。原始 redis 异常穿透后该处接不住,会破坏这条既有降级行为;改为 `except Exception` 则违反 P5 |
|
||||||
|
|
||||||
|
## 5. 行为审计(既有行为逐条标注)
|
||||||
|
|
||||||
|
| 既有行为 | 出处 | 处置 |
|
||||||
|
|---|---|---|
|
||||||
|
| 限流/熔断后端不可用 → 报错而非放行(fail-closed) | 库铁律 | **保留**,一字不改 |
|
||||||
|
| 记账路径后端故障降级为 warning | `middleware/retry.py:404` `_record_quietly` | **保留**。仅闸门路径需要到达调用方 |
|
||||||
|
| permit `release`/`settle` 失败降级 warning | `redis/limiter.py:133,151` | **保留**(§4 路线 D 因此被否决) |
|
||||||
|
| 遥测对后端故障发 `emit_terminal_failure` | `middleware/telemetry.py:250` | **保留**,父子关系后由父类分支承接,行为不变 |
|
||||||
|
| `except GovernanceBackendError: raise` 原样放行 | 包装器 9 处 | **保留** |
|
||||||
|
| "未知源"抛 `GovernanceBackendError` | `memory/limiter.py:92`、`redis/limiter.py:198` | **替换**为 `SourceNotConfiguredError`(§3.4) |
|
||||||
|
| `str(exc)` 为诊断串 | 22 处 | **保留**(§3.5 显式保全) |
|
||||||
|
|
||||||
|
## 6. 非功能维度
|
||||||
|
|
||||||
|
| 维度 | 回答 |
|
||||||
|
|---|---|
|
||||||
|
| **并发与取消** | 不适用于新增并发路径。异常构造是纯同步无状态操作,不引入共享状态。`CancelledError` 穿透路径完全不受影响——本设计不新增任何 `except` 子句,`_record_quietly` 中 `except asyncio.CancelledError`(`:402`)先于 `except GovernanceBackendError`(`:404`)的顺序不动 |
|
||||||
|
| **降级方向** | 不变。fail-closed 是本类存在的理由,本设计只改"它被归入哪一类",不改"它是否被抛出" |
|
||||||
|
| **幂等与重复** | 异常类型变更不涉及幂等性。需注意的是下游行为改变:同一次后端故障从"扣失败预算"变为"延期重投",重投次数由下游队列策略决定——这正是期望的变更,已在 CHANGELOG 行为变更段声明 |
|
||||||
|
| **持久化与原子性** | 无持久化改动。遥测落库路径(`emit_terminal_failure`)的字段与调用时机均不变 |
|
||||||
|
|
||||||
|
## 7. 错误处理与测试策略
|
||||||
|
|
||||||
|
新失败面只有一个:`SourceNotConfiguredError`,它落在四分类之外。这**不违反** CLAUDE.md §4.2「一切失败必须落入四分类」——该铁律的论域是 **transport 层翻译的调用失败**(`ARCHITECTURE.md` §6.2 的翻译规则表逐条对应 HTTP 状态码与解析失败),而本库已有一整族异常合法地处在四分类之外:`GatewayUnavailableError` / `CircuitOpenError` / `AllSourcesExhausted` 都不是四分类之一,`ARCHITECTURE.md` §6.1 把它们单列一行,因为它们回答的是另一个问题——"整个 scope 还能不能用",而非"这一次调用怎么失败的"。
|
||||||
|
|
||||||
|
`SourceNotConfiguredError` 属于第三个论域:**装配缺陷**(配置与治理循环不一致,正常不可达)。四分类决定重试/换源/熔断,而装配缺陷根本不该进入治理循环去被"决定",它应当立刻失败并让人看见。将其塞进四分类中的任何一类都会赋予它一份不该有的治理语义(如 `RequestRejectedError` 会让下游以为请求本身有问题、去修请求)。§9 Q1 保留了"复用 `RequestRejectedError`"作为备选供人类权衡。
|
||||||
|
|
||||||
|
| 测试 | 位置 | 先失败后通过的证据 |
|
||||||
|
|---|---|---|
|
||||||
|
| `GovernanceBackendError` 可被 `except GatewayUnavailableError` 接住 | `tests/unit/test_errors.py` | 改前 `pytest.raises(GatewayUnavailableError)` 必失败 |
|
||||||
|
| 闸门泄漏路径(五条,§1.1)抛出的异常携带正确 `scope` 与非零 `retry_after_s`;钉住 `try_acquire`/`try_enter`/`progress_age_s` 三条代表路径,余两条由同一注入机制覆盖 | `tests/unit/test_backpressure.py` — **三条桩都需新增**(Codex 审计划时核出: `:176-186` 是记账侧 `record_success`/`record_failure`/`mark_progress` 的降级桩,不是闸门路径;`progress_age_s` 仅 `:243-257` 覆盖包装行为、不验 scope) | 改前无 `scope` 属性,`AttributeError` |
|
||||||
|
| `str(exc)` 仍为原诊断串 | `tests/unit/test_errors.py` | 防 §3.5 回归 |
|
||||||
|
| 未知源抛 `SourceNotConfiguredError` 且**不是** `GatewayUnavailableError` | 改 `tests/unit/test_redis_key_layout.py:70-74`;内存版**当前无覆盖,需新增** | 改前抛 `GovernanceBackendError`,断言"不是 scope 级"必失败 |
|
||||||
|
| Redis 真实掉线时准入侧行为 | `tests/integration/test_redis_cross_connection.py:228-245`(真实 Redis,不 mock) | 断言由 `GovernanceBackendError` 收紧为"是 `GatewayUnavailableError` 且 `reason == governance_backend_down`" |
|
||||||
|
|
||||||
|
## 8. 影响面清单
|
||||||
|
|
||||||
|
| 类别 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| **源码** | `errors.py`(新常量+新类+继承变更)、`backends/redis/limiter.py`(7)、`backends/redis/breaker.py`(5)、`backends/memory/limiter.py`(1)、`middleware/breaker.py`(6:构造函数+5 处)、`middleware/ratelimit.py`(5)、`middleware/retry.py`/`ocr.py`/`embedding.py`(各 1 行装配)、`__init__.py`(导出新类) |
|
||||||
|
| **测试** | `tests/unit/test_errors.py`、`test_backpressure.py`、`test_redis_key_layout.py`、`tests/integration/test_redis_cross_connection.py` |
|
||||||
|
| **文档** | `README.md` §"错误模型"增两列表 + `GovernanceBackendError` 行;`ARCHITECTURE.md` §6.1 回补该类并记录本次归位;`migrations/chsanalyzer.md` G1 条目补注;`CHANGELOG.md` 1.1.0;按 `docs-convention.md` §2 同步 Gitea Wiki |
|
||||||
|
| **版本** | **1.1.0**。有行为变更(下游对后端故障的处置路线改变)但无 API 破坏(加父类是扩大),按语义化版本走 minor |
|
||||||
|
| **下游** | CHSAnalyzer3 当前在 1.0.1。升级后 `except GatewayUnavailableError` 即覆盖后端故障,其现有 `except GovernanceBackendError`(若有)继续有效,无需改代码即可获得修复 |
|
||||||
|
|
||||||
|
### 8.1 执行顺序(单一事实源纪律)
|
||||||
|
|
||||||
|
`ARCHITECTURE.md` 是架构单一事实源,`SCOPE_REASONS` 新增值域与 `GovernanceBackendError` 的归位都与其 §6.1 现状冲突。因此 **§6.1 的修订必须先于或同批于代码实现落地**,不得"先改代码、事后补文档"。具体为:人类批准本设计后,`writing-plans` 的第一项任务即为修订 `ARCHITECTURE.md` §6.1(补 `GovernanceBackendError` 与 `SourceNotConfiguredError` 行、scope 级 reason 值域增 `governance_backend_down`、记录本次归位的理由与日期),与实现同一分支、同批提交。
|
||||||
|
|
||||||
|
## 9. 待人类确认的决策点
|
||||||
|
|
||||||
|
(编号用 Q 前缀,避免与 `ARCHITECTURE.md` 的架构决策 D1–D14 混淆)
|
||||||
|
|
||||||
|
**三点均已由人类拍板(2026-08-06),全部采纳本文的选择:**
|
||||||
|
|
||||||
|
| # | 决策 | 裁定 | 被否决的备选及理由 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Q1 | "未知源"归到哪 | ✅ **拆为 `SourceNotConfiguredError`**,不在 `GatewayUnavailableError` 之下(§3.4) | ① 沿用 `GovernanceBackendError`——配置写错的任务将无限重投、永不进死信、无人发现;② 复用 `RequestRejectedError`——治理行为与选定方案**完全等价**,但名称误导:下游会去查 prompt 而非配置文件 |
|
||||||
|
| Q2 | `retry_after_s` 取值 | ✅ **常量 `5.0`**(§3.2) | 取 0 会让积压任务零延迟同时冲击已挂掉的后端,把一次故障放大成风暴 |
|
||||||
|
| Q3 | 新类是否公共导出 | ✅ **导出**(进 `__init__.py`) | 不导出则下游无法给"配置写错"单独接告警,而导出无成本 |
|
||||||
|
|
||||||
|
## 10. 审批记录
|
||||||
|
|
||||||
|
| 阶段 | 状态 |
|
||||||
|
|---|---|
|
||||||
|
| Claude 自审 | 已完成(全部结论对应本会话内 grep/read 输出;§3.5 的 `self.args` 保全机制经 conda 环境实跑验证) |
|
||||||
|
| Codex 独立审 | 已完成(2026-08-06),4 条意见逐条核验见下 |
|
||||||
|
| 人类审批 | ✅ **已批准(2026-08-06)**。方向 A′ 于设计前即由人类选定;Q1–Q3 三个决策点逐条拍板,全部采纳本文选择(见 §9)。可进入 `writing-plans` |
|
||||||
|
|
||||||
|
### 10.1 Codex 意见的核验结果
|
||||||
|
|
||||||
|
| 意见 | 判定 | 处置 |
|
||||||
|
|---|---|---|
|
||||||
|
| ARCHITECTURE §6.1 未同步前实施违反单一事实源(判为阻塞) | **实质成立**,但性质是执行顺序而非设计缺陷——§8 本已把 §6.1 回补列入影响面 | 新增 §8.1 明确"架构文档修订先于/同批于实现" |
|
||||||
|
| §6.1 错误分类表未承认 `GovernanceBackendError`(判为阻塞) | **与上条同源**,且 §1.1 已自陈此为根因 | 同上,由 §8.1 覆盖 |
|
||||||
|
| `SourceNotConfiguredError` 落在四分类外违反 §4.2 铁律(判为阻塞) | **部分成立**:铁律论域被误读——`GatewayUnavailableError` 族本就合法处在四分类之外(§6.1 单列一行)。但原文表述确会引起该疑虑 | §7 补写三个论域的划分论证;§9 Q1 增列"复用 `RequestRejectedError`"备选交人类权衡 |
|
||||||
|
| 硬编码常量与 §4.5 存在张力(建议性) | **成立** | §3.2 补写"非环境配置项"及 docstring 要求 |
|
||||||
|
| Q 编号与架构 D1–D14 混淆(建议性) | **成立** | §9 决策点编号由 `D` 改为 `Q` |
|
||||||
@@ -0,0 +1,254 @@
|
|||||||
|
# stall 判定改为非生产性等待口径设计(Issue #8)
|
||||||
|
|
||||||
|
- **日期**: 2026-08-06
|
||||||
|
- **来源**: Gitea Issue #8(本机全套件跑 391.67s,1 failed;失败源于单次 300s 超时耗尽 stall 窗口,基于 1.1.0 源码核查)
|
||||||
|
- **状态**: **已批准(2026-08-06)**,待 `writing-plans`
|
||||||
|
- **触发档位**: 强制(变更治理行为——判死条件的度量口径,是库对下游的承诺)
|
||||||
|
- **方案范围**: 人类明确要求单一方案,故本文不列平行备选,仅在 §5 记录被否决路线及否决理由(体例沿用 Issue #7 设计)
|
||||||
|
|
||||||
|
## 1. 目标与非目标
|
||||||
|
|
||||||
|
| | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| **G1** | 消除"单次超时即判 scope 级死亡"——`timeout_s` 与 `stall_window_s` 的隐式耦合彻底解除,重试预算在超时场景下真实可用 |
|
||||||
|
| **G2** | 使 stall 判定的度量对象与它的职责一致:**它治理的是无人治理的非生产性循环,不是已被重试预算治理的真实尝试** |
|
||||||
|
| **G3** | 三条治理循环(chat / embedding / ocr)口径一致,计时逻辑收敛为单一共享单元,杜绝第四次复制 |
|
||||||
|
| **G4** | 配置方不再需要心算 `stall_window > timeout × max_attempts`;`.env.example` 注释与实际语义对齐 |
|
||||||
|
| **非目标** | 不改 `progress_age_s()` 的 `inf` 语义(见 §3.4);不新增装配期校验(见 §5.2);不新增配置项;不改 `AllSourcesExhausted` 的字段与 `reason` 取值;不给 embedding/ocr 新增主循环判死路径(见 §5.4);不改 429 免预算、AIMD、选源、熔断任何既有行为 |
|
||||||
|
|
||||||
|
### 1.1 Issue 前提的三处修正(按 1.1.0 源码核实)
|
||||||
|
|
||||||
|
| Issue 原文 | 实际情况 |
|
||||||
|
|---|---|
|
||||||
|
| 失效点为 `retry.py:216` 一处 | **三处同构**:`retry.py:216`(主循环)、`retry.py:305` / `embedding.py:247` / `ocr.py:272`(`_on_no_runnable`)。四个判定点共用同一个墙钟 `entered_at`,故 embedding/ocr 在"先超时一次、再遇到无可用源"时同样误判——issue 只覆盖了 chat |
|
||||||
|
| 建议方向 1:装配期校验 `stall_window_s > max(timeout_s)` | **不采纳**。它把耦合固化成契约而非消除耦合,且约束值须为 `timeout × max_attempts`(本机即 900s),会让 stall 兜底迟钝到近乎失效。详见 §5.1 |
|
||||||
|
| 建议方向 2:`inf` 不参与判死 | **不采纳**。在新口径下 `inf` 从"有害恒真"变回"正确的保守默认";且它会反转已被测试钉住的既有行为。详见 §3.4 与 §5.3 |
|
||||||
|
|
||||||
|
## 2. 根因:两个预算重叠计费
|
||||||
|
|
||||||
|
`retry.py:214-215` 的注释自述这处判定是「429 免预算后的兜底,防饱和期无限循环」——它治理的对象是**非生产性循环**。但条件 A `now - entered_at > stall` 度量的是**墙钟总耗时**,无法区分两类性质相反的时间:
|
||||||
|
|
||||||
|
| 时间性质 | 构成 | 应由谁治理 | 耗尽后 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **生产性** | 一次尝试的完整生命周期(发请求、等响应含耗满 `timeout_s` 的超时/TTFT/流式读取,以及该次尝试的记账与遥测收尾) | `max_attempts`(重试预算) | `retry_exhausted` |
|
||||||
|
| **非生产性** | 429 退避、配额 wait 轮询、熔断冷却轮询、AIMD 排队 | **无人治理**(429 不计 `fails`)→ 正是 stall 的职责 | `stalled` |
|
||||||
|
|
||||||
|
**缺陷即:生产性时间同时向两个预算计费。** 而 stall 预算(默认 300s)远小于重试预算(`3 × 300s`),必然先耗尽,于是重试预算在超时场景下**永远用不上**——issue 观察到的"静默失效"就是这个重叠计费的直接后果。
|
||||||
|
|
||||||
|
`.env` 里 `TIMEOUT_S=300` 与 `_DEFAULT_STALL_WINDOW_S=300.0`(`config.py:60`)相等只是把它暴露得最快;只要 `timeout_s ≥ stall_window_s / 1`,一次超时就够。
|
||||||
|
|
||||||
|
### 2.1 两条佐证:`inf` 恒真是遗漏而非设计
|
||||||
|
|
||||||
|
| 证据 | 出处 | 含义 |
|
||||||
|
|---|---|---|
|
||||||
|
| `_PROGRESS_TTL_S = 3600 # 远大于任何 stall_window,防进度键过期造成假停滞` | `backends/redis/limiter.py:32` | 「无 progress 记录 ≠ 停滞」早已是设计共识,作者用超长 TTL 规避了"键过期"这一路径,但 TTL 再长也救不了"**从来没写过**"——冷启动是同类情形的漏网之鱼 |
|
||||||
|
| `test_global_stale_but_local_fresh_keeps_waiting` docstring 写「仅全局超窗(从未出餐 age=inf)」 | `tests/unit/test_backpressure.py:120-121` | 现有测试把 `inf` 当作"全局超窗成立"钉住了;`test_both_windows_exceeded_raises_stalled`(:89)更是**全靠 `inf` 恒真**才能触发判死 |
|
||||||
|
|
||||||
|
## 3. 选定方案:双预算正交模型
|
||||||
|
|
||||||
|
### 3.1 一句话
|
||||||
|
|
||||||
|
**stall 计时器只累计非生产性等待时间**:`stalled_s = (now − entered_at) − 真实尝试累计耗时`。
|
||||||
|
|
||||||
|
两个预算自此正交,各管一段,无缝覆盖调用的全部时间:
|
||||||
|
|
||||||
|
| 花在哪 | 烧哪个预算 |
|
||||||
|
|---|---|
|
||||||
|
| 真实尝试(`_attempt` 内),**429 除外** | 重试预算 `max_attempts` |
|
||||||
|
| 其余一切等待,**含 429 尝试本身** | stall 预算 `stall_window_s` |
|
||||||
|
|
||||||
|
> **划分依据是"谁消耗重试预算",不是"是否发出了请求"**(2026-08-06 实施期订正,见 §3.6)。初稿按后者划分,使 429 尝试两个预算都不烧。
|
||||||
|
|
||||||
|
**"生产性"的边界即 `_attempt` 的边界**——包含该次尝试的记账(`record_success`/`mark_progress`)与遥测收尾,而不止于"等响应"。这是有意的:这些收尾是"尝试已有结论"之后的动作,不是"在等待重试机会"的停滞;把它们计入 stall 会让遥测抖动参与判死,与「遥测写失败降级不冒泡」所守的"遥测不得影响主路径判决"同精神。其耗时本也在毫秒量级。
|
||||||
|
|
||||||
|
这与库内既有原则**同构**:429 不烧重试预算,所以 429 等待烧 stall 预算;真实尝试烧重试预算,所以它不烧 stall 预算。
|
||||||
|
|
||||||
|
### 3.2 为什么取补集,而不是逐处标记 sleep
|
||||||
|
|
||||||
|
两种实现都能达到 §3.1 的语义,选**取补集**(总时间减去 `_attempt` 耗时):
|
||||||
|
|
||||||
|
| 维度 | 取补集(选定) | 逐处标记 sleep(否决) |
|
||||||
|
|---|---|---|
|
||||||
|
| 埋点数量 | 每条循环 **1 处**(`_attempt` 调用点) | chat 3 处、embedding/ocr 各 2 处,共 7 处 |
|
||||||
|
| 演进安全性 | **默认安全**:将来新增任何等待路径自动计入 stall,兜底不会漏 | 默认危险:新增等待路径若忘记标记,即成新的 stall 盲区 |
|
||||||
|
| 语义可读性 | 「stall 时间 = 总时间 − 花在真实尝试上的时间」,一句话说清 | 需读者遍历全部标记点才能确认覆盖完整 |
|
||||||
|
|
||||||
|
`_attempt` 是纯生产性的:permit 获取、熔断准入、AIMD 判定全部在 `_pick_runnable` 内完成,`_attempt` 进入时已持 permit,内部只做"发请求 + 记账"。故补集口径不会把非生产性时间误算为生产性。
|
||||||
|
|
||||||
|
### 3.3 共享单元:`StallClock`
|
||||||
|
|
||||||
|
计时逻辑提取为 `middleware/retry.py` 的模块级小类,embedding/ocr 复用——沿用 `backoff_delay` 已被两者复用的既有手法(`tests/unit/test_backpressure.py:258` 记录该先例),不新建模块、不动依赖层次。
|
||||||
|
|
||||||
|
```python
|
||||||
|
class StallClock:
|
||||||
|
"""调用级 stall 计时器: 只累计非生产性等待(设计 §3.1)。
|
||||||
|
|
||||||
|
实例per调用创建, 严禁提升为实例属性——并发调用共享会互相污染。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, now: Callable[[], float]) -> None:
|
||||||
|
self._now = now
|
||||||
|
self._entered_at = now()
|
||||||
|
self._productive_s = 0.0
|
||||||
|
|
||||||
|
def stalled_s(self) -> float:
|
||||||
|
return self._now() - self._entered_at - self._productive_s
|
||||||
|
|
||||||
|
@contextlib.asynccontextmanager
|
||||||
|
async def attempting(self):
|
||||||
|
started = self._now()
|
||||||
|
try:
|
||||||
|
yield
|
||||||
|
finally:
|
||||||
|
# 只做算术, 不吞任何异常——CancelledError 照常穿透(库铁律)
|
||||||
|
self._productive_s += self._now() - started
|
||||||
|
```
|
||||||
|
|
||||||
|
调用点改动(三处循环同款):
|
||||||
|
|
||||||
|
```python
|
||||||
|
clock = StallClock(self._now) # 替换 entered_at = self._now()
|
||||||
|
...
|
||||||
|
if clock.stalled_s() > stall and await self._quota.progress_age_s() > stall:
|
||||||
|
raise AllSourcesExhausted(..., reason="stalled", ...)
|
||||||
|
...
|
||||||
|
async with clock.attempting(): # 包裹真实尝试
|
||||||
|
outcome = await self._attempt(request, *picked, reasons, attempt_fails)
|
||||||
|
```
|
||||||
|
|
||||||
|
`_on_no_runnable` 的形参由 `entered_at: float` 改为 `clock: StallClock`(三处同改)。
|
||||||
|
|
||||||
|
### 3.4 `inf` 语义为何不动(本设计的核心权衡)
|
||||||
|
|
||||||
|
新口径下第一象限的含义变为:「**非生产性排队已耗满 `stall_window_s`,且整个 scope 从未出餐**」。此时判死是正当的——真的没有任何证据表明这个 scope 还活着,而调用方已经白等了一整个窗口。`inf` 由此从"有害的恒真"回归为"正确的保守默认"。
|
||||||
|
|
||||||
|
反过来,若同时改 `inf` 语义:
|
||||||
|
|
||||||
|
- 冷启动窗口内 stall 判定**完全失效**,429 饱和场景下 chat 主循环重新暴露无限循环风险(429 不计 `fails`,无其他兜底);
|
||||||
|
- 会反转 `test_both_windows_exceeded_raises_stalled` 钉住的行为,并与 CHS 保真蓝本分叉。
|
||||||
|
|
||||||
|
**一次改动解决问题,优于两次改动互相牵制。** 这是本设计只动条件 A 的理由。
|
||||||
|
|
||||||
|
### 3.5 429 饱和场景下兜底仍然有效(正确性验证)
|
||||||
|
|
||||||
|
修改后必须确认 stall 兜底没有被削弱。429 免预算使 `fails` 恒为 0,`retry.py` 的 `max(fails, 1)` 令退避恒定在 `backoff_base_s` 档(或取 `Retry-After` 提示的较大值),不随轮次增长。每轮构成为「一次 429 往返」+「一段恒定退避 sleep」,后者非生产性且每轮累加,`stalled_s` 单调逼近 `stall_window_s`,兜底有效。
|
||||||
|
|
||||||
|
**但这个论证在初稿里依赖一个未加保护的假设**:「429 往返是快速失败,毫秒至秒级」。§3.6 处理它不成立的情形。
|
||||||
|
|
||||||
|
### 3.6 订正:429 尝试必须退还给 stall 账(2026-08-06 实施期,独立验证发现)
|
||||||
|
|
||||||
|
**缺陷**:初稿按"是否发出请求"划分两个预算,于是 429 尝试的耗时算生产性。但 429 **不消耗重试预算**——它于是**两个预算都不烧**,掉进缝隙。§3.1 初稿声称的"无缝覆盖调用的全部时间"因此不成立。
|
||||||
|
|
||||||
|
**后果实测**(排队型网关:持满 `timeout_s` 才回 429,`timeout=300 / stall=300 / backoff_base=2 / rng=0`):
|
||||||
|
|
||||||
|
| | 尝试次数 | 墙钟 |
|
||||||
|
|---|---|---|
|
||||||
|
| 修复前(main) | 1 | 301s |
|
||||||
|
| 初稿口径 | **301** | **90,601s ≈ 25.2 小时** |
|
||||||
|
| 订正后 | 1 | 301s |
|
||||||
|
|
||||||
|
即初稿把一个 bug 换成了一个更严重的 bug——25 小时的挂起。
|
||||||
|
|
||||||
|
**订正**:划分依据改为**"谁消耗重试预算"**。429 免重试预算 → 429 尝试的耗时归 stall 治理,由 `StallClock.attempting()` yield 的句柄 `refund()` 退还。缝隙就此闭合,且这条规则比初稿更本质:两个预算按"由谁治理"划分,而非按"是否发出请求"这个表象。
|
||||||
|
|
||||||
|
**影响范围仅 chat**:embedding/ocr 无 429 免预算(无条件 `fails += 1`),429 照常烧重试预算,不存在缝隙,无需改动(与 §5.4 的分析一致)。
|
||||||
|
|
||||||
|
## 4. 旧版行为审计(stall 子系统逐条)
|
||||||
|
|
||||||
|
| 既有行为 | 处置 | 说明 |
|
||||||
|
|---|---|---|
|
||||||
|
| 双条件判死(本地超窗 ∧ 全局无进展超窗) | **保留** | 结构不变,只改条件 A 的度量口径 |
|
||||||
|
| 条件 A = 调用级累计、循环内不重置(CHS `governance.py:207`) | **保留** | `StallClock` 同样每调用一个实例、循环内不重置 |
|
||||||
|
| 条件 A 计入真实尝试耗时 | **替换** | 本设计的唯一行为变更 |
|
||||||
|
| 条件 B `progress_age_s()`,`inf` = 从未进展 | **保留** | 见 §3.4 |
|
||||||
|
| 本地 monotonic 与后端时钟刻意不混用 | **保留** | `StallClock` 只用注入的 `self._now`,不读后端时钟 |
|
||||||
|
| poll jitter ∈ [0.5p, 1.0p] 防惊群 | **保留** | 不触碰 |
|
||||||
|
| `fail_fast` 不进入 stall 判定 | **保留** | 不触碰 |
|
||||||
|
| 429 免预算(chat 独有) | **保留** | 不触碰;embedding/ocr 无此逻辑,故无对应缺口(§5.4) |
|
||||||
|
| `AllSourcesExhausted(reason="stalled")` 及其 `retry_after_s` 取值 | **保留** | 错误面零变更,下游 `except` 写法不受影响 |
|
||||||
|
|
||||||
|
**有意放弃**: 无。本设计不删除任何既有行为。
|
||||||
|
|
||||||
|
## 5. 被否决的路线
|
||||||
|
|
||||||
|
### 5.1 装配期校验 `stall_window_s > max(timeout_s)`(Issue 建议方向 1)
|
||||||
|
|
||||||
|
否决理由三条:
|
||||||
|
|
||||||
|
1. **治标**。它把"两个预算重叠计费"这个缺陷固化成一条配置契约,要求配置方绕开它,而不是消除它。
|
||||||
|
2. **约束值不可接受**。要让重试预算真正可用,须 `stall_window > timeout × max_attempts`(本机 900s)。stall 兜底随之迟钝到 900s 才触发,饱和期无限循环的防护近乎失效——**修好一个洞,挖开另一个**。
|
||||||
|
3. **挡不住残余情形**。即便配到 1200s,一次调用若在 429 轮询与超时上累计超过 1200s,条件 B 的 `inf` 仍恒真,双条件仍退化为单条件。坑只是被推远。
|
||||||
|
|
||||||
|
新口径下 `stall_window_s` 与 `timeout_s` 不再有任何耦合,**这条校验没有存在的理由**——不加校验、而是消除掉需要校验的耦合。
|
||||||
|
|
||||||
|
### 5.2 既有校验 `stall_window_s ≥ max(ttft_timeout_s)` 的处置
|
||||||
|
|
||||||
|
`config.py:240-247` 的 `_validate_stall` 被 `ARCHITECTURE.md` §7.3 记为契约补强 G6。新口径下 TTFT 等待属生产性时间,其 docstring 的理由「防把正常慢首包误判为卡死」**已不成立**。
|
||||||
|
|
||||||
|
**人类已定夺:保留校验,改写 docstring 说明新口径**。校验本身无害(不会误拒任何合理配置),保留可避免改动 ARCHITECTURE.md 既有契约、把本次改动的影响面控制在最小。docstring 改为说明"该校验在新口径下为保守冗余,TTFT 已不计入 stall"。
|
||||||
|
|
||||||
|
### 5.3 `inf` 不参与判死(Issue 建议方向 2)
|
||||||
|
|
||||||
|
见 §3.4:新口径下 `inf` 已无害,单独改它会制造冷启动兜底真空并反转既有测试。
|
||||||
|
|
||||||
|
### 5.4 给 embedding/ocr 补主循环 stall 判定
|
||||||
|
|
||||||
|
设计过程中一度提出(前提是"429 饱和时它们没有防无限循环兜底"),**核实后前提不成立,故否决**:
|
||||||
|
|
||||||
|
| 循环路径 | embedding/ocr 的兜底 |
|
||||||
|
|---|---|
|
||||||
|
| `picked is None` → `_on_no_runnable` 轮询(不烧 `fails`) | `_on_no_runnable` 内已有 stall 判定(`embedding.py:247` / `ocr.py:272`)✓ |
|
||||||
|
| 尝试失败 → `fails += 1` | `max_attempts` ✓ |
|
||||||
|
|
||||||
|
`retry.py:233-234` 的 429 免预算分支是 chat **独有**的(`embedding.py:191`、`ocr.py:216` 均为无条件 `fails += 1`,两文件亦无 `pacer`),主循环判定正是为它打的补丁。embedding/ocr 两条路径均已封闭,补齐等于凭空新增一条判死路径,使其比 chat 更易判死——纯 gold-plating。
|
||||||
|
|
||||||
|
## 6. 非功能维度
|
||||||
|
|
||||||
|
| 维度 | 回答 |
|
||||||
|
|---|---|
|
||||||
|
| **并发** | `StallClock` **每次调用创建一个实例**,是调用级局部状态,与被替换的 `entered_at` 局部变量同性质。严禁提升为实例属性(并发调用会互相污染计时)——docstring 已写明,单测钉住并发两路调用互不干扰 |
|
||||||
|
| **取消** | `attempting()` 的 `finally` 只做浮点加法,不含 `await`、不捕获任何异常,`CancelledError` 逐字穿透。既有 `test_cancellation_pierces_wait_loop` 继续有效,并新增一条"取消发生在 `_attempt` 内"的用例 |
|
||||||
|
| **降级方向** | 不变。stall 判定读取的 `progress_age_s()` 属准入侧,后端故障仍 fail-closed 抛 `GovernanceBackendError`(scope 级),不放行 |
|
||||||
|
| **幂等与重复** | `stalled_s()` 是纯读,可任意次调用;`attempting()` 可重入多次(每次尝试一次),累加语义天然幂等于"总生产性时间" |
|
||||||
|
| **持久化与原子性** | 不适用。纯进程内计时,无落盘、无后端写入,不新增任何 Redis 往返 |
|
||||||
|
| **性能** | 每次尝试新增两次 `self._now()` 调用与一次浮点加法,可忽略 |
|
||||||
|
|
||||||
|
## 7. 错误处理与测试策略
|
||||||
|
|
||||||
|
**错误分类**: 无变更。判死仍抛 `AllSourcesExhausted(reason="stalled")`,属 scope 级不可用(`GatewayUnavailableError` 家族),下游延期重投语义不变。
|
||||||
|
|
||||||
|
### 7.1 回归证据(先失败后通过)
|
||||||
|
|
||||||
|
核心用例 `test_single_timeout_does_not_exhaust_stall_budget`:`stall_window_s == timeout_s == 300`,第一次尝试推进 `FakeClock` 超过 300s 后抛 `TransientError`,第二次返回成功。
|
||||||
|
|
||||||
|
- **改前**:第二次尝试发出前即被判死,抛 `AllSourcesExhausted(reason="stalled")` → **失败**
|
||||||
|
- **改后**:重试预算正常生效,返回成功响应 → **通过**
|
||||||
|
|
||||||
|
embedding / ocr 各一条同构用例(经"先超时一次、再遇到无可用源"触发 `_on_no_runnable`)。
|
||||||
|
|
||||||
|
### 7.2 其余用例
|
||||||
|
|
||||||
|
| 用例 | 钉住什么 |
|
||||||
|
|---|---|
|
||||||
|
| 四象限现有四条(`TestStallQuadrants`) | 非生产性路径行为逐字不变;`test_both_windows_exceeded` 全程无真实尝试,`stalled_s` 等价于旧墙钟,**应原样通过** |
|
||||||
|
| `test_productive_time_excluded_from_stall` | 直接断言:仅靠真实尝试耗时无论多久都不触发判死 |
|
||||||
|
| `test_nonproductive_wait_still_triggers_stall` | 反向:纯轮询等待累满窗口仍正常判死(兜底未被削弱) |
|
||||||
|
| `test_saturation_429_still_stalls` | §3.5 的正确性验证:429 连续拒绝 + 退避,最终仍判死而非无限循环 |
|
||||||
|
| `test_cancel_inside_attempt_pierces` | 取消穿透 `attempting()` 的 `finally` |
|
||||||
|
| `test_concurrent_calls_do_not_share_clock` | 两路并发调用,一路长尝试不影响另一路的 stall 账 |
|
||||||
|
|
||||||
|
Redis 后端无需新增用例:本设计不改后端接口与 `progress_age_s()` 语义。
|
||||||
|
|
||||||
|
## 8. 交付清单(供 `writing-plans` 展开)
|
||||||
|
|
||||||
|
| # | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| T1 | `middleware/retry.py` 新增 `StallClock`;主循环与 `_on_no_runnable` 改用之 |
|
||||||
|
| T2 | `embedding.py` / `ocr.py` 复用 `StallClock`,`_on_no_runnable` 形参改签名 |
|
||||||
|
| T3 | `config.py:240-247` `_validate_stall` docstring 改写(§5.2) |
|
||||||
|
| T4 | 测试:§7.1 回归三条 + §7.2 五条;`test_backpressure.py:121` docstring 订正 |
|
||||||
|
| T5 | `.env.example:41` 注释改写(删除误导性的"须 ≥ 最大源 TTFT",说明新口径);本机 `.env:37` 的临时缓解 `STALL_WINDOW_S=1200` 可回退默认(不入库,仅记录) |
|
||||||
|
| T6 | `ARCHITECTURE.md` §7.3 背压条目补记新口径与本设计指针;`CHANGELOG.md` 记治理行为变更 |
|
||||||
|
| T7 | Wiki 同步(`docs-convention.md` §2「治理行为变更」行):`解释-治理行为` + `指南-限流与熔断` |
|
||||||
|
|
||||||
|
**副作用提醒**: 修复后单次调用最坏耗时由 `stall_window_s` 抬升至 `max_attempts × timeout_s`(本机 900s)——这是重试预算恢复生效的**正确表现**,但 e2e 冒烟测试的最坏耗时随之变长,`tests/e2e` 的源 `timeout_s` 配置可能需要相应调小。
|
||||||
@@ -0,0 +1,238 @@
|
|||||||
|
# HTTP 错误响应体留存设计(Issue #10)
|
||||||
|
|
||||||
|
- **日期**: 2026-08-16
|
||||||
|
- **来源**: Gitea Issue #10(下游 1050 张医学影像批处理,1 张收到 400 被判确定性失败;事后无从查证原因。基于 1.1.2 源码核查)
|
||||||
|
- **状态**: **已批准(2026-08-16)**,待 `writing-plans`
|
||||||
|
- **触发档位**: 强制(`errors.py` 属最内层内核,新增公共字段即变更库对下游的承诺)
|
||||||
|
- **方案范围**: 人类明确要求单一方案(2026-08-16),故本文不列平行备选,仅在 §6 记录被否决路线及否决理由(体例沿用 Issue #7/#8 设计)
|
||||||
|
|
||||||
|
## 1. 目标与非目标
|
||||||
|
|
||||||
|
| | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| **G1** | 网关拒绝一次调用时,**它说了什么必须可事后查证**——库自己的遥测表里就能查到,不依赖下游额外埋点 |
|
||||||
|
| **G2** | 留存口径覆盖 transport 层**全部**非 2xx 分支与**全部** transport(chat / embedding / stream / OCR),杜绝"只修 400 → 下次 401 复发" |
|
||||||
|
| **G3** | 摘要文本单点规范化(折叠空白 + 截断 + 截断标记),message 与结构化字段**取同一份串**,两处永不打架 |
|
||||||
|
| **G4** | 不改变任何状态码 → 错误分类的映射(ARCHITECTURE §6.2 表原封不动),下游 `except` 写法零影响 |
|
||||||
|
| **非目标** | 不改 400 的治理语义(不重试不换源,见 §5.2);不新增遥测列(见 §6.1);不新增配置项;不做错误分类可插拔(见 §6.4);不顺手修 `_status_to_error` 的 `operation` 硬编码缺陷(见 §5.4) |
|
||||||
|
|
||||||
|
### 1.1 Issue 前提的两处修正(按 1.1.2 源码核实)
|
||||||
|
|
||||||
|
| Issue 原文 | 实际情况 |
|
||||||
|
|---|---|
|
||||||
|
| 建议方向一「让异常带上截断后的响应体……就能让下游把它记进日志和遥测」 | **只做这一半解决不了 Issue 自己陈述的痛点**。库的逐次遥测写的是 `error=str(exc)`(`middleware/retry.py:558` → `middleware/telemetry.py:89` → `telemetry/sqlite.py:43` 的 `error TEXT` 列),即**异常 message**。新增字段不会进库的遥测表;下游说的"写进遥测表"是他们自己的埋点。故本设计**两件都做,且以 message 为主**(§3.3) |
|
||||||
|
| 缺陷范围 = 400 分支 + 4xx 兜底 | 实为 **6 处同构**:`openai_compat._status_to_error` 的 400 / 4xx 兜底 / 401·403 / 5xx 四支,`_translate_429` 的两支(读了 body 判 `insufficient_quota`,但 message 仍不带),以及 `monkey_ocr._classify_status:74-88` 的**全部**分支(message 只有 `HTTP {status}`)。Issue 场景是"读表格",极可能正落在 OCR 路径 |
|
||||||
|
|
||||||
|
## 2. 根因:诊断信息在翻译层被丢弃,而遥测只看 message
|
||||||
|
|
||||||
|
`_status_to_error`(`transports/openai_compat.py:131-143`)手上握着 `body_text`,却只把它用于 429 的类型细分,翻出的异常与 message 都不携带它。响应体在这一层之后**不再存在于进程任何位置**:该模块无 logger(grep `logger|loguru` 零命中),异常类无字段,遥测只写 message。
|
||||||
|
|
||||||
|
三条留存通道同时为空,是"永久查不到"的完整解释:
|
||||||
|
|
||||||
|
| 通道 | 现状 | 本设计后 |
|
||||||
|
|---|---|---|
|
||||||
|
| 日志 | 模块无 logger | 仍无(§6.2:不加日志) |
|
||||||
|
| 异常字段 | 无承载处 | `body_text`(§3.1) |
|
||||||
|
| 库遥测 `error` 列 | 只有 `"{源名} 请求被拒: 400"` | message 携带摘要(§3.3) |
|
||||||
|
|
||||||
|
## 3. 选定方案
|
||||||
|
|
||||||
|
### 3.1 内核:`PolyGatewayError` 基类新增 `body_text`
|
||||||
|
|
||||||
|
```python
|
||||||
|
class PolyGatewayError(Exception):
|
||||||
|
def __init__(self, message, *, source_name=None, status_code=None,
|
||||||
|
operation=None, body_text: str = "") -> None:
|
||||||
|
```
|
||||||
|
|
||||||
|
**加在基类而非 `RequestRejectedError`**:这些错误全部由同一个 HTTP 响应翻译而来,"对方说了什么"与"它属于哪一类"正交。只给一个子类加,下次给 `SourceDeadError` 加又是一次公共 API 变更 + 一次人类门。
|
||||||
|
|
||||||
|
与既有 `ResultInvalidError.raw_text`(`errors.py:77`)的界限必须在 docstring 钉死,否则两个"原文字段"必然被混用:
|
||||||
|
|
||||||
|
| 字段 | 语义 | 来源 |
|
||||||
|
|---|---|---|
|
||||||
|
| `body_text` | **非 2xx** 的 HTTP 错误响应体摘要——对方**拒绝**你的理由 | transport 翻译层 |
|
||||||
|
| `raw_text` | **2xx** 但内容不可解析时的模型输出原文 | 结构化解析层 |
|
||||||
|
|
||||||
|
`GatewayUnavailableError` 一族继承到一个恒空的 `body_text` 不是噪音:scope 级失败本就"没有单一响应体可言",空串是对这件事的如实表达。
|
||||||
|
|
||||||
|
### 3.2 共享单元:`transports/_http_errors.py`(新建,~40 行)
|
||||||
|
|
||||||
|
两个 transport 各有自己的状态码分类逻辑(OCR 无 429 细分,有意保留,见 `monkey_ocr.py:53-54`),但**摘要口径必须同一份**,否则就是下一个"只修一半"。两函数:
|
||||||
|
|
||||||
|
| 函数 | 职责 | 关键防御 |
|
||||||
|
|---|---|---|
|
||||||
|
| `summarize_body(text) -> str` | 折叠空白 → 按 §3.4 的机械规则截断 | 空/空白入参返回 `""` |
|
||||||
|
| `response_body(response) -> str` | 从 `httpx.Response` 取已缓冲文本 | `ResponseNotRead` → 返回 `""`,**绝不触发网络读** |
|
||||||
|
|
||||||
|
- **折叠空白不是洁癖**:错误体常是缩进 JSON,直接拼进 message 会让一行日志炸成多行、遥测列不可读。
|
||||||
|
- **截断必须留标记**:不标记,读的人分不清"网关只说了这么多"和"库切的"。
|
||||||
|
- **`response_body` 的防御是硬要求**:`monkey_ocr._classify_status` 只拿得到 `httpx.HTTPStatusError`,若某天 OCR 走 stream 请求,`.text` 会抛 `ResponseNotRead`,把一次可分类的 4xx 变成泄漏的 httpx 异常——**违反"一切失败必须落入四分类"铁律**。诊断信息缺失绝不能升级为崩溃(降级方向,§4.2)。
|
||||||
|
|
||||||
|
放在 `transports/` 私有模块而非 `errors.py`:职责是"HTTP 响应 → 领域错误"的工具,放内核会稀释 `errors.py` 的单一职责(P3)。两个 transport 同 import 一个私有模块,不构成 transport 之间的互相依赖,import-linter 的 layers 契约(同层 `|` 独立性)不受影响。
|
||||||
|
|
||||||
|
### 3.3 翻译层:表驱动收口,message 与字段共用一份摘要
|
||||||
|
|
||||||
|
`_status_to_error` 现在是五个分支各拼各的 message,新增摘要意味着五处重复。改为**分类表 + 单点拼装**,代码反而变短:
|
||||||
|
|
||||||
|
```
|
||||||
|
summary = summarize_body(body_text) # 全函数只算一次
|
||||||
|
ctx = {..., "body_text": summary} # 字段
|
||||||
|
429 → _translate_429(source, body_text, headers, ctx) # 需原文判 type,单列
|
||||||
|
其余 → cls, label = _STATUS_MAP 查表 → cls(_compose(source, label, status, summary), **ctx)
|
||||||
|
```
|
||||||
|
|
||||||
|
message 形态:`"{源名} {标签}: {状态码} | {摘要}"`;**摘要为空时不拼后缀**,避免出现悬空的 ` | `。分隔符取 ` | ` 而非既有的 `: `,让"库的话"与"网关的话"一眼可分。
|
||||||
|
|
||||||
|
**429 也拼,不设例外**:例外就是下一个复发点。`insufficient_quota` 那支尤其需要(配额细节全在 body 里);普通限速 body 通常很短。代价是高频限速场景遥测 `error` 列变长,由 `_ERROR_BODY_CAP` 兜住。
|
||||||
|
|
||||||
|
`monkey_ocr._classify_status` 同款处理:`summary = summarize_body(response_body(exc.response))`,message 追加同一后缀,`ctx` 带上字段。
|
||||||
|
|
||||||
|
### 3.4 常量取值
|
||||||
|
|
||||||
|
**机械规则(实现与测试逐字照此)**:
|
||||||
|
|
||||||
|
```
|
||||||
|
_ERROR_BODY_CAP = 2048 # 字符(非字节),含省略标记在内的最终总长上限
|
||||||
|
_HEAD_CHARS = 1400
|
||||||
|
_TAIL_CHARS = 600
|
||||||
|
|
||||||
|
折叠空白后 len ≤ 2048 → 原样返回
|
||||||
|
否则 → s[:1400] + f"…(略 {len(s) - 2000} 字)…" + s[-600:]
|
||||||
|
```
|
||||||
|
|
||||||
|
> 规则必须写成算术而非叙述:"截断至 cap 并补标记"能同时被读成总长 2048 与 2049,两者会让测试断言与遥测长度承诺对不上(Codex 审查 2026-08-16 提出)。
|
||||||
|
|
||||||
|
**头尾保留而非头部硬切**(2026-08-16 调研决策)。截断的对象是**结构化 JSON 错误体**,信息分布头重尾也重:人话(`message`)在前,机器可判的 `type` / `code` / `param` / `request_id` 在后。Issue 给出的真实样本即 `"code":"invalid_parameter_error"` 收尾——头部硬切正好切掉向网关方追查时唯一有用的那部分。省略标记记下**被省略的字符数**,读的人才知道自己丢了多少,不会误以为网关只说了这么多。
|
||||||
|
|
||||||
|
按**字符**而非字节切:多字节字符不会被切成半个(Sentry 曾为按字节切开 issue #1691),且 `error TEXT` 列无定长约束,无需字节口径。
|
||||||
|
|
||||||
|
### 3.4.1 取值依据:同场景开源实践
|
||||||
|
|
||||||
|
| 项目 | 场景 | 上限 | 保留策略 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **Kubernetes client-go** `rest/request.go` | **读 HTTP 错误体生成错误信息**(与本设计同构) | `maxUnstructuredResponseTextBytes = 2048` | 头部硬切 |
|
||||||
|
| OpenAI Python SDK `_exceptions.py` | 异常对象持有 body | **不截断**(内存对象,不落库) | — |
|
||||||
|
| Sentry Python `strip_string` | 事件写入前 trim | `max_value_length`,2.34.0 前默认 1024 | 头部 + `...`,另用 metadata 记原长 |
|
||||||
|
| Elastic APM | 长字段 | keyword 1024 / long field 10000 | 截断带省略号 |
|
||||||
|
| Python 标准库 `reprlib` | 给人读的长字符串 | `maxstring` | **头 + 尾,中间省略** |
|
||||||
|
|
||||||
|
**2048 对齐 k8s client-go**——它是唯一与本设计同场景(读 HTTP 错误体做诊断)的成熟先例。初稿的 500 仅以 issue 的单个样本(约 160 字符)为据,是拿一个样本定上限,已废弃。头部硬切在 k8s/Sentry 成立是因为它们截的是任意文本;本设计截的是结构化 JSON,故取 `reprlib` 的头尾策略。
|
||||||
|
|
||||||
|
遥测代价:纯 ASCII 约 2KB/条,纯中文最多约 6KB/条;5xx 重试 3 次即一次调用最多约 18KB。批处理场景(1050 次调用、5% 失败)约 300KB,`TEXT` 列可忽略。
|
||||||
|
|
||||||
|
message 与 `body_text` **共用同一变量**,不设两个长度:两份不同长度会让"遥测里看到的"与"下游 catch 到的"对不上,排查时反而多一层困惑。
|
||||||
|
|
||||||
|
### 3.5 改动清单
|
||||||
|
|
||||||
|
| 文件 | 改动 |
|
||||||
|
|---|---|
|
||||||
|
| `errors.py` | 基类新增 `body_text` 字段 + 与 `raw_text` 的界限 docstring |
|
||||||
|
| `transports/_http_errors.py` | **新建**:`summarize_body` / `response_body` / `_ERROR_BODY_CAP` |
|
||||||
|
| `transports/openai_compat.py` | `_status_to_error` 表驱动重写;`_translate_429` 收 `ctx` |
|
||||||
|
| `transports/monkey_ocr.py` | `_classify_status` 带摘要 |
|
||||||
|
| `errors.py` docstring + `ARCHITECTURE.md` §6.2 | 中转拓扑下 400 的提醒(§5.2) |
|
||||||
|
| `README.md:34` | 安装 pin `==1.1.*` → `>=1.2,<2`(§5.3,发布前置,漏改则下游拿不到本修复) |
|
||||||
|
|
||||||
|
三个调用点(`openai_compat.py:402` embed、`:417` stream、`:509` 非流式)**签名不变**,无需改动。
|
||||||
|
|
||||||
|
## 4. 非功能维度
|
||||||
|
|
||||||
|
### 4.1 并发与取消
|
||||||
|
|
||||||
|
新增全部是纯函数与数据字段,无状态、无锁、无 IO、不引入 `await`。`response_body` 只读已缓冲字节,`ResponseNotRead` 时直接返回空串而**不发起网络读**——否则会在错误路径上凭空插入一次可能挂住的 IO。`CancelledError` 路径逐字不变。
|
||||||
|
|
||||||
|
### 4.2 降级方向
|
||||||
|
|
||||||
|
响应体不可得(未读缓冲 / 解码失败 / 空体)→ `body_text=""`,**静默降级,绝不报错**。诊断信息属可观测性,按库铁律与缓存/遥测同档:缺了降级,不得把一次本可正确分类的失败变成不可分类的崩溃。流式路径的 `(await resp.aread()).decode("utf-8", errors="replace")`(`:416`)已是这个口径,保持。
|
||||||
|
|
||||||
|
### 4.3 幂等与重复
|
||||||
|
|
||||||
|
纯函数,同输入同输出。`summarize_body` 对自身输出再调用一次是幂等的:输出总长恒为 `2000 + len(标记) ≤ 2048`(标记形如 `…(略 N 字)…`,8 + N 的位数,现实中远不足 48),且不含需折叠的空白,故第二次调用走"原样返回"分支,不会出现标记被反复嵌套。
|
||||||
|
|
||||||
|
### 4.4 持久化与原子性
|
||||||
|
|
||||||
|
不新增表、不改 DDL、不动遥测端口的 22 字段与列序。摘要经既有 `error TEXT` 列落盘,原子性由既有单行写入保证。
|
||||||
|
|
||||||
|
### 4.5 安全与体积
|
||||||
|
|
||||||
|
- **响应体可能回显请求内容**(部分网关的 `error.param` 会带违规字段值)。截断 + 空白折叠是主要止血手段;字段 docstring 须写明"可能包含请求回显,已截断"。库不做内容脱敏——库不知道下游哪些字段敏感,猜测式脱敏只会同时丢掉诊断价值与安全性。
|
||||||
|
- **本设计不放大既有的读取风险**:`_complete_stream:416` 的 `aread()` 对错误响应体无大小上限(超大错误体可打爆内存),该风险今天已经存在(读完即丢),留存后只是更显眼。**不夹带修复**,见 §5.4。
|
||||||
|
|
||||||
|
## 5. 错误处理、语义与边界
|
||||||
|
|
||||||
|
### 5.1 错误分类
|
||||||
|
|
||||||
|
不改任何映射。`body_text` 是**旁路数据**,不参与任何治理判定——不影响重试、换源、熔断计数、AIMD、限流结算。这是本设计能与 ARCHITECTURE §6.1/§6.2 零冲突的根本原因。
|
||||||
|
|
||||||
|
### 5.2 400 语义:不改行为,补文档
|
||||||
|
|
||||||
|
Issue 报告了一个有说服力的观察:同字节 15 次重发全部成功、`prompt_tokens=0`、耗时 2996ms 远低于同批 631 次成功调用的最快值 7366ms——说明那次 400 来自中转服务自身抖动,而非"你的输入有问题"。
|
||||||
|
|
||||||
|
**仍不改分类**:400 重试对直连供应商是纯浪费(确定性坏输入,重试只烧配额并拖延失败);"中转也回 400"是**部署拓扑**引入的信息损失,库从状态码无从分辨。默认改为可重试 = 让所有直连用户为一种部署形态买单,且推翻已冻结的公共契约。
|
||||||
|
|
||||||
|
**但本设计本身就是对这个观察最好的答复**:body 留存后,下游能自己区分——中转抖动的 400 体与供应商 `invalid_request_error` 体形态不同。库不替下游做判断,而是把判断所需的信息交出去。配套文档动作:`RequestRejectedError` docstring 与 ARCHITECTURE §6.2 各加一句"经中转部署时 400 可能源于中转自身抖动,批处理场景下游宜自备兜底分类"。
|
||||||
|
|
||||||
|
### 5.3 兼容性
|
||||||
|
|
||||||
|
`LLMResponse` 一族的"字段只增不删不改名"约束(ARCHITECTURE §5.1)同样适用于异常。本次是**纯新增关键字参数且带默认值**:既有构造点、既有 `except` 写法、既有 `str(exc)` 消费方全部不受影响。message 文本变化不构成破坏——现有测试对这些 message 无格式依赖(仅 `test_openai_compat.py:558` match 源名)。
|
||||||
|
|
||||||
|
版本 **1.2.0**(公共类型新增字段属 minor;2026-08-16 人类定夺)。
|
||||||
|
|
||||||
|
**发布时必须同步改 README 的安装 pin**:`README.md:34` 现为 `"polygateway[redis,postgres,structured]==1.1.*"`,发 1.2.0 后照此命令安装的下游会**静默停在 1.1.2**——无报错、无警告,与 CLAUDE.md §4.4.1 点名的"极易漏改"完全同款(registry 长期停在 1.0.5 即此类事故)。本次改为 **`>=1.2,<2`**,把"每发一个 minor 就要通知三个下游改 pin"这一反复出现的麻烦一次性消除。此项列入实现计划的发布前置步骤,不是发布日的临时动作。
|
||||||
|
|
||||||
|
### 5.4 有意不夹带的两项(建议单开 issue)
|
||||||
|
|
||||||
|
| 项 | 说明 |
|
||||||
|
|---|---|
|
||||||
|
| `_status_to_error` 的 `operation` 硬编码 `"chat"`(`:134`),而 `embed()` 也调它(`:402`) | embedding 的 HTTP 错误在遥测里被标成 `operation="chat"`,是既有数据正确性缺陷,与本 issue 无关 |
|
||||||
|
| `_complete_stream:416` 的 `aread()` 无大小上限 | 恶意/故障网关的超大错误体可打爆内存,属独立的健壮性问题 |
|
||||||
|
|
||||||
|
两项都在本次重构触及的函数附近,但修它们既不服务 G1-G4,也各自需要独立的行为讨论——按反 gold-plating 铁律留给独立 issue。
|
||||||
|
|
||||||
|
## 6. 被否决的路线
|
||||||
|
|
||||||
|
### 6.1 给遥测端口加一列(22 → 23 字段)
|
||||||
|
|
||||||
|
最"正统"的结构化留存,但成本极不相称:端口 Protocol 签名变更 + SQLite/Postgres 双后端 DDL 迁移 + 下游已有表的 ALTER + 列序契约测试全线改动——为一个诊断串付出一次跨三项目的迁移。而复用既有 `error TEXT` 列可达成同样的可查证性。
|
||||||
|
|
||||||
|
### 6.2 只在 `_status_to_error` 打一条 WARNING 日志(Issue 方向二)
|
||||||
|
|
||||||
|
不采纳为**主**手段:日志与遥测是两套留存,日志轮转后仍然查不到,而 Issue 的痛点恰是"事后"。且库铁律要求库不擅自向下游日志流写入高频内容(4xx/5xx 在批处理下可能极高频)。message 携带摘要已让 loguru 侧的下游在捕获点自然拿到同一份信息,再加一条独立日志属重复留存。
|
||||||
|
|
||||||
|
### 6.3 截断放在异常构造器内
|
||||||
|
|
||||||
|
构造器自动规范化更"防遗漏",但会让下游自建异常时传入的文本被悄悄改写,违反 P4;且 message 里的摘要仍需在翻译层单独算一次,反而出现两条规范化路径。选定方案在翻译层算一次、两处共用,更简且更显式。
|
||||||
|
|
||||||
|
### 6.4 错误分类映射可插拔(provider profile 注入 classifier)
|
||||||
|
|
||||||
|
Issue 的中转 400 场景确实指向这个方向,但当前只有一个使用方且他们已用自己的兜底分类解决。`ProviderProfile`(`providers.py:17-45`)目前也没有这个扩展点,加它是新子系统级的设计。YAGNI:等第二个使用方提出。
|
||||||
|
|
||||||
|
## 7. 测试策略(先失败后通过)
|
||||||
|
|
||||||
|
**验收主张**:一次 400 调用后,注入的 recorder 收到的 `error` 串含网关响应体摘要。这条端到端断言直接对应 Issue 的痛点,是本设计成立与否的唯一硬判据;其余为覆盖性用例。
|
||||||
|
|
||||||
|
| # | 用例 | 覆盖 |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | **端到端遥测**:mock transport 返回 400 + 真实样本体 → 断言 recorder 收到的 `error` 含摘要 | G1 |
|
||||||
|
| 2 | 参数化状态码(400 / 401 / 404 兜底 / 429 普通 / 429 `insufficient_quota` / 500)→ 断言 message 含摘要且 `exc.body_text` 非空,**分类与既有断言逐一不变** | G2, G4 |
|
||||||
|
| 3 | 超长体 → 前 1400 字符与原文头部逐字相同、**末 600 字符与原文尾部逐字相同**、中段为 `…(略 N 字)…` 且 N 等于实际省略数;`body_text` 与 message 中的摘要逐字相同 | G3 |
|
||||||
|
| 3b | 长度恰为 2048 / 2049 的体 → 前者原样无标记,后者走头尾保留(边界) | §3.4 |
|
||||||
|
| 3c | **尾部关键字段可见**:以 issue 的真实样本尾部 `"code":"invalid_parameter_error"}}` 构造超长体 → 断言该串出现在摘要中 | §3.4 头尾决策的验收 |
|
||||||
|
| 3d | 摘要对自身幂等(再摘要一次不嵌套标记) | §4.3 |
|
||||||
|
| 4 | 多行缩进 JSON → 折叠为单行 | G3 |
|
||||||
|
| 5 | 空体 / 纯空白体 → 不拼悬空分隔符,`body_text == ""` | §3.3 |
|
||||||
|
| 6 | 非 JSON 体、非 UTF-8 字节 → 不抛异常,分类不变 | §4.2 |
|
||||||
|
| 7 | 流式错误路径(`_complete_stream` 415-417)同样带摘要 | G2 |
|
||||||
|
| 8 | `monkey_ocr._classify_status` 同款(含 `ResponseNotRead` 时降级为空串而非抛出) | G2, §4.2 |
|
||||||
|
| 9 | embedding 路径(`:402`)HTTP 错误带摘要 | G2 |
|
||||||
|
|
||||||
|
`tests/unit/test_errors.py:29` 现有的"四类构造形态"参数化用例需扩展 `body_text` 默认值断言(默认 `""`、可传入、`GatewayUnavailableError` 一族恒空)。
|
||||||
|
|
||||||
|
## 8. 人类定夺记录(2026-08-16)
|
||||||
|
|
||||||
|
| 议题 | 定夺 |
|
||||||
|
|---|---|
|
||||||
|
| 摘要上限与保留策略 | 初稿 500 + 头部硬切被否:上限提至 **2048**(对齐 k8s client-go 同场景先例),策略改为**头 1400 + 尾 600 + 省略字数标记**——人类指出"有用的信息可能只在后半部分",经调研证实 JSON 错误体的 `code`/`request_id` 确实收尾(§3.4、§3.4.1) |
|
||||||
|
| 429 是否设例外 | **不设**,一律拼摘要(§3.3) |
|
||||||
|
| 版本 | **1.2.0**,并同步把 README pin 由 `==1.1.*` 改为 `>=1.2,<2`(§5.3) |
|
||||||
@@ -0,0 +1,44 @@
|
|||||||
|
---
|
||||||
|
type: design
|
||||||
|
node_id: design:governance-backend-error
|
||||||
|
title: "治理后端故障归位为 scope 级不可用(Issue #7)"
|
||||||
|
date: 2026-08-06
|
||||||
|
---
|
||||||
|
|
||||||
|
# 治理后端故障归位为 scope 级不可用(Issue #7)
|
||||||
|
|
||||||
|
全文见 `2026-08-06-governance-backend-error-design.md`。来源: Gitea Issue #7(下游 CHSAnalyzer3 按异常类型分流失败)。**状态: 已批准(2026-08-06,人类逐条拍板 Q1/Q2/Q3),待 `writing-plans`。**
|
||||||
|
|
||||||
|
问题: 限流/熔断状态后端故障时库 fail-closed,一个请求都发不出去——语义上就是 scope 级不可用,但 `GovernanceBackendError` 是 `PolyGatewayError` 的**直接子类**,只写 `except GatewayUnavailableError` 的调用方接不住,于是 Redis 抖一下,积压任务一批批消耗业务失败预算进死信,而那是运维重启就好的故障。
|
||||||
|
|
||||||
|
## 选定方案
|
||||||
|
|
||||||
|
| 决策 | 选定 | 关键理由 |
|
||||||
|
|---|---|---|
|
||||||
|
| A 类型树 | `GovernanceBackendError` 改继承 `GatewayUnavailableError`,`SCOPE_REASONS` 增 `governance_backend_down`,`reason` 恒为该值 | 加父类是**扩大**不是破坏(既有 `except GovernanceBackendError` 照旧命中);库内仅 `telemetry.py:250` 一处捕父类且已并列写两者,**零回归** |
|
||||||
|
| B `retry_after_s` | 模块常量 `GOVERNANCE_BACKEND_RETRY_AFTER_S = 5.0`,非环境配置项 | 后端恢复时间物理上不可知(不同于熔断冷却有确定到期时刻);取 0 会让积压任务零延迟批量重投,把一次故障放大成风暴 |
|
||||||
|
| C scope 来源 | 后端层用 `self._scope`;`QuotaGate`/`BreakerGate` 构造函数注入,三处装配(`retry.py`/`ocr.py`/`embedding.py`)各传一行 | 两个包装器是后端异常的唯一入口,注入点收敛;三处装配本就持有 `self._scope` |
|
||||||
|
| D 未知源拆分 | `_cfg()` 的 2 处改抛新增的 `SourceNotConfiguredError`,**有意不放在** `GatewayUnavailableError` 之下 | 那是装配缺陷不是后端故障;随整类归入"可重投"会让配置写错的任务永远重投、永不进死信——本 issue 要修的 bug 的镜像 |
|
||||||
|
| E message 保全 | `super().__init__()` 后覆写 `self.args = (message,)` | 父类会把 message 覆盖为 `f"{scope} 网关暂时不可用: {reason}"`,而 22 处构造点的诊断串是排障主线索。机制已实跑验证 |
|
||||||
|
|
||||||
|
## 被否决的备选
|
||||||
|
|
||||||
|
| 备选 | 否决原因 |
|
||||||
|
|---|---|
|
||||||
|
| B(issue 原议): 只补文档,类型树不动 | 正确性依赖每个下游都读到那句话;本 issue 本身就是"文档读不出来"引发的,同一失效模式不能用同一种药治 |
|
||||||
|
| C: 在 RetryMW 边界包成 `AllSourcesExhausted` | 比选定方案更具破坏性——下游现有 `except GovernanceBackendError` 直接失效 |
|
||||||
|
| D: 后端层不再构造该异常,原始异常穿透由包装器统一翻译 | 初评时倾向。`redis/limiter.py:133,151` 的 `RedisPermit.release/settle` 依赖 `except GovernanceBackendError` 实现**释放侧降级**,穿透后接不住会破坏该既有行为;改 `except Exception` 则违反 P5 |
|
||||||
|
| `retry_after_s` 复用 `BackpressureConfig.poll_interval_s` | 该值只有三个装配点持有,为此给后端加构造参数等于让状态存储层持有重投策略,违反 P7 |
|
||||||
|
| 新增配置项 `PGW_GOVERNANCE_BACKEND_RETRY_AFTER_S` | YAGNI;无下游表达过需要,真需要时下游可忽略该字段用自有退避 |
|
||||||
|
|
||||||
|
## 对 issue 前提的四处修正
|
||||||
|
|
||||||
|
泄漏路径是**五条**不是两条(判据: 该 gate 调用点是否被 `_record_quietly` 包裹——`QuotaGate` 的 try_acquire / stats / progress_age_s 与 `BreakerGate` 的 try_enter / retry_after_s 均未包裹,直达调用方);构造点 **22 处**;其中 2 处语义完全不同(未知源);`retry_after_s=0` 语义通但工程不通。
|
||||||
|
|
||||||
|
根因记录: `ARCHITECTURE.md` §6.1 错误分类表里 `GovernanceBackendError` **一次都没出现**——它是 M2 引入分布式后端时新增的,当时未回补架构表,于是它在"调用方视角的分类学"中从来没有位置,README 的遗漏是这个遗漏的下游后果。
|
||||||
|
|
||||||
|
## 独立审查修正(2026-08-06, Codex)
|
||||||
|
|
||||||
|
4 条意见逐条核验: 两条"架构文档未同步"实质成立但性质是执行顺序 → 新增 §8.1 钉死"`ARCHITECTURE.md` §6.1 修订先于/同批于实现";"新错误类违反四分类铁律"**部分成立**——铁律论域被误读(`GatewayUnavailableError` 族本就合法处在四分类之外),但原表述确会引起疑虑 → §7 补写三论域划分论证,并把"复用 `RequestRejectedError`"增列为待人类权衡的备选;两条建议性意见(常量非配置项的说明、决策编号 `D`→`Q` 防与架构 D1–D14 混淆)已采纳。
|
||||||
|
|
||||||
|
相关: [[m2-distributed]]、[[m1-core-design]]、[[m25-resilience]]
|
||||||
@@ -0,0 +1,62 @@
|
|||||||
|
---
|
||||||
|
type: design
|
||||||
|
node_id: design:issue10-error-body-retention
|
||||||
|
title: "HTTP 错误响应体留存(Issue #10)"
|
||||||
|
date: 2026-08-16
|
||||||
|
---
|
||||||
|
|
||||||
|
# HTTP 错误响应体留存(Issue #10)
|
||||||
|
|
||||||
|
**来源**: Gitea issue #10(CHSAnalyzer3 现场,1050 张影像批处理中 1 张 400 被判确定性失败、事后无从查证)|**范围**: `errors.py` + 两个 transport|**全文**: `designs/2026-08-16-issue10-error-body-retention-design.md`|**相关**: [[design:issue8-stall-budget]](同为下游实测反馈驱动的治理修正)
|
||||||
|
|
||||||
|
## 问题
|
||||||
|
|
||||||
|
网关拒绝一次调用时,它说的话在 transport 翻译层被丢弃,进程中不再有任何副本:该模块无 logger、异常类无承载字段、库遥测只写 message。三条留存通道同时为空,故"永久查不到"。
|
||||||
|
|
||||||
|
## 根因(Issue 前提的关键修正)
|
||||||
|
|
||||||
|
Issue 建议"给异常加 `body_text` 字段,下游就能记进遥测"——**只做这一半解决不了它自己陈述的痛点**。库的逐次遥测写的是 `error=str(exc)`(`retry.py:552` → `telemetry.py` → `sqlite.py` 的 `error TEXT` 列),即**异常 message**;新增字段不进库的遥测表。下游说的"写进遥测表"是他们自己的埋点。
|
||||||
|
|
||||||
|
缺陷范围也大于 issue 所述:实为 6 处同构——`_status_to_error` 的 400 / 4xx 兜底 / 401·403 / 5xx 四支,`_translate_429` 两支(读了 body 判类型却不带),以及 `monkey_ocr._classify_status` 全部分支(message 只有 `HTTP {status}`)。Issue 场景"读表格"极可能正落在 OCR 路径。
|
||||||
|
|
||||||
|
## 选定方案
|
||||||
|
|
||||||
|
**摘要在翻译层算一次,同一份串同时进 message 与新增的基类字段**——前者解决"事后可查"(走既有遥测列,零 DDL),后者解决下游结构化留存。
|
||||||
|
|
||||||
|
| 决策 | 理由 |
|
||||||
|
|---|---|
|
||||||
|
| 字段加在 `PolyGatewayError` 基类,非 `RequestRejectedError` | 这些错误全由同一个 HTTP 响应翻译而来,"对方说了什么"与"属于哪一类"正交;只加子类,下次给 `SourceDeadError` 加又是一次公共 API 变更 + 人类门 |
|
||||||
|
| 与 `ResultInvalidError.raw_text` 的界限写进 docstring | `body_text` = 非 2xx 的拒绝理由;`raw_text` = 2xx 但不可解析的模型输出。两个"原文字段"不钉死必被混用 |
|
||||||
|
| 新建 `transports/_http_errors.py` 共用摘要口径 | 两 transport 各有分类逻辑(OCR 无 429 细分,有意保留),但摘要必须同一份,否则就是下一个"只修一半" |
|
||||||
|
| `_status_to_error` 改表驱动 | 五分支各拼各的 message,加摘要即五处重复;查表 + 单点拼装后代码更短 |
|
||||||
|
| 摘要 = 折叠空白 + 总长 ≤ 2048,超出则**保留头 1400 + 尾 600**,中段记省略字数 | 折叠是因错误体常是缩进 JSON,拼进 message 会炸成多行。2048 对齐 k8s client-go 的 `maxUnstructuredResponseTextBytes`(唯一同场景先例);**头尾保留取自 `reprlib`**——JSON 错误体的 `code`/`request_id` 收尾,头部硬切正好切掉向网关追查唯一有用的部分(人类质疑 + 2026-08-16 调研,初稿的 500 + 头部硬切已废) |
|
||||||
|
| 429 不设例外 | 例外就是下一个复发点;`insufficient_quota` 那支的配额细节全在 body 里 |
|
||||||
|
| 400 治理语义不动,只补文档 | 见下 |
|
||||||
|
|
||||||
|
## 400 语义:不改行为,本次修复本身就是答复
|
||||||
|
|
||||||
|
Issue 给出有力证据(同字节 15 次重发全成功、`prompt_tokens=0`、2996ms 远低于同批成功最快的 7366ms),说明那次 400 来自中转服务抖动而非坏输入。**仍不改分类**:400 重试对直连供应商是纯浪费,而"中转也回 400"是部署拓扑引入的信息损失,库从状态码无从分辨;默认改可重试 = 让所有直连用户为一种部署形态买单。
|
||||||
|
|
||||||
|
但 body 留存后**下游能自己区分**——中转抖动体与供应商 `invalid_request_error` 体形态不同。库不替下游判断,把判断所需的信息交出去。配套在 docstring 与 ARCHITECTURE §6.2 加一句中转拓扑提醒。
|
||||||
|
|
||||||
|
## 被否决的备选
|
||||||
|
|
||||||
|
| 备选 | 否决理由 |
|
||||||
|
|---|---|
|
||||||
|
| 遥测端口加一列(22 → 23 字段) | 端口签名变更 + 双后端 DDL + 下游 ALTER + 列序契约全线改动,为一个诊断串付出跨三项目迁移;复用既有 `error` 列可达成同样可查证性 |
|
||||||
|
| 只打一条 WARNING 日志(issue 方向二) | 日志轮转后仍查不到,而痛点恰是"事后";且 4xx/5xx 在批处理下可能极高频 |
|
||||||
|
| 截断放进异常构造器 | 下游自建异常的文本被悄悄改写(违反 P4),且 message 侧仍需单独算一次,反出现两条规范化路径 |
|
||||||
|
| 错误分类映射可插拔 | Issue 场景确实指向它,但当前只有一个使用方且已用自己的兜底分类解决;`ProviderProfile` 无此扩展点,加它是子系统级设计。YAGNI |
|
||||||
|
|
||||||
|
## 有意不夹带(留独立 issue)
|
||||||
|
|
||||||
|
- `_status_to_error` 的 `operation` 硬编码 `"chat"`,而 `embed()` 也调它 → embedding 的 HTTP 错误在遥测里被标成 chat。
|
||||||
|
- `_complete_stream` 的 `aread()` 对错误响应体无大小上限,超大错误体可打爆内存(既有风险,留存后更显眼)。
|
||||||
|
|
||||||
|
两项都在本次触及的函数附近,但均不服务本 issue 目标,且各需独立行为讨论。
|
||||||
|
|
||||||
|
## 验收主张
|
||||||
|
|
||||||
|
一次 400 调用后,注入的 recorder 收到的 `error` 串含网关响应体摘要——这条端到端断言是本设计成立与否的唯一硬判据,其余用例为覆盖性(状态码参数化、截断边界 2048/2049、**尾部关键字段可见**、空白折叠、空体不拼悬空分隔符、非 UTF-8 不炸、流式路径、OCR 路径含 `ResponseNotRead` 降级)。
|
||||||
|
|
||||||
|
**发布约束**:版本 1.2.0,且 README 安装 pin 必须由 `==1.1.*` 改为 `>=1.2,<2`——否则照 README 安装的下游静默停在 1.1.2,拿不到本修复。
|
||||||
@@ -0,0 +1,46 @@
|
|||||||
|
---
|
||||||
|
type: design
|
||||||
|
node_id: design:issue8-stall-budget
|
||||||
|
title: "stall 判定改为非生产性等待口径"
|
||||||
|
date: 2026-08-06
|
||||||
|
---
|
||||||
|
|
||||||
|
# stall 判定改为非生产性等待口径
|
||||||
|
|
||||||
|
**全文**: `designs/2026-08-06-issue8-stall-budget-design.md`(已批准 2026-08-06)|**来源**: Gitea issue #8 |**实施**: [[plan:issue8-stall-budget-plan]]
|
||||||
|
|
||||||
|
## 问题
|
||||||
|
|
||||||
|
`timeout_s ≥ stall_window_s` 时,一次耗满超时的请求即判 scope 死,`max_attempts` **静默失效**(无报错无 warning)。`stall_window_s` 默认 300 极易被 `TIMEOUT_S` 追平,"只配 timeout 不配 stall"这种最常见写法正好踩中。
|
||||||
|
|
||||||
|
## 根因
|
||||||
|
|
||||||
|
**两个预算重叠计费**:真实尝试的耗时同时向重试预算(`max_attempts`)与 stall 预算(`stall_window_s`)计费,而后者更小,必然先耗尽。
|
||||||
|
|
||||||
|
## 选定方案
|
||||||
|
|
||||||
|
`StallClock` 让 stall 只累计非生产性等待。**划分依据是"谁消耗重试预算"**,不是"是否发出请求"——烧 `max_attempts` 的时间不烧 `stall_window_s`,不烧 `max_attempts` 的时间(含 429 尝试本身)归 stall 治理。
|
||||||
|
|
||||||
|
关键理由:
|
||||||
|
|
||||||
|
- **消除耦合而非守护耦合**。`stall_window_s` 与 `timeout_s` 自此无关系,配置方不必心算 `stall > timeout × retries`。
|
||||||
|
- **`inf` 语义因此不必改**。新口径下"非生产性排队耗满窗口且 scope 从未出餐"判死本就正当,`inf` 从"有害恒真"回归为"正确的保守默认"。一次改动解决问题,优于两次改动互相牵制。
|
||||||
|
- **取补集实现**(总时间减 `_attempt` 耗时)而非逐处标记 sleep:埋点 7 处降到 3 处,且将来新增等待路径自动计入 stall,默认安全。
|
||||||
|
|
||||||
|
## 被否决的备选
|
||||||
|
|
||||||
|
| 备选 | 否决理由 |
|
||||||
|
|---|---|
|
||||||
|
| **装配期校验 `stall_window_s > max(timeout_s)`**(issue 建议方向 1) | 治标:把缺陷固化成配置契约。且约束值须为 `timeout × max_attempts`(本机 900s),使 stall 兜底迟钝到近乎失效——修好一个洞挖开另一个。仍挡不住残余情形 |
|
||||||
|
| **`inf` 不参与判死**(issue 建议方向 2) | 新口径下 `inf` 已无害。单独改它会制造冷启动兜底真空(429 免预算无其他兜底),并反转 `test_both_windows_exceeded_raises_stalled` 钉住的行为、与 CHS 蓝本分叉 |
|
||||||
|
| **逐处标记 sleep** | 埋点 7 处且默认危险:新增等待路径忘记标记即成 stall 盲区 |
|
||||||
|
| **给 embedding/ocr 补主循环 stall 判定** | 前提不成立。429 免预算是 chat 独有,embedding/ocr 无条件 `fails += 1`,两条循环路径均已封闭,补齐等于凭空新增判死路径 |
|
||||||
|
| **删除既有 ttft 装配校验** | 其理由虽已消失(TTFT 属生产性时间),但校验无害且不误拒合理配置;删除需动 ARCHITECTURE §7.3 契约 G6,超出本 issue 范围(人类定夺:保留并改注释) |
|
||||||
|
|
||||||
|
## 实施期订正(§3.6)
|
||||||
|
|
||||||
|
初稿按"是否发出请求"划分,使 429 尝试**两个预算都不烧**(429 免重试预算,其耗时又算生产性)。排队型网关持满 timeout 才回 429 时实测挂 **25.2 小时**(301 次尝试),而改前只有 301s——**把一个 bug 换成了更严重的 bug**。由独立验证发现。订正为按"谁消耗重试预算"划分,429 尝试耗时退还 stall 账,实测回到 301s。
|
||||||
|
|
||||||
|
## 不变量
|
||||||
|
|
||||||
|
双条件结构、`progress_age_s()` 的 `inf` 语义、429 免预算、退避与 jitter 公式、`fail_fast` 分支、`AllSourcesExhausted` 字段与 `reason` 取值全部未动——**错误面零变更**。
|
||||||
@@ -0,0 +1,62 @@
|
|||||||
|
---
|
||||||
|
type: design
|
||||||
|
node_id: design:issue9-telemetry-ddl-probe
|
||||||
|
title: "建表前先探测,判死只认「确定写不进去」"
|
||||||
|
date: 2026-08-07
|
||||||
|
---
|
||||||
|
|
||||||
|
# 建表前先探测,判死只认「确定写不进去」
|
||||||
|
|
||||||
|
**来源**: Gitea issue #9(CHSAnalyzer3 现场)|**范围**: `telemetry/postgres.py` 单模块,无独立 plan(小改动自判)|**相关**: [[design:response-observability-fields]](issue #3 修的是同一个坑的另一半)
|
||||||
|
|
||||||
|
## 问题
|
||||||
|
|
||||||
|
应用账号有表级 `INSERT`、表也已存在,但没有 schema 的 `CREATE` 权限时,`_ensure_ready()` 的 `CREATE TABLE IF NOT EXISTS` 被拒 → `_failed = True` → **整个进程遥测永久 no-op**。业务调用一切正常,只留一行 warning,从外部完全看不出异常;下游 CHSAnalyzer3 首次端到端跑的 150+ 次调用数据因此全丢且无法补回。
|
||||||
|
|
||||||
|
## 根因
|
||||||
|
|
||||||
|
**PostgreSQL 对 schema 的 CREATE 权限检查早于 `IF NOT EXISTS` 的存在性判断**(`RangeVarGetAndCheckCreationNamespace()` 先 aclcheck 后查 relid)。这与 issue #3 里 `ALTER TABLE` 的 ownership 检查早于 `IF NOT EXISTS` 是同一类问题——当时只修了补列那一半,建表这一半原样留着,于是同一账号形态下"补列失败只丢一行日志接着干活,建表失败却把整个 recorder 判死"。
|
||||||
|
|
||||||
|
**实测(PostgreSQL 16.14,临时角色只授 `SELECT, INSERT ON llm_calls`)**:
|
||||||
|
|
||||||
|
| 语句 | 结果 |
|
||||||
|
|---|---|
|
||||||
|
| `SELECT to_regclass('llm_calls')` | 非 NULL(表就在那儿) |
|
||||||
|
| `CREATE TABLE IF NOT EXISTS llm_calls (...)` | **被拒 InsufficientPrivilegeError: permission denied for schema** |
|
||||||
|
| `INSERT INTO llm_calls ...` | 通过 |
|
||||||
|
| `ALTER TABLE ... ADD COLUMN IF NOT EXISTS` | 被拒 must be owner(即 issue #3 那条) |
|
||||||
|
|
||||||
|
## 选定方案
|
||||||
|
|
||||||
|
两条,第二条才是治本的那条:
|
||||||
|
|
||||||
|
1. **表存在就绝不发 DDL**。探测走 `to_regclass`(不需要任何权限,且与 `INSERT` 走同一套 search_path 解析——裸 `CREATE TABLE` 落在首个**可建**的 schema,可能与写入命中的不是同一张表,故探测优先反而更准)。表不存在才建;新建表列已齐全,顺带跳过补列。
|
||||||
|
2. **"结构性失能"的判据从「初始化时出过异常」收窄为「确定写不进去」**:
|
||||||
|
|
||||||
|
| 情形 | 处置 | 理由 |
|
||||||
|
|---|---|---|
|
||||||
|
| 建池失败 | 永久 no-op | 重试要在业务调用路径上内联吞掉 connect 超时 |
|
||||||
|
| 表存在 | 不发 DDL,只补列(失败仅 warning) | 本 issue 的直接修复 |
|
||||||
|
| 表不存在 → 建表成功 | 就绪,跳过补列 | 新建表列已齐 |
|
||||||
|
| 表不存在 → 建表失败 | 永久 no-op | 后续 INSERT 必然全败,重试无意义、日志纯噪音 |
|
||||||
|
| 探测/取连接失败 | 只跳过本条,下次调用重试 | 瞬时抖动,判死代价远大于多一次往返 |
|
||||||
|
|
||||||
|
**SQLite 侧有意不对称**:实测其对已存在的表在**解析期**就把 `CREATE TABLE IF NOT EXISTS` 短路掉——另一连接持 `BEGIN EXCLUSIVE`、或文件 `chmod 444` 时该语句均通过(同条件下 `INSERT` 与新表名建表分别报 database is locked / readonly database),既不抢写锁也不检查可写性。故 PG 侧的坑在此不存在,加探测零收益。**需要对称的是保证(表存在就不该因建表失败而失能),不是代码**;结论已钉进 `sqlite.py` 模块 docstring,防止后人为"对称"加回来。
|
||||||
|
|
||||||
|
## 被否决的备选
|
||||||
|
|
||||||
|
| 备选 | 否决理由 |
|
||||||
|
|---|---|
|
||||||
|
| 只加探测,`_failed` 语义不动(issue 原方案) | 治标。初始化瞬间的 DB 抖动、一次 `pool.acquire` 失败、search_path 配错仍会让整个进程永久失遥测——同一个开关,换个触发口 |
|
||||||
|
| 除建池外一律不判死 | 方向最统一,但表真的不存在时每次调用都发一条注定失败的 INSERT + 一条 warning(150 次调用 = 150 行噪音),而这种情形是**可确定判定**的,没必要留活路 |
|
||||||
|
| 捕获 `InsufficientPrivilegeError` 特判放行 | 按异常类型打补丁,漏一种错误码就复发;探测是把"该不该发这条 DDL"判断在前,与错误面无关 |
|
||||||
|
| SQLite 侧同步加探测 | 实测证明零收益,属为对称而对称的 gold-plating |
|
||||||
|
|
||||||
|
## 遗留
|
||||||
|
|
||||||
|
**SQLite 的窄缝**:表不存在 + 构造瞬间库被排他锁(多进程共库)→ `__init__` 里的建表失败 → recorder 永久失能。修它要把 SQLite 也改成 lazy 重试结构,超出本 issue 范围,记此备查。
|
||||||
|
|
||||||
|
## 测试证据
|
||||||
|
|
||||||
|
- 单测 `TestPostgresTableProbe`(5 例,fake conn):表存在不发 DDL / DDL 被拒仍照常 INSERT 且 `_failed` 不置位 / 表缺失则建表且不补列 / 表缺失且建不出来才判死 / 探测失败下次重试。
|
||||||
|
- 集成 `TestLeastPrivilegeDeployment`(真实 PG,临时 schema + 临时角色,teardown 删净):先钉死"该角色确实建不了表"这条库外事实,再验两行记录照常落库。**修复前该用例复现 issue 原文那行 warning 并失败**。
|
||||||
@@ -0,0 +1,195 @@
|
|||||||
|
---
|
||||||
|
type: finding
|
||||||
|
node_id: finding:2026-08-02-thinking-switch-and-reasoning-tokens
|
||||||
|
title: "推理开关与 reasoning_tokens: 供应商实测与业界做法"
|
||||||
|
date: 2026-08-02
|
||||||
|
---
|
||||||
|
|
||||||
|
# 推理开关与 reasoning_tokens:供应商实测与业界做法
|
||||||
|
|
||||||
|
> 类型:findings(事实基础)|日期:2026-08-02|来源:issue #5 / #6 调研
|
||||||
|
> 本文只记录**已验证的事实与其证据**,设计取舍见 `designs/2026-08-02-thinking-capability-design.md`。
|
||||||
|
> 本文的价值不限于这两条 issue——「同一语义、形态因模型而异」是本库长期要面对的一类问题,此处的结论与方法可复用。
|
||||||
|
|
||||||
|
## 1. 实验环境与方法
|
||||||
|
|
||||||
|
| 项 | 值 |
|
||||||
|
|---|---|
|
||||||
|
| 端点 | 自建 new-api 中转(`newapi.iomgaa.online/v1`,OpenAI 兼容) |
|
||||||
|
| 参数 | `temperature=0`、`max_tokens=800`、非流式为主,流式单独验证 |
|
||||||
|
| 题目 | 固定一道鸡兔同笼题,要求"只输出两个数字" |
|
||||||
|
| 判据 | `usage.completion_tokens_details.reasoning_tokens`(**唯一可靠的判别量**,见 §2.5) |
|
||||||
|
| 旁证 | `prompt_tokens` 变化——注入生效的参数会改变模型侧模板,输入侧 token 数随之变化 |
|
||||||
|
|
||||||
|
**方法论要点(可复用)**:判断一个参数"是否被上游真正消费",`prompt_tokens` 比输出长度可靠得多。输出长度受采样影响、方差大;而输入侧 token 数在同一请求体下是确定的,一旦变化就说明服务端换了模板,即参数确实到达了模型。本次三条关键结论全部由这个旁证锁定。
|
||||||
|
|
||||||
|
## 2. MiniMax:真开关是 `reasoning_effort`
|
||||||
|
|
||||||
|
### 2.1 M3 参数矩阵(非流式)
|
||||||
|
|
||||||
|
| 注入参数 | prompt | completion | reasoning_tokens | 判定 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| 默认(不传) | 194 | 4 | 无 ctd | 不推理 |
|
||||||
|
| `reasoning_effort=none` | 194 | 10 | 无 ctd | 不推理 |
|
||||||
|
| `reasoning_effort=minimal` | **207** | 129 | 123 | 推理 |
|
||||||
|
| `reasoning_effort=low` | **207** | 98 | 93 | 推理 |
|
||||||
|
| `reasoning_effort=medium` | **207** | 183 | 177 | 推理 |
|
||||||
|
| `reasoning_effort=high` | **207** | 158 | 142 | 推理 |
|
||||||
|
| `thinking={"type":"enabled"}` | 194 | 5 | 无 ctd | **被静默丢弃** |
|
||||||
|
| `thinking={"type":"disabled"}` | 194 | 4 | 无 ctd | **被静默丢弃** |
|
||||||
|
| `enable_thinking=true` | 194 | 5 | 无 ctd | **被静默丢弃** |
|
||||||
|
| `enable_thinking=false` | 194 | 5 | 无 ctd | **被静默丢弃** |
|
||||||
|
|
||||||
|
`prompt_tokens` 194→207 的 13 token 差是硬证据:`reasoning_effort` 被消费时模型注入了推理指令;另四种写法 prompt 恒为 194,参数根本没到达模型。
|
||||||
|
|
||||||
|
### 2.2 `none` 是被识别的真值,不是被当非法值丢弃
|
||||||
|
|
||||||
|
这是一个必须排除的伪解释——若中转把不认识的值直接丢掉,`none` 的表现会与"不传"无异,我们就会误以为它生效。
|
||||||
|
|
||||||
|
反证实验:传乱码值 `reasoning_effort="xyzzy"` → 返回 200、prompt=207、reasoning_tokens=180。**未知值不但没被丢弃,反而开启了推理。** 既然无效值的行为是"开推理",而 `none` 的行为是"不推理",两者不同,`none` 就必然是被识别的枚举值。
|
||||||
|
|
||||||
|
对照组:完全未知的**键** `zzz_bogus_param=1` → prompt=194、无 ctd、无报错,确认未知**键**才会被静默吞掉。
|
||||||
|
|
||||||
|
### 2.3 M2.7 / M2.5 的推理关不掉
|
||||||
|
|
||||||
|
三种参数形态各 3 次,`completion_tokens` 全部落在推理区间:
|
||||||
|
|
||||||
|
| 模型 | 默认(基线) | `reasoning_effort=none` | `thinking:{disabled}` | `thinking:{adaptive}` |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| MiniMax-M2.7 | 372/283/285 | 275/301/248 | 310/190/219 | 299/269/246 |
|
||||||
|
| MiniMax-M2.5 | 273/–/256 | 363/353/264 | 286/278/320 | 278/228/259 |
|
||||||
|
|
||||||
|
真关闭应为 5–10("23 12" 两个数字),实测无一接近。
|
||||||
|
|
||||||
|
**三个独立外部来源与实测完全吻合**:
|
||||||
|
|
||||||
|
| 来源 | M3 | M2.7 / M2.5 |
|
||||||
|
|---|---|---|
|
||||||
|
| OpenRouter `/api/v1/models` 的 `reasoning` 描述符 | `mandatory: false` | **`mandatory: true`** |
|
||||||
|
| models.dev 的 `reasoning_options` | `[{"type":"toggle"}]`(二元可控) | `[]`(有推理但无控制手段) |
|
||||||
|
| MiniMax 官方仓库 issue #121 | — | "M2.7 不允许关闭思考",无官方回复 |
|
||||||
|
|
||||||
|
**结论:M2.x 的推理是模型固有属性,不是参数没找对。** 任何库层改动都无法让它关闭;唯一诚实的做法是如实报错。
|
||||||
|
|
||||||
|
### 2.4 M3 的稳定性
|
||||||
|
|
||||||
|
同一请求打 10 次,`(prompt_tokens, 是否上报 ctd)` 全部为 `(194, False)`,零跳变——`enable_thinking=False` 的修复可以建立在 M3 上。
|
||||||
|
|
||||||
|
### 2.5 输出长度不是有效判别量(2026-08-02 e2e 补测,各 15 轮)
|
||||||
|
|
||||||
|
初版判据用 `completion_tokens` 阈值区分推理开关,被自己的数据证伪:
|
||||||
|
|
||||||
|
| 档位 | `completion_tokens` 观测范围 | `reasoning_tokens` |
|
||||||
|
|---|---|---|
|
||||||
|
| 关闭(`reasoning_effort=none`) | 4 – **46** | 15/15 轮为 `None` |
|
||||||
|
| 开启(`medium`) | **13** – 186 | 15/15 轮 > 0 |
|
||||||
|
|
||||||
|
**两档的输出长度分布重叠**:关闭档偶尔到 46(模型没照做「只输出两个数字」,把解题过程写进了正文——那是正文不是推理);开启档最低到 13(medium 档想得少的轮次)。按长度阈值判,两个方向都会误判。
|
||||||
|
|
||||||
|
而 `reasoning_tokens` 在同一批 30 轮里干净分开。**这条对下游同样成立**:想判断某次调用是否发生了推理,只能看 `reasoning_tokens`,不能看输出长度。
|
||||||
|
|
||||||
|
另有一个不含魔数的确定性锚点:同一模型上关闭档的 `prompt_tokens` 严格小于开启档(实测 194 < 207),因为供应商在开启时向模板注入了推理指令。这是相对比较,供应商改模板也不会失效。
|
||||||
|
|
||||||
|
## 3. qwen / deepseek:现有 profile 正确
|
||||||
|
|
||||||
|
| 模型 | `enable_thinking=false` | `thinking:{disabled}` | `reasoning_effort=none` | 现有 profile |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| qwen3.7-plus | ✅ 关闭(compl 5) | ✅ 关闭 | ✅ 关闭 | `enable_thinking` — **正确** |
|
||||||
|
| deepseek-v4-pro | ❌ 无效(仍推理 198) | ✅ 关闭(compl 3) | ✅ 关闭 | `thinking:{type}` — **正确** |
|
||||||
|
|
||||||
|
两点附带事实:
|
||||||
|
|
||||||
|
- **`reasoning_effort=none` 在三家都有效**,但这很可能是中转做了参数归一化。**不可据此认为可以统一发一个参数**——下游若直连供应商官方端点,该假设大概率不成立。翻译表必须一家一行。
|
||||||
|
- **qwen 的 `strip_think_tags=True` 已过时**:实测 qwen 走 `reasoning_content` 字段,正文中无 `<think>` 标签。无害,但属于死代码。
|
||||||
|
- **非流式没有 400**:DashScope 系"`enable_thinking` 仅支持流式"的限制经中转不存在。直连时是否仍存在未验证。
|
||||||
|
|
||||||
|
## 4. new-api 中转的三个行为(会污染观测)
|
||||||
|
|
||||||
|
这一节对任何经中转做实测的场景都适用,值得单独记住。
|
||||||
|
|
||||||
|
**(a)不校验参数值。** `reasoning_effort="xyzzy"` 返回 200 并当作"开推理"处理。**意味着"靠上游报错兜底"的设计模式在此失效**——Bedrock 式的"最小交集 + 裸逃生口"在这里等于零保护。
|
||||||
|
|
||||||
|
**(b)静默丢弃未知键。** 默认路径是 struct round-trip(`ConvertRequest` 返回 struct 再 `json.Marshal`),未知键在第一次序列化就消失。new-api 有 per-channel 的 `pass_through_body_enabled` 开关可改变此行为。
|
||||||
|
|
||||||
|
**(c)上游不返回 usage 时用本地 tokenizer 补算并整体替换。** 补算出的 usage 只有三个标量,`completion_tokens_details` 为零值。这直接解释了实测中的双峰现象:
|
||||||
|
|
||||||
|
| 现象 | 解释 |
|
||||||
|
|---|---|
|
||||||
|
| 同一请求 10 次:`prompt=74` 者 6 次不上报 `reasoning_tokens`,`prompt=72` 者 4 次上报,从不交叉 | `74` = 本地估算值,`72` = 上游真值;补算路径吃掉了 ctd |
|
||||||
|
|
||||||
|
**这不是多渠道路由**(MiniMax 侧为单渠道单密钥),也不是配置错误,而是上游偶发不返回 usage 时的兜底逻辑。中转日志中的 `local_count_tokens` 标志可现场确认。
|
||||||
|
|
||||||
|
**对库的直接影响**:`reasoning_tokens` 缺失**不能**解释为"该源不上报这个字段",只能解释为"**本次调用未上报**"。下游若按前者建立统计口径会算错。
|
||||||
|
|
||||||
|
## 5. 业界如何建模"同一语义、形态因模型而异"
|
||||||
|
|
||||||
|
调研覆盖 LiteLLM、OpenRouter、models.dev、LangChain、Vercel AI SDK、AWS Bedrock Converse、Portkey、Helicone、LlamaIndex、new-api/one-api。
|
||||||
|
|
||||||
|
### 5.1 核心共识:形态按 provider,能力按 model
|
||||||
|
|
||||||
|
| 概念 | 变化频率 | 应归属层次 |
|
||||||
|
|---|---|---|
|
||||||
|
| **形态**:参数长什么样(`enable_thinking` / `thinking.type` / `reasoning_effort`) | 协议方言,一个供应商数年不变 | provider 级 |
|
||||||
|
| **能力**:能否关闭、有几档、默认开不开 | 模型属性,同一供应商每代都变 | **model 级** |
|
||||||
|
|
||||||
|
注册单位的分布很能说明问题:LiteLLM(2986 条目)、models.dev(5949 条)、LangChain、OpenRouter(细到 endpoint)、Helicone 全部下沉到 model 级;**仍停在 provider 级的只有 Portkey 与 LlamaIndex,而这两家恰是失败语义最差的两家(均静默丢弃)**。二者相关不是偶然:注册单位不够细,就只能靠"表里没有 = 不发"来兜底,而这正是静默失效的成因。
|
||||||
|
|
||||||
|
### 5.2 失败语义的四种谱系
|
||||||
|
|
||||||
|
| 语义 | 代表 | 适用前提 |
|
||||||
|
|---|---|---|
|
||||||
|
| 默认报错 + 可配置降级开关 | LiteLLM(`UnsupportedParamsError` + `drop_params`) | 有 model 级能力表可依据 |
|
||||||
|
| 软降级 + 显式 warning 通道 | Vercel AI SDK(丢弃参数并 push `warnings[]`) | 调用方愿意读 warning |
|
||||||
|
| 静默忽略 + 可选路由过滤 | OpenRouter(默认忽略;`require_parameters:true` 改为排除不支持的上游) | 网关自己拥有路由权 |
|
||||||
|
| 硬失败(透传给上游报错) | Bedrock(`inferenceConfig` 4 字段交集 + `additionalModelRequestFields` 裸透传) | **上游会诚实报错** |
|
||||||
|
|
||||||
|
**选型时先问"我的上游会不会诚实报错"**。若不会(如本项目的中转),最后一种直接出局,静默类也不能选。
|
||||||
|
|
||||||
|
### 5.3 表会过期,这是公理
|
||||||
|
|
||||||
|
LiteLLM 有过真实事故(issue #27351:`gpt-5.1-mini` 漏登记导致 `temperature` 被误拒)。它的应对是**两种相反极性**,值得直接借鉴:
|
||||||
|
|
||||||
|
- **opt-in 能力**(用错会 400 或悄悄花钱):未登记 → 视作不支持 → 拒绝
|
||||||
|
- **opt-out 能力**(多半支持,误拒代价大):未登记 → 放行 → 只有表里显式写 `false` 才拒
|
||||||
|
|
||||||
|
维护方式上,LiteLLM/models.dev 靠社区 PR + CI 校验,LangChain 靠"上游拉取 + 本地增补 + 代码生成"。**对内部库而言唯一现实的答案是:谁实测出来谁登记,登记必须附实测证据与日期。**
|
||||||
|
|
||||||
|
### 5.4 「布尔开关 → 多档旋钮」无语义共识
|
||||||
|
|
||||||
|
| 系统 | effort → 预算的换算 |
|
||||||
|
|---|---|
|
||||||
|
| LiteLLM | 一组 2 的幂(1024/2048/4096/8192/16384),全部可用环境变量覆盖;gemini 各型号还另有分叉 |
|
||||||
|
| OpenRouter | `max_tokens` 的百分比(≈80%/50%/20%) |
|
||||||
|
| Helicone | 一律 `max_tokens/2`,完全不看档位 |
|
||||||
|
| LangChain | 明确不保证跨 provider 可比 |
|
||||||
|
|
||||||
|
**唯一对齐的是"关"**:`none` / `disabled` / `thinking:{type:"disabled"}` / OpenRouter `effort:"none"` 语义一致。"开"那一端没有任何标准。
|
||||||
|
|
||||||
|
**工程共识只有一条:这个映射必须是可覆盖的常量,不是可推导的公式。** 业界所有人都在拍脑袋,区别只在拍完让不让调用方改。
|
||||||
|
|
||||||
|
### 5.5 Vercel AI SDK 的一处设计值得单记
|
||||||
|
|
||||||
|
它的推理档位枚举里有一个 `'provider-default'`,与 `'none'`(明确关闭)严格区分。这与本库 `enable_thinking` 的三态(`None` 不干预 / `True` / `False`)是同一思想——**"调用方不表态"必须是一个独立的值,不能与任何具体档位混同**。本库这一点原本就做对了,应保持。
|
||||||
|
|
||||||
|
## 6. 附带发现(不属本次范围,建议另立 issue)
|
||||||
|
|
||||||
|
**kimi-k3 拒绝 `temperature=0`**:返回 `400 invalid temperature: only 1 is supported`(另有渠道回 `only 0.6`)。本库把 400 归入 `RequestRejectedError`——不重试、不换源。若下游统一下发 `temperature=0`,此类源会 100% 硬失败。这与本次两条 issue 同源:**供应商能力差异未被建模**。
|
||||||
|
|
||||||
|
**中转渠道可用性会波动**:kimi 渠道在 429 后被中转下线,随后返回 `404 Model not supported by any channel`。任何依赖真实 API 的测试都必须容忍源不可用(跳过并给出明确原因),而不是失败。
|
||||||
|
|
||||||
|
## 7. 未能证实
|
||||||
|
|
||||||
|
1. **MiniMax 官方文档对 `reasoning_effort` 的一手定义**:官方文档站三次抓取均失败。M2.x 关不掉有三处佐证,但官方原文未取得。另有二手来源称 MiniMax 原生开关是 `thinking:{type:"adaptive"/"disabled"}`——**该说法已被本次实测证伪**(M2.7/M2.5 上两种写法均无效),但"中转是否对 `reasoning_effort` 做了改写"仍未排除。直连官方端点复测可彻底澄清。
|
||||||
|
2. **qwen 直连 DashScope 时非流式 `enable_thinking` 是否仍报 400**:仅验证了经中转的行为。
|
||||||
|
3. **new-api 走本地补算的确切触发条件**:读到了补算分支与 `local_count_tokens` 标记,未逐条比对所有渠道类型。双峰现象与该解释高度吻合,但未在日志中直接验证。
|
||||||
|
4. **能力表条目对非本次实测模型的正确性**:qwen / deepseek 只测了各一个型号,同系其他型号未验证。
|
||||||
|
|
||||||
|
## 8. 对后续开发的指导
|
||||||
|
|
||||||
|
1. **判定参数是否生效,优先看 `prompt_tokens` 而非输出长度**(§1)。
|
||||||
|
2. **排除"无效值被静默丢弃"必须做反证实验**:传一个乱码值,看它的行为是否与目标值不同(§2.2)。
|
||||||
|
3. **经中转做的任何实测都要标注"经中转,直连未验证"**,并写进注释(§3、§7)。
|
||||||
|
4. **新增供应商或模型前,先查 OpenRouter `/api/v1/models` 与 models.dev**——它们的登记与本次实测 100% 吻合,可作为低成本预判,但不可作为运行时依赖。
|
||||||
|
5. **能力表条目必须附实测证据与日期**;表过期是必然事件,退化路径与漂移检测要一起设计(§5.3)。
|
||||||
|
6. **`reasoning_tokens` 缺失只能记 `None`,绝不可记 `0`**(§4c)——"观测不到"与"没发生"是两件事。
|
||||||
|
7. **判断"是否发生了推理"只能看 `reasoning_tokens`,不能看输出长度**(§2.5)——两档的 `completion_tokens` 分布是重叠的,长度阈值两个方向都会误判。
|
||||||
@@ -130,6 +130,36 @@
|
|||||||
"id": "plan:sampling-params-plan",
|
"id": "plan:sampling-params-plan",
|
||||||
"label": "采样参数透传实现计划(issue #4)",
|
"label": "采样参数透传实现计划(issue #4)",
|
||||||
"type": "plan"
|
"type": "plan"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "design:governance-backend-error",
|
||||||
|
"label": "治理后端故障归位为 scope 级不可用(Issue #7)",
|
||||||
|
"type": "design"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "plan:governance-backend-error",
|
||||||
|
"label": "实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)",
|
||||||
|
"type": "plan"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "design:issue8-stall-budget",
|
||||||
|
"label": "stall 判定改为非生产性等待口径",
|
||||||
|
"type": "design"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "plan:issue8-stall-budget-plan",
|
||||||
|
"label": "issue #8 实施计划: stall 非生产性等待口径",
|
||||||
|
"type": "plan"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "design:issue10-error-body-retention",
|
||||||
|
"label": "HTTP 错误响应体留存(Issue #10)",
|
||||||
|
"type": "design"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"id": "plan:issue10-error-body-retention-plan",
|
||||||
|
"label": "实现计划: HTTP 错误响应体留存(Issue #10)",
|
||||||
|
"type": "plan"
|
||||||
}
|
}
|
||||||
],
|
],
|
||||||
"links": [
|
"links": [
|
||||||
@@ -223,6 +253,41 @@
|
|||||||
"relation": "implements",
|
"relation": "implements",
|
||||||
"evidence": "11 个任务逐条覆盖设计的决策 A-G 与 §5 的 14 条测试清单",
|
"evidence": "11 个任务逐条覆盖设计的决策 A-G 与 §5 的 14 条测试清单",
|
||||||
"added": "2026-07-31T16:59:35.657367+00:00"
|
"added": "2026-07-31T16:59:35.657367+00:00"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"source": "finding:2026-08-02-thinking-switch-and-reasoning-tokens",
|
||||||
|
"target": "design:2026-08-02-thinking-capability-design",
|
||||||
|
"relation": "supports",
|
||||||
|
"evidence": "供应商实测与业界调研为该设计的形态/能力分层与失败语义提供事实依据",
|
||||||
|
"added": "2026-08-02T09:38:57.033054+00:00"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"source": "plan:2026-08-02-thinking-capability",
|
||||||
|
"target": "design:2026-08-02-thinking-capability-design",
|
||||||
|
"relation": "implements",
|
||||||
|
"evidence": "T1-T10 逐条实现设计的 D1-D6 六个决策与 §11 九条验收标准",
|
||||||
|
"added": "2026-08-02T09:49:48.126539+00:00"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"source": "plan:governance-backend-error",
|
||||||
|
"target": "design:governance-backend-error",
|
||||||
|
"relation": "implements",
|
||||||
|
"evidence": "T1-T5 逐任务实现设计 §3 的五项决策与 §8 影响面清单",
|
||||||
|
"added": "2026-08-06T08:08:51.865565+00:00"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"source": "plan:issue8-stall-budget-plan",
|
||||||
|
"target": "design:issue8-stall-budget",
|
||||||
|
"relation": "implements",
|
||||||
|
"evidence": "T1-T6 实施该设计,含 §3.6 订正",
|
||||||
|
"added": "2026-08-06T14:58:01.673693+00:00"
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"source": "plan:issue10-error-body-retention-plan",
|
||||||
|
"target": "design:issue10-error-body-retention",
|
||||||
|
"relation": "implements",
|
||||||
|
"evidence": "7 任务覆盖设计 G1-G4 与 §7 全部验收用例",
|
||||||
|
"added": "2026-08-16T09:50:57.830855+00:00"
|
||||||
}
|
}
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
+21
-5
@@ -1,8 +1,8 @@
|
|||||||
# Research Wiki 索引
|
# Research Wiki 索引
|
||||||
|
|
||||||
> 自动生成,更新时间:2026-08-01 01:58 UTC
|
> 自动生成,更新时间:2026-08-16 09:50 UTC
|
||||||
|
|
||||||
## design (20)
|
## design (28)
|
||||||
- [2026-07-20-m1-core-design](designs/2026-07-20-m1-core-design.md) `design:2026-07-20-m1-core-design`
|
- [2026-07-20-m1-core-design](designs/2026-07-20-m1-core-design.md) `design:2026-07-20-m1-core-design`
|
||||||
- [2026-07-20-m2-distributed-design](designs/2026-07-20-m2-distributed-design.md) `design:2026-07-20-m2-distributed-design`
|
- [2026-07-20-m2-distributed-design](designs/2026-07-20-m2-distributed-design.md) `design:2026-07-20-m2-distributed-design`
|
||||||
- [2026-07-21-m25-resilience-design](designs/2026-07-21-m25-resilience-design.md) `design:2026-07-21-m25-resilience-design`
|
- [2026-07-21-m25-resilience-design](designs/2026-07-21-m25-resilience-design.md) `design:2026-07-21-m25-resilience-design`
|
||||||
@@ -13,18 +13,26 @@
|
|||||||
- [2026-07-30-settings-invariants-round-2-design](designs/2026-07-30-settings-invariants-round-2-design.md) `design:2026-07-30-settings-invariants-round-2-design`
|
- [2026-07-30-settings-invariants-round-2-design](designs/2026-07-30-settings-invariants-round-2-design.md) `design:2026-07-30-settings-invariants-round-2-design`
|
||||||
- [2026-07-31-response-observability-fields-design](designs/2026-07-31-response-observability-fields-design.md) `design:2026-07-31-response-observability-fields-design`
|
- [2026-07-31-response-observability-fields-design](designs/2026-07-31-response-observability-fields-design.md) `design:2026-07-31-response-observability-fields-design`
|
||||||
- [2026-07-31-sampling-params-design](designs/2026-07-31-sampling-params-design.md) `design:2026-07-31-sampling-params-design`
|
- [2026-07-31-sampling-params-design](designs/2026-07-31-sampling-params-design.md) `design:2026-07-31-sampling-params-design`
|
||||||
|
- [2026-08-06-governance-backend-error-design](designs/2026-08-06-governance-backend-error-design.md) `design:2026-08-06-governance-backend-error-design`
|
||||||
|
- [2026-08-06-issue8-stall-budget-design](designs/2026-08-06-issue8-stall-budget-design.md) `design:2026-08-06-issue8-stall-budget-design`
|
||||||
|
- [2026-08-16-issue10-error-body-retention-design](designs/2026-08-16-issue10-error-body-retention-design.md) `design:2026-08-16-issue10-error-body-retention-design`
|
||||||
- [est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)](designs/est-tokens-decoupling.md) `design:est-tokens-decoupling`
|
- [est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)](designs/est-tokens-decoupling.md) `design:est-tokens-decoupling`
|
||||||
- [GatewaySettings 装配校验补齐(第二轮)](designs/settings-invariants-round-2.md) `design:settings-invariants-round-2`
|
- [GatewaySettings 装配校验补齐(第二轮)](designs/settings-invariants-round-2.md) `design:settings-invariants-round-2`
|
||||||
- [GatewaySettings 跨字段不变量守卫的生效范围](designs/settings-invariant-guards.md) `design:settings-invariant-guards`
|
- [GatewaySettings 跨字段不变量守卫的生效范围](designs/settings-invariant-guards.md) `design:settings-invariant-guards`
|
||||||
|
- [HTTP 错误响应体留存(Issue #10)](designs/issue10-error-body-retention.md) `design:issue10-error-body-retention`
|
||||||
- [M1 核心里程碑设计:公共签名冻结与治理栈落地](designs/m1-core-design.md) `design:m1-core-design`
|
- [M1 核心里程碑设计:公共签名冻结与治理栈落地](designs/m1-core-design.md) `design:m1-core-design`
|
||||||
- [M2 分布式:Redis 治理后端+背压+Postgres 遥测+pricing+Embedding+压测 harness](designs/m2-distributed.md) `design:m2-distributed`
|
- [M2 分布式:Redis 治理后端+背压+Postgres 遥测+pricing+Embedding+压测 harness](designs/m2-distributed.md) `design:m2-distributed`
|
||||||
- [M2.5 治理韧性: 半死源隔离与健康感知调度](designs/m25-resilience.md) `design:m25-resilience`
|
- [M2.5 治理韧性: 半死源隔离与健康感知调度](designs/m25-resilience.md) `design:m25-resilience`
|
||||||
- [M3 OCR 端口族设计](designs/m3-ocr.md) `design:m3-ocr`
|
- [M3 OCR 端口族设计](designs/m3-ocr.md) `design:m3-ocr`
|
||||||
- [M4 迁移验证设计(GovDoc→CHS,发 v1.0)](designs/m4-migration.md) `design:m4-migration`
|
- [M4 迁移验证设计(GovDoc→CHS,发 v1.0)](designs/m4-migration.md) `design:m4-migration`
|
||||||
|
- [stall 判定改为非生产性等待口径](designs/issue8-stall-budget.md) `design:issue8-stall-budget`
|
||||||
- [响应可观测字段扩展(Issue #3)](designs/response-observability-fields.md) `design:response-observability-fields`
|
- [响应可观测字段扩展(Issue #3)](designs/response-observability-fields.md) `design:response-observability-fields`
|
||||||
|
- [建表前先探测,判死只认「确定写不进去」](designs/issue9-telemetry-ddl-probe.md) `design:issue9-telemetry-ddl-probe`
|
||||||
|
- [推理开关能力建模与 reasoning_tokens 采集(issue #5 + #6)](designs/2026-08-02-thinking-capability-design.md) `design:2026-08-02-thinking-capability-design`
|
||||||
|
- [治理后端故障归位为 scope 级不可用(Issue #7)](designs/governance-backend-error.md) `design:governance-backend-error`
|
||||||
- [采样参数透传设计(issue #4)](designs/sampling-params.md) `design:sampling-params`
|
- [采样参数透传设计(issue #4)](designs/sampling-params.md) `design:sampling-params`
|
||||||
|
|
||||||
## finding (11)
|
## finding (12)
|
||||||
- [2026-07-20-m2-soak-workload](findings/2026-07-20-m2-soak-workload.md) `finding:2026-07-20-m2-soak-workload`
|
- [2026-07-20-m2-soak-workload](findings/2026-07-20-m2-soak-workload.md) `finding:2026-07-20-m2-soak-workload`
|
||||||
- [2026-07-21-m25-acceptance](findings/2026-07-21-m25-acceptance.md) `finding:2026-07-21-m25-acceptance`
|
- [2026-07-21-m25-acceptance](findings/2026-07-21-m25-acceptance.md) `finding:2026-07-21-m25-acceptance`
|
||||||
- [2026-07-21-p6-soak-baseline](findings/2026-07-21-p6-soak-baseline.md) `finding:2026-07-21-p6-soak-baseline`
|
- [2026-07-21-p6-soak-baseline](findings/2026-07-21-p6-soak-baseline.md) `finding:2026-07-21-p6-soak-baseline`
|
||||||
@@ -36,8 +44,9 @@
|
|||||||
- [M4 迁移验收(GovDoc+CHS)](findings/m4-acceptance.md) `finding:m4-acceptance`
|
- [M4 迁移验收(GovDoc+CHS)](findings/m4-acceptance.md) `finding:m4-acceptance`
|
||||||
- [P6 混合浸泡首跑基线与记分板三重伪击穿修复](findings/p6-soak-baseline.md) `finding:p6-soak-baseline`
|
- [P6 混合浸泡首跑基线与记分板三重伪击穿修复](findings/p6-soak-baseline.md) `finding:p6-soak-baseline`
|
||||||
- [P7 OCR soak 验收: 99.73% 与 13 不变量全 PASS](findings/p7-ocr-soak.md) `finding:p7-ocr-soak`
|
- [P7 OCR soak 验收: 99.73% 与 13 不变量全 PASS](findings/p7-ocr-soak.md) `finding:p7-ocr-soak`
|
||||||
|
- [推理开关与 reasoning_tokens: 供应商实测与业界做法](findings/2026-08-02-thinking-switch-and-reasoning-tokens.md) `finding:2026-08-02-thinking-switch-and-reasoning-tokens`
|
||||||
|
|
||||||
## plan (16)
|
## plan (23)
|
||||||
- [2026-07-20-m1-core-plan](plans/2026-07-20-m1-core-plan.md) `plan:2026-07-20-m1-core-plan`
|
- [2026-07-20-m1-core-plan](plans/2026-07-20-m1-core-plan.md) `plan:2026-07-20-m1-core-plan`
|
||||||
- [2026-07-20-m2-distributed-plan](plans/2026-07-20-m2-distributed-plan.md) `plan:2026-07-20-m2-distributed-plan`
|
- [2026-07-20-m2-distributed-plan](plans/2026-07-20-m2-distributed-plan.md) `plan:2026-07-20-m2-distributed-plan`
|
||||||
- [2026-07-21-m25-resilience-plan](plans/2026-07-21-m25-resilience-plan.md) `plan:2026-07-21-m25-resilience-plan`
|
- [2026-07-21-m25-resilience-plan](plans/2026-07-21-m25-resilience-plan.md) `plan:2026-07-21-m25-resilience-plan`
|
||||||
@@ -46,17 +55,24 @@
|
|||||||
- [2026-07-30-est-tokens-decoupling-plan](plans/2026-07-30-est-tokens-decoupling-plan.md) `plan:2026-07-30-est-tokens-decoupling-plan`
|
- [2026-07-30-est-tokens-decoupling-plan](plans/2026-07-30-est-tokens-decoupling-plan.md) `plan:2026-07-30-est-tokens-decoupling-plan`
|
||||||
- [2026-07-31-response-observability-fields](plans/2026-07-31-response-observability-fields.md) `plan:2026-07-31-response-observability-fields`
|
- [2026-07-31-response-observability-fields](plans/2026-07-31-response-observability-fields.md) `plan:2026-07-31-response-observability-fields`
|
||||||
- [2026-07-31-sampling-params](plans/2026-07-31-sampling-params.md) `plan:2026-07-31-sampling-params`
|
- [2026-07-31-sampling-params](plans/2026-07-31-sampling-params.md) `plan:2026-07-31-sampling-params`
|
||||||
|
- [2026-08-06-governance-backend-error-plan](plans/2026-08-06-governance-backend-error-plan.md) `plan:2026-08-06-governance-backend-error-plan`
|
||||||
|
- [2026-08-06-issue8-stall-budget](plans/2026-08-06-issue8-stall-budget.md) `plan:2026-08-06-issue8-stall-budget`
|
||||||
|
- [2026-08-16-issue10-error-body-retention](plans/2026-08-16-issue10-error-body-retention.md) `plan:2026-08-16-issue10-error-body-retention`
|
||||||
- [est_tokens 解耦实施计划](plans/est-tokens-decoupling.md) `plan:est-tokens-decoupling`
|
- [est_tokens 解耦实施计划](plans/est-tokens-decoupling.md) `plan:est-tokens-decoupling`
|
||||||
|
- [issue #8 实施计划: stall 非生产性等待口径](plans/issue8-stall-budget-plan.md) `plan:issue8-stall-budget-plan`
|
||||||
- [M1 核心里程碑实现计划](plans/m1-core-plan.md) `plan:m1-core-plan`
|
- [M1 核心里程碑实现计划](plans/m1-core-plan.md) `plan:m1-core-plan`
|
||||||
- [M2 分布式实现计划](plans/m2-distributed.md) `plan:m2-distributed`
|
- [M2 分布式实现计划](plans/m2-distributed.md) `plan:m2-distributed`
|
||||||
- [M2.5 治理韧性实现计划](plans/m25-resilience.md) `plan:m25-resilience`
|
- [M2.5 治理韧性实现计划](plans/m25-resilience.md) `plan:m25-resilience`
|
||||||
- [M3 OCR 实现计划](plans/m3-ocr.md) `plan:m3-ocr`
|
- [M3 OCR 实现计划](plans/m3-ocr.md) `plan:m3-ocr`
|
||||||
- [M4 迁移实现计划(T0-T14)](plans/m4-migration.md) `plan:m4-migration`
|
- [M4 迁移实现计划(T0-T14)](plans/m4-migration.md) `plan:m4-migration`
|
||||||
- [响应可观测字段扩展实现计划](plans/response-observability-fields.md) `plan:response-observability-fields`
|
- [响应可观测字段扩展实现计划](plans/response-observability-fields.md) `plan:response-observability-fields`
|
||||||
|
- [实现计划: HTTP 错误响应体留存(Issue #10)](plans/issue10-error-body-retention-plan.md) `plan:issue10-error-body-retention-plan`
|
||||||
|
- [实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)](plans/governance-backend-error.md) `plan:governance-backend-error`
|
||||||
|
- [推理开关能力建模与 reasoning_tokens 采集实施计划(issue #5 + #6)](plans/2026-08-02-thinking-capability.md) `plan:2026-08-02-thinking-capability`
|
||||||
- [采样参数透传实现计划(issue #4)](plans/sampling-params-plan.md) `plan:sampling-params-plan`
|
- [采样参数透传实现计划(issue #4)](plans/sampling-params-plan.md) `plan:sampling-params-plan`
|
||||||
|
|
||||||
## schema (1)
|
## schema (1)
|
||||||
- [表结构: llm_calls(遥测 21 字段)](schemas/llm-calls.md) `schema:llm-calls`
|
- [表结构: llm_calls(遥测 22 字段)](schemas/llm-calls.md) `schema:llm-calls`
|
||||||
|
|
||||||
## metric (2)
|
## metric (2)
|
||||||
- [OCR 治理调用成功率与错误分类分布](metrics/ocr-call-success.md) `metric:ocr-call-success`
|
- [OCR 治理调用成功率与错误分类分布](metrics/ocr-call-success.md) `metric:ocr-call-success`
|
||||||
|
|||||||
@@ -73,3 +73,32 @@
|
|||||||
- [2026-07-31 16:59 UTC] 重建索引: 50 篇页面
|
- [2026-07-31 16:59 UTC] 重建索引: 50 篇页面
|
||||||
- [2026-07-31 17:01 UTC] 重建索引: 50 篇页面
|
- [2026-07-31 17:01 UTC] 重建索引: 50 篇页面
|
||||||
- [2026-08-01 01:58 UTC] 重建索引: 50 篇页面
|
- [2026-08-01 01:58 UTC] 重建索引: 50 篇页面
|
||||||
|
- [2026-08-02 09:38 UTC] 重建索引: 52 篇页面
|
||||||
|
- [2026-08-02 09:38 UTC] 新增边: finding:2026-08-02-thinking-switch-and-reasoning-tokens --supports--> design:2026-08-02-thinking-capability-design
|
||||||
|
- [2026-08-02 09:38 UTC] 新增 finding: 推理开关与 reasoning_tokens 供应商实测与业界做法 (finding:2026-08-02-thinking-switch-and-reasoning-tokens)
|
||||||
|
- [2026-08-02 09:38 UTC] 新增 design: 推理开关能力建模与 reasoning_tokens 采集 issue #5+#6 (design:2026-08-02-thinking-capability-design)
|
||||||
|
- [2026-08-02 09:39 UTC] 重建 Query Pack: 29 字符
|
||||||
|
- [2026-08-02 09:49 UTC] 重建索引: 53 篇页面
|
||||||
|
- [2026-08-02 09:49 UTC] 新增边: plan:2026-08-02-thinking-capability --implements--> design:2026-08-02-thinking-capability-design
|
||||||
|
- [2026-08-02 09:49 UTC] 新增 plan: 推理开关能力建模与 reasoning_tokens 采集实施计划 (plan:2026-08-02-thinking-capability)
|
||||||
|
- [2026-08-02 09:49 UTC] 重建索引: 53 篇页面
|
||||||
|
- [2026-08-02 10:55 UTC] 重建索引: 53 篇页面
|
||||||
|
- [2026-08-02 10:55 UTC] 更新 finding: 补 §2.5 输出长度不是有效判别量(e2e 各 15 轮实测)
|
||||||
|
- [2026-08-06 06:37 UTC] 新增 design: 治理后端故障归位为 scope 级不可用(Issue #7) (design:governance-backend-error)
|
||||||
|
- [2026-08-06 06:38 UTC] 重建索引: 55 篇页面
|
||||||
|
- [2026-08-06 08:08 UTC] 新增 plan: 实现计划: 治理后端故障归位为 scope 级不可用(Issue #7) (plan:governance-backend-error)
|
||||||
|
- [2026-08-06 08:08 UTC] 新增边: plan:governance-backend-error --implements--> design:governance-backend-error
|
||||||
|
- [2026-08-06 08:08 UTC] 重建索引: 57 篇页面
|
||||||
|
- [2026-08-06 08:11 UTC] 重建索引: 57 篇页面
|
||||||
|
- [2026-08-06 14:57 UTC] 新增 design: stall 判定改为非生产性等待口径 (design:issue8-stall-budget)
|
||||||
|
- [2026-08-06 14:58 UTC] 新增 plan: issue #8 实施计划: stall 非生产性等待口径 (plan:issue8-stall-budget-plan)
|
||||||
|
- [2026-08-06 14:58 UTC] 新增边: plan:issue8-stall-budget-plan --implements--> design:issue8-stall-budget
|
||||||
|
- [2026-08-06 14:58 UTC] 重建索引: 61 篇页面
|
||||||
|
- [2026-08-07 15:20 UTC] 新增 design: 建表前先探测,判死只认「确定写不进去」 (design:issue9-telemetry-ddl-probe)
|
||||||
|
- [2026-08-07 15:20 UTC] 重建索引: 62 篇页面
|
||||||
|
- [2026-08-07 15:11 UTC] 重建索引: 62 篇页面
|
||||||
|
- [2026-08-16 09:09 UTC] 新增 design: HTTP 错误响应体留存(Issue #10) (design:issue10-error-body-retention)
|
||||||
|
- [2026-08-16 09:12 UTC] 重建索引: 64 篇页面
|
||||||
|
- [2026-08-16 09:50 UTC] 新增 plan: 实现计划: HTTP 错误响应体留存(Issue #10) (plan:issue10-error-body-retention-plan)
|
||||||
|
- [2026-08-16 09:50 UTC] 新增边: plan:issue10-error-body-retention-plan --implements--> design:issue10-error-body-retention
|
||||||
|
- [2026-08-16 09:50 UTC] 重建索引: 66 篇页面
|
||||||
|
|||||||
@@ -181,7 +181,7 @@ stack = ExtractionProviderStack(
|
|||||||
| R4 | 六道闸+契约 5 条、服务器时钟窗口、settle 落 acquire 窗口、transient 按 est 保守结算 | M2 | §7.3 大体覆盖 |
|
| R4 | 六道闸+契约 5 条、服务器时钟窗口、settle 落 acquire 窗口、transient 按 est 保守结算 | M2 | §7.3 大体覆盖 |
|
||||||
| R5 | RequestRejected 二分(真实响应记成功/本地拒绝释放探针);换源重试跨源计数口径 | M2 | 须进 M2 设计 |
|
| R5 | RequestRejected 二分(真实响应记成功/本地拒绝释放探针);换源重试跨源计数口径 | M2 | 须进 M2 设计 |
|
||||||
| R6 | OCR ZIP 协议 + bbox 数值防御下沉;OCR Usage=0;glm 白名单预留 | M3 | §7.10 已覆盖 |
|
| R6 | OCR ZIP 协议 + bbox 数值防御下沉;OCR Usage=0;glm 白名单预留 | M3 | §7.10 已覆盖 |
|
||||||
| **G1** | ✅ 已闭(M3 核实): 库 `GatewayUnavailableError` 一族自 M1 起携 `scope/reason/retry_after_s/per_source_reasons`(errors.py:74-105),chat/embedding/OCR 三循环抛出点均已填充且有契约测试钉住;项目侧仅剩约 10 行翻译 shim(库异常 → ProviderUnavailableError)或 tracking.py 直接 except 库异常 | M2 | 已闭 |
|
| **G1** | ✅ 已闭(M3 核实): 库 `GatewayUnavailableError` 一族自 M1 起携 `scope/reason/retry_after_s/per_source_reasons`(errors.py:74-105),chat/embedding/OCR 三循环抛出点均已填充且有契约测试钉住;项目侧仅剩约 10 行翻译 shim(库异常 → ProviderUnavailableError)或 tracking.py 直接 except 库异常。**2026-08-06 补(issue #7,库 1.1.0)**: 治理后端故障(`GovernanceBackendError`,Redis 挂等 fail-closed 情形)此前**不在**该族内,`except GatewayUnavailableError` 接不住,会落进 `_TERMINAL` 兜底而消耗业务失败预算;现已归入该族(`reason=governance_backend_down`,`retry_after_s` 默认 5.0),tracking.py 一条 except 即覆盖完整,**无需为它单列分支**。同批新增的 `SourceNotConfiguredError`(源名与配置不匹配的装配缺陷)**有意在族外**,应当落进 `_TERMINAL` 让配置错误浮出水面 | M2 | 已闭 |
|
||||||
| **G2** | ✅ 已闭(2026-07-30 核实): `est_tokens` 已进 ARCH §7.7 SourceConfig 字段清单,且 §7.3 `try_acquire` 的 est 来源已定义为 `SourceConfig.effective_est_tokens()`(显式值优先,否则按 `tpm // 60` 派生)。双职责一并拆开:该字段只剩 TPM 预扣的可选调优覆盖,usage 缺失不再由它兜底(见 §7 行 151 的推翻判定) | M2 | 已闭 |
|
| **G2** | ✅ 已闭(2026-07-30 核实): `est_tokens` 已进 ARCH §7.7 SourceConfig 字段清单,且 §7.3 `try_acquire` 的 est 来源已定义为 `SourceConfig.effective_est_tokens()`(显式值优先,否则按 `tpm // 60` 派生)。双职责一并拆开:该字段只剩 TPM 预扣的可选调优覆盖,usage 缺失不再由它兜底(见 §7 行 151 的推翻判定) | M2 | 已闭 |
|
||||||
| **G3** | ⚠️ §4.3 层序图文矛盾:图示 熔断→限流→重试(重试最内),但理由要求"每次重试重新过限流闸"且熔断/限流是 per-source 的、选源在重试循环内(governance.py:120-167 实践为每次尝试执行 选源→冷却备忘→permit→熔断门)。洋葱不澄清"逐次准入"机制则多源语义无法成立 | M2 | **架构缺口**,澄清 §4.3/§4.4 |
|
| **G3** | ⚠️ §4.3 层序图文矛盾:图示 熔断→限流→重试(重试最内),但理由要求"每次重试重新过限流闸"且熔断/限流是 per-source 的、选源在重试循环内(governance.py:120-167 实践为每次尝试执行 选源→冷却备忘→permit→熔断门)。洋葱不澄清"逐次准入"机制则多源语义无法成立 | M2 | **架构缺口**,澄清 §4.3/§4.4 |
|
||||||
| **G4** | ⚠️ per-scope 韧性配置命名(`{SCOPE}__RETRY__*`/`BREAKER__*`/`BACKPRESSURE__*`/`SELECTOR`/`GLOBAL__*`)未进 ARCH §9,现文只有平铺 `LLM_*` 键;CHSAnalyzer 的 VLM/OCR 两 scope 参数各异,平铺键无法表达 | M2 | **架构缺口**,修订 §9 |
|
| **G4** | ⚠️ per-scope 韧性配置命名(`{SCOPE}__RETRY__*`/`BREAKER__*`/`BACKPRESSURE__*`/`SELECTOR`/`GLOBAL__*`)未进 ARCH §9,现文只有平铺 `LLM_*` 键;CHSAnalyzer 的 VLM/OCR 两 scope 参数各异,平铺键无法表达 | M2 | **架构缺口**,修订 §9 |
|
||||||
|
|||||||
@@ -0,0 +1,276 @@
|
|||||||
|
---
|
||||||
|
type: plan
|
||||||
|
node_id: plan:2026-08-02-thinking-capability
|
||||||
|
title: "推理开关能力建模与 reasoning_tokens 采集实施计划(issue #5 + #6)"
|
||||||
|
date: 2026-08-02
|
||||||
|
---
|
||||||
|
|
||||||
|
# 推理开关能力建模与 reasoning_tokens 采集实施计划(issue #5 + #6)
|
||||||
|
|
||||||
|
**目标**:让 `enable_thinking` 对每个源要么真实生效、要么显式报错,并采集 `reasoning_tokens` 以区分推理开销与生成开销。
|
||||||
|
|
||||||
|
**方案概述**:`ProviderProfile` 保留为「形态层」(参数长什么样,按 provider),新增 model 级「能力层」声明该模型能否关闭推理;两层在单一判定函数 `resolve_thinking` 相遇,装配期与请求期共用。同时照搬 issue #3 的 `_coerce_cached_tokens` 采集 `reasoning_tokens`,并把 `enable_thinking` 纳入缓存指纹。
|
||||||
|
|
||||||
|
**涉及技术**:Python 3.11 frozen dataclass、`MappingProxyType` 只读注册表、httpx、SQLite/PostgreSQL DDL 迁移、pytest。
|
||||||
|
|
||||||
|
**依据文档**:设计 `designs/2026-08-02-thinking-capability-design.md`;事实基础 `findings/2026-08-02-thinking-switch-and-reasoning-tokens.md`。
|
||||||
|
|
||||||
|
**保真校验**:本计划不涉及 `reference/` 参考实现迁移,保真校验不适用。
|
||||||
|
|
||||||
|
## 文件结构
|
||||||
|
|
||||||
|
| 文件 | 动作 | 职责 |
|
||||||
|
|---|---|---|
|
||||||
|
| `src/polygateway/types.py` | 修改 | `LLMResponse` / `TransportResult` 尾部各加 `reasoning_tokens` |
|
||||||
|
| `src/polygateway/providers.py` | 修改 | 形态层放宽为 `dict \| None`;新增能力层与 `resolve_thinking` |
|
||||||
|
| `src/polygateway/transports/openai_compat.py` | 修改 | 采集 `reasoning_tokens`;`_build_payload` 接入 `resolve_thinking`;收 `capabilities` |
|
||||||
|
| `src/polygateway/middleware/retry.py` | 修改 | `_build_response` 透传 `reasoning_tokens` |
|
||||||
|
| `src/polygateway/ports.py` | 修改 | `record_llm_call` 21 → 22 字段 |
|
||||||
|
| `src/polygateway/telemetry/sqlite.py` | 修改 | 建表列 + `_BACKFILL_COLUMNS` + `_COLUMNS`(新列排末尾) |
|
||||||
|
| `src/polygateway/telemetry/postgres.py` | 修改 | 同上 |
|
||||||
|
| `src/polygateway/middleware/telemetry.py` | 修改 | `_record` + 三个 `emit_*` 入口 |
|
||||||
|
| `src/polygateway/client.py` | 修改 | `capabilities` 参数贯通;装配守卫;缓存指纹纳入 `enable_thinking` |
|
||||||
|
| `tests/e2e/test_thinking_live.py` | 新建 | 真实 API 矩阵 L1–L9 |
|
||||||
|
| `CHANGELOG.md` / `research-wiki/schemas/llm-calls.md` | 修改 | 行为变更说明与字段表 21 → 22 |
|
||||||
|
|
||||||
|
**任务顺序不可调换**:T1–T3 先把 `reasoning_tokens` 打通(#6 是 #5 的验收仪器),T4–T7 再改推理开关,T8 用真实 API 验证,T9 收尾文档。
|
||||||
|
|
||||||
|
## 关键接口(跨任务消费,此处给出实际代码)
|
||||||
|
|
||||||
|
`providers.py` 新增:
|
||||||
|
|
||||||
|
```python
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ThinkingCapability:
|
||||||
|
"""某个具体模型的推理能力(model 级);登记必须附实测证据与日期。"""
|
||||||
|
|
||||||
|
can_disable: bool
|
||||||
|
evidence: str
|
||||||
|
|
||||||
|
|
||||||
|
def get_capability(
|
||||||
|
model: str, *, table: Mapping[str, ThinkingCapability] | None = None
|
||||||
|
) -> ThinkingCapability | None:
|
||||||
|
"""按模型名精确查找;未登记返回 None(= 能力未知,由调用方决定退化)。"""
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
def register_capability(
|
||||||
|
model: str,
|
||||||
|
capability: ThinkingCapability,
|
||||||
|
*,
|
||||||
|
base: Mapping[str, ThinkingCapability] | None = None,
|
||||||
|
) -> dict[str, ThinkingCapability]:
|
||||||
|
"""纯函数注册: 返回 base(缺省 DEFAULT_CAPABILITIES)+ 新条目的新表,同名覆盖。"""
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_thinking(
|
||||||
|
profile: ProviderProfile,
|
||||||
|
capability: ThinkingCapability | None,
|
||||||
|
enable_thinking: bool | None,
|
||||||
|
*,
|
||||||
|
model: str,
|
||||||
|
) -> Mapping[str, Any]:
|
||||||
|
"""三态 + 两层能力 → 注入片段;不可满足时 ValueError(调用点翻译为领域错误)。
|
||||||
|
|
||||||
|
model 只用于错误与告警文案: 报错必须能定位到具体模型才有可操作性,
|
||||||
|
而 capability 为 None(未登记)时无从从别处取得模型名。
|
||||||
|
"""
|
||||||
|
```
|
||||||
|
|
||||||
|
`resolve_thinking` 的判定顺序(**顺序即语义,不可调换**):
|
||||||
|
|
||||||
|
| 步 | 条件 | 行为 |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | `enable_thinking is None` | 返回 `{}`(不干预) |
|
||||||
|
| 2 | 对应档 `slot is None` | `ValueError`:形态未知,指路 `register_provider` / `extra_body` |
|
||||||
|
| 3 | `capability is None` | `loguru.warning` 后返回 `slot`(能力未登记,从宽放行) |
|
||||||
|
| 4 | `enable_thinking is False` 且 `capability.can_disable is False` | `ValueError`:该模型无法关闭推理 |
|
||||||
|
| 5 | 其余 | 返回 `slot` |
|
||||||
|
|
||||||
|
第 2 步必须先于第 4 步:形态未知时无从注入,能力如何无关紧要。第 3 步先于第 4 步:未登记模型无 `can_disable` 可读。
|
||||||
|
|
||||||
|
`transports/openai_compat.py` 新增:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _coerce_reasoning_tokens(usage: Any) -> int | None:
|
||||||
|
"""取 usage.completion_tokens_details.reasoning_tokens(issue #6);形态异常一律 None。"""
|
||||||
|
```
|
||||||
|
|
||||||
|
## 任务清单
|
||||||
|
|
||||||
|
### T1 — `reasoning_tokens` 进入类型与采集路径
|
||||||
|
|
||||||
|
- [ ] **文件**:改 `src/polygateway/types.py`、`src/polygateway/transports/openai_compat.py`、`src/polygateway/middleware/retry.py`;改测 `tests/unit/test_types.py`、`tests/unit/test_openai_compat.py`、`tests/unit/test_retry.py`
|
||||||
|
|
||||||
|
**行为**:`LLMResponse` 与 `TransportResult` **尾部**各加 `reasoning_tokens: int | None = None`(字段顺序是公共承诺,见 `types.py:1-5`,只增不删不改名)。新增 `_coerce_reasoning_tokens`,语义与 `_coerce_cached_tokens`(`openai_compat.py:161-177`)逐条对齐:非 `dict` 返回 `None`;`completion_tokens_details` 非 `dict` 返回 `None`;`bool` 显式排除(`isinstance(True, int)` 为真,放行会把 `True` 记成 1);负数返回 `None`;`0` 如实保留。流式(`:401` 附近)取 `sink.get("usage")`、非流式(`:485` 附近)取 `body.get("usage")`,与 `cached_prompt_tokens` 同处填值。`retry.py:_build_response` 透传。
|
||||||
|
|
||||||
|
**docstring 措辞**(必须逐字,理由见 findings §4c):`None` = **本次调用**未上报,**不可**写「该源未上报」——中转在上游不返回 usage 时会本地补算并吃掉该字段。
|
||||||
|
|
||||||
|
**验收**:非流式与流式响应含 `completion_tokens_details.reasoning_tokens: 7` → `reasoning_tokens == 7`;该键为 `0` → `0`(不与 `None` 混同);`completion_tokens_details` 缺失 / 非 dict / 值为 `True` / 值为 `-1` → 均为 `None`;`missing_done="salvage"` 打捞路径(无 usage 帧)→ `None` 而非 `0`。
|
||||||
|
|
||||||
|
**测试证据**:先加断言 → 失败(字段不存在)→ 实现 → 通过。
|
||||||
|
|
||||||
|
**验证**:`conda run -n PolyGateway pytest tests/unit/test_types.py tests/unit/test_openai_compat.py tests/unit/test_retry.py -v` → 全部 PASS。
|
||||||
|
|
||||||
|
### T2 — 遥测端口 21 → 22 字段与两后端落库
|
||||||
|
|
||||||
|
- [ ] **文件**:改 `src/polygateway/ports.py`、`src/polygateway/telemetry/sqlite.py`、`src/polygateway/telemetry/postgres.py`、`src/polygateway/middleware/telemetry.py`;改测 `tests/unit/test_ports.py`、`tests/unit/test_telemetry.py`、`tests/integration/test_postgres_telemetry.py`
|
||||||
|
|
||||||
|
**行为**:`record_llm_call` 在 `sampling` 之后追加 `reasoning_tokens: int | None`(不设默认值——库外无第三方实现者,见 `ports.py:250` 注释)。两个后端在建表 DDL、`_BACKFILL_COLUMNS`(sqlite)/ 迁移语句列表(postgres)、`_COLUMNS` 三处各加一项,**新列必须排在末尾**(两文件均有明文注释:旧表只能 ALTER 追加,新建库若插在前面会与迁移路径的物理列序分叉)。`middleware/telemetry.py` 的 `_record` 加参数,三个 `emit_*` 入口按 `cached_prompt_tokens` 的既有形态填值:`emit_attempt` 用 `response.reasoning_tokens if response else None`,`emit_cache_hit` 原样回放,`emit_terminal_failure` 填 `None`。
|
||||||
|
|
||||||
|
**不改 `pricing.py`**:推理 token 已含在 `completion_tokens` 内,单列计价即重复计费。
|
||||||
|
|
||||||
|
**验收**:新建库与经 ALTER 迁移的旧库物理列序一致;`reasoning_tokens=7` / `0` / `None` 三种值各自如实落库(`0` 与 `NULL` 可区分);遥测写失败仍降级为 warning 不冒泡。
|
||||||
|
|
||||||
|
**测试证据**:先扩字段清单断言 → 失败 → 实现 → 通过。
|
||||||
|
|
||||||
|
**验证**:`conda run -n PolyGateway pytest tests/unit/test_ports.py tests/unit/test_telemetry.py -v` → PASS;`conda run -n PolyGateway pytest tests/integration/test_postgres_telemetry.py -v` → PASS 或按既有约定 SKIP(无 PG 凭据时)。
|
||||||
|
|
||||||
|
### T3 — 提交点:issue #6 完整可用
|
||||||
|
|
||||||
|
- [ ] 运行 `conda run -n PolyGateway make ci`,确认全绿后提交。提交信息类型 `feat`,正文说明 `LLMResponse` 新增字段与遥测端口 21 → 22。此处独立成一个提交,便于 #6 单独回滚。
|
||||||
|
|
||||||
|
**测试证据**:本任务不引入新行为,证据即 T1 与 T2 各自的「先失败后通过」记录;提交前需确认这两组记录都已产生,不得以 `make ci` 全绿代替。
|
||||||
|
|
||||||
|
### T4 — 形态层放宽与 profile 修正
|
||||||
|
|
||||||
|
- [ ] **文件**:改 `src/polygateway/providers.py`;改测 `tests/unit/test_providers.py`
|
||||||
|
|
||||||
|
**行为**:`ProviderProfile.thinking_on` / `thinking_off` 类型由 `dict[str, Any]` 改为 `Mapping[str, Any] | None`。三值语义写进类 docstring:`{...}` = 已知注入片段;`{}` = 已知无需注入即处于该档;`None` = **未知**(库不知道该 provider 如何表达)。删除现有 docstring 里「两档皆空 ⇒ 不产生任何效果」那段(`providers.py:23-25` 与 `:50-51`)——它正是把「不支持」与「未知」编码成同一个值的根因。
|
||||||
|
|
||||||
|
`minimax` 填 `thinking_on={"reasoning_effort": "medium"}`、`thinking_off={"reasoning_effort": "none"}`;`openai` 两档改 `None`。qwen / deepseek **不动**(实测正确)。两处均加注释写明:取值依据 2026-08-02 经自建 new-api 中转的实测,直连官方端点未验证。
|
||||||
|
|
||||||
|
**验收**:`get_provider("minimax").thinking_off == {"reasoning_effort": "none"}`;`get_provider("openai").thinking_on is None`;qwen / deepseek 两档与改动前逐字相同。
|
||||||
|
|
||||||
|
**测试证据**:`tests/unit/test_providers.py:29,34` 现有断言锁的是空字典,先改成新期望 → 失败 → 实现 → 通过。
|
||||||
|
|
||||||
|
**验证**:`conda run -n PolyGateway pytest tests/unit/test_providers.py -v` → PASS。
|
||||||
|
|
||||||
|
### T5 — 能力层与 `resolve_thinking`
|
||||||
|
|
||||||
|
- [ ] **文件**:改 `src/polygateway/providers.py`;改测 `tests/unit/test_providers.py`
|
||||||
|
|
||||||
|
**行为**:按「关键接口」一节的签名实现 `ThinkingCapability`、`DEFAULT_CAPABILITIES`、`get_capability`、`register_capability`、`resolve_thinking`。注册表用 `MappingProxyType` 只读,注册走纯函数返回新表(不修改共享状态,纯 asyncio 中立铁律),与既有 `register_provider`(`providers.py:84-90`)同形。
|
||||||
|
|
||||||
|
`DEFAULT_CAPABILITIES` 首发五条,`evidence` 逐条写明实测日期与样本量:
|
||||||
|
|
||||||
|
| 键 | `can_disable` | `evidence` 要点 |
|
||||||
|
|---|---|---|
|
||||||
|
| `MiniMax-M3` | `True` | 2026-08-02 实测 N=10,`reasoning_effort=none` 稳定关闭 |
|
||||||
|
| `MiniMax-M2.7` | `False` | 三形态各 N=3 全无效;OpenRouter 登记 `mandatory:true` |
|
||||||
|
| `MiniMax-M2.5` | `False` | 同上 |
|
||||||
|
| `qwen3.7-plus` | `True` | 实测 `enable_thinking=false` 关闭 |
|
||||||
|
| `deepseek-v4-pro` | `True` | 实测 `thinking:{"type":"disabled"}` 关闭 |
|
||||||
|
|
||||||
|
**验收**:`resolve_thinking` 五条判定各有一例;未登记模型返回 `slot` 并产生一条 warning(**loguru 不经标准 logging,pytest 的 `caplog` 抓不到**——必须复用项目既有写法 `logger.add(messages.append, level="WARNING")`,见 `tests/unit/test_config.py:29-31`);`enable_thinking=False` + `MiniMax-M2.7` 抛 `ValueError` 且消息含模型名与"无法关闭"字样;`enable_thinking` 任意非 `None` + `openai` profile 抛 `ValueError` 且消息含 `register_provider` 与 `extra_body` 两个指路词;`register_capability` 不修改 `DEFAULT_CAPABILITIES`。
|
||||||
|
|
||||||
|
**测试证据**:先写五条判定的参数化测试 → 失败(函数不存在)→ 实现 → 通过。
|
||||||
|
|
||||||
|
**验证**:`conda run -n PolyGateway pytest tests/unit/test_providers.py -v` → PASS。
|
||||||
|
|
||||||
|
### T6 — transport 接入与请求期兜底
|
||||||
|
|
||||||
|
- [ ] **文件**:改 `src/polygateway/transports/openai_compat.py`;改测 `tests/unit/test_openai_compat.py`
|
||||||
|
|
||||||
|
**行为**:`OpenAICompatTransport.__init__` 增加 `capabilities: Mapping[str, ThinkingCapability] | None = None`,与既有 `registry` 参数同形并存于 `self._capabilities`。`_build_payload` 的两个 `if` 分支(`:293-296`)收敛为两行——先 `capability = get_capability(source.model, table=self._capabilities)`,再 `payload.update(resolve_thinking(profile, capability, source.enable_thinking, model=source.model))`;其后 `payload.update(source.extra_body)` 与 `payload.update(overlay)` 两行**顺序不变**(顺序即优先级,issue #4 决策 A)。`complete()` 中把构造 payload 的 `ValueError` 翻译为 `RequestRejectedError`(四分类之一,不重试不换源)。
|
||||||
|
|
||||||
|
**绝不在 `_build_payload` 内抛裸 `ValueError` 让它冒泡**:该处位于 RetryMW 内侧,裸异常不属错误四分类、`TelemetryMW` 也不捕,会导致一行遥测都没有就逃出 `chat()`。
|
||||||
|
|
||||||
|
**验收**:`enable_thinking=True` + minimax 源 → 请求体含 `reasoning_effort: "medium"`;`False` → `"none"`;`None` → 请求体无 `reasoning_effort` 键;`extra_body={"reasoning_effort":"high"}` 时实发 `high`(覆盖 profile);`enable_thinking=False` + M2.7 源经 transport 调用 → `RequestRejectedError` 而非裸 `ValueError`。
|
||||||
|
|
||||||
|
**测试证据**:扩 `tests/unit/test_openai_compat.py:418-431` 的三态参数化,加 minimax 用例 → 失败 → 实现 → 通过。
|
||||||
|
|
||||||
|
**验证**:`conda run -n PolyGateway pytest tests/unit/test_openai_compat.py -v` → PASS。
|
||||||
|
|
||||||
|
### T7 — 装配守卫、参数贯通与缓存指纹
|
||||||
|
|
||||||
|
- [ ] **文件**:改 `src/polygateway/client.py`;改测 `tests/unit/test_cache.py`(指纹相关)、新增装配守卫测试至 `tests/unit/test_config.py`
|
||||||
|
|
||||||
|
**行为**(三件事,同一文件):
|
||||||
|
|
||||||
|
其一,`from_settings` 与 `from_env` 各增加 `capabilities` 参数并透传给 `OpenAICompatTransport`;`from_settings` 在已解析 `profiles` 之后(`client.py:248`)加装配守卫:对 `zip(sources, profiles, strict=True)` 的每一对,先 `get_capability(src.model, table=capabilities)` 取能力,再调用一次 `resolve_thinking(prof, cap, src.enable_thinking, model=src.model)` 并丢弃返回值——只为让配置错误在装配期即抛 `ValueError`。守卫与 transport 内的判定共用同一函数,不复制逻辑——这与 `get_provider` 在 `client.py:248` 与 `openai_compat.py:313` 双点调用的既有形态一致。
|
||||||
|
|
||||||
|
其二,`build_model_fingerprint`(`client.py:63-80`)把 `enable_thinking` 纳入摘要。实现必须保持既有不变量——**全源不配 `enable_thinking` 时指纹字面量与改动前逐字相同**:
|
||||||
|
|
||||||
|
```python
|
||||||
|
def _fingerprint_mark(s: SourceConfig) -> str:
|
||||||
|
parts: list[Any] = [s.model, dict(s.extra_body)]
|
||||||
|
if s.enable_thinking is not None: # 仅在表态时追加,保证存量指纹字面量不变
|
||||||
|
parts.append(s.enable_thinking)
|
||||||
|
return json.dumps(parts, sort_keys=True, ensure_ascii=False)
|
||||||
|
```
|
||||||
|
|
||||||
|
筛选条件由 `if s.extra_body` 扩为 `if s.extra_body or s.enable_thinking is not None`。
|
||||||
|
|
||||||
|
其三,为守卫补测:`enable_thinking=False` + `provider=minimax` + `model=MiniMax-M2.7` 的 `GatewaySettings` 经 `from_settings` → `ValueError`;`provider=openai` + 任意非 `None` 的 `enable_thinking` → `ValueError`。
|
||||||
|
|
||||||
|
**验收**:装配期报错两例;改 `enable_thinking` → 指纹变化;只配 `extra_body`、不配 `enable_thinking` 的源 → 指纹与改动前逐字相同(用硬编码的历史字面量断言,防回归)。
|
||||||
|
|
||||||
|
**测试证据**:先写三条断言 → 失败 → 实现 → 通过。
|
||||||
|
|
||||||
|
**验证**:`conda run -n PolyGateway pytest tests/unit/test_cache.py tests/unit/test_config.py -v` → PASS;随后 `conda run -n PolyGateway make ci` → 全绿。
|
||||||
|
|
||||||
|
### T8 — 真实 API e2e 矩阵
|
||||||
|
|
||||||
|
- [ ] **文件**:新建 `tests/e2e/test_thinking_live.py`
|
||||||
|
|
||||||
|
**行为**:沿用既有 e2e 约定(`tests/e2e/test_smoke_gateway.py:19-26`)——`dotenv_values(".env")` 读凭据、`skipif(not _HAS_SOURCE, ...)`、结构化报告写入 `tests/outputs/e2e/`。**不新造开关机制**:另加项目既有的 `slow` 标记,靠 `pyproject.toml` 的 `addopts = "-m 'not slow'"` 把本组挡在 `make ci` 之外(137 次真实调用、约 7 分钟,且判据是统计性的,网络抖动会造成假红——执行期实测撞到过一次 `network_error` 耗尽源)。合并前用 `pytest -m slow tests/e2e/test_thinking_live.py` 显式真跑。
|
||||||
|
|
||||||
|
**源映射**:L1–L5、L8 的 MiniMax 行用现有的 `LLM__MINIMAX__1__*`(`MODEL=MiniMax-M3`);M2.7 / M2.5 行经 `dataclasses.replace(source, model=...)` 派生,不新增 `.env` 键。**L6 / L7 目前无对应源**——`.env` 里只有 MINIMAX 与 MONKEY 两类;需新增 `{SCOPE}__QWEN__1__*` 与 `{SCOPE}__DEEPSEEK__1__*`(同一中转 `BASE_URL` 与密钥,仅 `MODEL` 不同)。未配置时按既有 `skipif` 约定跳过,并在报告中记为「未覆盖」,**不得静默计入通过**。
|
||||||
|
|
||||||
|
覆盖矩阵(轮数经环境变量可调,默认值如下):
|
||||||
|
|
||||||
|
| # | 场景 | 源 | 轮数 | 判据 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| L1 | `enable_thinking=False` | MiniMax-M3 | 10 | **每轮** `completion_tokens < 30`(主判据)且 `reasoning_tokens in (None, 0)`(辅判据,与下游口径一致) |
|
||||||
|
| L2 | `enable_thinking=True` | MiniMax-M3 | 10 | 多数轮 `completion_tokens > 100`;请求体实发 `reasoning_effort=medium` |
|
||||||
|
| L3 | `enable_thinking=None` | MiniMax-M3 | 10 | 请求体无 `reasoning_effort` 键 |
|
||||||
|
| L4 | `extra_body` 覆盖 profile | MiniMax-M3 | 5 | 实发 `high` |
|
||||||
|
| L5 | L1 / L2 的**流式**重跑 | MiniMax-M3 | 各 10 | 同 L1 / L2(`stream=True` 是库的默认主路径) |
|
||||||
|
| L6 | `enable_thinking=False` | qwen | 10 | 每轮 `completion_tokens < 30` |
|
||||||
|
| L7 | `enable_thinking=False` | deepseek | 10 | 每轮 `completion_tokens < 30` |
|
||||||
|
| L8 | 能力表漂移哨兵 | 全部登记模型 | 各 5 | 实测行为与 `can_disable` 声明一致 |
|
||||||
|
| L9 | M2.7 + `enable_thinking=False` → 装配期报错 | — | — | 纯本地,无需真实调用 |
|
||||||
|
|
||||||
|
**三条必须遵守的测试纪律**:
|
||||||
|
|
||||||
|
其一,**判别量只能是 `reasoning_tokens`**。(执行时按 e2e 实测修正:本条初稿写的是「主判据用 `completion_tokens`」,被数据推翻——两档的输出长度分布**重叠**,关闭档实测最高 46、开启档最低 13,按长度阈值判两个方向都会误判。)`completion_tokens` 仅作 `reasoning_tokens` 被中转吃掉时的退路(findings §4c、§2.5)。
|
||||||
|
|
||||||
|
其二,**关闭方向要求每轮满足,开启方向只要求多数轮满足**。中转吃掉 ctd 时开启方向可能偶尔观测不到,关闭方向不受影响。
|
||||||
|
|
||||||
|
其四,**必须有不依赖输出侧噪声的锚点**:L2b 比较两档的 `prompt_tokens`(相对比较,无魔数),L3b 用非法值反证 `none` 是被识别而非被静默丢弃——后者正是 issue #5 的原始故障形态,不排除它,关闭方向的证据就只到「未回归」,够不到「已生效」。
|
||||||
|
|
||||||
|
其三,**源不可用必须跳过并在报告中显式记为「未覆盖」**,不得静默计入通过(实测中 kimi 渠道 429 后被中转下线并返回 404)。报告要能一眼看出哪些矩阵行没跑到。
|
||||||
|
|
||||||
|
**报告内容**(`tests/outputs/e2e/test_thinking_live_<ts>.md`):逐轮记录实际注入的 thinking 片段、`prompt_tokens` / `completion_tokens` / `reasoning_tokens`、单轮判定结果;逐行记录矩阵编号、通过或跳过及其原因;文末给出总调用次数与时间戳。原始数字必须落盘——结论可以复核,才算证据。
|
||||||
|
|
||||||
|
**验收**:矩阵九行全部有结论(通过 / 明确跳过),报告落盘 `tests/outputs/e2e/`。
|
||||||
|
|
||||||
|
**验证**:`conda run -n PolyGateway pytest tests/e2e/test_thinking_live.py -v -s` → PASS,人工核对报告。
|
||||||
|
|
||||||
|
### T9 — 文档同步与收尾
|
||||||
|
|
||||||
|
- [ ] **文件**:改 `CHANGELOG.md`、`research-wiki/schemas/llm-calls.md`
|
||||||
|
|
||||||
|
**行为**:CHANGELOG 必须醒目标注这是**行为变更而非纯修复**——MiniMax 源的 `ENABLE_THINKING` 从「无效」变为「生效」,且配了该项的 scope 会有一次性缓存冷启动。同时写明 `reasoning_tokens` 的语义:`None` = 本次调用未上报,下游判据须为 `in (None, 0)`,写 `== 0` 永远不成立。`schemas/llm-calls.md` 的字段表由 21 改 22,新增行说明该列。
|
||||||
|
|
||||||
|
Gitea Wiki(独立仓库)**本任务内必须同步**:按 `docs-convention.md` §2「新公共 API / 新能力」一行,需改 `参考-公共API`(`LLMResponse` 新字段)与相关指南页;该表把同步绑定在**变更**上而非发版上,不可推迟。`Home.md` 的版本号与安装命令等发版项不在本计划范围。
|
||||||
|
|
||||||
|
**验收**:CHANGELOG 含行为变更与冷启动两处提示;schema 文档字段数与 `_COLUMNS` 长度一致。
|
||||||
|
|
||||||
|
**验证**:人工核对;`conda run -n PolyGateway make ci` → 全绿。
|
||||||
|
|
||||||
|
## 完成后的独立验证
|
||||||
|
|
||||||
|
按 `verification-before-completion` 的强制档,本计划跨多文件,合并前须派**全新上下文**的 verifier subagent 逐条核对设计 §11 的九条验收标准与本计划各任务的测试证据,不得自审代替。
|
||||||
|
|
||||||
|
### T10 — 同步结论给 dissect
|
||||||
|
|
||||||
|
- [ ] **动作**:在本分支合并时,向 dissect 提一条 issue 或在其 `ROADMAP` 风险表中记录下述结论,并确认对方已读。
|
||||||
|
|
||||||
|
**验收**:dissect 侧存在可追溯的记录(issue 编号或文档行号),不以口头告知为准。
|
||||||
|
|
||||||
|
## 需要同步给下游的结论
|
||||||
|
|
||||||
|
`MiniMax-M2.7` / `M2.5` 的推理**关不掉**是模型固有属性,任何库层改动都无法改变。dissect 的 Phase-0 若要做「开思考 vs 关思考」对照,只能在 M3 上做,或把因子改为「高档 vs 低档」。此结论须在本分支合并时同步给 dissect。
|
||||||
@@ -0,0 +1,290 @@
|
|||||||
|
# 实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)
|
||||||
|
|
||||||
|
- **设计**: `research-wiki/designs/2026-08-06-governance-backend-error-design.md`(已批准 2026-08-06,Q1/Q2/Q3 逐条拍板)
|
||||||
|
- **分支**: `feat/issue-7-governance-backend-error`
|
||||||
|
- **目标**: 让"限流/熔断后端故障"在类型上落入 `GatewayUnavailableError`,使调用方一条 `except` 覆盖完整;同时把混在同一类里的装配缺陷拆出去,避免配置写错的任务永远重投。
|
||||||
|
- **方案概述**: `GovernanceBackendError` 改继承 `GatewayUnavailableError`(新 reason `governance_backend_down`,`retry_after_s` 默认 5.0);两处"未知源"改抛新增的 `SourceNotConfiguredError`(**不**在 scope 级家族内);`scope` 由后端层 `self._scope` 与两个 gate 包装器注入。
|
||||||
|
- **涉及技术**: Python 3.11+,pytest(含真实 Redis 的 integration),radon/ruff 门禁。
|
||||||
|
|
||||||
|
## 保真校验(适用)
|
||||||
|
|
||||||
|
本计划触及 ARCHITECTURE.md §1.4 索引的移植蓝本:错误分类(`reference/CHSAnalyzer/app/domain/errors.py`)与限流/熔断(`reference/CHSAnalyzer/app/coordination/`)。
|
||||||
|
|
||||||
|
本次**有意变更**的语义只有一条,已在设计 §3.1 声明:`GovernanceBackendError` 的类型归属(CHS 的 `LimiterError` 是独立异常,本库将其提升为 scope 级不可用的一员)。除此之外,下列承自 CHS 的语义**不得被顺带改动**,每个任务完成前逐条自查:
|
||||||
|
|
||||||
|
| 不得改动 | 出处 |
|
||||||
|
|---|---|
|
||||||
|
| `retry_after_s` 非可选、`0 = 可立即重试` | `errors.py:74-78` |
|
||||||
|
| `SCOPE_REASONS` 既有 5 值与 `SOURCE_REASONS` 既有 7 值 | `errors.py:7-20` |
|
||||||
|
| fail-closed 降级方向(限流/熔断后端挂 → 报错而非放行) | 库铁律 |
|
||||||
|
| 记账路径降级为 warning、闸门路径上抛的分工 | `middleware/retry.py:404` |
|
||||||
|
| `RedisPermit.release/settle` 的释放侧降级 | `backends/redis/limiter.py:133,151` |
|
||||||
|
|
||||||
|
## 文件结构
|
||||||
|
|
||||||
|
| 文件 | 职责 | 动作 |
|
||||||
|
|---|---|---|
|
||||||
|
| `research-wiki/ARCHITECTURE.md` | 架构单一事实源 §6.1 错误分类表 | 修改(**必须先行**,见设计 §8.1) |
|
||||||
|
| `src/polygateway/errors.py` | 错误类型树内核 | 修改: 新常量、新 reason、新类、继承变更 |
|
||||||
|
| `src/polygateway/__init__.py` | 公共 API 面 | 修改: 导出新类 |
|
||||||
|
| `src/polygateway/backends/redis/limiter.py` | Redis 限流后端 | 修改: 6 处补 scope、1 处换新类 |
|
||||||
|
| `src/polygateway/backends/redis/breaker.py` | Redis 熔断后端 | 修改: 5 处补 scope |
|
||||||
|
| `src/polygateway/backends/memory/limiter.py` | 内存限流后端 | 修改: 1 处换新类 |
|
||||||
|
| `src/polygateway/middleware/ratelimit.py` | `QuotaGate` 包装器 | 修改: 构造增 scope、4 处补 scope |
|
||||||
|
| `src/polygateway/middleware/breaker.py` | `BreakerGate` 包装器 | 修改: 构造增 scope、5 处补 scope |
|
||||||
|
| `src/polygateway/middleware/retry.py` / `ocr.py` / `embedding.py` | 三处 gate 装配 | 修改: 各 2 行传 scope |
|
||||||
|
| `tests/unit/test_errors.py` | 错误类型契约 | 修改 |
|
||||||
|
| `tests/unit/test_backpressure.py` | 后端故障传播 | 修改 |
|
||||||
|
| `tests/unit/test_redis_key_layout.py` | 未知源行为 | 修改 |
|
||||||
|
| `tests/integration/test_redis_cross_connection.py` | 真实 Redis 掉线 | 修改 |
|
||||||
|
| `README.md` / `research-wiki/migrations/chsanalyzer.md` / `CHANGELOG.md` / `pyproject.toml` | 文档与版本 | 修改 |
|
||||||
|
|
||||||
|
## 关键接口(跨任务消费,此处写死)
|
||||||
|
|
||||||
|
`errors.py` 新增与变更部分:
|
||||||
|
|
||||||
|
```python
|
||||||
|
GOVERNANCE_BACKEND_RETRY_AFTER_S = 5.0
|
||||||
|
"""治理后端故障的建议重投间隔(秒)。
|
||||||
|
|
||||||
|
**不是环境配置项**——后端恢复时间物理上不可知(不同于熔断冷却有确定到期
|
||||||
|
时刻),故取一个保守固定值;下游有自己的退避策略时可忽略本字段。取 0 会让
|
||||||
|
积压任务零延迟同时冲击已挂掉的后端(issue #7 §3.2)。
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
class SourceNotConfiguredError(PolyGatewayError):
|
||||||
|
"""源名不在限流后端的配置字典中: 装配缺陷,正常不可达。
|
||||||
|
|
||||||
|
**有意不在** `GatewayUnavailableError` 之下: 它不是"暂时不可用"而是
|
||||||
|
"配置写错了",必须消耗失败预算进死信让人看见;归入可重投家族会让配置
|
||||||
|
错误的任务永远重投、永不告警(issue #7 §3.4)。
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
class GovernanceBackendError(GatewayUnavailableError):
|
||||||
|
"""限流/熔断状态后端自身故障: 必须报错而非放行(防击穿网关,降级方向铁律)。
|
||||||
|
|
||||||
|
继承 `GatewayUnavailableError`: fail-closed 时一个请求都发不出去,语义
|
||||||
|
上即 scope 级不可用,调用方一条 except 即可覆盖(issue #7)。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
message: str,
|
||||||
|
*,
|
||||||
|
scope: str,
|
||||||
|
retry_after_s: float = GOVERNANCE_BACKEND_RETRY_AFTER_S,
|
||||||
|
source_name: str | None = None,
|
||||||
|
) -> None:
|
||||||
|
super().__init__(
|
||||||
|
scope=scope,
|
||||||
|
reason="governance_backend_down",
|
||||||
|
retry_after_s=retry_after_s,
|
||||||
|
source_name=source_name,
|
||||||
|
)
|
||||||
|
# 父类会把 message 覆写为 "{scope} 网关暂时不可用: {reason}",而各构造点
|
||||||
|
# 携带的诊断串是排障主线索,必须保住(设计 §3.5,机制已实跑验证)
|
||||||
|
self.args = (message,)
|
||||||
|
```
|
||||||
|
|
||||||
|
两个 gate 包装器的构造签名(`scope` 为 keyword-only 必填):
|
||||||
|
|
||||||
|
```python
|
||||||
|
class QuotaGate:
|
||||||
|
def __init__(self, limiter: RateLimiter, *, scope: str) -> None:
|
||||||
|
self._limiter = limiter
|
||||||
|
self._scope = scope
|
||||||
|
|
||||||
|
|
||||||
|
class BreakerGate:
|
||||||
|
def __init__(self, gate: ProviderGate, *, scope: str) -> None:
|
||||||
|
self._gate = gate
|
||||||
|
self._scope = scope
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 任务清单
|
||||||
|
|
||||||
|
### - [x] T1: ARCHITECTURE §6.1 回补(必须先行)
|
||||||
|
|
||||||
|
**文件**: `research-wiki/ARCHITECTURE.md`(§6.1,约 372-380 行)
|
||||||
|
|
||||||
|
**行为**: 在错误分类表补两行——`GovernanceBackendError`(scope 级不可用,reason 恒为 `governance_backend_down`)与 `SourceNotConfiguredError`(装配缺陷,不重试不换源,消耗失败预算);scope 级 `reason` 值域由 5 值扩为 6 值,增 `governance_backend_down`。同时记录本次归位的理由与日期,并说明根因(该类是 M2 引入分布式后端时新增,当时未回补本表)。
|
||||||
|
|
||||||
|
**为什么先行**: `ARCHITECTURE.md` 是单一事实源,新 reason 值域与其现状冲突;先改代码后补文档等于让实现与事实源脱节(设计 §8.1)。
|
||||||
|
|
||||||
|
**验收**: §6.1 表格含上述两行;reason 值域文字与 `errors.py` 将要写入的 `SCOPE_REASONS` 逐字一致。
|
||||||
|
|
||||||
|
**测试要求**: 纯文档,无测试证据要求。
|
||||||
|
|
||||||
|
**验证**: `grep -n "governance_backend_down\|SourceNotConfiguredError" research-wiki/ARCHITECTURE.md` → 至少各 1 处命中。
|
||||||
|
|
||||||
|
**提交**: `docs: admit governance backend failures into the scope-level error model`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### - [x] T2: errors.py 纯增量(新常量、新 reason、新类)+ 导出
|
||||||
|
|
||||||
|
**文件**: 改 `src/polygateway/errors.py`、`src/polygateway/__init__.py`;改 `tests/unit/test_errors.py`
|
||||||
|
|
||||||
|
**行为**:
|
||||||
|
1. 加模块级常量 `GOVERNANCE_BACKEND_RETRY_AFTER_S = 5.0`(docstring 逐字见上文"关键接口");
|
||||||
|
2. `SCOPE_REASONS` 增 `"governance_backend_down"`;
|
||||||
|
3. 新增 `SourceNotConfiguredError(PolyGatewayError)`(定义逐字见上文);
|
||||||
|
4. `__init__.py` 的 import 块与 `__all__` 各增 `SourceNotConfiguredError`(`__all__` 保持字母序: `SourceDeadError` → **`SourceNotConfiguredError`** → `TransientError`,即插在 `SourceDeadError` **之后**)。
|
||||||
|
|
||||||
|
**本任务不动 `GovernanceBackendError`**——它是纯增量,不破坏任何既有调用点,可独立提交且全套件保持通过。
|
||||||
|
|
||||||
|
**测试要求(先失败后通过)**:
|
||||||
|
- 新增用例断言 `SourceNotConfiguredError` **不是** `GatewayUnavailableError` 的子类,且是 `PolyGatewayError` 的子类。改前该类不存在 → `ImportError`;改后 PASS。
|
||||||
|
- 新增用例断言 `"governance_backend_down" in SCOPE_REASONS`,且 `GatewayUnavailableError(scope="llm", reason="governance_backend_down", retry_after_s=0.0)` 可构造。改前 `reason` 校验抛 `ValueError` → 用例失败;改后 PASS。
|
||||||
|
- 新增用例断言 `from polygateway import SourceNotConfiguredError` 可用。
|
||||||
|
|
||||||
|
**验证**: `conda run -n PolyGateway pytest tests/unit/test_errors.py -v` → 全 PASS;`conda run -n PolyGateway pytest tests/ -q` → 与改动前同样全绿(纯增量不应影响任何既有用例)。
|
||||||
|
|
||||||
|
**提交**: `feat: add SourceNotConfiguredError and the governance backend reason`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### - [x] T3: `GovernanceBackendError` 归位 + 22 处构造点 + scope 注入(原子)
|
||||||
|
|
||||||
|
**文件**: 改 `src/polygateway/errors.py`、`backends/redis/limiter.py`、`backends/redis/breaker.py`、`backends/memory/limiter.py`、`middleware/ratelimit.py`、`middleware/breaker.py`、`middleware/retry.py`、`ocr.py`、`embedding.py`;改 `tests/unit/test_errors.py`、`tests/unit/test_backpressure.py`、`tests/unit/test_redis_key_layout.py`、`tests/integration/test_redis_cross_connection.py`
|
||||||
|
|
||||||
|
**为什么必须原子**: `scope` 是必填 keyword,继承变更与全部构造点若分批提交,中间状态会 `TypeError`,门禁跑不过。
|
||||||
|
|
||||||
|
**行为**:
|
||||||
|
|
||||||
|
1. `errors.py`: `GovernanceBackendError` 改继承 `GatewayUnavailableError` 并覆写 `__init__`(逐字见上文"关键接口")。
|
||||||
|
|
||||||
|
2. **两处未知源改抛新类**(设计 §3.4,Q1 已拍板):
|
||||||
|
|
||||||
|
| 位置 | 改为 |
|
||||||
|
|---|---|
|
||||||
|
| `backends/redis/limiter.py:198` | `raise SourceNotConfiguredError(f"未知源 {source_key!r}(scope={self._scope})")` |
|
||||||
|
| `backends/memory/limiter.py:92` | 同上 |
|
||||||
|
|
||||||
|
3. **后端层 11 处补 `scope=self._scope`**(该属性已存在: redis limiter `:170`、redis breaker `:291`、memory limiter 同名字段):
|
||||||
|
- `backends/redis/limiter.py` 的 `:250 / :268 / :275 / :286 / :298 / :305`(6 处)
|
||||||
|
- `backends/redis/breaker.py` 的 `:370 / :388 / :410 / :422 / :432`(5 处)
|
||||||
|
|
||||||
|
4. **两个 gate 包装器**: 构造函数改为上文"关键接口"的签名;`QuotaGate` 4 处(`ratelimit.py:30/38/46/54`)与 `BreakerGate` 5 处(`breaker.py:26/36/46/54/62`)的 `raise` 补 `scope=self._scope`。
|
||||||
|
- ~~各方法开头的 `except GovernanceBackendError: raise` **保持不变**~~ **← 这条是错的,2026-08-06 独立验证时炸出(见 §T6)**。正确做法: 该放行必须扩为 `except (GovernanceBackendError, SourceNotConfiguredError): raise`,否则新增的兄弟类型会落进下一行的 `except Exception` 被**重新包成** `GovernanceBackendError`,使 Q1 的拆分在唯一的生产路径上完全失效。
|
||||||
|
|
||||||
|
5. **三处装配各传 scope**(三处的 `self._scope` 均已在装配前赋值,无需调整顺序):
|
||||||
|
|
||||||
|
| 文件 | 行 | 改为 |
|
||||||
|
|---|---|---|
|
||||||
|
| `middleware/retry.py` | 186-187 | `QuotaGate(limiter, scope=self._scope)` / `BreakerGate(gate, scope=self._scope)` |
|
||||||
|
| `ocr.py` | 122-123 | 同款 |
|
||||||
|
| `embedding.py` | 123-124 | 同款 |
|
||||||
|
|
||||||
|
**测试要求(先失败后通过,逐条对应)**:
|
||||||
|
|
||||||
|
| 用例 | 文件 | 改前为何失败 |
|
||||||
|
|---|---|---|
|
||||||
|
| `GovernanceBackendError` 可被 `except GatewayUnavailableError` 接住,且 `reason == "governance_backend_down"`、`retry_after_s == 5.0` | `tests/unit/test_errors.py` | 改前非其子类,`pytest.raises(GatewayUnavailableError)` 不匹配 |
|
||||||
|
| `str(exc)` 仍为构造时的诊断串(防 §3.5 回归) | `tests/unit/test_errors.py` | 改前无该风险但改后若漏写 `self.args` 即失败,是回归护栏 |
|
||||||
|
| 闸门泄漏路径(共五条,见设计 §1.1)抛出的异常带正确 `scope`、且可被 `except GatewayUnavailableError` 接住;钉住 `try_acquire` / `try_enter` / `progress_age_s` 三条代表路径 | `tests/unit/test_backpressure.py` — **三条都要新增桩**。现状: `progress_age_s` 只有 `TestQuotaGateProgressAge`(`:243-257`)覆盖包装行为、不验 scope;`try_acquire`(`QuotaGate`)与 `try_enter`(`BreakerGate`)**完全无桩** | 改前异常无 `scope` 属性 → `AttributeError`;两条新路径改前无覆盖 |
|
||||||
|
| 未知源抛 `SourceNotConfiguredError`,且断言它**不是** `GatewayUnavailableError` | 改 `tests/unit/test_redis_key_layout.py:70-74`(`test_unknown_source_rejected`,现断言 `GovernanceBackendError`);内存版**当前无对应用例,需新增**一条同款(`backends/memory/limiter.py:92` 的 `_cfg("nope")`) | 改前 redis 版类型断言失败;内存版改前无覆盖(该分支从未被测过) |
|
||||||
|
| Redis 真实掉线时准入侧抛 scope 级异常且 `reason == "governance_backend_down"` | `tests/integration/test_redis_cross_connection.py:228-245`(真实 Redis,不 mock) | 改前无 `reason` 属性 |
|
||||||
|
|
||||||
|
**必须同批更新的既有测试构造点**(新签名为 keyword-only 必填,漏改即 `TypeError: missing required keyword-only argument`,门禁直接红):
|
||||||
|
|
||||||
|
| 位置 | 现状 | 改为 |
|
||||||
|
|---|---|---|
|
||||||
|
| `tests/unit/test_backpressure.py:176 / :181 / :186` | `raise GovernanceBackendError("redis 抖动")` | 补 `scope=`(任意测试 scope,如 `"llm"`) |
|
||||||
|
| `tests/unit/test_errors.py:89` | `exc = GovernanceBackendError("redis down")` | 同上;该用例现断言它**不属于**可重试分类,须一并改为断言它**是** `GatewayUnavailableError` |
|
||||||
|
| `tests/unit/test_backpressure.py:255 / :257` | `QuotaGate(_L())` / `QuotaGate(_Broken())` | `QuotaGate(_L(), scope="llm")` 等 |
|
||||||
|
|
||||||
|
**保真校验检查点**: 提交前对照上文"保真校验"五条逐条自查,确认无一被顺带改动。特别核对 `RedisPermit.release/settle`(`redis/limiter.py:133,151`)的 `except GovernanceBackendError` 仍能接住释放侧失败——该处是设计 §4 否决"让原始异常穿透"路线的直接原因。
|
||||||
|
|
||||||
|
**验证**:
|
||||||
|
- `conda run -n PolyGateway pytest tests/unit tests/contracts -v` → 全 PASS
|
||||||
|
- `conda run -n PolyGateway pytest tests/integration -v` → 全 PASS(需真实 Redis)
|
||||||
|
- `conda run -n PolyGateway pytest tests/ -q` → `0 failed`
|
||||||
|
- `conda run -n PolyGateway radon cc src -n C -s` → 无输出
|
||||||
|
- `make lint` → import-linter 契约全绿(本次不新增跨层依赖,应无变化)
|
||||||
|
|
||||||
|
**提交**: `fix: reparent governance backend failures under GatewayUnavailableError (issue #7)`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### - [x] T4: 公开错误面文档(issue #7 第二诉求)
|
||||||
|
|
||||||
|
**文件**: 改 `README.md`(§"错误模型(四分类)",约 114-125 行)、`research-wiki/migrations/chsanalyzer.md`
|
||||||
|
|
||||||
|
**行为**:
|
||||||
|
1. README 增一张两列表,明确区分**会到达调用方**与**库内吸收**:
|
||||||
|
|
||||||
|
| 会到达调用方 | 库内吸收 |
|
||||||
|
|---|---|
|
||||||
|
| `GatewayUnavailableError` 族(`CircuitOpenError` / `AllSourcesExhausted` / `GovernanceBackendError`) | `TransientError` |
|
||||||
|
| `RequestRejectedError` | `SourceDeadError` |
|
||||||
|
| `ResultInvalidError` | |
|
||||||
|
| `SourceNotConfiguredError` | |
|
||||||
|
|
||||||
|
2. 在该表下补一句说明: `TransientError` / `SourceDeadError` 的 docstring 描述的是**库内治理行为**,它们被 `middleware/retry.py:365` 接住并在预算耗尽时包成 `AllSourcesExhausted`,**不会**到达调用方——issue #7 记载下游曾据此写错整段设计文档。
|
||||||
|
3. `migrations/chsanalyzer.md` 的 G1 条目补注:后端故障现已并入 `GatewayUnavailableError`,项目侧 `except GatewayUnavailableError` 一条即覆盖完整,无需为 `GovernanceBackendError` 单列分支。
|
||||||
|
|
||||||
|
**验收**: 调用方仅读 README 即可判断该 catch 什么,无需读 `middleware/retry.py`。
|
||||||
|
|
||||||
|
**测试要求**: 纯文档,无测试证据要求。
|
||||||
|
|
||||||
|
**验证**: `grep -n "库内吸收" README.md` → 命中。
|
||||||
|
|
||||||
|
**提交**: `docs: publish which errors reach callers and which the library absorbs`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### - [x] T5: 版本 1.1.0 + CHANGELOG + Wiki 同步
|
||||||
|
|
||||||
|
**文件**: 改 `pyproject.toml`(version)、`src/polygateway/__init__.py`(`__version__`)、`CHANGELOG.md`;按 `research-wiki/docs-convention.md` §2 同步 Gitea Wiki
|
||||||
|
|
||||||
|
**行为**: 版本 `1.0.6` → `1.1.0`(有行为变更但无 API 破坏:加父类是扩大)。CHANGELOG 需写明:
|
||||||
|
|
||||||
|
- **行为变更**: 后端故障从"落入调用方兜底分支"变为"被 `except GatewayUnavailableError` 捕获";下游据此把它按"延期重投、不消耗失败预算"处置,这正是修复目标,但**处置路线确实变了**,升级前须确认下游的兜底分支没有依赖它。
|
||||||
|
- **新增**: `SourceNotConfiguredError`(公共导出)、`GOVERNANCE_BACKEND_RETRY_AFTER_S`、scope 级 reason `governance_backend_down`。
|
||||||
|
- **下游请读**: `GovernanceBackendError` 现携带 `scope` / `reason` / `retry_after_s`(默认 5.0)/`per_source_reasons`;`str(exc)` 仍是原诊断串,结构化字段并存。配置写错(源名不匹配)现在抛 `SourceNotConfiguredError` 而非 `GovernanceBackendError`,它**不**属于可重投家族——这是有意的,目的是让装配缺陷进死信而不是永远重投。
|
||||||
|
|
||||||
|
**验收**: 版本三处一致(`pyproject.toml` / `__init__.py` / CHANGELOG 标题);CLAUDE.md §6 要求"版本 bump 提交不得裸发",故本任务必须与 wiki 同步同批。
|
||||||
|
|
||||||
|
**测试要求**: 无行为变更,`pytest tests/ -q` 保持全绿即可。
|
||||||
|
|
||||||
|
**验证**: `grep -n "1.1.0" pyproject.toml src/polygateway/__init__.py CHANGELOG.md` → 三处命中。
|
||||||
|
|
||||||
|
**提交**: `chore: release 1.1.0`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
### - [x] T6: 修复独立验证炸出的阻塞缺陷(计划外,2026-08-06)
|
||||||
|
|
||||||
|
T1–T5 全绿、全部门禁通过之后,全新上下文的 verifier 用一个**走 `QuotaGate` 的**端到端用例炸出:装配缺陷在唯一的生产路径上根本没有拆出去。
|
||||||
|
|
||||||
|
**缺陷**: `QuotaGate`/`BreakerGate` 的 `except GovernanceBackendError: raise` 只放行了旧类型,新增的 `SourceNotConfiguredError` 落进下一行 `except Exception` 被重新包成 `GovernanceBackendError`(`reason=governance_backend_down`、`retry_after_s=5.0`)。实证:
|
||||||
|
|
||||||
|
```
|
||||||
|
RAISED: GovernanceBackendError | isGatewayUnavailable=True | isSourceNotConfigured=False
|
||||||
|
| 限流后端故障(source_stats): 未知源 's1'(scope=llm)
|
||||||
|
```
|
||||||
|
|
||||||
|
即配置写错的任务照样落进"可延期重投"家族,**永远重投、永不进死信、无人告警**——正是 Q1 要防的镜像 bug,G2 等于没做。
|
||||||
|
|
||||||
|
**为什么原有测试测不出来**: T3 写的两条用例(`test_backpressure.py`、`test_redis_key_layout.py`)都直接打私有 `_cfg()`,绕过了包装器;而治理循环只经包装器访问后端。**盲区在于测试打的层次比生产路径低一层。**
|
||||||
|
|
||||||
|
**修复**(三处):
|
||||||
|
|
||||||
|
| 文件 | 改动 |
|
||||||
|
|---|---|
|
||||||
|
| `middleware/ratelimit.py` | 4 个方法的放行扩为 `except (GovernanceBackendError, SourceNotConfiguredError): raise` |
|
||||||
|
| `middleware/breaker.py` | 同上,5 个方法 |
|
||||||
|
| `middleware/telemetry.py:254` | 终态捕获元组加 `SourceNotConfiguredError`。**连带坑**: 放行生效后该异常不再是 `GovernanceBackendError`,而它在任何 attempt 之前抛出,若不显式捕获则 `emit_terminal_failure` 不触发、该路径**遥测归零**,违反"遥测必录"铁律 |
|
||||||
|
|
||||||
|
**回归测试**: `test_backpressure.py::TestUnknownSourceIsAssemblyDefect::test_survives_the_quota_gate_wrapper`(参数化覆盖 `try_acquire` / `stats`),**走包装器而非私有方法**。修前 2 failed,修后 PASS。
|
||||||
|
|
||||||
|
**同批文档订正**: 泄漏路径由"三条"改为**五条**(遗漏了 `QuotaGate.stats` 与 `BreakerGate.retry_after_s`,判据是该调用点是否被 `_record_quietly` 包裹);CHANGELOG 的 `per_source_reasons` 表述改为"属性存在但恒为 `{}`"。
|
||||||
|
|
||||||
|
## 完成后
|
||||||
|
|
||||||
|
按 CLAUDE.md §3 Phase 2,合并前须派**全新上下文**的 verifier subagent 做独立验证(`verification-before-completion`),并按新规则**前台运行**。随后走 `finishing-a-development-branch` 决定合并方式,并在 Gitea 关闭 issue #7。
|
||||||
@@ -0,0 +1,276 @@
|
|||||||
|
# 实施计划: stall 判定改为非生产性等待口径(Issue #8)
|
||||||
|
|
||||||
|
- **依据设计**: `research-wiki/designs/2026-08-06-issue8-stall-budget-design.md`(**已批准 2026-08-06**)
|
||||||
|
- **分支**: `feat/issue-8-stall-budget`(已建,已含设计提交 `bfe423d` + `ce2dda7`)
|
||||||
|
- **目标**: 让 stall 计时器只累计非生产性等待,解除 `timeout_s` 与 `stall_window_s` 的隐式耦合,使重试预算在超时场景下真实可用。
|
||||||
|
- **方案概述**: 新增调用级 `StallClock`(总时间减去 `_attempt` 耗时),替换三条治理循环里的墙钟 `entered_at`。判死双条件的结构、`inf` 语义、错误面、429 免预算全部不动。
|
||||||
|
- **涉及技术**: Python 3.11 asyncio、`contextlib.asynccontextmanager`、pytest + `FakeClock`。
|
||||||
|
|
||||||
|
## 保真校验适用性
|
||||||
|
|
||||||
|
**适用**。三条治理循环均为 `reference/CHSAnalyzer app/providers/governance.py:200-285` 的移植物(ARCHITECTURE.md §1.4 关键资产)。但 **`reference/` 当前不在工作区**,无法逐段比对源码,故保真基准改为两处已入库的等价证据:
|
||||||
|
|
||||||
|
1. 设计文档 §4「旧版行为审计」表——9 条既有行为逐条标注保留/替换,实施时逐条核对;
|
||||||
|
2. 代码内既有的 CHS 行号注释(`retry.py:212`「调用级累计计时,循环内不重置(CHS governance.py:207)」、`:303-304`「双条件 stall 判死(CHS governance.py:270-281)」、`:315`「jitter 防惊群(CHS governance.py:283-285)」)与 `tests/unit/test_backpressure.py:1-6` 的蓝本 docstring。
|
||||||
|
|
||||||
|
**唯一允许的语义变更是条件 A 的度量口径**(设计 §4 中标"替换"的那一行)。其余任何条件分支、退避公式、jitter 区间、状态迁移若发生行为改变,即为违规,必须回退。
|
||||||
|
|
||||||
|
## 文件结构
|
||||||
|
|
||||||
|
| 文件 | 动作 | 职责 |
|
||||||
|
|---|---|---|
|
||||||
|
| `src/polygateway/middleware/retry.py` | 修改 | 新增模块级 `StallClock`(共享单元);主循环与 `_on_no_runnable` 改用之 |
|
||||||
|
| `src/polygateway/embedding.py` | 修改 | 复用 `StallClock`;`_on_no_runnable` 改签名 |
|
||||||
|
| `src/polygateway/ocr.py` | 修改 | 同上 |
|
||||||
|
| `src/polygateway/config.py` | 修改 | `_validate_stall` docstring 改写(仅注释,不改逻辑) |
|
||||||
|
| `tests/unit/test_backpressure.py` | 修改 | 新增 `TestStallBudget` 类;订正 `:121` docstring |
|
||||||
|
| `tests/unit/test_embedding.py` | 修改 | 新增 embedding 回归用例 |
|
||||||
|
| `tests/unit/test_ocr_client.py` | 修改 | 新增 ocr 回归用例 |
|
||||||
|
| `.env.example` | 修改 | 第 41 行注释改写 |
|
||||||
|
| `research-wiki/ARCHITECTURE.md` | 修改 | §7.3 背压条目补记新口径 |
|
||||||
|
| `CHANGELOG.md` | 修改 | 记治理行为变更 |
|
||||||
|
|
||||||
|
**不创建任何新模块**。`StallClock` 放在 `retry.py`,沿用 `backoff_delay` 已被 embedding/ocr 复用的既有手法(依赖方向不变:`embedding.py:42`、`ocr.py:38` 已在 import 该模块)。
|
||||||
|
|
||||||
|
## 关键接口(跨任务消费,此处给出实际代码)
|
||||||
|
|
||||||
|
`StallClock` 由 T1 落地,T2/T3 直接消费,签名以此为准:
|
||||||
|
|
||||||
|
```python
|
||||||
|
class StallClock:
|
||||||
|
"""调用级 stall 计时器: 只累计非生产性等待(设计 §3.1)。
|
||||||
|
|
||||||
|
stall 预算治理的是"无人治理的等待"(429 退避、配额轮询、熔断冷却),
|
||||||
|
真实尝试已由重试预算 max_attempts 治理,故须从 stall 账里扣除——
|
||||||
|
两者重叠计费正是 issue #8 的根因。
|
||||||
|
|
||||||
|
每次调用创建一个实例。严禁提升为实例属性: 并发调用共享会互相污染计时。
|
||||||
|
"""
|
||||||
|
|
||||||
|
__slots__ = ("_now", "_entered_at", "_productive_s")
|
||||||
|
|
||||||
|
def __init__(self, now: Callable[[], float]) -> None:
|
||||||
|
self._now = now
|
||||||
|
self._entered_at = now()
|
||||||
|
self._productive_s = 0.0
|
||||||
|
|
||||||
|
def stalled_s(self) -> float:
|
||||||
|
"""非生产性等待累计秒数 = 总耗时 - 真实尝试耗时。"""
|
||||||
|
return self._now() - self._entered_at - self._productive_s
|
||||||
|
|
||||||
|
@contextlib.asynccontextmanager
|
||||||
|
async def attempting(self) -> AsyncIterator[None]:
|
||||||
|
"""包裹一次真实尝试, 其耗时记为生产性(边界即 _attempt 的边界)。"""
|
||||||
|
started = self._now()
|
||||||
|
try:
|
||||||
|
yield
|
||||||
|
finally:
|
||||||
|
# 只做算术, 不吞任何异常——CancelledError 逐字穿透(库铁律)
|
||||||
|
self._productive_s += self._now() - started
|
||||||
|
```
|
||||||
|
|
||||||
|
需在 `retry.py` 新增 `import contextlib`;`AsyncIterator` 从 `collections.abc` 引入(该文件已有 `from __future__ import annotations`,类型注解延迟求值,若 `TYPE_CHECKING` 块中已有 `Callable` 则复用)。
|
||||||
|
|
||||||
|
三条循环的改造模式一致:
|
||||||
|
|
||||||
|
```python
|
||||||
|
clock = StallClock(self._now) # 替换 entered_at = self._now()
|
||||||
|
...
|
||||||
|
await self._on_no_runnable(gate_rejections, reasons, clock) # 形参改类型
|
||||||
|
...
|
||||||
|
async with clock.attempting():
|
||||||
|
outcome = await self._attempt(...) # 原调用不变, 仅被包裹
|
||||||
|
```
|
||||||
|
|
||||||
|
判定式由 `self._now() - entered_at > stall` 改为 `clock.stalled_s() > stall`,**条件 B 与 `and` 结构逐字不动**。
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## T1 — `StallClock` 落地与 chat 路径改造
|
||||||
|
|
||||||
|
- [ ] **文件**: `src/polygateway/middleware/retry.py`(修改)、`tests/unit/test_backpressure.py`(修改)
|
||||||
|
|
||||||
|
### 行为与验收标准
|
||||||
|
|
||||||
|
1. 按上文「关键接口」实现 `StallClock`,置于模块级(建议紧邻既有 `backoff_delay` 纯函数,便于 embedding/ocr 一并 import)。
|
||||||
|
2. `RetryMW.__call__`:`entered_at = self._now()`(`retry.py:212`)改为 `clock = StallClock(self._now)`;主循环判定(`:217`)改为 `clock.stalled_s() > stall`;`_attempt` 调用(`:228`)用 `async with clock.attempting():` 包裹。
|
||||||
|
3. `_on_no_runnable`(`:286-288`)形参 `entered_at: float` 改为 `clock: StallClock`,其内判定(`:306`)同步改为 `clock.stalled_s() > stall`。
|
||||||
|
4. **不得改动**:429 免预算分支(`:233-234`)、`max(fails, 1)` 退避(`:243`)、jitter 公式(`:315`)、`fail_fast` 分支、条件 B `await self._quota.progress_age_s() > stall`、`AllSourcesExhausted` 的任何字段。
|
||||||
|
5. 保留 `retry.py:212` 的 CHS 行号注释并补记新口径(说明"调用级累计、循环内不重置"仍然成立,变的只是不再计入真实尝试)。
|
||||||
|
|
||||||
|
### 测试要求(先失败后通过)
|
||||||
|
|
||||||
|
在 `tests/unit/test_backpressure.py` 新增 `class TestStallBudget`:
|
||||||
|
|
||||||
|
| 用例 | 构造 | 断言 |
|
||||||
|
|---|---|---|
|
||||||
|
| `test_single_timeout_does_not_exhaust_stall_budget` | `_STALL=300`,源 `timeout_s` 等价;脚本 `[TransientError(耗时 350s), _ok()]`——用 `FakeTransport` 配合在尝试中推进 `FakeClock` 350s | 返回成功响应。**改前**:抛 `AllSourcesExhausted(reason="stalled")` |
|
||||||
|
| `test_productive_time_excluded_from_stall` | 连续两次尝试各推进时钟 `_STALL+100`,第三次成功 | 返回成功;全程不触发 `stalled` |
|
||||||
|
| `test_nonproductive_wait_still_triggers_stall` | 沿用 `_blocked_limiter`,轮询中推进时钟超窗且不 `mark_progress` | 抛 `stalled`(兜底未被削弱) |
|
||||||
|
| `test_saturation_429_still_stalls` | 源持续抛 429(`TransientError(status_code=429)`),退避 sleep 中推进时钟 | 抛 `stalled` 而非无限循环(`BoundedSleep` 上限内)。钉住设计 §3.5 |
|
||||||
|
| `test_cancel_inside_attempt_pierces` | 在 `_attempt` 内挂起后 `task.cancel()` | 抛 `CancelledError`(`attempting()` 的 finally 不吞) |
|
||||||
|
| `test_concurrent_calls_do_not_share_clock` | 两路并发调用,一路长尝试、一路正常 | 两路互不影响;钉住 `StallClock` 不得为实例属性 |
|
||||||
|
| `test_telemetry_time_counts_as_productive` | 注入一个在 `emit_attempt` 中推进 `FakeClock` 超过 `_STALL` 的慢 emitter,transport 正常成功 | 返回成功响应,不触发 `stalled`。**钉住设计 §3.1 的边界声明**:遥测收尾属生产性,遥测抖动不得参与判死。若将来有人把 `attempting()` 的包裹范围收窄到只包 transport 调用,该不变式会被悄悄破坏而其余用例抓不到 |
|
||||||
|
|
||||||
|
**既有四象限用例(`TestStallQuadrants` 四条)必须原样通过,不得修改断言**——它们全程无真实尝试或真实尝试耗时为 0,`stalled_s()` 与旧墙钟等价。若其中任何一条需要改断言才能通过,说明实现越界,停下来复核。
|
||||||
|
|
||||||
|
同时订正 `tests/unit/test_backpressure.py:121` 的 docstring:「仅全局超窗(从未出餐 age=inf)」保持不变(该语义确实不变),但补一句说明本地口径已是非生产性等待。
|
||||||
|
|
||||||
|
### 验证命令
|
||||||
|
|
||||||
|
```bash
|
||||||
|
conda run -n PolyGateway pytest tests/unit/test_backpressure.py -v
|
||||||
|
conda run -n PolyGateway pytest tests/unit/test_retry.py -v
|
||||||
|
```
|
||||||
|
|
||||||
|
预期:全部 PASS。先在实现前跑新增用例,记录 `test_single_timeout_does_not_exhaust_stall_budget` 的 FAILED 输出作为红证据。
|
||||||
|
|
||||||
|
- [ ] **提交点**: `fix: bill only non-productive waiting against the chat stall budget`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## T2 — embedding 路径改造
|
||||||
|
|
||||||
|
- [ ] **文件**: `src/polygateway/embedding.py`(修改)、`tests/unit/test_embedding.py`(修改)
|
||||||
|
|
||||||
|
### 行为与验收标准
|
||||||
|
|
||||||
|
1. `embedding.py:42` 的 import 增加 `StallClock`(该行已 import `_failure_reason, backoff_delay`)。
|
||||||
|
2. `_embed_batch`(`:182`):`entered_at = self._now()` 改为 `clock = StallClock(self._now)`;`_attempt` 调用(`:188`)用 `async with clock.attempting():` 包裹;`_on_no_runnable` 传参(`:186`)改为 `clock`。
|
||||||
|
3. `_on_no_runnable`(`:230-232`)形参改 `clock: StallClock`,判定(`:248`)改 `clock.stalled_s() > stall`。
|
||||||
|
4. **不得新增主循环 stall 判定**(设计 §5.4:embedding 无 429 免预算,`fails += 1` 无条件,缺口不存在;新增等于凭空多一条判死路径)。
|
||||||
|
5. **不得改动**:`fails += 1` 的无条件性(`:191`)、`max_attempts` 判定、退避调用(`:200`)。
|
||||||
|
|
||||||
|
### 测试要求(先失败后通过)
|
||||||
|
|
||||||
|
在 `tests/unit/test_embedding.py` 新增 `test_single_timeout_does_not_exhaust_stall_budget`。
|
||||||
|
|
||||||
|
**构造方式(已核实可行,不必绕过既有 helper)**:`_embed_client`(`:219`)的 `**overrides` 直通 `EmbeddingClient.__init__`,而后者接受 `now`/`sleep`/`rng`(`embedding.py:109-111`),故可写 `_embed_client([src], script, now=clock, sleep=<推进时钟的 fake>)`。制造一轮 `_on_no_runnable` 沿用 `test_backpressure.py:75-85` `_blocked_limiter` 的手法:源 `max_concurrency=1`,测试先 `try_acquire` 占满 permit,在 fake sleep 回调里释放。helper 内的 `InMemoryLimiter` 未注入 `now` 不影响本用例——判定要的是 `progress_age_s()` 返回 `inf`(从未 `mark_progress`),与 limiter 时钟无关。
|
||||||
|
|
||||||
|
**断言**:第一次尝试推进 `FakeClock` 超过 `stall_window_s` 后抛 `TransientError`,随后经一轮 `_on_no_runnable` 再恢复,最终返回成功的 `EmbeddingResponse`。改前应抛 `AllSourcesExhausted(reason="stalled")`。
|
||||||
|
|
||||||
|
**取消穿透验收点**:既有 `test_cancel_releases_permit`(`test_embedding.py:331-339`)的取消路径**将被新的 `async with clock.attempting()` 包住**,故它是本任务的必过回归项,不得因改动而修改其断言。若它转红,说明 `attempting()` 的 `finally` 吞了 `CancelledError` 或泄漏了 permit,停下来复核而非改测试。
|
||||||
|
|
||||||
|
### 验证命令
|
||||||
|
|
||||||
|
```bash
|
||||||
|
conda run -n PolyGateway pytest tests/unit/test_embedding.py -v
|
||||||
|
```
|
||||||
|
|
||||||
|
预期:全部 PASS(含既有取消与遥测用例)。
|
||||||
|
|
||||||
|
- [ ] **提交点**: `fix: apply the non-productive stall budget to the embedding loop`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## T3 — ocr 路径改造
|
||||||
|
|
||||||
|
- [ ] **文件**: `src/polygateway/ocr.py`(修改)、`tests/unit/test_ocr_client.py`(修改)
|
||||||
|
|
||||||
|
### 行为与验收标准
|
||||||
|
|
||||||
|
与 T2 同构,对应行号:import(`:38`)、`_call` 的 `entered_at`(`:207`)、`_on_no_runnable` 传参(`:211`)、`_attempt` 调用(`:213`)、`_on_no_runnable` 签名(`:255-257`)与判定(`:273`)。同样**不得新增主循环 stall 判定**,不得改动 `fails += 1`(`:216`)与退避(`:225`)。
|
||||||
|
|
||||||
|
### 测试要求(先失败后通过)
|
||||||
|
|
||||||
|
`tests/unit/test_ocr_client.py` 新增与 T2 同构的 `test_single_timeout_does_not_exhaust_stall_budget`,覆盖 `recognize_text` 或 `parse_layout` 任一端点即可(两者共用 `_call`)。
|
||||||
|
|
||||||
|
**构造方式**:同 T2——经该文件既有的 `_client(...)` helper 传 `now=clock` 与推进时钟的 fake `sleep`;`_on_no_runnable` 一轮用"源 `max_concurrency=1` + 测试预先占满 permit + 在 fake sleep 回调里释放"制造。
|
||||||
|
|
||||||
|
**取消穿透验收点**:既有 `test_cancel_during_transport_releases_permit`(`test_ocr_client.py:339-348`)与其上方的退避期取消用例同样会被新包裹覆盖,均为必过回归项,不得修改断言。
|
||||||
|
|
||||||
|
### 验证命令
|
||||||
|
|
||||||
|
```bash
|
||||||
|
conda run -n PolyGateway pytest tests/unit/test_ocr_client.py tests/unit/test_monkey_ocr.py -v
|
||||||
|
```
|
||||||
|
|
||||||
|
预期:全部 PASS。
|
||||||
|
|
||||||
|
- [ ] **提交点**: `fix: apply the non-productive stall budget to the ocr loop`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## T4 — 配置侧注释对齐(无逻辑变更)
|
||||||
|
|
||||||
|
- [ ] **文件**: `src/polygateway/config.py`(修改)、`.env.example`(修改)
|
||||||
|
|
||||||
|
### 行为与验收标准
|
||||||
|
|
||||||
|
1. `config.py:240-241` `_validate_stall` 的 docstring 由「stall 窗口须 ≥ 最慢源 TTFT 上限,防把正常慢首包误判为卡死」改写为:说明该校验在新口径下**属保守冗余**——TTFT 等待是生产性时间,已不计入 stall;保留校验是为不改动 ARCHITECTURE.md §7.3 契约 G6(人类 2026-08-06 定夺)。**校验逻辑本身一字不改**。
|
||||||
|
2. `.env.example:41` 注释由「stall 双条件判死窗口;须 ≥ 最大源 TTFT」改写为说明它度量的是**非生产性等待**(429 退避/配额轮询/熔断冷却)累计,与 `TIMEOUT_S` 无耦合、无需按 `timeout × retries` 放大。
|
||||||
|
3. 不新增、不改名任何配置键(`_DEFAULT_STALL_WINDOW_S = 300.0` 保持不变)。
|
||||||
|
|
||||||
|
### 验证命令
|
||||||
|
|
||||||
|
```bash
|
||||||
|
conda run -n PolyGateway pytest tests/unit/test_config.py -v
|
||||||
|
conda run -n PolyGateway make lint
|
||||||
|
```
|
||||||
|
|
||||||
|
预期:全部 PASS(本任务不改逻辑,`test_config.py` 应零变化通过)。
|
||||||
|
|
||||||
|
- [ ] **提交点**: `docs: align the stall window comments with the new metering`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## T5 — 全套件回归与文档同步
|
||||||
|
|
||||||
|
- [ ] **文件**: `research-wiki/ARCHITECTURE.md`(修改)、`CHANGELOG.md`(修改)
|
||||||
|
|
||||||
|
### 行为与验收标准
|
||||||
|
|
||||||
|
1. 跑全套件确认零回归。**命令末尾不得接管道**(CLAUDE.md 执行模式:管道会掩盖真实退出码),需要后台跑时用 `wait`/轮询 PID 判完成。
|
||||||
|
2. `ARCHITECTURE.md` §7.3 背压条目(第 429-431 行区域)补记:stall 双条件的条件 A 现为**非生产性等待累计**,并给出本设计文档指针。既有 G6 契约行保留,补注其在新口径下为保守冗余。
|
||||||
|
3. `CHANGELOG.md` 记治理行为变更(属公共行为变更,须显式列出:单次调用最坏耗时由 `stall_window_s` 抬升至 `max_attempts × timeout_s`)。
|
||||||
|
4. 核对设计 §4 行为审计表 9 条,逐条确认实现与标注一致(保真校验检查点)。
|
||||||
|
|
||||||
|
### 副作用处置
|
||||||
|
|
||||||
|
修复后单次调用最坏耗时变为 `max_attempts × timeout_s`(本机 900s)。跑 e2e 前先评估 `tests/e2e` 的源 `timeout_s` 是否需调小,以免冒烟耗时失控。本机 `.env:37` 的临时缓解 `STALL_WINDOW_S=1200` 可回退默认值(`.env` 不入库,仅在本任务记录该动作)。
|
||||||
|
|
||||||
|
### 验证命令
|
||||||
|
|
||||||
|
```bash
|
||||||
|
conda run -n PolyGateway make lint
|
||||||
|
conda run -n PolyGateway pytest tests/unit tests/integration -v
|
||||||
|
conda run -n PolyGateway make test
|
||||||
|
```
|
||||||
|
|
||||||
|
预期:lint 通过(含 import-linter 依赖契约);单元与集成全绿;覆盖率不低于既有水平。
|
||||||
|
|
||||||
|
- [ ] **提交点**: `docs: record the stall metering change in architecture and changelog`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## T6 — 独立验证与 Wiki 同步
|
||||||
|
|
||||||
|
- [ ] **文件**: Gitea Wiki(独立仓库)
|
||||||
|
|
||||||
|
### 行为与验收标准
|
||||||
|
|
||||||
|
1. **派全新上下文 verifier subagent**(`verification-before-completion`,MANDATORY:跨 3 模块属里程碑级),**前台运行**(`run_in_background: false`,CLAUDE.md 执行模式)。核验对象:设计 §1 的 G1-G4 是否逐条兑现、§4 行为审计表 9 条是否与实现一致、是否出现设计未声明的语义变更、测试是否真的覆盖"先失败后通过"。
|
||||||
|
2. 按 `docs-convention.md` §2「治理行为变更」行同步 Wiki:`解释-治理行为`(stall 判定口径)、`指南-限流与熔断`(配置说明中删除"须按 timeout×retries 放大 stall"一类误导)。
|
||||||
|
3. 在 Gitea issue #8 下回帖:根因、方案、被否决的两个原建议方向及理由、影响面。
|
||||||
|
4. Wiki 注册:
|
||||||
|
```bash
|
||||||
|
.claude/tools/research_wiki.py add_entity research-wiki/ --type plan --id issue8-stall-budget --title "stall 判定改为非生产性等待口径"
|
||||||
|
.claude/tools/research_wiki.py add_edge research-wiki/ --from "plan:issue8-stall-budget" --to "design:issue8-stall-budget" --type implements --evidence "本计划实施该设计的 T1-T6"
|
||||||
|
.claude/tools/research_wiki.py rebuild_index research-wiki/
|
||||||
|
```
|
||||||
|
|
||||||
|
### 验证命令
|
||||||
|
|
||||||
|
```bash
|
||||||
|
conda run -n PolyGateway make ci
|
||||||
|
```
|
||||||
|
|
||||||
|
预期:只读验证全绿。verifier 报告须逐条对应本会话内的工具输出(证据化声明,禁止虚报)。
|
||||||
|
|
||||||
|
- [ ] **提交点**: `chore: register the issue #8 plan and sync the wiki`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 任务依赖
|
||||||
|
|
||||||
|
T1 → (T2 ‖ T3) → T4 → T5 → T6。T2 与 T3 相互独立,但都依赖 T1 落地的 `StallClock`。
|
||||||
@@ -0,0 +1,259 @@
|
|||||||
|
# 实现计划: HTTP 错误响应体留存(Issue #10)
|
||||||
|
|
||||||
|
- **设计**: `research-wiki/designs/2026-08-16-issue10-error-body-retention-design.md`(**已批准 2026-08-16**)
|
||||||
|
- **分支**: `feat/issue-10-error-body-retention`
|
||||||
|
- **目标**: 网关拒绝一次调用时,它说的话必须能在库自己的遥测表里被事后查到。
|
||||||
|
- **方案概述**: transport 翻译层把 HTTP 错误响应体折叠空白并按头尾策略摘要,**同一份串**同时拼进异常 message(经既有 `error` 列落遥测)与新增的基类字段 `body_text`(供下游结构化留存)。覆盖两个 transport 的全部非 2xx 分支。不改任何状态码→分类的映射。
|
||||||
|
- **涉及技术**: Python 3.11 / httpx / pytest。无新增依赖。
|
||||||
|
- **保真校验**: 本计划**不涉及** `reference/` 参考实现迁移——错误分类映射逐条不变,保真体现为"既有分类断言全部保留、无一条被改写"(Task 3/4 验收项)。
|
||||||
|
|
||||||
|
## 文件结构
|
||||||
|
|
||||||
|
| 文件 | 动作 | 职责 |
|
||||||
|
|---|---|---|
|
||||||
|
| `src/polygateway/errors.py` | 改 | 基类 `PolyGatewayError` 新增 `body_text` 字段;与 `raw_text` 的界限 docstring;`RequestRejectedError` 补中转拓扑提醒 |
|
||||||
|
| `src/polygateway/transports/_http_errors.py` | **新建** | 摘要口径单一实现:`summarize_body` / `compose_message` / `response_body` + 三个常量 |
|
||||||
|
| `src/polygateway/transports/openai_compat.py` | 改 | `_status_to_error` 表驱动重写;`_translate_429` 收 `ctx` |
|
||||||
|
| `src/polygateway/transports/monkey_ocr.py` | 改 | `_classify_status` 带摘要 |
|
||||||
|
| `tests/unit/test_http_error_body.py` | **新建** | 摘要单元的纯函数用例(截断边界、头尾保留、折叠、幂等) |
|
||||||
|
| `tests/unit/test_openai_compat.py` | 改 | 状态码参数化断言 message + 字段;超长 `insufficient_quota` 回归 |
|
||||||
|
| `tests/unit/test_monkey_ocr.py` | 改 | OCR 分支同款 + `ResponseNotRead` 降级 |
|
||||||
|
| `tests/unit/test_errors.py` | 改 | `body_text` 默认值与可传性 |
|
||||||
|
| `tests/integration/test_governance_stack.py` | 改 | **端到端验收**:400 调用后 SQLite `error` 列含摘要 |
|
||||||
|
| `README.md` / `CHANGELOG.md` / `pyproject.toml` / `src/polygateway/__init__.py` / `research-wiki/ARCHITECTURE.md` | 改 | 文档与 1.2.0 版本号 |
|
||||||
|
|
||||||
|
## 关键接口(跨任务消费,必须逐字一致)
|
||||||
|
|
||||||
|
```python
|
||||||
|
# src/polygateway/transports/_http_errors.py
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import httpx # response_body 的类型与 ResponseNotRead 都来自它
|
||||||
|
|
||||||
|
_ERROR_BODY_CAP = 2048 # 字符(非字节),含省略标记在内的最终总长上限
|
||||||
|
_HEAD_CHARS = 1400
|
||||||
|
_TAIL_CHARS = 600
|
||||||
|
|
||||||
|
|
||||||
|
def summarize_body(text: str) -> str:
|
||||||
|
"""折叠空白后按头尾策略摘要;空/空白入参返回空串。"""
|
||||||
|
collapsed = " ".join(text.split())
|
||||||
|
if len(collapsed) <= _ERROR_BODY_CAP:
|
||||||
|
return collapsed
|
||||||
|
omitted = len(collapsed) - _HEAD_CHARS - _TAIL_CHARS
|
||||||
|
return f"{collapsed[:_HEAD_CHARS]}…(略 {omitted} 字)…{collapsed[-_TAIL_CHARS:]}"
|
||||||
|
|
||||||
|
|
||||||
|
def compose_message(message: str, summary: str) -> str:
|
||||||
|
"""摘要非空才拼后缀,避免悬空分隔符。"""
|
||||||
|
return f"{message} | {summary}" if summary else message
|
||||||
|
|
||||||
|
|
||||||
|
def response_body(response: httpx.Response) -> str:
|
||||||
|
"""取已缓冲的响应文本;未读缓冲一律降级空串,绝不触发网络读。"""
|
||||||
|
try:
|
||||||
|
return response.text
|
||||||
|
except httpx.ResponseNotRead:
|
||||||
|
return ""
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
# src/polygateway/errors.py
|
||||||
|
class PolyGatewayError(Exception):
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
message: str,
|
||||||
|
*,
|
||||||
|
source_name: str | None = None,
|
||||||
|
status_code: int | None = None,
|
||||||
|
operation: str | None = None,
|
||||||
|
body_text: str = "",
|
||||||
|
) -> None:
|
||||||
|
```
|
||||||
|
|
||||||
|
```python
|
||||||
|
# src/polygateway/transports/openai_compat.py
|
||||||
|
# 既有 errors 导入(:18-23)须补入 PolyGatewayError —— 当前只导了四个子类,
|
||||||
|
# 直接写 _classify 的返回注解会让 ruff 报 F821 未定义名。
|
||||||
|
from polygateway.errors import (
|
||||||
|
PolyGatewayError, # ← 新增
|
||||||
|
RequestRejectedError,
|
||||||
|
ResultInvalidError,
|
||||||
|
SourceDeadError,
|
||||||
|
TransientError,
|
||||||
|
)
|
||||||
|
from polygateway.transports._http_errors import compose_message, summarize_body
|
||||||
|
|
||||||
|
|
||||||
|
def _classify(status: int) -> tuple[type[PolyGatewayError], str]:
|
||||||
|
"""状态码 → (错误类, message 标签);映射与 1.1.2 逐条相同。"""
|
||||||
|
if status in (401, 403):
|
||||||
|
return SourceDeadError, "凭据失效/欠费"
|
||||||
|
if status == 400:
|
||||||
|
return RequestRejectedError, "请求被拒"
|
||||||
|
if status >= 500:
|
||||||
|
return TransientError, "瞬时错误"
|
||||||
|
return RequestRejectedError, "客户端错误"
|
||||||
|
|
||||||
|
|
||||||
|
def _status_to_error(
|
||||||
|
source: SourceConfig, status: int, body_text: str, headers: Mapping[str, str]
|
||||||
|
) -> Exception:
|
||||||
|
summary = summarize_body(body_text) # 全函数只算一次
|
||||||
|
ctx: dict[str, Any] = {
|
||||||
|
"source_name": source.name,
|
||||||
|
"status_code": status,
|
||||||
|
"operation": "chat",
|
||||||
|
"body_text": summary,
|
||||||
|
}
|
||||||
|
if status == 429:
|
||||||
|
return _translate_429(source, body_text, headers, ctx) # 传**原文**,见下
|
||||||
|
cls, label = _classify(status)
|
||||||
|
return cls(compose_message(f"{source.name} {label}: {status}", summary), **ctx)
|
||||||
|
```
|
||||||
|
|
||||||
|
> **实现红线**:`_translate_429` 判 `insufficient_quota` 必须解析**未截断的原文** `body_text`,不得改用 `summary`。摘要会破坏 JSON 结构,超长体一旦改用摘要解析,配额耗尽的源将不再 `force_open`——那是把一个诊断改进变成治理 bug。Task 3 有专门的回归用例钉死这条。
|
||||||
|
|
||||||
|
## 任务清单
|
||||||
|
|
||||||
|
### - [ ] Task 1: 内核新增 `body_text` 字段
|
||||||
|
|
||||||
|
**改**: `src/polygateway/errors.py`
|
||||||
|
|
||||||
|
- `PolyGatewayError.__init__` 按上文签名新增 `body_text: str = ""`,存为实例属性。
|
||||||
|
- 类 docstring 增补与 `ResultInvalidError.raw_text` 的界限:`body_text` = 非 2xx 的 HTTP 错误响应体摘要(对方拒绝的理由);`raw_text` = 2xx 但内容不可解析时的模型输出。并写明"可能包含请求回显,已截断"。
|
||||||
|
- **不动** `TransientError` / `SourceDeadError` / `RequestRejectedError` / `ResultInvalidError` / `GatewayUnavailableError` 的任何既有签名与行为。
|
||||||
|
|
||||||
|
**测试**(`tests/unit/test_errors.py`,扩展 `:29` 的四类构造形态参数化):
|
||||||
|
- 四个 transport 错误类默认 `body_text == ""`;显式传入后可读回。
|
||||||
|
- `GatewayUnavailableError` / `CircuitOpenError` / `AllSourcesExhausted` / `GovernanceBackendError` 的 `body_text` 恒为 `""`(它们不经 HTTP 响应翻译)。
|
||||||
|
|
||||||
|
**验收**: 新增字段不改变任何既有异常的 `str()` 输出。
|
||||||
|
**验证**: `conda run -n PolyGateway pytest tests/unit/test_errors.py -v` → 全 PASS。
|
||||||
|
|
||||||
|
### - [ ] Task 2: 共享摘要单元
|
||||||
|
|
||||||
|
**新建**: `src/polygateway/transports/_http_errors.py`(按上文"关键接口"逐字实现,**含其中的 `import httpx`**,加中文模块/函数 docstring 解释**为什么**折叠空白、为什么头尾保留、为什么 `response_body` 必须降级)
|
||||||
|
|
||||||
|
**新建测试**: `tests/unit/test_http_error_body.py`
|
||||||
|
|
||||||
|
| 用例 | 断言 |
|
||||||
|
|---|---|
|
||||||
|
| 短体原样 | `summarize_body('{"a":1}') == '{"a":1}'` |
|
||||||
|
| 空白折叠 | 多行缩进 JSON → 单行,无连续空格 |
|
||||||
|
| 空 / 纯空白入参 | 返回 `""` |
|
||||||
|
| 长度恰 2048 | 原样返回,无标记 |
|
||||||
|
| 长度 2049 | 走头尾策略 |
|
||||||
|
| 超长体头尾 | 前 1400 字符 == 原文前 1400;**末 600 字符 == 原文末 600**;中段标记内 N == `len(原文) - 2000` |
|
||||||
|
| **尾部关键字段可见**(设计 §7 用例 3c) | 以 issue 真实样本尾部 `"code":"invalid_parameter_error"}}` 收尾构造超长体 → 断言该串出现在摘要中 |
|
||||||
|
| 幂等 | `summarize_body(summarize_body(x)) == summarize_body(x)`(标记不嵌套) |
|
||||||
|
| `compose_message` | 摘要为空时返回原 message 不变;非空时以竖线分隔符拼接 |
|
||||||
|
| `response_body` 降级 | `httpx.Response(400, stream=<未读 SyncByteStream>)` → 返回 `""` 且不抛(构造法见下) |
|
||||||
|
|
||||||
|
未读响应的构造(已实测可用):
|
||||||
|
|
||||||
|
```python
|
||||||
|
class _Unread(httpx.SyncByteStream):
|
||||||
|
def __iter__(self):
|
||||||
|
yield b"body"
|
||||||
|
|
||||||
|
resp = httpx.Response(400, stream=_Unread()) # 未 read → .text 抛 ResponseNotRead
|
||||||
|
```
|
||||||
|
|
||||||
|
**验收**: 摘要总长恒 ≤ `2000 + len(标记)`;头尾各自与原文逐字对应。
|
||||||
|
**验证**: `conda run -n PolyGateway pytest tests/unit/test_http_error_body.py -v` → 全 PASS。
|
||||||
|
|
||||||
|
### - [ ] Task 3: openai_compat 翻译层收口
|
||||||
|
|
||||||
|
**改**: `src/polygateway/transports/openai_compat.py`
|
||||||
|
|
||||||
|
- 新增 `_classify`,`_status_to_error` 按上文骨架重写(五分支各拼各的 message → 查表 + 单点拼装)。
|
||||||
|
- `_translate_429` 签名改为 `(source, body_text, headers, ctx)`,两支 message 各自追加 `compose_message` 后缀,构造改用 `**ctx`;**`json.loads` 仍读原文 `body_text`**。
|
||||||
|
- message 主体逐字保持 1.1.2 原样(`凭据失效/欠费: {status}` / `请求被拒: 400` / `瞬时错误: {status}` / `客户端错误: {status}` / `配额耗尽(insufficient_quota)` / `限速: 429`),只在末尾追加 ` | {摘要}`。
|
||||||
|
- 三个调用点(`:402` embed、`:417` stream、`:509` 非流式)签名不变,**不改动**。
|
||||||
|
- **补 import**:`PolyGatewayError`(errors)与 `compose_message` / `summarize_body`(`._http_errors`),见上文关键接口——漏补则 `make lint` 报 F821(Codex 审查 2026-08-16 提出)。
|
||||||
|
- **不改** `operation` 硬编码 `"chat"`(设计 §5.4 有意留给独立 issue)。
|
||||||
|
|
||||||
|
**测试**(`tests/unit/test_openai_compat.py`,沿用既有 `_transport_for(handler)` + `httpx.MockTransport`):
|
||||||
|
|
||||||
|
| # | 用例 | 断言 |
|
||||||
|
|---|---|---|
|
||||||
|
| 3.1 | 状态码参数化 400 / 401 / 403 / 404 / 500 / 503,handler 返回带真实样本体 | 异常类型与 1.1.2 **逐条相同**;message 含摘要;`exc.body_text` == 摘要 |
|
||||||
|
| 3.2 | 429 普通限速(body 无 `insufficient_quota`) | `TransientError`,message 含摘要,`retry_after_s` 解析不受影响 |
|
||||||
|
| 3.3 | 429 + `insufficient_quota` | `SourceDeadError`,message 含摘要 |
|
||||||
|
| 3.4 | **回归红线**:429 + `insufficient_quota` 且 body 长度 > 2048(前置大量填充字段) | 仍判 `SourceDeadError`——证明类型判定读的是原文而非摘要 |
|
||||||
|
| 3.5 | 空 body 的 400 | message 无悬空分隔符,`body_text == ""` |
|
||||||
|
| 3.6 | 非 JSON body、非 UTF-8 字节 body | 不抛额外异常,分类不变 |
|
||||||
|
| 3.7 | 流式路径(handler 对 stream 请求返回 400 + body) | 经 `_complete_stream:415-417` 抛出的异常同样带摘要 |
|
||||||
|
| 3.8 | embedding 路径(`transport.embed(...)` 遇 400) | 同样带摘要 |
|
||||||
|
|
||||||
|
**验收**: 既有测试零修改通过(除 3.x 新增外);`test_openai_compat.py:558`(match 源名)仍 PASS。
|
||||||
|
**验证**: `conda run -n PolyGateway pytest tests/unit/test_openai_compat.py -v` → 全 PASS。
|
||||||
|
|
||||||
|
### - [ ] Task 4: monkey_ocr 同款收口
|
||||||
|
|
||||||
|
**改**: `src/polygateway/transports/monkey_ocr.py`
|
||||||
|
|
||||||
|
- `_classify_status`:`summary = summarize_body(response_body(exc.response))`,三支 message 统一经 `compose_message` 追加后缀,`ctx` 带 `body_text=summary`。
|
||||||
|
- message 主体保持 `f"{source_name} OCR {operation} HTTP {status}"` 不变。
|
||||||
|
- 429/5xx → `TransientError`、401/403 → `SourceDeadError`、其余 → `RequestRejectedError` 的映射**逐条不变**(OCR 无 429 细分是设计有意保留,见模块 docstring `:53-54`)。
|
||||||
|
|
||||||
|
**测试**(`tests/unit/test_monkey_ocr.py`):
|
||||||
|
- 扩展 `:300` 的状态码参数化:各分支 message 含摘要且 `body_text` 非空,分类不变。
|
||||||
|
- `ResponseNotRead` 降级:`exc.response` 为未读流 → `body_text == ""`,message 无悬空分隔符,**分类仍正确**(不得因取 body 失败而改变错误类型或抛出 httpx 异常)。
|
||||||
|
|
||||||
|
**验收**: `:195` 与 `:300` 既有断言不被改写。
|
||||||
|
**验证**: `conda run -n PolyGateway pytest tests/unit/test_monkey_ocr.py -v` → 全 PASS。
|
||||||
|
|
||||||
|
### - [ ] Task 5: 端到端遥测验收(**本计划的硬判据**)
|
||||||
|
|
||||||
|
**改**: `tests/integration/test_governance_stack.py`
|
||||||
|
|
||||||
|
新增用例,沿用既有 `_full_client(handler, telemetry=SQLiteRecorder(...))` 与 `:135` 的 `SELECT error FROM llm_calls` 断言模式:
|
||||||
|
|
||||||
|
- handler 对 chat 请求返回 `httpx.Response(400, content=<issue #10 真实样本体>)`。
|
||||||
|
- `client.chat(...)` 抛 `RequestRejectedError`(400 不重试不换源,行为不变)。
|
||||||
|
- `recorder.close()` 后查 `SELECT error FROM llm_calls`:该行 `error` 串**含样本体里的 `InvalidParameter` 与结尾的 `invalid_parameter_error`**。
|
||||||
|
|
||||||
|
真实样本体(取自 issue #10 原文,一字不改):
|
||||||
|
|
||||||
|
```json
|
||||||
|
{"error":{"message":"<400> ***.***.InvalidParameter: The image format is illegal and cannot be opened","type":"invalid_request_error","param":"","code":"invalid_parameter_error"}}
|
||||||
|
```
|
||||||
|
|
||||||
|
**验收**: 这条断言在 Task 1-4 之前**必然失败**(1.1.2 的 `error` 列只有 `"qwen_1 请求被拒: 400"`),之后通过——这就是本 issue 的"先失败后通过"证据主体,执行时须保留失败输出截图/文本进提交说明。
|
||||||
|
**验证**: `conda run -n PolyGateway pytest tests/integration/test_governance_stack.py -v` → 全 PASS。
|
||||||
|
|
||||||
|
### - [ ] Task 6: 文档与版本
|
||||||
|
|
||||||
|
**改**:
|
||||||
|
|
||||||
|
| 文件 | 内容 |
|
||||||
|
|---|---|
|
||||||
|
| `src/polygateway/errors.py` | `RequestRejectedError` docstring 加一句:经中转部署时 400 可能源于中转自身抖动,批处理场景下游宜自备兜底分类(设计 §5.2) |
|
||||||
|
| `research-wiki/ARCHITECTURE.md` §6.2 | 同一提醒 + 注明四分类错误自 1.2.0 起携带 `body_text` |
|
||||||
|
| `CHANGELOG.md` | 新增 `## 1.2.0(2026-08-16)` 段:行为变更(message 追加摘要 → 遥测 `error` 列变长)、新增字段、不变项(分类映射零变更、错误面零变更) |
|
||||||
|
| `README.md:34` | 安装 pin `==1.1.*` → **`>=1.2,<2`**(2026-08-16 人类定夺;漏改则下游静默停在 1.1.2) |
|
||||||
|
| `pyproject.toml` + `src/polygateway/__init__.py` | 版本号 `1.1.2` → `1.2.0`,**两处必须一致** |
|
||||||
|
|
||||||
|
**验收**: `grep -rn "1\.1\.\*" README.md` 零命中;两处版本号一致。
|
||||||
|
**验证**: `conda run -n PolyGateway python -c "import polygateway; print(polygateway.__version__)"` → `1.2.0`。
|
||||||
|
|
||||||
|
### - [ ] Task 7: 合并前全量门
|
||||||
|
|
||||||
|
1. `make lint`(ruff + import-linter)→ 零违规,**重点确认新建 `transports/_http_errors.py` 未触发洋葱分层契约**。
|
||||||
|
2. `make test` 全套件 → 全 PASS,覆盖率不低于既有水平。
|
||||||
|
3. 派**全新上下文** verifier subagent 独立验证(CLAUDE.md §3 Phase 2 硬门):逐条核对 Task 1-6 验收项与本会话工具输出。
|
||||||
|
4. `finishing-a-development-branch` 合并回 main(`--no-ff`),合并后在 main 上重跑 `make lint` 与全套件。
|
||||||
|
|
||||||
|
**发布**(合并后)严格按 CLAUDE.md §4.4.1 九步执行,不在本计划展开;其中步骤 1(更新 README)已在 Task 6 前置完成,**构建前须再次确认 pin 已是 `>=1.2,<2`**。
|
||||||
|
|
||||||
|
## 执行顺序与提交点
|
||||||
|
|
||||||
|
```
|
||||||
|
Task 1 ──┐
|
||||||
|
├── Task 3 ──┐
|
||||||
|
Task 2 ──┴── Task 4 ──┴── Task 5 ── Task 6 ── Task 7
|
||||||
|
```
|
||||||
|
|
||||||
|
Task 1 与 2 可并行(互不依赖);Task 3、4 都依赖 1+2;Task 5 依赖 3;Task 6 独立于代码但须在 Task 7 之前。每个 Task 一次语义化提交(`commit` skill),Task 5 的提交说明须附"修复前失败、修复后通过"的实际输出。
|
||||||
@@ -0,0 +1,45 @@
|
|||||||
|
---
|
||||||
|
type: plan
|
||||||
|
node_id: plan:governance-backend-error
|
||||||
|
title: "实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)"
|
||||||
|
date: 2026-08-06
|
||||||
|
---
|
||||||
|
|
||||||
|
# 实现计划: 治理后端故障归位为 scope 级不可用(Issue #7)
|
||||||
|
|
||||||
|
全文见 `2026-08-06-governance-backend-error-plan.md`。实现设计 [[governance-backend-error]](已批准 2026-08-06)。
|
||||||
|
|
||||||
|
## 五个任务
|
||||||
|
|
||||||
|
| # | 任务 | 关键约束 |
|
||||||
|
|---|---|---|
|
||||||
|
| T1 | `ARCHITECTURE.md` §6.1 回补两行 + reason 值域扩为 6 值 | **必须先行**——单一事实源纪律,先改代码后补文档等于让实现与事实源脱节(设计 §8.1) |
|
||||||
|
| T2 | `errors.py` 纯增量: 新常量、新 reason、`SourceNotConfiguredError` + 顶层导出 | 刻意不动 `GovernanceBackendError`,故全套件保持通过,可独立提交 |
|
||||||
|
| T3 | `GovernanceBackendError` 归位 + 22 处构造点 + gate scope 注入 + 全部受影响测试 | **必须原子**: `scope` 是必填 keyword,分批提交的中间状态会 `TypeError` |
|
||||||
|
| T4 | README 公开错误面两列表 + 迁移文档补注 | issue #7 的第二诉求,作者认为比第一条更值得改 |
|
||||||
|
| T5 | 版本 1.1.0 + CHANGELOG + Wiki 同步 | 加父类是扩大不是破坏,故 minor 而非 major |
|
||||||
|
|
||||||
|
## 保真校验(适用)
|
||||||
|
|
||||||
|
触及 ARCHITECTURE §1.4 的移植蓝本(CHS `app/domain/errors.py` 与 `app/coordination/`)。本次**有意变更**的语义仅一条(`GovernanceBackendError` 的类型归属);`retry_after_s` 非可选语义、两个 reason 值域、fail-closed 方向、记账/闸门分工、`RedisPermit` 释放侧降级五条**不得被顺带改动**,每任务完成前逐条自查。
|
||||||
|
|
||||||
|
## 独立审查修正(2026-08-06, Codex)
|
||||||
|
|
||||||
|
4 条意见全部核实属实并已折回:
|
||||||
|
|
||||||
|
1. **T3 测试证据定位错误**(重要)——原写"复用 `test_backpressure.py:176-186` 的注入桩"覆盖三条泄漏路径,实测那三个桩是 `record_success`/`record_failure`/`mark_progress` 的**记账侧降级**,与闸门路径无关;`try_acquire`/`try_enter` 全无覆盖。已改为"三条桩都要新增"并写明现状。
|
||||||
|
2. **T3 漏了既有测试构造点**(重要)——新签名 keyword-only 必填,`test_backpressure.py:176/181/186`、`test_errors.py:89` 的裸 `GovernanceBackendError("...")` 与 `:255/:257` 的 `QuotaGate(_L())` 漏改即 `TypeError`。已补一张同批更新清单。
|
||||||
|
3. **`__all__` 插入位置写反**(次要)——按字母序应在 `SourceDeadError` **之后**而非之前。已改。
|
||||||
|
4. **设计中 telemetry 行号过时**(次要)——`:210` → `:250`,系本分支加 `_AttemptUsage` 造成的漂移。设计与摘要页已同步更新。
|
||||||
|
|
||||||
|
Codex 同时独立核实了计划的可执行性锚点: 22 处构造点、三处 gate 装配、后端层 `self._scope` 位置、README/ARCH 章节行号,均与 `src/` 现状相符。
|
||||||
|
|
||||||
|
## 独立验证炸出的阻塞缺陷(2026-08-06,全新上下文 verifier)
|
||||||
|
|
||||||
|
T1–T5 全绿、四道门禁全过之后,verifier 用一个**走 `QuotaGate` 的**端到端用例证明: 装配缺陷在唯一的生产路径上根本没拆出去——包装器的 `except GovernanceBackendError: raise` 只放行旧类型,`SourceNotConfiguredError` 落进下一行 `except Exception` 被重新包回去,配置写错照样永远重投。**盲区在于 T3 写的两条用例都直接打私有 `_cfg()`,比生产路径低一层。**
|
||||||
|
|
||||||
|
修复见正文 §T6(9 处放行 + 遥测终态捕获 + 走包装器的回归测试)。复核时 verifier 又指出一颗雷: 新放行让该异常能穿透 `_record_quietly`,而那层降级的存在理由是"调用已真实完成,写回失败不该丢弃成功响应"——同批把三处 `_record_quietly` 一并放宽并加了回归断言。
|
||||||
|
|
||||||
|
两轮都订正了同一处事实错误: 闸门泄漏路径是**五条**不是三条(`QuotaGate.stats` 与 `BreakerGate.retry_after_s` 同样未被 `_record_quietly` 包裹)。
|
||||||
|
|
||||||
|
相关: [[governance-backend-error]](design)、[[m2-distributed]]
|
||||||
@@ -0,0 +1,37 @@
|
|||||||
|
---
|
||||||
|
type: plan
|
||||||
|
node_id: plan:issue10-error-body-retention-plan
|
||||||
|
title: "实现计划: HTTP 错误响应体留存(Issue #10)"
|
||||||
|
date: 2026-08-16
|
||||||
|
---
|
||||||
|
|
||||||
|
# 实现计划: HTTP 错误响应体留存(Issue #10)
|
||||||
|
|
||||||
|
**全文**: `plans/2026-08-16-issue10-error-body-retention.md`|**实现**: [[design:issue10-error-body-retention]]|**分支**: `feat/issue-10-error-body-retention`
|
||||||
|
|
||||||
|
## 任务分解
|
||||||
|
|
||||||
|
| # | 任务 | 产出 |
|
||||||
|
|---|---|---|
|
||||||
|
| 1 | 内核字段 | `PolyGatewayError.body_text`(默认空串)+ 与 `raw_text` 的界限 docstring |
|
||||||
|
| 2 | 共享摘要单元 | 新建 `transports/_http_errors.py`:`summarize_body` / `compose_message` / `response_body` |
|
||||||
|
| 3 | openai_compat 收口 | `_status_to_error` 表驱动;429 判类型仍读原文 |
|
||||||
|
| 4 | monkey_ocr 收口 | `_classify_status` 同款,含 `ResponseNotRead` 降级 |
|
||||||
|
| 5 | **端到端验收** | 400 调用后 SQLite `error` 列含摘要——本计划的硬判据 |
|
||||||
|
| 6 | 文档与版本 | 1.2.0、README pin `>=1.2,<2`、CHANGELOG、ARCHITECTURE §6.2 |
|
||||||
|
| 7 | 合并前门 | lint + 全套件 + 全新上下文 verifier |
|
||||||
|
|
||||||
|
依赖:1‖2 → 3‖4 → 5 → 6 → 7。
|
||||||
|
|
||||||
|
## 计划期钉死的两条实现红线
|
||||||
|
|
||||||
|
1. **429 判 `insufficient_quota` 必须解析未截断原文**,不得改用摘要——摘要会破坏 JSON,超长体一旦改用摘要解析,配额耗尽的源将不再 `force_open`,把诊断改进变成治理 bug。Task 3.4 有专门回归用例。
|
||||||
|
2. **message 主体逐字保持 1.1.2 原样**,只在末尾追加摘要后缀;状态码→分类映射逐条不变,验收要求既有分类断言零改写。
|
||||||
|
|
||||||
|
## 保真校验
|
||||||
|
|
||||||
|
不涉及 `reference/` 迁移。错误分类映射不变,保真体现为"既有分类断言全部保留"。
|
||||||
|
|
||||||
|
## 执行期观察
|
||||||
|
|
||||||
|
Task 计划提交时 pre-commit 钩子的全套件跑出现一次 `tests/e2e/test_compat_projects.py::TestGovDocOnboarding::test_call_site_shape_runs_governed` 失败,单独跑与 e2e 全目录跑(7 passed / 23.30s,每例 2-3.5s)均通过,重跑全套件亦通过 → 判为真实网关抖动,非回归。该用例正是 [[design:issue8-stall-budget]] 当年记录的三个漂移用例之一,e2e 打真实网关的固有 flaky 面仍在。
|
||||||
@@ -0,0 +1,35 @@
|
|||||||
|
---
|
||||||
|
type: plan
|
||||||
|
node_id: plan:issue8-stall-budget-plan
|
||||||
|
title: "issue #8 实施计划: stall 非生产性等待口径"
|
||||||
|
date: 2026-08-06
|
||||||
|
---
|
||||||
|
|
||||||
|
# issue #8 实施计划: stall 非生产性等待口径
|
||||||
|
|
||||||
|
**全文**: `plans/2026-08-06-issue8-stall-budget.md` |**实现**: [[design:issue8-stall-budget]] |**分支**: `feat/issue-8-stall-budget`
|
||||||
|
|
||||||
|
## 交付
|
||||||
|
|
||||||
|
| 任务 | 内容 | 提交 |
|
||||||
|
|---|---|---|
|
||||||
|
| T1 | `StallClock` 落地 + chat 路径改造 + 8 条测试 | `02c3d06` |
|
||||||
|
| T2 | embedding 路径复用 | `6d0f3c9` |
|
||||||
|
| T3 | ocr 路径复用 | `0477d95` |
|
||||||
|
| T4 | `config.py` docstring 与 `.env.example` 注释对齐(无逻辑变更) | `d05114e` |
|
||||||
|
| T5 | 全套件回归 + `ARCHITECTURE.md` §7.3 与 `CHANGELOG` 同步 | `3645e57` |
|
||||||
|
| T6 | 独立验证(全新上下文 verifier)+ 三个问题的修复 | `bc4683d`、`a0a5cf7` |
|
||||||
|
|
||||||
|
## 测试证据(先失败后通过)
|
||||||
|
|
||||||
|
三条路径的失效链条各有一条回归用例,改前均转红于 `reason="stalled"`:`retry.py:218`、`embedding.py:250`、`ocr.py:275`。issue 只记录了 chat 路径,embedding/ocr 两条为本次核出。
|
||||||
|
|
||||||
|
## 独立验证发现的三个问题(均已修)
|
||||||
|
|
||||||
|
1. **429 缝隙(中)**: 初稿使 429 尝试两个预算都不烧,慢 429 场景实测挂 25.2 小时——**修复引入的回归**。见设计 §3.6。
|
||||||
|
2. **测试假证据(中)**: 并发用例用了两个 `RetryMW` 实例,实例级共享被对象隔离掩盖,clock 提升为实例属性时 7 条用例全部逃逸。改为复用同一 `mw` 并补"两次调用间空转超窗"用例,变异测试确认可抓。
|
||||||
|
3. **文档遗漏(轻)**: 计划要求的 `test_backpressure.py` docstring 订正漏做。
|
||||||
|
|
||||||
|
## 保真校验
|
||||||
|
|
||||||
|
治理主循环为 CHS `governance.py:200-285` 移植物,但 `reference/` 不在工作区,故以设计 §4 行为审计表 9 条 + 代码内 CHS 行号注释为基准。核对结果:标"保留"的 8 条在 `git diff` 中零出现,唯一"替换"项为条件 A 度量口径。
|
||||||
@@ -1,3 +1,3 @@
|
|||||||
# Query Pack
|
# Query Pack
|
||||||
|
|
||||||
> 尚无数据。运行 research-lit 或 idea-creator 后自动生成。
|
> 自动生成,请勿手动编辑。
|
||||||
|
|||||||
@@ -1,11 +1,11 @@
|
|||||||
---
|
---
|
||||||
type: schema
|
type: schema
|
||||||
node_id: schema:llm-calls
|
node_id: schema:llm-calls
|
||||||
title: "表结构: llm_calls(遥测 21 字段)"
|
title: "表结构: llm_calls(遥测 22 字段)"
|
||||||
date: 2026-07-20
|
date: 2026-07-20
|
||||||
---
|
---
|
||||||
|
|
||||||
# 表结构: llm_calls(遥测 21 字段)
|
# 表结构: llm_calls(遥测 22 字段)
|
||||||
|
|
||||||
|
|
||||||
## 列定义(冻结,M1 设计 §4.4 / ARCH §7.8)
|
## 列定义(冻结,M1 设计 §4.4 / ARCH §7.8)
|
||||||
@@ -27,6 +27,7 @@ date: 2026-07-20
|
|||||||
| cached_prompt_tokens | INTEGER | 供应商 prompt cache 命中的输入 token(2026-07-31,issue #3);NULL = 该源未上报,`0` = 上报了真实零命中,两者不可混同 |
|
| cached_prompt_tokens | INTEGER | 供应商 prompt cache 命中的输入 token(2026-07-31,issue #3);NULL = 该源未上报,`0` = 上报了真实零命中,两者不可混同 |
|
||||||
| model_reported | TEXT | API 响应体实际返回的 model;NULL = 未上报。与 `model`(配置别名)可能分叉 |
|
| model_reported | TEXT | API 响应体实际返回的 model;NULL = 未上报。与 `model`(配置别名)可能分叉 |
|
||||||
| sampling | TEXT | 本次调用的采样参数 canonical JSON(2026-07-31,issue #4);NULL = 未传。见下方口径 |
|
| sampling | TEXT | 本次调用的采样参数 canonical JSON(2026-07-31,issue #4);NULL = 未传。见下方口径 |
|
||||||
|
| reasoning_tokens | INTEGER | 推理消耗的输出 token(2026-08-02,issue #6);**含在 completion_tokens 内**,不影响成本总额,只补归因。NULL = **本次调用**未上报 |
|
||||||
|
|
||||||
## usage/成本口径(2026-07-30,est_tokens 解耦)
|
## usage/成本口径(2026-07-30,est_tokens 解耦)
|
||||||
|
|
||||||
@@ -53,6 +54,8 @@ FROM llm_calls WHERE cache_hit = false AND cached_prompt_tokens IS NOT NULL;
|
|||||||
|
|
||||||
## 采样参数口径(2026-07-31,issue #4)
|
## 采样参数口径(2026-07-31,issue #4)
|
||||||
|
|
||||||
|
`reasoning_tokens` 的 NULL 语义与 `cached_prompt_tokens` **不同**: 后者的 NULL 是"该源不报这个数",前者只能读作"**本次调用**未上报"——中转在上游不返回 usage 时会用本地 tokenizer 补算并整体替换 usage 对象,把 `completion_tokens_details` 一并吃掉(实测同一请求 10 轮呈 6:4 双峰)。故统计口径须为 `IS NULL OR = 0` 才算"未推理",写 `= 0` 的条件永远不成立——实测三家供应商在未推理时都是整个 details 缺失,无人上报字面 `0`。**不可用 `completion_tokens` 反推是否推理**: 两档的输出长度分布重叠(关闭档实测最高 46,开启档最低 13)。
|
||||||
|
|
||||||
`sampling` 列 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化输出注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。补列纪律与 issue #3 两列逐字相同(排在末尾、先探测再 ALTER、失败只逐行降级)。
|
`sampling` 列 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化输出注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。补列纪律与 issue #3 两列逐字相同(排在末尾、先探测再 ALTER、失败只逐行降级)。
|
||||||
|
|
||||||
三个 emit 入口的取值必须各自定死,否则同一列在不同行含义不同:
|
三个 emit 入口的取值必须各自定死,否则同一列在不同行含义不同:
|
||||||
|
|||||||
@@ -17,6 +17,7 @@ from polygateway.errors import (
|
|||||||
RequestRejectedError,
|
RequestRejectedError,
|
||||||
ResultInvalidError,
|
ResultInvalidError,
|
||||||
SourceDeadError,
|
SourceDeadError,
|
||||||
|
SourceNotConfiguredError,
|
||||||
TransientError,
|
TransientError,
|
||||||
)
|
)
|
||||||
from polygateway.ocr import OcrClient
|
from polygateway.ocr import OcrClient
|
||||||
@@ -31,7 +32,7 @@ from polygateway.types import (
|
|||||||
SourceConfig,
|
SourceConfig,
|
||||||
)
|
)
|
||||||
|
|
||||||
__version__ = "1.0.5"
|
__version__ = "1.2.0"
|
||||||
|
|
||||||
__all__ = [
|
__all__ = [
|
||||||
"DEFAULT_PROFILES",
|
"DEFAULT_PROFILES",
|
||||||
@@ -58,6 +59,7 @@ __all__ = [
|
|||||||
"ResultInvalidError",
|
"ResultInvalidError",
|
||||||
"SourceConfig",
|
"SourceConfig",
|
||||||
"SourceDeadError",
|
"SourceDeadError",
|
||||||
|
"SourceNotConfiguredError",
|
||||||
"TransientError",
|
"TransientError",
|
||||||
"__version__",
|
"__version__",
|
||||||
"gather_bounded",
|
"gather_bounded",
|
||||||
|
|||||||
@@ -15,7 +15,7 @@ import time
|
|||||||
import uuid
|
import uuid
|
||||||
from typing import TYPE_CHECKING
|
from typing import TYPE_CHECKING
|
||||||
|
|
||||||
from polygateway.errors import GovernanceBackendError
|
from polygateway.errors import SourceNotConfiguredError
|
||||||
from polygateway.types import GlobalLimits, SourceConfig, SourceStats
|
from polygateway.types import GlobalLimits, SourceConfig, SourceStats
|
||||||
|
|
||||||
if TYPE_CHECKING:
|
if TYPE_CHECKING:
|
||||||
@@ -89,7 +89,7 @@ class InMemoryLimiter:
|
|||||||
def _cfg(self, source_key: str) -> SourceConfig:
|
def _cfg(self, source_key: str) -> SourceConfig:
|
||||||
cfg = self._sources.get(source_key)
|
cfg = self._sources.get(source_key)
|
||||||
if cfg is None:
|
if cfg is None:
|
||||||
raise GovernanceBackendError(f"未知源 {source_key!r}(scope={self._scope})")
|
raise SourceNotConfiguredError(f"未知源 {source_key!r}(scope={self._scope})")
|
||||||
return cfg
|
return cfg
|
||||||
|
|
||||||
def _window(self) -> int:
|
def _window(self) -> int:
|
||||||
|
|||||||
@@ -367,7 +367,7 @@ class RedisGate:
|
|||||||
keys=[self._key(source_name)], args=[owner, self._probe_ttl_ms]
|
keys=[self._key(source_name)], args=[owner, self._probe_ttl_ms]
|
||||||
)
|
)
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端 try_enter 失败: {exc}") from exc
|
raise GovernanceBackendError(f"熔断后端 try_enter 失败: {exc}", scope=self._scope) from exc
|
||||||
return self._decision(source_name, result)
|
return self._decision(source_name, result)
|
||||||
|
|
||||||
async def record_success(
|
async def record_success(
|
||||||
@@ -385,7 +385,7 @@ class RedisGate:
|
|||||||
try:
|
try:
|
||||||
result = await self._success_lua(keys=[self._key(entry.source_name)], args=args)
|
result = await self._success_lua(keys=[self._key(entry.source_name)], args=args)
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端 record_success 失败: {exc}") from exc
|
raise GovernanceBackendError(f"熔断后端 record_success 失败: {exc}", scope=self._scope) from exc
|
||||||
return self._update(result)
|
return self._update(result)
|
||||||
|
|
||||||
async def record_failure(
|
async def record_failure(
|
||||||
@@ -407,7 +407,7 @@ class RedisGate:
|
|||||||
try:
|
try:
|
||||||
result = await self._failure_lua(keys=[self._key(entry.source_name)], args=args)
|
result = await self._failure_lua(keys=[self._key(entry.source_name)], args=args)
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端 record_failure 失败: {exc}") from exc
|
raise GovernanceBackendError(f"熔断后端 record_failure 失败: {exc}", scope=self._scope) from exc
|
||||||
return self._update(result)
|
return self._update(result)
|
||||||
|
|
||||||
async def release_probe(self, entry: GateDecision) -> GateUpdate:
|
async def release_probe(self, entry: GateDecision) -> GateUpdate:
|
||||||
@@ -419,7 +419,7 @@ class RedisGate:
|
|||||||
keys=[self._key(entry.source_name)], args=[entry.epoch, entry.probe_owner]
|
keys=[self._key(entry.source_name)], args=[entry.epoch, entry.probe_owner]
|
||||||
)
|
)
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端 release_probe 失败: {exc}") from exc
|
raise GovernanceBackendError(f"熔断后端 release_probe 失败: {exc}", scope=self._scope) from exc
|
||||||
return self._update(result)
|
return self._update(result)
|
||||||
|
|
||||||
async def retry_after_s(self, sources: tuple[str, ...]) -> float:
|
async def retry_after_s(self, sources: tuple[str, ...]) -> float:
|
||||||
@@ -429,7 +429,7 @@ class RedisGate:
|
|||||||
try:
|
try:
|
||||||
result = await self._retry_after_lua(keys=[self._key(s) for s in sources])
|
result = await self._retry_after_lua(keys=[self._key(s) for s in sources])
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端 retry_after_s 失败: {exc}") from exc
|
raise GovernanceBackendError(f"熔断后端 retry_after_s 失败: {exc}", scope=self._scope) from exc
|
||||||
return int(result) / 1000.0
|
return int(result) / 1000.0
|
||||||
|
|
||||||
async def aclose(self) -> None:
|
async def aclose(self) -> None:
|
||||||
|
|||||||
@@ -22,7 +22,7 @@ from typing import TYPE_CHECKING
|
|||||||
from loguru import logger
|
from loguru import logger
|
||||||
from redis.exceptions import RedisError
|
from redis.exceptions import RedisError
|
||||||
|
|
||||||
from polygateway.errors import GovernanceBackendError
|
from polygateway.errors import GovernanceBackendError, SourceNotConfiguredError
|
||||||
from polygateway.types import GlobalLimits, SourceConfig, SourceStats
|
from polygateway.types import GlobalLimits, SourceConfig, SourceStats
|
||||||
|
|
||||||
if TYPE_CHECKING:
|
if TYPE_CHECKING:
|
||||||
@@ -195,7 +195,7 @@ class RedisLimiter:
|
|||||||
def _cfg(self, source_key: str) -> SourceConfig:
|
def _cfg(self, source_key: str) -> SourceConfig:
|
||||||
cfg = self._sources.get(source_key)
|
cfg = self._sources.get(source_key)
|
||||||
if cfg is None:
|
if cfg is None:
|
||||||
raise GovernanceBackendError(f"未知源 {source_key!r}(scope={self._scope})")
|
raise SourceNotConfiguredError(f"未知源 {source_key!r}(scope={self._scope})")
|
||||||
return cfg
|
return cfg
|
||||||
|
|
||||||
def _lease_keys(self, source_key: str) -> tuple[str, str]:
|
def _lease_keys(self, source_key: str) -> tuple[str, str]:
|
||||||
@@ -247,7 +247,7 @@ class RedisLimiter:
|
|||||||
],
|
],
|
||||||
)
|
)
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"限流后端 try_acquire 失败: {exc}") from exc
|
raise GovernanceBackendError(f"限流后端 try_acquire 失败: {exc}", scope=self._scope) from exc
|
||||||
if ok != 1:
|
if ok != 1:
|
||||||
return None
|
return None
|
||||||
return _RedisPermit(self, source_key, lease_id, est_tokens, window)
|
return _RedisPermit(self, source_key, lease_id, est_tokens, window)
|
||||||
@@ -265,14 +265,14 @@ class RedisLimiter:
|
|||||||
try:
|
try:
|
||||||
await self._release_lua(keys=[gl, sl], args=[lease_id])
|
await self._release_lua(keys=[gl, sl], args=[lease_id])
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"限流后端 release 失败: {exc}") from exc
|
raise GovernanceBackendError(f"限流后端 release 失败: {exc}", scope=self._scope) from exc
|
||||||
|
|
||||||
async def _settle_tpm(self, source_key: str, delta: int, window: int) -> None:
|
async def _settle_tpm(self, source_key: str, delta: int, window: int) -> None:
|
||||||
wk = self._window_keys(source_key, window)
|
wk = self._window_keys(source_key, window)
|
||||||
try:
|
try:
|
||||||
await self._settle_lua(keys=[wk["g_tpm"], wk["s_tpm"]], args=[delta, _WINDOW_TTL_S])
|
await self._settle_lua(keys=[wk["g_tpm"], wk["s_tpm"]], args=[delta, _WINDOW_TTL_S])
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"限流后端 settle 失败: {exc}") from exc
|
raise GovernanceBackendError(f"限流后端 settle 失败: {exc}", scope=self._scope) from exc
|
||||||
|
|
||||||
async def source_stats(self, source_key: str) -> SourceStats:
|
async def source_stats(self, source_key: str) -> SourceStats:
|
||||||
"""当前窗口快照;读侧 clamp ≥0(展示口径,存储保留负值)。"""
|
"""当前窗口快照;读侧 clamp ≥0(展示口径,存储保留负值)。"""
|
||||||
@@ -283,7 +283,7 @@ class RedisLimiter:
|
|||||||
wk = self._window_keys(source_key, window)
|
wk = self._window_keys(source_key, window)
|
||||||
res = await self._stats_lua(keys=[sl, wk["s_rpm"], wk["s_tpm"]])
|
res = await self._stats_lua(keys=[sl, wk["s_rpm"], wk["s_tpm"]])
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"限流后端 source_stats 失败: {exc}") from exc
|
raise GovernanceBackendError(f"限流后端 source_stats 失败: {exc}", scope=self._scope) from exc
|
||||||
return SourceStats(
|
return SourceStats(
|
||||||
inflight=int(res[0]),
|
inflight=int(res[0]),
|
||||||
rpm_used=max(0, int(res[1])),
|
rpm_used=max(0, int(res[1])),
|
||||||
@@ -295,14 +295,14 @@ class RedisLimiter:
|
|||||||
try:
|
try:
|
||||||
await self._progress_mark_lua(keys=[self._progress_key()], args=[_PROGRESS_TTL_S])
|
await self._progress_mark_lua(keys=[self._progress_key()], args=[_PROGRESS_TTL_S])
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"限流后端 mark_progress 失败: {exc}") from exc
|
raise GovernanceBackendError(f"限流后端 mark_progress 失败: {exc}", scope=self._scope) from exc
|
||||||
|
|
||||||
async def progress_age_s(self) -> float:
|
async def progress_age_s(self) -> float:
|
||||||
"""距上次全局成功的秒数;仅键缺失(-1)= 从未进展 → inf(CHS limiter.py:208)。"""
|
"""距上次全局成功的秒数;仅键缺失(-1)= 从未进展 → inf(CHS limiter.py:208)。"""
|
||||||
try:
|
try:
|
||||||
res = await self._progress_age_lua(keys=[self._progress_key()])
|
res = await self._progress_age_lua(keys=[self._progress_key()])
|
||||||
except RedisError as exc:
|
except RedisError as exc:
|
||||||
raise GovernanceBackendError(f"限流后端 progress_age_s 失败: {exc}") from exc
|
raise GovernanceBackendError(f"限流后端 progress_age_s 失败: {exc}", scope=self._scope) from exc
|
||||||
return float("inf") if int(res) == -1 else int(res) / 1000.0
|
return float("inf") if int(res) == -1 else int(res) / 1000.0
|
||||||
|
|
||||||
async def aclose(self) -> None:
|
async def aclose(self) -> None:
|
||||||
|
|||||||
+46
-12
@@ -25,7 +25,7 @@ from polygateway.middleware.retry import RetryMW
|
|||||||
from polygateway.middleware.structured import StructuredMW
|
from polygateway.middleware.structured import StructuredMW
|
||||||
from polygateway.middleware.telemetry import TelemetryEmitter, TelemetryMW
|
from polygateway.middleware.telemetry import TelemetryEmitter, TelemetryMW
|
||||||
from polygateway.pricing import PricingTable
|
from polygateway.pricing import PricingTable
|
||||||
from polygateway.providers import get_provider
|
from polygateway.providers import get_capability, get_provider, resolve_thinking
|
||||||
from polygateway.sources import (
|
from polygateway.sources import (
|
||||||
AdaptivePacer,
|
AdaptivePacer,
|
||||||
HealthAwareSelector,
|
HealthAwareSelector,
|
||||||
@@ -51,7 +51,7 @@ if TYPE_CHECKING:
|
|||||||
TelemetryRecorder,
|
TelemetryRecorder,
|
||||||
Transport,
|
Transport,
|
||||||
)
|
)
|
||||||
from polygateway.providers import ProviderProfile
|
from polygateway.providers import ProviderProfile, ThinkingCapability
|
||||||
from polygateway.types import (
|
from polygateway.types import (
|
||||||
BackpressurePolicy,
|
BackpressurePolicy,
|
||||||
RetryPolicy,
|
RetryPolicy,
|
||||||
@@ -61,22 +61,52 @@ if TYPE_CHECKING:
|
|||||||
_T = TypeVar("_T")
|
_T = TypeVar("_T")
|
||||||
|
|
||||||
|
|
||||||
|
def _guard_thinking(
|
||||||
|
sources: list[SourceConfig],
|
||||||
|
profiles: list[ProviderProfile],
|
||||||
|
capabilities: Mapping[str, ThinkingCapability] | None,
|
||||||
|
) -> None:
|
||||||
|
"""装配期把不可满足的推理开关炸掉,而不是留到运行时(issue #5)。
|
||||||
|
|
||||||
|
与 transport 内的同一次判定不是重复: 那里兜的是"构造函数全量注入"这条路
|
||||||
|
(CLAUDE.md §4.5 的第二条装配路),而工厂路占 90% 场景,配置错误应当在装配期
|
||||||
|
就带着指路信息炸掉。`get_provider` 现在就是同一形态的双点调用。
|
||||||
|
"""
|
||||||
|
for source, profile in zip(sources, profiles, strict=True):
|
||||||
|
resolve_thinking(
|
||||||
|
profile,
|
||||||
|
get_capability(source.model, table=capabilities),
|
||||||
|
source.enable_thinking,
|
||||||
|
model=source.model,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _fingerprint_mark(source: SourceConfig) -> str:
|
||||||
|
"""单源的指纹标记;`enable_thinking` 仅在**表态时**追加。
|
||||||
|
|
||||||
|
只在表态时追加不是省事: 这样只配了 `extra_body` 的存量源字面量与 issue #4
|
||||||
|
时期逐字相同,升级本版本不会给它们平白来一次全量缓存冷启动。
|
||||||
|
"""
|
||||||
|
parts: list[Any] = [source.model, dict(source.extra_body)]
|
||||||
|
if source.enable_thinking is not None:
|
||||||
|
parts.append(source.enable_thinking)
|
||||||
|
return json.dumps(parts, sort_keys=True, ensure_ascii=False)
|
||||||
|
|
||||||
|
|
||||||
def build_model_fingerprint(sources: Iterable[SourceConfig]) -> str:
|
def build_model_fingerprint(sources: Iterable[SourceConfig]) -> str:
|
||||||
"""缓存 key 的模型身份: 多源 scope = 排序去重的 model 合集。
|
"""缓存 key 的模型身份: 多源 scope = 排序去重的 model 合集。
|
||||||
|
|
||||||
配置级采样参数(`extra_body`)必须参与,否则把 temperature 从 0 改成 1
|
配置级采样参数(`extra_body`)必须参与,否则把 temperature 从 0 改成 1
|
||||||
后重启仍会读到旧缓存(issue #4 设计决策 C)。全源 `extra_body` 皆空时
|
后重启仍会读到旧缓存(issue #4 设计决策 C)。`enable_thinking` 同理
|
||||||
字面量与历史实现逐字相同,不触发存量缓存冷启动。
|
(issue #5): 它一旦真正改变请求体,"关掉推理后重启"就会读到开着推理时
|
||||||
|
缓存的旧响应。全源两者皆未表态时字面量与历史实现逐字相同,不触发存量
|
||||||
|
缓存冷启动。
|
||||||
"""
|
"""
|
||||||
fingerprint = ",".join(sorted({s.model for s in sources}))
|
fingerprint = ",".join(sorted({s.model for s in sources}))
|
||||||
# 按 (model, extra_body) 而非源名摘要: 语义是"本 scope 会用哪些
|
# 按 (model, extra_body[, enable_thinking]) 而非源名摘要: 语义是"本 scope
|
||||||
# (模型, 解码参数)组合",改源名不该误触全量冷启动
|
# 会用哪些(模型, 请求形态)组合",改源名不该误触全量冷启动
|
||||||
marks = sorted(
|
marks = sorted(
|
||||||
{
|
{_fingerprint_mark(s) for s in sources if s.extra_body or s.enable_thinking is not None}
|
||||||
json.dumps([s.model, dict(s.extra_body)], sort_keys=True, ensure_ascii=False)
|
|
||||||
for s in sources
|
|
||||||
if s.extra_body
|
|
||||||
}
|
|
||||||
)
|
)
|
||||||
if marks:
|
if marks:
|
||||||
digest = hashlib.sha256("".join(marks).encode("utf-8")).hexdigest()
|
digest = hashlib.sha256("".join(marks).encode("utf-8")).hexdigest()
|
||||||
@@ -241,11 +271,13 @@ class GatewayClient:
|
|||||||
cache: CacheBackend | None = None,
|
cache: CacheBackend | None = None,
|
||||||
telemetry: TelemetryRecorder | None = None,
|
telemetry: TelemetryRecorder | None = None,
|
||||||
registry: Mapping[str, ProviderProfile] | None = None,
|
registry: Mapping[str, ProviderProfile] | None = None,
|
||||||
|
capabilities: Mapping[str, ThinkingCapability] | None = None,
|
||||||
rng: Any = random.random,
|
rng: Any = random.random,
|
||||||
) -> GatewayClient:
|
) -> GatewayClient:
|
||||||
"""按配置装配;显式传入的后端实例即共享(None 项按配置自建私有实例)。"""
|
"""按配置装配;显式传入的后端实例即共享(None 项按配置自建私有实例)。"""
|
||||||
sources = list(settings.sources)
|
sources = list(settings.sources)
|
||||||
profiles = [get_provider(s.provider, registry=registry) for s in sources]
|
profiles = [get_provider(s.provider, registry=registry) for s in sources]
|
||||||
|
_guard_thinking(sources, profiles, capabilities)
|
||||||
strategy, escalation = _build_structured(profiles)
|
strategy, escalation = _build_structured(profiles)
|
||||||
return cls(
|
return cls(
|
||||||
scope=settings.scope,
|
scope=settings.scope,
|
||||||
@@ -253,7 +285,7 @@ class GatewayClient:
|
|||||||
selector=_build_selector(settings.selector, rng=rng),
|
selector=_build_selector(settings.selector, rng=rng),
|
||||||
limiter=limiter or _build_limiter(settings, sources),
|
limiter=limiter or _build_limiter(settings, sources),
|
||||||
breaker=breaker or _build_breaker(settings),
|
breaker=breaker or _build_breaker(settings),
|
||||||
transport=OpenAICompatTransport(registry=registry),
|
transport=OpenAICompatTransport(registry=registry, capabilities=capabilities),
|
||||||
retry=settings.retry,
|
retry=settings.retry,
|
||||||
backpressure=settings.backpressure,
|
backpressure=settings.backpressure,
|
||||||
quota_full=settings.quota_full,
|
quota_full=settings.quota_full,
|
||||||
@@ -279,6 +311,7 @@ class GatewayClient:
|
|||||||
cache: CacheBackend | None = None,
|
cache: CacheBackend | None = None,
|
||||||
telemetry: TelemetryRecorder | None = None,
|
telemetry: TelemetryRecorder | None = None,
|
||||||
registry: Mapping[str, ProviderProfile] | None = None,
|
registry: Mapping[str, ProviderProfile] | None = None,
|
||||||
|
capabilities: Mapping[str, ThinkingCapability] | None = None,
|
||||||
env: Mapping[str, str] | None = None,
|
env: Mapping[str, str] | None = None,
|
||||||
) -> GatewayClient:
|
) -> GatewayClient:
|
||||||
"""从 .env/环境变量装配一个 scope 的 client(键名清单见 .env.example)。"""
|
"""从 .env/环境变量装配一个 scope 的 client(键名清单见 .env.example)。"""
|
||||||
@@ -289,6 +322,7 @@ class GatewayClient:
|
|||||||
cache=cache,
|
cache=cache,
|
||||||
telemetry=telemetry,
|
telemetry=telemetry,
|
||||||
registry=registry,
|
registry=registry,
|
||||||
|
capabilities=capabilities,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -149,9 +149,11 @@ class GatewaySettings:
|
|||||||
def _normalize(self) -> None:
|
def _normalize(self) -> None:
|
||||||
"""把 `from_env` 一直在做的规范化补到构造路上,两条路必须产出同一个值。
|
"""把 `from_env` 一直在做的规范化补到构造路上,两条路必须产出同一个值。
|
||||||
|
|
||||||
`scope` 最要紧: 它直接进 Redis key(`pgw:limit:{scope}:…`/`pgw:gate:{scope}:…`)。
|
`scope` 的 strip 才是要紧的那一半: 它进 Redis key(`pgw:limit:{scope}:…`
|
||||||
一个进程走 `from_env("LLM")` 拿到 "llm"、另一个直接构造传 "LLM",同一逻辑
|
/`pgw:gate:{scope}:…`),而两个 Redis 后端在构造函数里只 `.lower()` **不 strip**
|
||||||
scope 的限流与熔断状态会分裂到两套命名空间,各记各的,治理静默失效且不报错。
|
——`"llm "` 会产出 `pgw:limit:llm :…`,与 `from_env` 路的进程分裂成两套命名空间。
|
||||||
|
大小写则不会: 后端自 v1.0.0 起各自 lower,`from_settings` 传 "LLM" 也落在同一
|
||||||
|
套 key 上(此处 lower 只为让 `GatewaySettings.scope` 属性两路取值一致)。
|
||||||
|
|
||||||
空串归 None 同理: 留着空串会骗过 `is None` 判断,把错误推迟到 redis 客户端
|
空串归 None 同理: 留着空串会骗过 `is None` 判断,把错误推迟到 redis 客户端
|
||||||
抛连接串解析异常。`telemetry_pg_dsn` 的驱动后缀因为要看 backend 且需告警,
|
抛连接串解析异常。`telemetry_pg_dsn` 的驱动后缀因为要看 backend 且需告警,
|
||||||
@@ -236,7 +238,14 @@ class GatewaySettings:
|
|||||||
)
|
)
|
||||||
|
|
||||||
def _validate_stall(self) -> None:
|
def _validate_stall(self) -> None:
|
||||||
"""stall 窗口须 ≥ 最慢源 TTFT 上限,防把正常慢首包误判为卡死。"""
|
"""stall 窗口须 ≥ 最慢源 TTFT 上限(保守冗余,见下)。
|
||||||
|
|
||||||
|
原理由是"防把正常慢首包误判为卡死"。issue #8 起 stall 只累计**非
|
||||||
|
生产性等待**(429 退避、配额轮询、熔断冷却),TTFT 等待属生产性时间、
|
||||||
|
已不计入 stall 账,该误判在机制上不再可能。校验本身无害且不会误拒
|
||||||
|
任何合理配置,故保留——删除它需同步改动 ARCHITECTURE.md §7.3 的契约
|
||||||
|
补强 G6,超出 issue #8 的范围(2026-08-06 人类定夺)。
|
||||||
|
"""
|
||||||
ttfts = [s.ttft_timeout_s for s in self.sources if s.ttft_timeout_s is not None]
|
ttfts = [s.ttft_timeout_s for s in self.sources if s.ttft_timeout_s is not None]
|
||||||
if ttfts and self.backpressure.stall_window_s < max(ttfts):
|
if ttfts and self.backpressure.stall_window_s < max(ttfts):
|
||||||
raise ValueError(
|
raise ValueError(
|
||||||
|
|||||||
@@ -34,11 +34,12 @@ from polygateway.errors import (
|
|||||||
RequestRejectedError,
|
RequestRejectedError,
|
||||||
ResultInvalidError,
|
ResultInvalidError,
|
||||||
SourceDeadError,
|
SourceDeadError,
|
||||||
|
SourceNotConfiguredError,
|
||||||
TransientError,
|
TransientError,
|
||||||
)
|
)
|
||||||
from polygateway.middleware.breaker import BreakerGate
|
from polygateway.middleware.breaker import BreakerGate
|
||||||
from polygateway.middleware.ratelimit import QuotaGate
|
from polygateway.middleware.ratelimit import QuotaGate
|
||||||
from polygateway.middleware.retry import _failure_reason, backoff_delay
|
from polygateway.middleware.retry import StallClock, _failure_reason, backoff_delay
|
||||||
from polygateway.middleware.telemetry import TelemetryEmitter
|
from polygateway.middleware.telemetry import TelemetryEmitter
|
||||||
from polygateway.sources import SourceCooldownMemo
|
from polygateway.sources import SourceCooldownMemo
|
||||||
from polygateway.types import (
|
from polygateway.types import (
|
||||||
@@ -120,8 +121,8 @@ class EmbeddingClient:
|
|||||||
# 否则遥测会记录一个从未发出的采样参数(issue #4 决策 G)
|
# 否则遥测会记录一个从未发出的采样参数(issue #4 决策 G)
|
||||||
self._sources = strip_unsupported_extra_body(list(sources), path="embedding")
|
self._sources = strip_unsupported_extra_body(list(sources), path="embedding")
|
||||||
self._selector = selector
|
self._selector = selector
|
||||||
self._quota = QuotaGate(limiter)
|
self._quota = QuotaGate(limiter, scope=self._scope)
|
||||||
self._breaker = BreakerGate(breaker)
|
self._breaker = BreakerGate(breaker, scope=self._scope)
|
||||||
self._transport = transport
|
self._transport = transport
|
||||||
self._retry = retry
|
self._retry = retry
|
||||||
self._bp = backpressure
|
self._bp = backpressure
|
||||||
@@ -178,12 +179,14 @@ class EmbeddingClient:
|
|||||||
) -> _BatchOutcome:
|
) -> _BatchOutcome:
|
||||||
fails = 0
|
fails = 0
|
||||||
reasons: dict[str, str] = {}
|
reasons: dict[str, str] = {}
|
||||||
entered_at = self._now()
|
# 只计非生产性等待(issue #8): 真实尝试由重试预算治理,不重复烧 stall 预算
|
||||||
|
clock = StallClock(self._now)
|
||||||
while True:
|
while True:
|
||||||
picked, gate_rejections = await self._pick_runnable(reasons)
|
picked, gate_rejections = await self._pick_runnable(reasons)
|
||||||
if picked is None:
|
if picked is None:
|
||||||
await self._on_no_runnable(gate_rejections, reasons, entered_at)
|
await self._on_no_runnable(gate_rejections, reasons, clock)
|
||||||
continue
|
continue
|
||||||
|
async with clock.attempting():
|
||||||
outcome = await self._attempt(batch, *picked, reasons, session_id, parent_call_id)
|
outcome = await self._attempt(batch, *picked, reasons, session_id, parent_call_id)
|
||||||
if isinstance(outcome, _BatchOutcome):
|
if isinstance(outcome, _BatchOutcome):
|
||||||
return outcome
|
return outcome
|
||||||
@@ -227,7 +230,7 @@ class EmbeddingClient:
|
|||||||
return None, gate_rejections
|
return None, gate_rejections
|
||||||
|
|
||||||
async def _on_no_runnable(
|
async def _on_no_runnable(
|
||||||
self, gate_rejections: int, reasons: dict[str, str], entered_at: float
|
self, gate_rejections: int, reasons: dict[str, str], clock: StallClock
|
||||||
) -> None:
|
) -> None:
|
||||||
if gate_rejections == len(self._sources):
|
if gate_rejections == len(self._sources):
|
||||||
names = tuple(s.name for s in self._sources)
|
names = tuple(s.name for s in self._sources)
|
||||||
@@ -244,7 +247,7 @@ class EmbeddingClient:
|
|||||||
per_source_reasons=reasons,
|
per_source_reasons=reasons,
|
||||||
)
|
)
|
||||||
stall = self._bp.stall_window_s
|
stall = self._bp.stall_window_s
|
||||||
if self._now() - entered_at > stall and await self._quota.progress_age_s() > stall:
|
if clock.stalled_s() > stall and await self._quota.progress_age_s() > stall:
|
||||||
names = tuple(s.name for s in self._sources)
|
names = tuple(s.name for s in self._sources)
|
||||||
raise AllSourcesExhausted(
|
raise AllSourcesExhausted(
|
||||||
scope=self._scope,
|
scope=self._scope,
|
||||||
@@ -326,7 +329,7 @@ class EmbeddingClient:
|
|||||||
await write_back
|
await write_back
|
||||||
except asyncio.CancelledError:
|
except asyncio.CancelledError:
|
||||||
raise
|
raise
|
||||||
except GovernanceBackendError as exc:
|
except (GovernanceBackendError, SourceNotConfiguredError) as exc:
|
||||||
logger.warning("embedding 治理记账写回降级(不冒泡): {}", exc)
|
logger.warning("embedding 治理记账写回降级(不冒泡): {}", exc)
|
||||||
|
|
||||||
async def _settle_and_release(self, permit: Permit, actual: int) -> None:
|
async def _settle_and_release(self, permit: Permit, actual: int) -> None:
|
||||||
|
|||||||
@@ -5,8 +5,22 @@
|
|||||||
"""
|
"""
|
||||||
|
|
||||||
SCOPE_REASONS = frozenset(
|
SCOPE_REASONS = frozenset(
|
||||||
{"circuit_open", "retry_exhausted", "stalled", "quota_exhausted", "no_sources"}
|
{
|
||||||
|
"circuit_open",
|
||||||
|
"retry_exhausted",
|
||||||
|
"stalled",
|
||||||
|
"quota_exhausted",
|
||||||
|
"no_sources",
|
||||||
|
"governance_backend_down", # issue #7: 限流/熔断后端故障(fail-closed → 整个 scope 发不出请求)
|
||||||
|
}
|
||||||
)
|
)
|
||||||
|
GOVERNANCE_BACKEND_RETRY_AFTER_S = 5.0
|
||||||
|
"""治理后端故障的建议重投间隔(秒)。
|
||||||
|
|
||||||
|
**不是环境配置项**——后端恢复时间物理上不可知(不同于熔断冷却有确定到期时刻),
|
||||||
|
故取一个保守固定值;下游有自己的退避策略时可忽略本字段。取 0 会让积压任务零延迟
|
||||||
|
同时冲击已挂掉的后端,把一次故障放大成一场风暴(issue #7 §3.2)。
|
||||||
|
"""
|
||||||
SOURCE_REASONS = frozenset(
|
SOURCE_REASONS = frozenset(
|
||||||
{
|
{
|
||||||
"network_error",
|
"network_error",
|
||||||
@@ -21,7 +35,26 @@ SOURCE_REASONS = frozenset(
|
|||||||
|
|
||||||
|
|
||||||
class PolyGatewayError(Exception):
|
class PolyGatewayError(Exception):
|
||||||
"""库内一切领域错误的基类,携带来源上下文便于遥测与日志定位。"""
|
"""库内一切领域错误的基类,携带来源上下文便于遥测与日志定位。
|
||||||
|
|
||||||
|
`body_text` 是**非 2xx 响应体的摘要**——网关拒绝这次调用时说的话(issue #10)。
|
||||||
|
它与 `ResultInvalidError.raw_text` 是两回事,严禁混用:
|
||||||
|
|
||||||
|
============== ==================================================
|
||||||
|
``body_text`` **非 2xx** 的 HTTP 错误响应体: 对方**拒绝**的理由
|
||||||
|
``raw_text`` **2xx** 但内容不可解析时的模型输出原文
|
||||||
|
============== ==================================================
|
||||||
|
|
||||||
|
加在基类而非某个子类,是因为这些错误全部由同一个 HTTP 响应翻译而来——
|
||||||
|
"对方说了什么"与"它属于哪一类"正交。scope 级错误(`GatewayUnavailableError`
|
||||||
|
一族)继承到的恒空值不是噪音,而是"没有单一响应体可言"的如实表达。
|
||||||
|
|
||||||
|
**内容已由 transport 层截断**(`transports/_http_errors.summarize_body`),
|
||||||
|
且可能包含网关对请求的回显——库不做脱敏: 它不知道下游哪些字段敏感,
|
||||||
|
猜测式脱敏只会同时丢掉诊断价值与安全性。
|
||||||
|
|
||||||
|
本字段是**旁路数据**,不参与任何治理判定(重试/换源/熔断计数/限流结算)。
|
||||||
|
"""
|
||||||
|
|
||||||
def __init__(
|
def __init__(
|
||||||
self,
|
self,
|
||||||
@@ -30,11 +63,13 @@ class PolyGatewayError(Exception):
|
|||||||
source_name: str | None = None,
|
source_name: str | None = None,
|
||||||
status_code: int | None = None,
|
status_code: int | None = None,
|
||||||
operation: str | None = None,
|
operation: str | None = None,
|
||||||
|
body_text: str = "",
|
||||||
) -> None:
|
) -> None:
|
||||||
super().__init__(message)
|
super().__init__(message)
|
||||||
self.source_name = source_name
|
self.source_name = source_name
|
||||||
self.status_code = status_code
|
self.status_code = status_code
|
||||||
self.operation = operation
|
self.operation = operation
|
||||||
|
self.body_text = body_text
|
||||||
|
|
||||||
|
|
||||||
class TransientError(PolyGatewayError):
|
class TransientError(PolyGatewayError):
|
||||||
@@ -50,7 +85,16 @@ class SourceDeadError(PolyGatewayError):
|
|||||||
|
|
||||||
|
|
||||||
class RequestRejectedError(PolyGatewayError):
|
class RequestRejectedError(PolyGatewayError):
|
||||||
"""请求被拒(400/坏输入): 不重试不换源,直接上抛。"""
|
"""请求被拒(400/坏输入): 不重试不换源,直接上抛。
|
||||||
|
|
||||||
|
**经中转部署时请注意**(issue #10 下游实测): 第三方 API 中转服务自身抖动
|
||||||
|
时也会回 400,从状态码上与供应商说"你的输入有问题"无法区分。下游曾观测到
|
||||||
|
同一份字节(sha256 一致)重发 15 次全部成功,且失败那次 `prompt_tokens=0`、
|
||||||
|
耗时远低于任何成功调用——请求在推理开始前就被挡了。本库仍按确定性失败处理
|
||||||
|
(对直连供应商而言重试只会白烧配额),批处理场景的下游宜自备兜底分类;
|
||||||
|
`body_text` 即为此提供判据: 中转抖动的响应体与供应商的 `invalid_request_error`
|
||||||
|
形态不同。
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
class ResultInvalidError(PolyGatewayError):
|
class ResultInvalidError(PolyGatewayError):
|
||||||
@@ -127,5 +171,38 @@ class AllSourcesExhausted(GatewayUnavailableError): # noqa: N818 — ARCH §6.1
|
|||||||
"""重试预算耗尽 / 无可用源 / 配额 fail-fast 等 scope 级失败。"""
|
"""重试预算耗尽 / 无可用源 / 配额 fail-fast 等 scope 级失败。"""
|
||||||
|
|
||||||
|
|
||||||
class GovernanceBackendError(PolyGatewayError):
|
class SourceNotConfiguredError(PolyGatewayError):
|
||||||
"""限流/熔断状态后端自身故障: 必须报错而非放行(防击穿网关,降级方向铁律)。"""
|
"""源名不在限流后端的配置字典中: 装配缺陷,正常不可达。
|
||||||
|
|
||||||
|
**有意不在** `GatewayUnavailableError` 之下: 它不是"暂时不可用"而是"配置写
|
||||||
|
错了",必须消耗失败预算进死信让人看见;归入可重投家族会让配置错误的任务永远
|
||||||
|
重投、永不告警——正是 issue #7 要修的那个 bug 的镜像(§3.4)。
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
class GovernanceBackendError(GatewayUnavailableError):
|
||||||
|
"""限流/熔断状态后端自身故障: 必须报错而非放行(防击穿网关,降级方向铁律)。
|
||||||
|
|
||||||
|
继承 `GatewayUnavailableError`(issue #7): fail-closed 意味着整个 scope 一个
|
||||||
|
请求都发不出去,语义上即 scope 级不可用。此前它是 `PolyGatewayError` 的直接
|
||||||
|
子类,只写 `except GatewayUnavailableError` 的调用方接不住,后果是"Redis 抖
|
||||||
|
一下 → 积压任务消耗业务失败预算 → 进死信",而那是运维重启即可恢复的故障。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
message: str,
|
||||||
|
*,
|
||||||
|
scope: str,
|
||||||
|
retry_after_s: float = GOVERNANCE_BACKEND_RETRY_AFTER_S,
|
||||||
|
source_name: str | None = None,
|
||||||
|
) -> None:
|
||||||
|
super().__init__(
|
||||||
|
scope=scope,
|
||||||
|
reason="governance_backend_down",
|
||||||
|
retry_after_s=retry_after_s,
|
||||||
|
source_name=source_name,
|
||||||
|
)
|
||||||
|
# 父类会把 message 覆写为 "{scope} 网关暂时不可用: {reason}",而各构造点
|
||||||
|
# 携带的诊断串(如"限流后端 try_acquire 失败: ...")是排障主线索,必须保住
|
||||||
|
self.args = (message,)
|
||||||
|
|||||||
@@ -4,7 +4,7 @@ from __future__ import annotations
|
|||||||
|
|
||||||
from typing import TYPE_CHECKING
|
from typing import TYPE_CHECKING
|
||||||
|
|
||||||
from polygateway.errors import GovernanceBackendError
|
from polygateway.errors import GovernanceBackendError, SourceNotConfiguredError
|
||||||
|
|
||||||
if TYPE_CHECKING:
|
if TYPE_CHECKING:
|
||||||
from polygateway.ports import GateDecision, GateUpdate, ProviderGate
|
from polygateway.ports import GateDecision, GateUpdate, ProviderGate
|
||||||
@@ -14,49 +14,61 @@ if TYPE_CHECKING:
|
|||||||
class BreakerGate:
|
class BreakerGate:
|
||||||
"""RetryMW 面向熔断后端的唯一入口;包装一切后端异常。"""
|
"""RetryMW 面向熔断后端的唯一入口;包装一切后端异常。"""
|
||||||
|
|
||||||
def __init__(self, gate: ProviderGate) -> None:
|
def __init__(self, gate: ProviderGate, *, scope: str) -> None:
|
||||||
self._gate = gate
|
self._gate = gate
|
||||||
|
# 后端故障即 scope 级不可用,异常须携 scope 供调用方定位(issue #7 §3.3)
|
||||||
|
self._scope = scope
|
||||||
|
|
||||||
async def try_enter(self, source: SourceConfig, owner: str) -> GateDecision:
|
async def try_enter(self, source: SourceConfig, owner: str) -> GateDecision:
|
||||||
try:
|
try:
|
||||||
return await self._gate.try_enter(source.name, owner)
|
return await self._gate.try_enter(source.name, owner)
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端故障(try_enter): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"熔断后端故障(try_enter): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|
||||||
async def record_success(
|
async def record_success(
|
||||||
self, entry: GateDecision, *, count_attempt: bool = True
|
self, entry: GateDecision, *, count_attempt: bool = True
|
||||||
) -> GateUpdate:
|
) -> GateUpdate:
|
||||||
try:
|
try:
|
||||||
return await self._gate.record_success(entry, count_attempt=count_attempt)
|
return await self._gate.record_success(entry, count_attempt=count_attempt)
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端故障(record_success): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"熔断后端故障(record_success): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|
||||||
async def record_failure(
|
async def record_failure(
|
||||||
self, entry: GateDecision, reason: str, force_open: bool
|
self, entry: GateDecision, reason: str, force_open: bool
|
||||||
) -> GateUpdate:
|
) -> GateUpdate:
|
||||||
try:
|
try:
|
||||||
return await self._gate.record_failure(entry, reason, force_open)
|
return await self._gate.record_failure(entry, reason, force_open)
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端故障(record_failure): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"熔断后端故障(record_failure): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|
||||||
async def release_probe(self, entry: GateDecision) -> GateUpdate:
|
async def release_probe(self, entry: GateDecision) -> GateUpdate:
|
||||||
try:
|
try:
|
||||||
return await self._gate.release_probe(entry)
|
return await self._gate.release_probe(entry)
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端故障(release_probe): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"熔断后端故障(release_probe): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|
||||||
async def retry_after_s(self, sources: tuple[str, ...]) -> float:
|
async def retry_after_s(self, sources: tuple[str, ...]) -> float:
|
||||||
try:
|
try:
|
||||||
return await self._gate.retry_after_s(sources)
|
return await self._gate.retry_after_s(sources)
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"熔断后端故障(retry_after_s): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"熔断后端故障(retry_after_s): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|||||||
@@ -8,7 +8,7 @@ from __future__ import annotations
|
|||||||
|
|
||||||
from typing import TYPE_CHECKING
|
from typing import TYPE_CHECKING
|
||||||
|
|
||||||
from polygateway.errors import GovernanceBackendError
|
from polygateway.errors import GovernanceBackendError, SourceNotConfiguredError
|
||||||
|
|
||||||
if TYPE_CHECKING:
|
if TYPE_CHECKING:
|
||||||
from polygateway.ports import Permit, RateLimiter
|
from polygateway.ports import Permit, RateLimiter
|
||||||
@@ -18,37 +18,47 @@ if TYPE_CHECKING:
|
|||||||
class QuotaGate:
|
class QuotaGate:
|
||||||
"""RetryMW 面向限流后端的唯一入口;包装一切后端异常。"""
|
"""RetryMW 面向限流后端的唯一入口;包装一切后端异常。"""
|
||||||
|
|
||||||
def __init__(self, limiter: RateLimiter) -> None:
|
def __init__(self, limiter: RateLimiter, *, scope: str) -> None:
|
||||||
self._limiter = limiter
|
self._limiter = limiter
|
||||||
|
# 后端故障即 scope 级不可用,异常须携 scope 供调用方定位(issue #7 §3.3)
|
||||||
|
self._scope = scope
|
||||||
|
|
||||||
async def try_acquire(self, source: SourceConfig) -> Permit | None:
|
async def try_acquire(self, source: SourceConfig) -> Permit | None:
|
||||||
try:
|
try:
|
||||||
return await self._limiter.try_acquire(source.name, source.effective_est_tokens())
|
return await self._limiter.try_acquire(source.name, source.effective_est_tokens())
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"限流后端故障(try_acquire): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"限流后端故障(try_acquire): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|
||||||
async def stats(self, source: SourceConfig) -> SourceStats:
|
async def stats(self, source: SourceConfig) -> SourceStats:
|
||||||
try:
|
try:
|
||||||
return await self._limiter.source_stats(source.name)
|
return await self._limiter.source_stats(source.name)
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"限流后端故障(source_stats): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"限流后端故障(source_stats): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|
||||||
async def mark_progress(self) -> None:
|
async def mark_progress(self) -> None:
|
||||||
try:
|
try:
|
||||||
await self._limiter.mark_progress()
|
await self._limiter.mark_progress()
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"限流后端故障(mark_progress): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"限流后端故障(mark_progress): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|
||||||
async def progress_age_s(self) -> float:
|
async def progress_age_s(self) -> float:
|
||||||
try:
|
try:
|
||||||
return await self._limiter.progress_age_s()
|
return await self._limiter.progress_age_s()
|
||||||
except GovernanceBackendError:
|
except (GovernanceBackendError, SourceNotConfiguredError):
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
raise GovernanceBackendError(f"限流后端故障(progress_age_s): {exc}") from exc
|
raise GovernanceBackendError(
|
||||||
|
f"限流后端故障(progress_age_s): {exc}", scope=self._scope
|
||||||
|
) from exc
|
||||||
|
|||||||
@@ -11,6 +11,7 @@ httpx 是库的核心依赖而非实现层内部件,不违反"middleware 只依
|
|||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
|
|
||||||
import asyncio
|
import asyncio
|
||||||
|
import contextlib
|
||||||
import random
|
import random
|
||||||
import time
|
import time
|
||||||
import uuid
|
import uuid
|
||||||
@@ -28,6 +29,7 @@ from polygateway.errors import (
|
|||||||
RequestRejectedError,
|
RequestRejectedError,
|
||||||
ResultInvalidError,
|
ResultInvalidError,
|
||||||
SourceDeadError,
|
SourceDeadError,
|
||||||
|
SourceNotConfiguredError,
|
||||||
TransientError,
|
TransientError,
|
||||||
)
|
)
|
||||||
from polygateway.middleware.breaker import BreakerGate
|
from polygateway.middleware.breaker import BreakerGate
|
||||||
@@ -38,7 +40,7 @@ from polygateway.streaming import StreamLivenessTimeout
|
|||||||
from polygateway.types import LLMResponse
|
from polygateway.types import LLMResponse
|
||||||
|
|
||||||
if TYPE_CHECKING:
|
if TYPE_CHECKING:
|
||||||
from collections.abc import Awaitable, Callable
|
from collections.abc import AsyncIterator, Awaitable, Callable
|
||||||
|
|
||||||
from polygateway.ports import (
|
from polygateway.ports import (
|
||||||
GateDecision,
|
GateDecision,
|
||||||
@@ -73,6 +75,65 @@ def backoff_delay(
|
|||||||
return max(delay, retry_after)
|
return max(delay, retry_after)
|
||||||
|
|
||||||
|
|
||||||
|
class _Attempt:
|
||||||
|
"""一次尝试的计时句柄;`refund()` 把它退还给 stall 账(见 `StallClock`)。"""
|
||||||
|
|
||||||
|
__slots__ = ("productive",)
|
||||||
|
|
||||||
|
def __init__(self) -> None:
|
||||||
|
self.productive = True
|
||||||
|
|
||||||
|
def refund(self) -> None:
|
||||||
|
"""该次尝试不消耗重试预算(429),故其耗时归 stall 治理而非重试治理。"""
|
||||||
|
self.productive = False
|
||||||
|
|
||||||
|
|
||||||
|
class StallClock:
|
||||||
|
"""调用级 stall 计时器: 只累计非生产性等待(issue #8 设计 §3.1)。
|
||||||
|
|
||||||
|
**划分依据是"谁消耗重试预算"**,不是"是否发出了请求"。消耗 `max_attempts`
|
||||||
|
的时间已被重试预算治理,从 stall 账扣除;不消耗它的时间无人治理,归 stall。
|
||||||
|
两者重叠计费正是 issue #8 的根因: stall 预算(默认 300s)小于重试预算
|
||||||
|
(3 × timeout_s),必然先耗尽,于是重试预算在超时场景下永远用不上。
|
||||||
|
|
||||||
|
"生产性"的边界即 `_attempt` 的边界,含该次尝试的记账与遥测收尾——它们是
|
||||||
|
"尝试已有结论"之后的动作,不是在等待重试机会;把它们计入 stall 会让遥测
|
||||||
|
抖动参与判死。
|
||||||
|
|
||||||
|
**例外: 429 尝试须 `refund()`**。429 免重试预算(饱和期等待而非死亡),若其
|
||||||
|
耗时又算生产性,就掉进两个预算的缝隙——排队型网关持满 timeout 才回 429 时,
|
||||||
|
每轮只有退避那一两秒进 stall 账,调用可挂满 `stall_window/backoff_base` 轮
|
||||||
|
(实测 timeout=300/base=2 时达 25 小时)。退还后缝隙闭合。
|
||||||
|
|
||||||
|
每次调用创建一个实例。严禁提升为实例属性: `_entered_at` 会固定在进程启动
|
||||||
|
时刻,使 `stalled_s()` 随进程运行时长单调增长,最终所有调用被误判 stalled。
|
||||||
|
模块级共享单元, EmbeddingClient 与 OcrClient 复用(同 `backoff_delay`)。
|
||||||
|
"""
|
||||||
|
|
||||||
|
__slots__ = ("_now", "_entered_at", "_productive_s")
|
||||||
|
|
||||||
|
def __init__(self, now: Callable[[], float]) -> None:
|
||||||
|
self._now = now
|
||||||
|
self._entered_at = now()
|
||||||
|
self._productive_s = 0.0
|
||||||
|
|
||||||
|
def stalled_s(self) -> float:
|
||||||
|
"""非生产性等待累计秒数 = 调用总耗时 - 消耗重试预算的时间。"""
|
||||||
|
return self._now() - self._entered_at - self._productive_s
|
||||||
|
|
||||||
|
@contextlib.asynccontextmanager
|
||||||
|
async def attempting(self) -> AsyncIterator[_Attempt]:
|
||||||
|
"""包裹一次真实尝试,其耗时默认记为生产性(除非被 `refund()`)。"""
|
||||||
|
handle = _Attempt()
|
||||||
|
started = self._now()
|
||||||
|
try:
|
||||||
|
yield handle
|
||||||
|
finally:
|
||||||
|
# 只做算术与取值, 不吞任何异常——CancelledError 逐字穿透(库铁律)
|
||||||
|
if handle.productive:
|
||||||
|
self._productive_s += self._now() - started
|
||||||
|
|
||||||
|
|
||||||
def _demote_call_failures(
|
def _demote_call_failures(
|
||||||
ordered: list[SourceConfig],
|
ordered: list[SourceConfig],
|
||||||
attempt_fails: dict[str, int],
|
attempt_fails: dict[str, int],
|
||||||
@@ -156,6 +217,15 @@ class _Failed:
|
|||||||
immediate: bool
|
immediate: bool
|
||||||
|
|
||||||
|
|
||||||
|
def _is_rate_limited(outcome: LLMResponse | _Failed) -> bool:
|
||||||
|
"""429 = 服务端调度指令(gRPC pushback 语义,迭代 5): 按 Retry-After 退避但
|
||||||
|
**不消耗重试预算**——饱和窗口里等待而非死亡;其余失败照常计数。
|
||||||
|
|
||||||
|
因其免重试预算,该次尝试的耗时必须归 stall 治理(`StallClock` 的 refund)。
|
||||||
|
"""
|
||||||
|
return isinstance(outcome, _Failed) and _failure_reason(outcome.exc) == "rate_limited"
|
||||||
|
|
||||||
|
|
||||||
class RetryMW:
|
class RetryMW:
|
||||||
"""尝试编排器;时钟/睡眠/随机全部注入,纯确定性可测(P6)。"""
|
"""尝试编排器;时钟/睡眠/随机全部注入,纯确定性可测(P6)。"""
|
||||||
|
|
||||||
@@ -183,8 +253,8 @@ class RetryMW:
|
|||||||
self._scope = scope
|
self._scope = scope
|
||||||
self._sources = list(sources)
|
self._sources = list(sources)
|
||||||
self._selector = selector
|
self._selector = selector
|
||||||
self._quota = QuotaGate(limiter)
|
self._quota = QuotaGate(limiter, scope=self._scope)
|
||||||
self._breaker = BreakerGate(gate)
|
self._breaker = BreakerGate(gate, scope=self._scope)
|
||||||
self._transport = transport
|
self._transport = transport
|
||||||
self._retry = retry
|
self._retry = retry
|
||||||
self._bp = backpressure
|
self._bp = backpressure
|
||||||
@@ -208,12 +278,12 @@ class RetryMW:
|
|||||||
reasons: dict[str, str] = {}
|
reasons: dict[str, str] = {}
|
||||||
# 调用内失败计数(设计 §3.3): 局部状态,调用结束即弃;严禁实例属性(并发共享)
|
# 调用内失败计数(设计 §3.3): 局部状态,调用结束即弃;严禁实例属性(并发共享)
|
||||||
attempt_fails: dict[str, int] = {}
|
attempt_fails: dict[str, int] = {}
|
||||||
entered_at = self._now() # 调用级累计计时,循环内不重置(CHS governance.py:207)
|
# 调用级累计计时,循环内不重置(CHS governance.py:207);issue #8 起只计
|
||||||
|
# 非生产性等待——真实尝试由重试预算治理,不再重复烧 stall 预算
|
||||||
|
clock = StallClock(self._now)
|
||||||
while True:
|
while True:
|
||||||
# 调用级时间上限(迭代 5): 429 免预算后的兜底,防饱和期无限循环。
|
# 调用级时间上限(迭代 5): 429 免预算后的兜底,防饱和期无限循环
|
||||||
# 与 _on_no_runnable 同款双条件(CHS 口径): 本地超窗且全局无进展才判死
|
if await self._stalled(clock):
|
||||||
stall = self._bp.stall_window_s
|
|
||||||
if self._now() - entered_at > stall and await self._quota.progress_age_s() > stall:
|
|
||||||
raise AllSourcesExhausted(
|
raise AllSourcesExhausted(
|
||||||
scope=self._scope,
|
scope=self._scope,
|
||||||
reason="stalled",
|
reason="stalled",
|
||||||
@@ -222,14 +292,17 @@ class RetryMW:
|
|||||||
)
|
)
|
||||||
picked, gate_rejections = await self._pick_runnable(reasons, attempt_fails)
|
picked, gate_rejections = await self._pick_runnable(reasons, attempt_fails)
|
||||||
if picked is None:
|
if picked is None:
|
||||||
await self._on_no_runnable(gate_rejections, reasons, entered_at)
|
await self._on_no_runnable(gate_rejections, reasons, clock)
|
||||||
continue
|
continue
|
||||||
|
async with clock.attempting() as attempt:
|
||||||
outcome = await self._attempt(request, *picked, reasons, attempt_fails)
|
outcome = await self._attempt(request, *picked, reasons, attempt_fails)
|
||||||
|
rate_limited = _is_rate_limited(outcome)
|
||||||
|
if rate_limited:
|
||||||
|
# 免了重试预算就得进 stall 账,否则这段耗时无人治理(见 StallClock)
|
||||||
|
attempt.refund()
|
||||||
if isinstance(outcome, LLMResponse):
|
if isinstance(outcome, LLMResponse):
|
||||||
return outcome
|
return outcome
|
||||||
# 429 = 服务端调度指令(gRPC pushback 语义,迭代 5): 按 Retry-After
|
if not rate_limited:
|
||||||
# 退避但不消耗重试预算——饱和窗口里等待而非死亡;其余失败照常计数
|
|
||||||
if _failure_reason(outcome.exc) != "rate_limited":
|
|
||||||
fails += 1
|
fails += 1
|
||||||
if fails >= self._retry.max_attempts:
|
if fails >= self._retry.max_attempts:
|
||||||
raise AllSourcesExhausted(
|
raise AllSourcesExhausted(
|
||||||
@@ -282,8 +355,20 @@ class RetryMW:
|
|||||||
await self._settle_and_release(permit, 0)
|
await self._settle_and_release(permit, 0)
|
||||||
return None, gate_rejections
|
return None, gate_rejections
|
||||||
|
|
||||||
|
# —— 背压与 stall 判死(CHS governance.py:270-285)——
|
||||||
|
|
||||||
|
async def _stalled(self, clock: StallClock) -> bool:
|
||||||
|
"""双条件 stall 判死(CHS governance.py:270-281): 本地累计等待与全局
|
||||||
|
无进展**同时**超窗才判死——本地 monotonic 与后端时钟刻意不混用。
|
||||||
|
|
||||||
|
本地一侧只计非生产性等待(issue #8,见 `StallClock`)。短路顺序有意为之:
|
||||||
|
本地未超窗就不问后端,省一次 Redis 往返。
|
||||||
|
"""
|
||||||
|
stall = self._bp.stall_window_s
|
||||||
|
return clock.stalled_s() > stall and await self._quota.progress_age_s() > stall
|
||||||
|
|
||||||
async def _on_no_runnable(
|
async def _on_no_runnable(
|
||||||
self, gate_rejections: int, reasons: dict[str, str], entered_at: float
|
self, gate_rejections: int, reasons: dict[str, str], clock: StallClock
|
||||||
) -> None:
|
) -> None:
|
||||||
if gate_rejections == len(self._sources):
|
if gate_rejections == len(self._sources):
|
||||||
names = tuple(s.name for s in self._sources)
|
names = tuple(s.name for s in self._sources)
|
||||||
@@ -299,10 +384,7 @@ class RetryMW:
|
|||||||
retry_after_s=self._bp.poll_interval_s,
|
retry_after_s=self._bp.poll_interval_s,
|
||||||
per_source_reasons=reasons,
|
per_source_reasons=reasons,
|
||||||
)
|
)
|
||||||
# 双条件 stall 判死(CHS governance.py:270-281): 本地累计等待与全局
|
if await self._stalled(clock):
|
||||||
# 无进展**同时**超窗才判死——本地 monotonic 与后端时钟刻意不混用。
|
|
||||||
stall = self._bp.stall_window_s
|
|
||||||
if self._now() - entered_at > stall and await self._quota.progress_age_s() > stall:
|
|
||||||
names = tuple(s.name for s in self._sources)
|
names = tuple(s.name for s in self._sources)
|
||||||
raise AllSourcesExhausted(
|
raise AllSourcesExhausted(
|
||||||
scope=self._scope,
|
scope=self._scope,
|
||||||
@@ -401,7 +483,7 @@ class RetryMW:
|
|||||||
await write_back
|
await write_back
|
||||||
except asyncio.CancelledError:
|
except asyncio.CancelledError:
|
||||||
raise
|
raise
|
||||||
except GovernanceBackendError as exc:
|
except (GovernanceBackendError, SourceNotConfiguredError) as exc:
|
||||||
logger.warning("治理记账写回降级(不冒泡): {}", exc)
|
logger.warning("治理记账写回降级(不冒泡): {}", exc)
|
||||||
|
|
||||||
def _feed_outcome(self, source_name: str, ok: bool) -> None:
|
def _feed_outcome(self, source_name: str, ok: bool) -> None:
|
||||||
@@ -438,6 +520,7 @@ class RetryMW:
|
|||||||
usage_source=result.usage_source,
|
usage_source=result.usage_source,
|
||||||
cached_prompt_tokens=result.cached_prompt_tokens,
|
cached_prompt_tokens=result.cached_prompt_tokens,
|
||||||
model_reported=result.model_reported,
|
model_reported=result.model_reported,
|
||||||
|
reasoning_tokens=result.reasoning_tokens,
|
||||||
)
|
)
|
||||||
|
|
||||||
async def _settle_and_release(self, permit: Permit, actual: int) -> None:
|
async def _settle_and_release(self, permit: Permit, actual: int) -> None:
|
||||||
|
|||||||
@@ -12,11 +12,16 @@ import asyncio
|
|||||||
import json
|
import json
|
||||||
import time
|
import time
|
||||||
import uuid
|
import uuid
|
||||||
|
from dataclasses import dataclass
|
||||||
from typing import TYPE_CHECKING
|
from typing import TYPE_CHECKING
|
||||||
|
|
||||||
from loguru import logger
|
from loguru import logger
|
||||||
|
|
||||||
from polygateway.errors import GatewayUnavailableError, GovernanceBackendError
|
from polygateway.errors import (
|
||||||
|
GatewayUnavailableError,
|
||||||
|
GovernanceBackendError,
|
||||||
|
SourceNotConfiguredError,
|
||||||
|
)
|
||||||
from polygateway.middleware.cache import digest_messages
|
from polygateway.middleware.cache import digest_messages
|
||||||
from polygateway.types import canonical_sampling_json, merge_sampling
|
from polygateway.types import canonical_sampling_json, merge_sampling
|
||||||
|
|
||||||
@@ -28,6 +33,44 @@ if TYPE_CHECKING:
|
|||||||
from polygateway.types import ChatRequest, LLMResponse, SourceConfig
|
from polygateway.types import ChatRequest, LLMResponse, SourceConfig
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class _AttemptUsage:
|
||||||
|
"""一次尝试的用量视图;默认值即"失败尝试"档(无用量可言,记 0 并标 unavailable)。
|
||||||
|
|
||||||
|
存在的理由是把 `emit_attempt` 里逐字段重复的 `X if response else Y` 收敛为
|
||||||
|
一处判定——十处三元把该方法推到圈复杂度 C,而它们表达的是同一件事。
|
||||||
|
"""
|
||||||
|
|
||||||
|
response_text: str = ""
|
||||||
|
thinking: str = ""
|
||||||
|
prompt_tokens: int = 0
|
||||||
|
completion_tokens: int = 0
|
||||||
|
usage_source: str = "unavailable"
|
||||||
|
ttft_ms: float | None = None
|
||||||
|
max_inter_token_ms: float | None = None
|
||||||
|
cached_prompt_tokens: int | None = None
|
||||||
|
model_reported: str | None = None
|
||||||
|
reasoning_tokens: int | None = None
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def of(cls, response: LLMResponse | None) -> _AttemptUsage:
|
||||||
|
"""从响应取用量;`None`(失败尝试)返回全默认视图。"""
|
||||||
|
if response is None:
|
||||||
|
return cls()
|
||||||
|
return cls(
|
||||||
|
response_text=response.content,
|
||||||
|
thinking=response.thinking,
|
||||||
|
prompt_tokens=response.prompt_tokens,
|
||||||
|
completion_tokens=response.completion_tokens,
|
||||||
|
usage_source=response.usage_source,
|
||||||
|
ttft_ms=response.ttft_ms,
|
||||||
|
max_inter_token_ms=response.max_inter_token_ms,
|
||||||
|
cached_prompt_tokens=response.cached_prompt_tokens,
|
||||||
|
model_reported=response.model_reported,
|
||||||
|
reasoning_tokens=response.reasoning_tokens,
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
class TelemetryEmitter:
|
class TelemetryEmitter:
|
||||||
"""从请求与结果组装 21 字段并写入 recorder;一切写失败降级 warning。"""
|
"""从请求与结果组装 21 字段并写入 recorder;一切写失败降级 warning。"""
|
||||||
|
|
||||||
@@ -46,24 +89,26 @@ class TelemetryEmitter:
|
|||||||
error: str | None,
|
error: str | None,
|
||||||
) -> None:
|
) -> None:
|
||||||
"""逐次尝试记录(RetryMW 调用);失败尝试无用量可言,记 0 并标 unavailable。"""
|
"""逐次尝试记录(RetryMW 调用);失败尝试无用量可言,记 0 并标 unavailable。"""
|
||||||
|
usage = _AttemptUsage.of(response)
|
||||||
await self._record(
|
await self._record(
|
||||||
request=request,
|
request=request,
|
||||||
call_id=call_id,
|
call_id=call_id,
|
||||||
model=source.model,
|
model=source.model,
|
||||||
provider=source.provider,
|
provider=source.provider,
|
||||||
source_name=source.name,
|
source_name=source.name,
|
||||||
response_text=response.content if response else "",
|
response_text=usage.response_text,
|
||||||
thinking=response.thinking if response else "",
|
thinking=usage.thinking,
|
||||||
prompt_tokens=response.prompt_tokens if response else 0,
|
prompt_tokens=usage.prompt_tokens,
|
||||||
completion_tokens=response.completion_tokens if response else 0,
|
completion_tokens=usage.completion_tokens,
|
||||||
usage_source=response.usage_source if response else "unavailable",
|
usage_source=usage.usage_source,
|
||||||
latency_ms=latency_ms,
|
latency_ms=latency_ms,
|
||||||
ttft_ms=response.ttft_ms if response else None,
|
ttft_ms=usage.ttft_ms,
|
||||||
max_inter_token_ms=response.max_inter_token_ms if response else None,
|
max_inter_token_ms=usage.max_inter_token_ms,
|
||||||
cache_hit=False,
|
cache_hit=False,
|
||||||
error=error,
|
error=error,
|
||||||
cached_prompt_tokens=response.cached_prompt_tokens if response else None,
|
cached_prompt_tokens=usage.cached_prompt_tokens,
|
||||||
model_reported=response.model_reported if response else None,
|
model_reported=usage.model_reported,
|
||||||
|
reasoning_tokens=usage.reasoning_tokens,
|
||||||
# 唯一有"生效源"的入口,故是唯一能并上 extra_body 的(设计决策 D)
|
# 唯一有"生效源"的入口,故是唯一能并上 extra_body 的(设计决策 D)
|
||||||
sampling=canonical_sampling_json(merge_sampling(source.extra_body, request.sampling)),
|
sampling=canonical_sampling_json(merge_sampling(source.extra_body, request.sampling)),
|
||||||
)
|
)
|
||||||
@@ -90,6 +135,7 @@ class TelemetryEmitter:
|
|||||||
# 统计供应商缓存命中率必须带 WHERE cache_hit = false,否则重复计数。
|
# 统计供应商缓存命中率必须带 WHERE cache_hit = false,否则重复计数。
|
||||||
cached_prompt_tokens=response.cached_prompt_tokens,
|
cached_prompt_tokens=response.cached_prompt_tokens,
|
||||||
model_reported=response.model_reported,
|
model_reported=response.model_reported,
|
||||||
|
reasoning_tokens=response.reasoning_tokens,
|
||||||
# 由最外层 TelemetryMW 调用,手上没有 source。缓存命中行无损:
|
# 由最外层 TelemetryMW 调用,手上没有 source。缓存命中行无损:
|
||||||
# sampling 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同
|
# sampling 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同
|
||||||
sampling=canonical_sampling_json(request.sampling),
|
sampling=canonical_sampling_json(request.sampling),
|
||||||
@@ -117,6 +163,7 @@ class TelemetryEmitter:
|
|||||||
error=error,
|
error=error,
|
||||||
cached_prompt_tokens=None,
|
cached_prompt_tokens=None,
|
||||||
model_reported=None,
|
model_reported=None,
|
||||||
|
reasoning_tokens=None,
|
||||||
# 无具体源,与 model/provider/source_name 置空同一先例(设计决策 D)
|
# 无具体源,与 model/provider/source_name 置空同一先例(设计决策 D)
|
||||||
sampling=canonical_sampling_json(request.sampling),
|
sampling=canonical_sampling_json(request.sampling),
|
||||||
)
|
)
|
||||||
@@ -142,6 +189,7 @@ class TelemetryEmitter:
|
|||||||
cached_prompt_tokens: int | None,
|
cached_prompt_tokens: int | None,
|
||||||
model_reported: str | None,
|
model_reported: str | None,
|
||||||
sampling: str | None,
|
sampling: str | None,
|
||||||
|
reasoning_tokens: int | None,
|
||||||
) -> None:
|
) -> None:
|
||||||
try:
|
try:
|
||||||
# 成本换算(M2 §6): 成功行按单价换算;缓存命中 0.0(未产生新调用);
|
# 成本换算(M2 §6): 成功行按单价换算;缓存命中 0.0(未产生新调用);
|
||||||
@@ -182,6 +230,7 @@ class TelemetryEmitter:
|
|||||||
cached_prompt_tokens=cached_prompt_tokens,
|
cached_prompt_tokens=cached_prompt_tokens,
|
||||||
model_reported=model_reported,
|
model_reported=model_reported,
|
||||||
sampling=sampling,
|
sampling=sampling,
|
||||||
|
reasoning_tokens=reasoning_tokens,
|
||||||
)
|
)
|
||||||
except asyncio.CancelledError:
|
except asyncio.CancelledError:
|
||||||
raise
|
raise
|
||||||
@@ -202,7 +251,7 @@ class TelemetryMW:
|
|||||||
started = self._now()
|
started = self._now()
|
||||||
try:
|
try:
|
||||||
response = await call_next(request)
|
response = await call_next(request)
|
||||||
except (GatewayUnavailableError, GovernanceBackendError) as exc:
|
except (GatewayUnavailableError, GovernanceBackendError, SourceNotConfiguredError) as exc:
|
||||||
await self._emitter.emit_terminal_failure(
|
await self._emitter.emit_terminal_failure(
|
||||||
request=request,
|
request=request,
|
||||||
call_id=str(uuid.uuid4()),
|
call_id=str(uuid.uuid4()),
|
||||||
|
|||||||
+14
-9
@@ -30,11 +30,12 @@ from polygateway.errors import (
|
|||||||
RequestRejectedError,
|
RequestRejectedError,
|
||||||
ResultInvalidError,
|
ResultInvalidError,
|
||||||
SourceDeadError,
|
SourceDeadError,
|
||||||
|
SourceNotConfiguredError,
|
||||||
TransientError,
|
TransientError,
|
||||||
)
|
)
|
||||||
from polygateway.middleware.breaker import BreakerGate
|
from polygateway.middleware.breaker import BreakerGate
|
||||||
from polygateway.middleware.ratelimit import QuotaGate
|
from polygateway.middleware.ratelimit import QuotaGate
|
||||||
from polygateway.middleware.retry import _failure_reason, backoff_delay
|
from polygateway.middleware.retry import StallClock, _failure_reason, backoff_delay
|
||||||
from polygateway.middleware.telemetry import TelemetryEmitter
|
from polygateway.middleware.telemetry import TelemetryEmitter
|
||||||
from polygateway.ports import OutcomeAwareSelector
|
from polygateway.ports import OutcomeAwareSelector
|
||||||
from polygateway.sources import SourceCooldownMemo
|
from polygateway.sources import SourceCooldownMemo
|
||||||
@@ -119,8 +120,8 @@ class OcrClient:
|
|||||||
self._sources = strip_unsupported_extra_body(list(sources), path="OCR")
|
self._sources = strip_unsupported_extra_body(list(sources), path="OCR")
|
||||||
self._selector = selector
|
self._selector = selector
|
||||||
self._feed_health = isinstance(selector, OutcomeAwareSelector)
|
self._feed_health = isinstance(selector, OutcomeAwareSelector)
|
||||||
self._quota = QuotaGate(limiter)
|
self._quota = QuotaGate(limiter, scope=self._scope)
|
||||||
self._breaker = BreakerGate(breaker)
|
self._breaker = BreakerGate(breaker, scope=self._scope)
|
||||||
self._transport = transport
|
self._transport = transport
|
||||||
self._retry = retry
|
self._retry = retry
|
||||||
self._bp = backpressure
|
self._bp = backpressure
|
||||||
@@ -203,13 +204,17 @@ class OcrClient:
|
|||||||
raise AllSourcesExhausted(scope=self._scope, reason="no_sources", retry_after_s=0.0)
|
raise AllSourcesExhausted(scope=self._scope, reason="no_sources", retry_after_s=0.0)
|
||||||
fails = 0
|
fails = 0
|
||||||
reasons: dict[str, str] = {}
|
reasons: dict[str, str] = {}
|
||||||
entered_at = self._now()
|
# 只计非生产性等待(issue #8): 真实尝试由重试预算治理,不重复烧 stall 预算
|
||||||
|
clock = StallClock(self._now)
|
||||||
while True:
|
while True:
|
||||||
picked, gate_rejections = await self._pick_runnable(reasons)
|
picked, gate_rejections = await self._pick_runnable(reasons)
|
||||||
if picked is None:
|
if picked is None:
|
||||||
await self._on_no_runnable(gate_rejections, reasons, entered_at)
|
await self._on_no_runnable(gate_rejections, reasons, clock)
|
||||||
continue
|
continue
|
||||||
outcome = await self._attempt(kind, image, *picked, reasons, session_id, parent_call_id)
|
async with clock.attempting():
|
||||||
|
outcome = await self._attempt(
|
||||||
|
kind, image, *picked, reasons, session_id, parent_call_id
|
||||||
|
)
|
||||||
if isinstance(outcome, _AttemptOutcome):
|
if isinstance(outcome, _AttemptOutcome):
|
||||||
return outcome
|
return outcome
|
||||||
fails += 1
|
fails += 1
|
||||||
@@ -252,7 +257,7 @@ class OcrClient:
|
|||||||
return None, gate_rejections
|
return None, gate_rejections
|
||||||
|
|
||||||
async def _on_no_runnable(
|
async def _on_no_runnable(
|
||||||
self, gate_rejections: int, reasons: dict[str, str], entered_at: float
|
self, gate_rejections: int, reasons: dict[str, str], clock: StallClock
|
||||||
) -> None:
|
) -> None:
|
||||||
if gate_rejections == len(self._sources):
|
if gate_rejections == len(self._sources):
|
||||||
names = tuple(s.name for s in self._sources)
|
names = tuple(s.name for s in self._sources)
|
||||||
@@ -269,7 +274,7 @@ class OcrClient:
|
|||||||
per_source_reasons=reasons,
|
per_source_reasons=reasons,
|
||||||
)
|
)
|
||||||
stall = self._bp.stall_window_s
|
stall = self._bp.stall_window_s
|
||||||
if self._now() - entered_at > stall and await self._quota.progress_age_s() > stall:
|
if clock.stalled_s() > stall and await self._quota.progress_age_s() > stall:
|
||||||
names = tuple(s.name for s in self._sources)
|
names = tuple(s.name for s in self._sources)
|
||||||
raise AllSourcesExhausted(
|
raise AllSourcesExhausted(
|
||||||
scope=self._scope,
|
scope=self._scope,
|
||||||
@@ -360,7 +365,7 @@ class OcrClient:
|
|||||||
await write_back
|
await write_back
|
||||||
except asyncio.CancelledError:
|
except asyncio.CancelledError:
|
||||||
raise
|
raise
|
||||||
except GovernanceBackendError as exc:
|
except (GovernanceBackendError, SourceNotConfiguredError) as exc:
|
||||||
logger.warning("OCR 治理记账写回降级(不冒泡): {}", exc)
|
logger.warning("OCR 治理记账写回降级(不冒泡): {}", exc)
|
||||||
|
|
||||||
async def _settle_and_release(self, permit: Permit) -> None:
|
async def _settle_and_release(self, permit: Permit) -> None:
|
||||||
|
|||||||
@@ -245,7 +245,7 @@ class StructuredOutputStrategy(Protocol):
|
|||||||
|
|
||||||
@runtime_checkable
|
@runtime_checkable
|
||||||
class TelemetryRecorder(Protocol):
|
class TelemetryRecorder(Protocol):
|
||||||
"""遥测后端;20 字段冻结(M1 设计 §4.4 + issue #3),唯一调用点是 TelemetryEmitter。
|
"""遥测后端;22 字段冻结(M1 设计 §4.4 + issue #3/#4),唯一调用点是 TelemetryEmitter。
|
||||||
|
|
||||||
新增参数不设默认值: 库外无第三方实现者(三项目迁移时删除了各自的同名
|
新增参数不设默认值: 库外无第三方实现者(三项目迁移时删除了各自的同名
|
||||||
Protocol),完整签名的成本为零,而少写一列会被 emitter 的降级吞成 warning。
|
Protocol),完整签名的成本为零,而少写一列会被 emitter 的降级吞成 warning。
|
||||||
@@ -275,4 +275,5 @@ class TelemetryRecorder(Protocol):
|
|||||||
cached_prompt_tokens: int | None,
|
cached_prompt_tokens: int | None,
|
||||||
model_reported: str | None,
|
model_reported: str | None,
|
||||||
sampling: str | None,
|
sampling: str | None,
|
||||||
|
reasoning_tokens: int | None,
|
||||||
) -> None: ...
|
) -> None: ...
|
||||||
|
|||||||
+174
-16
@@ -10,24 +10,37 @@ from dataclasses import dataclass
|
|||||||
from types import MappingProxyType
|
from types import MappingProxyType
|
||||||
from typing import Any
|
from typing import Any
|
||||||
|
|
||||||
|
from loguru import logger
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
class ProviderProfile:
|
class ProviderProfile:
|
||||||
"""单个 provider 的能力与差异声明。
|
"""单个 provider 的能力与差异声明。
|
||||||
|
|
||||||
thinking_on/thinking_off 分别是 `SourceConfig.enable_thinking` 为
|
thinking_on/thinking_off 分别是 `SourceConfig.enable_thinking` 为
|
||||||
True/False 时并入请求体的参数片段(None 时二者都不注入,用模型默认);
|
True/False 时并入请求体的参数片段(`enable_thinking` 为 None 时二者都不
|
||||||
strip_think_tags 声明响应 content 需剥离 ``<think>`` 标签(qwen 系);
|
注入,用模型默认);strip_think_tags 声明响应 content 需剥离 ``<think>``
|
||||||
supports_native_schema 供 D14 阶梯选择原生 response_format 策略。
|
标签(qwen 系);supports_native_schema 供 D14 阶梯选择原生 response_format。
|
||||||
|
|
||||||
注: 某个 provider 的两档若皆为空字典(如 openai/minimax),说明该 provider
|
两档各有三种取值,**语义互不重叠**(issue #5):
|
||||||
无已知的推理开关参数——此时 `enable_thinking` 对它**不产生任何效果**,
|
|
||||||
而非静默生效。需要下发自定义参数时用 `SourceConfig.extra_body`。
|
========== ==========================================================
|
||||||
|
``{...}`` 已知的注入片段
|
||||||
|
``{}`` 已知**无需注入**任何参数即处于该档
|
||||||
|
``None`` **未知**: 本库不知道该 provider 如何表达这一档
|
||||||
|
========== ==========================================================
|
||||||
|
|
||||||
|
`None` 与 `{}` 必须分开: 二者曾同为空字典,导致 `enable_thinking=False`
|
||||||
|
对 minimax/openai 源静默失效——调用方以为关掉了推理,实际什么都没发生。
|
||||||
|
现在 `None` 会在装配期显式报错并指路 `register_provider` / `extra_body`。
|
||||||
|
|
||||||
|
注: 本类只声明**形态**(参数长什么样,按 provider 变);某个具体模型能否
|
||||||
|
关闭推理属**能力**(按 model 变),见 `ThinkingCapability`。
|
||||||
"""
|
"""
|
||||||
|
|
||||||
name: str
|
name: str
|
||||||
thinking_on: dict[str, Any]
|
thinking_on: Mapping[str, Any] | None
|
||||||
thinking_off: dict[str, Any]
|
thinking_off: Mapping[str, Any] | None
|
||||||
strip_think_tags: bool
|
strip_think_tags: bool
|
||||||
supports_native_schema: bool = False
|
supports_native_schema: bool = False
|
||||||
|
|
||||||
@@ -47,26 +60,171 @@ DEFAULT_PROFILES: Mapping[str, ProviderProfile] = MappingProxyType(
|
|||||||
thinking_off={"thinking": {"type": "disabled"}},
|
thinking_off={"thinking": {"type": "disabled"}},
|
||||||
strip_think_tags=False,
|
strip_think_tags=False,
|
||||||
),
|
),
|
||||||
# 两档皆空 ⇒ `enable_thinking` 对本 provider **不产生任何效果**(调用方
|
# OpenAI 兼容基线段名: 实践中被复用为**任意**兼容厂商的兜底(下游把
|
||||||
# 以为关掉了实际没关)。真需要控制推理时经 `SourceConfig.extra_body` 下发
|
# kimi-k3 挂在 provider=openai 下),故不能下发任何厂商方言参数——发给
|
||||||
|
# 不认识它的厂商会 400。两档标 None(未知): 配了 enable_thinking 即在
|
||||||
|
# 装配期报错并指路,真 OpenAI 推理模型的用户走 register_provider
|
||||||
"openai": ProviderProfile(
|
"openai": ProviderProfile(
|
||||||
name="openai",
|
name="openai",
|
||||||
thinking_on={},
|
thinking_on=None,
|
||||||
thinking_off={},
|
thinking_off=None,
|
||||||
strip_think_tags=False,
|
strip_think_tags=False,
|
||||||
),
|
),
|
||||||
# OpenAI 兼容基线,无已知注入差异;reasoning_content 由 transport 通用处理。
|
# 注入形态出处: 2026-08-02 经自建 new-api 中转实测(findings §2),
|
||||||
# 同上: 两档皆空 ⇒ `enable_thinking` 对 MiniMax 源不产生任何效果
|
# **直连官方端点未验证**。实测 enable_thinking / thinking 两种写法均被
|
||||||
|
# 静默丢弃(prompt_tokens 恒定不变),reasoning_effort 才是真开关。
|
||||||
|
# "开"取 medium: qwen 的 enable_thinking:true 与 deepseek 的
|
||||||
|
# thinking:{enabled} 都不指定预算、由模型自定,medium 是五档里语义最接近
|
||||||
|
# "厂商正常强度"的一档;取 high 等于替下游做"加钱换质量"的业务判断。
|
||||||
|
# 要精确控制档位经 `SourceConfig.extra_body`(优先级高于本片段)
|
||||||
"minimax": ProviderProfile(
|
"minimax": ProviderProfile(
|
||||||
name="minimax",
|
name="minimax",
|
||||||
thinking_on={},
|
thinking_on={"reasoning_effort": "medium"},
|
||||||
thinking_off={},
|
thinking_off={"reasoning_effort": "none"},
|
||||||
strip_think_tags=False,
|
strip_think_tags=False,
|
||||||
),
|
),
|
||||||
}
|
}
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class ThinkingUnsupportedError(ValueError):
|
||||||
|
"""推理开关无法满足: 形态未知或该模型不支持该方向(issue #5)。
|
||||||
|
|
||||||
|
是 `ValueError` 的子类而非 `errors.py` 四分类之一——它描述的是**配置**
|
||||||
|
不可满足(装配期就该炸),不是一次调用的运行时失败。transport 在请求期
|
||||||
|
捕获它并翻译为 `RequestRejectedError` 再进四分类。单列一个类型是为了让
|
||||||
|
捕获点能精确到它,而不是宽catch 整个 `ValueError`(那会把序列化等无关
|
||||||
|
错误误贴成"推理开关无法满足")。
|
||||||
|
"""
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class ThinkingCapability:
|
||||||
|
"""某个**具体模型**能否关闭推理(issue #5);登记必须附实测证据与日期。
|
||||||
|
|
||||||
|
与 `ProviderProfile` 的分工: 后者声明**形态**(参数长什么样,按 provider 变,
|
||||||
|
数年不变一次),本类声明**能力**(按 model 变,同一 provider 每代都变)。二者
|
||||||
|
合一在 provider 级表达不了代际差异——实测 MiniMax-M3 可关闭推理,而同厂的
|
||||||
|
M2.7/M2.5 三种参数形态全部无效(findings §2.3),profile 一格管不住三个模型。
|
||||||
|
|
||||||
|
`evidence` 不是装饰: 能力表过期是必然事件,没有出处就无从判断该不该信它。
|
||||||
|
"""
|
||||||
|
|
||||||
|
can_disable: bool
|
||||||
|
evidence: str
|
||||||
|
|
||||||
|
|
||||||
|
DEFAULT_CAPABILITIES: Mapping[str, ThinkingCapability] = MappingProxyType(
|
||||||
|
{
|
||||||
|
"MiniMax-M3": ThinkingCapability(
|
||||||
|
can_disable=True,
|
||||||
|
evidence="2026-08-02 经 new-api 中转实测 N=10: reasoning_effort=none 稳定关闭,零跳变",
|
||||||
|
),
|
||||||
|
"MiniMax-M2.7": ThinkingCapability(
|
||||||
|
can_disable=False,
|
||||||
|
evidence=(
|
||||||
|
"2026-08-02 实测 reasoning_effort=none / thinking:{disabled} / thinking:{adaptive} "
|
||||||
|
"各 N=3 全部无效;OpenRouter 注册表登记 mandatory:true,models.dev 登记无控制手段"
|
||||||
|
),
|
||||||
|
),
|
||||||
|
"MiniMax-M2.5": ThinkingCapability(
|
||||||
|
can_disable=False,
|
||||||
|
evidence="2026-08-02 实测同 M2.7: 三种形态各 N=3 全部无效;外部注册表同样登记为强制推理",
|
||||||
|
),
|
||||||
|
"qwen3.7-plus": ThinkingCapability(
|
||||||
|
can_disable=True,
|
||||||
|
evidence="2026-08-02 实测 enable_thinking=false 关闭(completion 5 token,无推理)",
|
||||||
|
),
|
||||||
|
"deepseek-v4-pro": ThinkingCapability(
|
||||||
|
can_disable=True,
|
||||||
|
evidence="2026-08-02 实测 thinking:{type:disabled} 关闭(completion 3 token,无推理)",
|
||||||
|
),
|
||||||
|
}
|
||||||
|
)
|
||||||
|
"""在用模型的推理能力登记(YAGNI: 不覆盖全世界,未登记走 `resolve_thinking` 退化)。"""
|
||||||
|
|
||||||
|
|
||||||
|
def get_capability(
|
||||||
|
model: str, *, table: Mapping[str, ThinkingCapability] | None = None
|
||||||
|
) -> ThinkingCapability | None:
|
||||||
|
"""按模型名精确查找;未登记返回 None(= 能力未知,由调用方决定如何退化)。
|
||||||
|
|
||||||
|
与 `get_provider` 未注册即报错不同: provider 是配置里写死的少数几个值,
|
||||||
|
写错就是配置错误;而模型名千变万化,新模型上线不该被库挡住(设计 §5 R4)。
|
||||||
|
"""
|
||||||
|
return (DEFAULT_CAPABILITIES if table is None else table).get(model)
|
||||||
|
|
||||||
|
|
||||||
|
def register_capability(
|
||||||
|
model: str,
|
||||||
|
capability: ThinkingCapability,
|
||||||
|
*,
|
||||||
|
base: Mapping[str, ThinkingCapability] | None = None,
|
||||||
|
) -> dict[str, ThinkingCapability]:
|
||||||
|
"""纯函数注册: 返回 base(缺省 DEFAULT_CAPABILITIES)+ 新条目的新表,同名覆盖。"""
|
||||||
|
table = dict(DEFAULT_CAPABILITIES if base is None else base)
|
||||||
|
table[model] = capability
|
||||||
|
return table
|
||||||
|
|
||||||
|
|
||||||
|
def resolve_thinking(
|
||||||
|
profile: ProviderProfile,
|
||||||
|
capability: ThinkingCapability | None,
|
||||||
|
enable_thinking: bool | None,
|
||||||
|
*,
|
||||||
|
model: str,
|
||||||
|
warn_unregistered: bool = True,
|
||||||
|
) -> Mapping[str, Any]:
|
||||||
|
"""三态 + 两层能力 → 请求体注入片段;不可满足时 ValueError。
|
||||||
|
|
||||||
|
调用点负责翻译: 装配期直接冒泡(配置错误),transport 内翻译为
|
||||||
|
`RequestRejectedError`(四分类之一)。判定顺序即语义,不可调换——形态未知时
|
||||||
|
无从注入,能力如何无关紧要,故 Phase 2 必须先于 Phase 4;未登记模型没有
|
||||||
|
`can_disable` 可读,故 Phase 3 必须先于 Phase 4。
|
||||||
|
|
||||||
|
`model` 只用于错误与告警文案: 报错能定位到具体模型才有可操作性,而
|
||||||
|
`capability` 为 None(未登记)时无从从别处取得模型名。
|
||||||
|
|
||||||
|
`warn_unregistered=False` 供请求热路径去重用: 装配期已经喊过一次,逐次
|
||||||
|
调用再喊只会刷屏。判定结果不受此参数影响。
|
||||||
|
"""
|
||||||
|
# Phase 1: 调用方不表态 —— 与 False 严格区分,用模型默认档
|
||||||
|
if enable_thinking is None:
|
||||||
|
return {}
|
||||||
|
slot = profile.thinking_on if enable_thinking else profile.thinking_off
|
||||||
|
direction = "thinking_on" if enable_thinking else "thinking_off"
|
||||||
|
# Phase 2: 形态未知 —— 提供了开关却不知道怎么发,静默放行就是欺骗调用方
|
||||||
|
if slot is None:
|
||||||
|
raise ThinkingUnsupportedError(
|
||||||
|
f"provider {profile.name!r} 的 {direction} 形态未知(模型 {model!r}): "
|
||||||
|
f"本库不知道该 provider 如何表达这一档。请用 register_provider 注册形态,"
|
||||||
|
f"或改用 SourceConfig.extra_body 直接下发供应商参数"
|
||||||
|
)
|
||||||
|
# Phase 3: 能力未登记 —— 新模型上线不该被库挡住,但也不该假装成功
|
||||||
|
if capability is None:
|
||||||
|
if warn_unregistered:
|
||||||
|
_warn_unregistered(model, profile, slot)
|
||||||
|
return slot
|
||||||
|
# Phase 4: 明确不支持关闭 —— 调用方要的是"不推理"的语义保证,给不了必须说
|
||||||
|
if enable_thinking is False and not capability.can_disable:
|
||||||
|
raise ThinkingUnsupportedError(
|
||||||
|
f"模型 {model!r} 无法关闭推理,enable_thinking=False 无法满足: "
|
||||||
|
f"{capability.evidence}。该模型的推理是固有属性,任何参数都关不掉——"
|
||||||
|
f"需要关闭思维链请换用支持关闭的模型"
|
||||||
|
)
|
||||||
|
return slot
|
||||||
|
|
||||||
|
|
||||||
|
def _warn_unregistered(model: str, profile: ProviderProfile, slot: Mapping[str, Any]) -> None:
|
||||||
|
logger.warning(
|
||||||
|
"模型 {} 的推理能力未登记,按 provider {} 的形态尽力注入 {};"
|
||||||
|
"若该模型实际不支持这一档,本次设置将静默失效。实测后请用 register_capability 登记",
|
||||||
|
model,
|
||||||
|
profile.name,
|
||||||
|
dict(slot),
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
def get_provider(
|
def get_provider(
|
||||||
name: str, *, registry: Mapping[str, ProviderProfile] | None = None
|
name: str, *, registry: Mapping[str, ProviderProfile] | None = None
|
||||||
) -> ProviderProfile:
|
) -> ProviderProfile:
|
||||||
|
|||||||
@@ -1,12 +1,17 @@
|
|||||||
"""Postgres 遥测后端(M2 设计 §5): asyncpg lazy 池 + 两级降级。
|
"""Postgres 遥测后端(M2 设计 §5): asyncpg lazy 池 + 两级降级。
|
||||||
|
|
||||||
参考仓无先例(三项目遥测全 SQLite);asyncpg 工程写法取 GovDoc
|
参考仓无先例(三项目遥测全 SQLite);asyncpg 工程写法取 GovDoc
|
||||||
`taskrun/postgres_store.py`($n 占位、`CREATE TABLE IF NOT EXISTS`、
|
`taskrun/postgres_store.py`($n 占位、`ON CONFLICT DO NOTHING`),但其
|
||||||
`ON CONFLICT DO NOTHING`),但其"失败冒泡"方向按遥测铁律**有意反转**:
|
"失败冒泡"方向按遥测铁律**有意反转**:
|
||||||
① 结构性失败(建池/建表)→ warning 一次后永久降级(池置 None 短路);
|
① 结构性失败 → warning 一次后永久降级(所有写入短路);
|
||||||
② 运行时单条写失败 → 逐条 warning 丢弃,不降级不重试(连接抖动由
|
② 运行时单条写失败 → 逐条 warning 丢弃,不降级不重试(连接抖动由
|
||||||
asyncpg 池自恢复;避免浸泡开头一次抖动导致后续全程失遥测)。
|
asyncpg 池自恢复;避免浸泡开头一次抖动导致后续全程失遥测)。
|
||||||
构造不连库(lazy),20 列 schema 与 SQLite 版同名同序。
|
构造不连库(lazy),22 列 schema 与 SQLite 版同名同序。
|
||||||
|
|
||||||
|
**"结构性"的判据是「确定写不进去」,不是「初始化时出过错」**(issue #9):
|
||||||
|
只有建池失败(重试要在业务路径上内联吞掉 connect 超时)与"表确定不存在
|
||||||
|
且建不出来"(后续 INSERT 必然全败)才判死;探测失败、补列失败、取连接
|
||||||
|
失败一律只 warning,让写入照常尝试或下次调用重试。
|
||||||
"""
|
"""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
@@ -42,7 +47,8 @@ CREATE TABLE IF NOT EXISTS llm_calls (
|
|||||||
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
||||||
cached_prompt_tokens INTEGER,
|
cached_prompt_tokens INTEGER,
|
||||||
model_reported TEXT,
|
model_reported TEXT,
|
||||||
sampling TEXT
|
sampling TEXT,
|
||||||
|
reasoning_tokens INTEGER
|
||||||
);
|
);
|
||||||
"""
|
"""
|
||||||
|
|
||||||
@@ -51,8 +57,12 @@ _BACKFILL = (
|
|||||||
("cached_prompt_tokens", "ALTER TABLE llm_calls ADD COLUMN cached_prompt_tokens INTEGER"),
|
("cached_prompt_tokens", "ALTER TABLE llm_calls ADD COLUMN cached_prompt_tokens INTEGER"),
|
||||||
("model_reported", "ALTER TABLE llm_calls ADD COLUMN model_reported TEXT"),
|
("model_reported", "ALTER TABLE llm_calls ADD COLUMN model_reported TEXT"),
|
||||||
("sampling", "ALTER TABLE llm_calls ADD COLUMN sampling TEXT"),
|
("sampling", "ALTER TABLE llm_calls ADD COLUMN sampling TEXT"),
|
||||||
|
("reasoning_tokens", "ALTER TABLE llm_calls ADD COLUMN reasoning_tokens INTEGER"),
|
||||||
)
|
)
|
||||||
|
|
||||||
|
# 探测表是否存在;不需要任何权限,且与 INSERT 走同一套 search_path 解析
|
||||||
|
_TABLE_EXISTS = "SELECT to_regclass('llm_calls')"
|
||||||
|
|
||||||
# 探测现有列;尊重 search_path(to_regclass 按当前 search_path 解析)
|
# 探测现有列;尊重 search_path(to_regclass 按当前 search_path 解析)
|
||||||
_EXISTING_COLUMNS = (
|
_EXISTING_COLUMNS = (
|
||||||
"SELECT attname FROM pg_attribute "
|
"SELECT attname FROM pg_attribute "
|
||||||
@@ -81,6 +91,7 @@ _COLUMNS = (
|
|||||||
"cached_prompt_tokens",
|
"cached_prompt_tokens",
|
||||||
"model_reported",
|
"model_reported",
|
||||||
"sampling",
|
"sampling",
|
||||||
|
"reasoning_tokens",
|
||||||
)
|
)
|
||||||
|
|
||||||
_INSERT = (
|
_INSERT = (
|
||||||
@@ -108,30 +119,81 @@ class PostgresRecorder:
|
|||||||
self._init_lock = asyncio.Lock()
|
self._init_lock = asyncio.Lock()
|
||||||
|
|
||||||
async def _ensure_ready(self) -> asyncpg.Pool | None:
|
async def _ensure_ready(self) -> asyncpg.Pool | None:
|
||||||
"""lazy 建池+建表;结构性失败 warning 一次后永久降级(设计 §5 两级之一)。"""
|
"""lazy 建池+备表;判死只认「确定写不进去」(issue #9),其余失败都留活路。"""
|
||||||
if self._failed:
|
if self._failed:
|
||||||
return None
|
return None
|
||||||
if self._schema_ready:
|
if self._schema_ready:
|
||||||
return self._pool
|
return self._pool
|
||||||
async with self._init_lock:
|
async with self._init_lock:
|
||||||
if self._failed or self._schema_ready:
|
if self._failed:
|
||||||
return None if self._failed else self._pool
|
return None
|
||||||
|
if self._schema_ready:
|
||||||
|
return self._pool
|
||||||
|
pool = await self._open_pool()
|
||||||
|
if pool is None:
|
||||||
|
return None
|
||||||
|
return await self._prepare_schema(pool)
|
||||||
|
|
||||||
|
async def _open_pool(self) -> asyncpg.Pool | None:
|
||||||
|
"""建池;失败即永久降级(唯一一处「无条件判死」)。"""
|
||||||
|
if self._pool is not None:
|
||||||
|
return self._pool
|
||||||
try:
|
try:
|
||||||
if self._pool is None:
|
|
||||||
import asyncpg
|
import asyncpg
|
||||||
|
|
||||||
self._pool = await asyncpg.create_pool(self._dsn, timeout=10)
|
self._pool = await asyncpg.create_pool(self._dsn, timeout=10)
|
||||||
async with self._pool.acquire() as conn:
|
|
||||||
await conn.execute(_DDL)
|
|
||||||
await self._backfill_columns(conn)
|
|
||||||
self._schema_ready = True
|
|
||||||
return self._pool
|
|
||||||
except asyncio.CancelledError:
|
except asyncio.CancelledError:
|
||||||
raise
|
raise
|
||||||
except Exception as exc:
|
except Exception as exc:
|
||||||
|
# 池建不出来 = 确定写不进去;且每次调用重试都要内联吞掉 connect
|
||||||
|
# 超时,而遥测是业务路径上的 await —— 此处必须永久降级
|
||||||
self._failed = True
|
self._failed = True
|
||||||
logger.warning("Postgres 遥测初始化失败,后续记录降级为 no-op: {}", exc)
|
logger.warning("Postgres 遥测建池失败,后续记录降级为 no-op: {}", exc)
|
||||||
return None
|
return None
|
||||||
|
return self._pool
|
||||||
|
|
||||||
|
async def _prepare_schema(self, pool: asyncpg.Pool) -> asyncpg.Pool | None:
|
||||||
|
"""备好表并交回可用的池;瞬时失败只跳过本次,确定写不进去才判死。"""
|
||||||
|
try:
|
||||||
|
async with pool.acquire() as conn:
|
||||||
|
writable = await self._prepare_table(conn)
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
# 池已在手,取连接/探测失败多为瞬时抖动: 不判死也不标就绪,
|
||||||
|
# 只跳过本次记录,下次调用重新准备
|
||||||
|
logger.warning("Postgres 遥测建表探测失败(跳过本条,下次重试): {}", exc)
|
||||||
|
return None
|
||||||
|
if not writable:
|
||||||
|
self._failed = True
|
||||||
|
return None
|
||||||
|
self._schema_ready = True
|
||||||
|
return pool
|
||||||
|
|
||||||
|
async def _prepare_table(self, conn: object) -> bool:
|
||||||
|
"""备好 `llm_calls`;**表存在就绝不发 DDL**。返回 False 仅表示表确定不存在。
|
||||||
|
|
||||||
|
`CREATE TABLE IF NOT EXISTS` 不能无条件发: PostgreSQL 对 schema 的
|
||||||
|
CREATE 权限检查**早于** `IF NOT EXISTS` 的存在性判断(PG 16.14 实测:
|
||||||
|
只授 `SELECT, INSERT ON llm_calls` 的角色,表明明在、也写得进去,这一句
|
||||||
|
照样被拒 `permission denied for schema`)。这与 `_backfill_columns` 撞的
|
||||||
|
是同一类问题(issue #3/#9),故守卫也必须同款: 先探测,后 DDL。
|
||||||
|
探测走 `to_regclass`,不需要任何权限,且与 INSERT 的 search_path 解析
|
||||||
|
口径一致——比裸 DDL 更准(裸 `CREATE TABLE` 落在首个**可建**的 schema,
|
||||||
|
可能与 INSERT 命中的不是同一张表)。
|
||||||
|
"""
|
||||||
|
exists = await conn.fetchval(_TABLE_EXISTS) is not None # type: ignore[attr-defined]
|
||||||
|
if exists:
|
||||||
|
await self._backfill_columns(conn) # 旧表可能缺列;失败只逐行降级
|
||||||
|
return True
|
||||||
|
try:
|
||||||
|
await conn.execute(_DDL) # type: ignore[attr-defined]
|
||||||
|
except asyncio.CancelledError:
|
||||||
|
raise
|
||||||
|
except Exception as exc:
|
||||||
|
logger.warning("Postgres 遥测建表失败(表不存在,记录无处可落): {}", exc)
|
||||||
|
return False
|
||||||
|
return True # 新建表列已齐全,无需再走补列
|
||||||
|
|
||||||
async def _backfill_columns(self, conn: object) -> None:
|
async def _backfill_columns(self, conn: object) -> None:
|
||||||
"""给已存在的旧表补新列(issue #3);**先探测再 ALTER,失败绝不置 `_failed`**。
|
"""给已存在的旧表补新列(issue #3);**先探测再 ALTER,失败绝不置 `_failed`**。
|
||||||
|
|||||||
@@ -3,6 +3,14 @@
|
|||||||
蓝本 VT `adapters/telemetry.py`: 构造期建连接与表,失败降级为 no-op
|
蓝本 VT `adapters/telemetry.py`: 构造期建连接与表,失败降级为 no-op
|
||||||
(记录基础设施不得拖垮业务调用);`INSERT OR IGNORE` 幂等(call_id 主键);
|
(记录基础设施不得拖垮业务调用);`INSERT OR IGNORE` 幂等(call_id 主键);
|
||||||
写入经 threading.Lock 串行化后由 `asyncio.to_thread` 执行,不阻塞事件循环。
|
写入经 threading.Lock 串行化后由 `asyncio.to_thread` 执行,不阻塞事件循环。
|
||||||
|
|
||||||
|
**这里不做 postgres.py 那样的建表前探测,是实测后的有意不对称**(issue #9):
|
||||||
|
SQLite 对已存在的表在**解析期**就把 `CREATE TABLE IF NOT EXISTS` 短路掉,
|
||||||
|
既不抢写锁也不检查可写性——实测同一时刻另一连接持 `BEGIN EXCLUSIVE`、或
|
||||||
|
文件 `chmod 444`,该语句均通过,而同条件下的 `INSERT` 与新表名建表分别报
|
||||||
|
database is locked / readonly database。故 PG 侧"权限检查早于存在性判断"
|
||||||
|
的坑在此不存在,加探测零收益。**别为了代码对称把它加回来**;需要对称的是
|
||||||
|
保证(表存在就不该因建表失败而失能),这一条两侧都已满足。
|
||||||
"""
|
"""
|
||||||
|
|
||||||
from __future__ import annotations
|
from __future__ import annotations
|
||||||
@@ -37,7 +45,8 @@ CREATE TABLE IF NOT EXISTS llm_calls (
|
|||||||
created_at TEXT NOT NULL DEFAULT (datetime('now')),
|
created_at TEXT NOT NULL DEFAULT (datetime('now')),
|
||||||
cached_prompt_tokens INTEGER,
|
cached_prompt_tokens INTEGER,
|
||||||
model_reported TEXT,
|
model_reported TEXT,
|
||||||
sampling TEXT
|
sampling TEXT,
|
||||||
|
reasoning_tokens INTEGER
|
||||||
);
|
);
|
||||||
"""
|
"""
|
||||||
|
|
||||||
@@ -47,6 +56,7 @@ _BACKFILL_COLUMNS = (
|
|||||||
("cached_prompt_tokens", "INTEGER"),
|
("cached_prompt_tokens", "INTEGER"),
|
||||||
("model_reported", "TEXT"),
|
("model_reported", "TEXT"),
|
||||||
("sampling", "TEXT"),
|
("sampling", "TEXT"),
|
||||||
|
("reasoning_tokens", "INTEGER"),
|
||||||
)
|
)
|
||||||
|
|
||||||
_COLUMNS = (
|
_COLUMNS = (
|
||||||
@@ -71,6 +81,7 @@ _COLUMNS = (
|
|||||||
"cached_prompt_tokens",
|
"cached_prompt_tokens",
|
||||||
"model_reported",
|
"model_reported",
|
||||||
"sampling",
|
"sampling",
|
||||||
|
"reasoning_tokens",
|
||||||
)
|
)
|
||||||
|
|
||||||
_INSERT = (
|
_INSERT = (
|
||||||
|
|||||||
@@ -0,0 +1,63 @@
|
|||||||
|
"""HTTP 错误响应体的取用与摘要(issue #10 设计 §3.2)。
|
||||||
|
|
||||||
|
两个 transport 各有自己的状态码分类逻辑(OCR 有意不做 429 细分),但**摘要口径
|
||||||
|
必须是同一份**——issue #10 的教训正是"只有一个分支用了响应体",一处例外就是
|
||||||
|
下一次事后查不到原因。故本模块是全库唯一的摘要实现,不得在别处复制。
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
|
||||||
|
_ERROR_BODY_CAP = 2048
|
||||||
|
"""摘要总长上限(**字符**,含省略标记在内)。
|
||||||
|
|
||||||
|
取值对齐 Kubernetes client-go `rest/request.go` 的 `maxUnstructuredResponseTextBytes
|
||||||
|
= 2048`——它是唯一与本设计同场景(读 HTTP 错误体做诊断)的成熟先例。按字符而非
|
||||||
|
字节切,多字节字符不会被切成半个;`error` 列是 TEXT,无定长约束,不需要字节口径。
|
||||||
|
"""
|
||||||
|
|
||||||
|
_HEAD_CHARS = 1400
|
||||||
|
_TAIL_CHARS = 600
|
||||||
|
|
||||||
|
|
||||||
|
def summarize_body(text: str) -> str:
|
||||||
|
"""折叠空白后按头尾策略摘要;空/空白入参返回空串。
|
||||||
|
|
||||||
|
**折叠空白**不是洁癖: 错误体常是缩进 JSON,原样拼进 message 会把一行日志
|
||||||
|
炸成多行、把遥测列变得不可读。
|
||||||
|
|
||||||
|
**保头保尾**而非头部硬切: 截断的对象是结构化 JSON,信息分布头重尾也重——
|
||||||
|
人话(`message`)在前,机器可判的 `type`/`code`/`param`/`request_id` 在后。
|
||||||
|
k8s/Sentry 用头部硬切是因为它们截的是任意文本;本函数截的是错误 JSON,
|
||||||
|
头部硬切正好切掉向网关方追查时唯一有用的那部分。策略取自标准库 `reprlib`
|
||||||
|
对"给人读的长字符串"的处置。
|
||||||
|
|
||||||
|
**标记记下省略字数**,读的人才知道自己丢了多少,不会误以为网关只说了这么多。
|
||||||
|
"""
|
||||||
|
collapsed = " ".join(text.split())
|
||||||
|
if len(collapsed) <= _ERROR_BODY_CAP:
|
||||||
|
return collapsed
|
||||||
|
omitted = len(collapsed) - _HEAD_CHARS - _TAIL_CHARS
|
||||||
|
return f"{collapsed[:_HEAD_CHARS]}…(略 {omitted} 字)…{collapsed[-_TAIL_CHARS:]}"
|
||||||
|
|
||||||
|
|
||||||
|
def compose_message(message: str, summary: str) -> str:
|
||||||
|
"""摘要非空才拼后缀,避免留下悬空的分隔符。
|
||||||
|
|
||||||
|
分隔符取 ` | ` 而非既有的 `: `,让"库说的话"与"网关说的话"一眼可分。
|
||||||
|
"""
|
||||||
|
return f"{message} | {summary}" if summary else message
|
||||||
|
|
||||||
|
|
||||||
|
def response_body(response: httpx.Response) -> str:
|
||||||
|
"""取**已缓冲**的响应文本;未读缓冲一律降级空串。
|
||||||
|
|
||||||
|
绝不在此触发网络读: 那会在错误路径上凭空插入一次可能挂住的 IO。降级方向
|
||||||
|
与缓存/遥测同档(库铁律)——诊断信息缺失不得把一次本可正确分类的失败变成
|
||||||
|
不可分类的崩溃,那正是 `ResponseNotRead` 泄漏出四分类之外的后果。
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
return response.text
|
||||||
|
except httpx.ResponseNotRead:
|
||||||
|
return ""
|
||||||
@@ -24,6 +24,11 @@ from polygateway.errors import (
|
|||||||
SourceDeadError,
|
SourceDeadError,
|
||||||
TransientError,
|
TransientError,
|
||||||
)
|
)
|
||||||
|
from polygateway.transports._http_errors import (
|
||||||
|
compose_message,
|
||||||
|
response_body,
|
||||||
|
summarize_body,
|
||||||
|
)
|
||||||
from polygateway.types import (
|
from polygateway.types import (
|
||||||
OcrLayoutElement,
|
OcrLayoutElement,
|
||||||
OcrLayoutTransportResult,
|
OcrLayoutTransportResult,
|
||||||
@@ -74,13 +79,20 @@ def _translate_http_errors(source_name: str, operation: str) -> Iterator[None]:
|
|||||||
def _classify_status(
|
def _classify_status(
|
||||||
exc: httpx.HTTPStatusError, source_name: str, operation: str
|
exc: httpx.HTTPStatusError, source_name: str, operation: str
|
||||||
) -> TransientError | SourceDeadError | RequestRejectedError:
|
) -> TransientError | SourceDeadError | RequestRejectedError:
|
||||||
|
"""HTTP 状态码 → 错误四分类,**全部分支**携带响应体摘要(issue #10)。
|
||||||
|
|
||||||
|
分类映射本身零变更;摘要口径与 chat 侧共用同一实现,不得在此另起一份——
|
||||||
|
"只有一个分支用了响应体"正是 issue #10 的成因。
|
||||||
|
"""
|
||||||
status = exc.response.status_code
|
status = exc.response.status_code
|
||||||
|
summary = summarize_body(response_body(exc.response))
|
||||||
ctx: dict[str, Any] = {
|
ctx: dict[str, Any] = {
|
||||||
"source_name": source_name,
|
"source_name": source_name,
|
||||||
"status_code": status,
|
"status_code": status,
|
||||||
"operation": operation,
|
"operation": operation,
|
||||||
|
"body_text": summary,
|
||||||
}
|
}
|
||||||
message = f"{source_name} OCR {operation} HTTP {status}"
|
message = compose_message(f"{source_name} OCR {operation} HTTP {status}", summary)
|
||||||
if status >= 500 or status == 429:
|
if status >= 500 or status == 429:
|
||||||
return TransientError(message, **ctx)
|
return TransientError(message, **ctx)
|
||||||
if status in (401, 403):
|
if status in (401, 403):
|
||||||
|
|||||||
@@ -16,13 +16,22 @@ from typing import TYPE_CHECKING, Any
|
|||||||
import httpx
|
import httpx
|
||||||
|
|
||||||
from polygateway.errors import (
|
from polygateway.errors import (
|
||||||
|
PolyGatewayError,
|
||||||
RequestRejectedError,
|
RequestRejectedError,
|
||||||
ResultInvalidError,
|
ResultInvalidError,
|
||||||
SourceDeadError,
|
SourceDeadError,
|
||||||
TransientError,
|
TransientError,
|
||||||
)
|
)
|
||||||
from polygateway.providers import ProviderProfile, get_provider
|
from polygateway.providers import (
|
||||||
|
ProviderProfile,
|
||||||
|
ThinkingCapability,
|
||||||
|
ThinkingUnsupportedError,
|
||||||
|
get_capability,
|
||||||
|
get_provider,
|
||||||
|
resolve_thinking,
|
||||||
|
)
|
||||||
from polygateway.streaming import StreamLivenessTimeout, stream_with_liveness_timeouts
|
from polygateway.streaming import StreamLivenessTimeout, stream_with_liveness_timeouts
|
||||||
|
from polygateway.transports._http_errors import compose_message, summarize_body
|
||||||
from polygateway.types import EmbeddingTransportResult, SourceConfig, TransportResult
|
from polygateway.types import EmbeddingTransportResult, SourceConfig, TransportResult
|
||||||
|
|
||||||
if TYPE_CHECKING:
|
if TYPE_CHECKING:
|
||||||
@@ -100,40 +109,59 @@ def _parse_retry_after(raw: str | None) -> float | None:
|
|||||||
return seconds if seconds > 0 else None
|
return seconds if seconds > 0 else None
|
||||||
|
|
||||||
|
|
||||||
def _translate_429(source: SourceConfig, body_text: str, headers: Mapping[str, str]) -> Exception:
|
def _translate_429(
|
||||||
|
source: SourceConfig, body_text: str, headers: Mapping[str, str], ctx: dict[str, Any]
|
||||||
|
) -> Exception:
|
||||||
|
"""429 细分。`body_text` 必须是**未截断的原文**——`ctx["body_text"]` 是摘要,
|
||||||
|
头尾保留会破坏 JSON 结构,拿它解析会让超长 body 的配额耗尽退化成普通限速
|
||||||
|
(该源不再 force_open),把一个诊断改进变成治理 bug(issue #10 实现红线)。
|
||||||
|
"""
|
||||||
try:
|
try:
|
||||||
err_type = json.loads(body_text).get("error", {}).get("type", "")
|
err_type = json.loads(body_text).get("error", {}).get("type", "")
|
||||||
except (json.JSONDecodeError, AttributeError):
|
except (json.JSONDecodeError, AttributeError):
|
||||||
err_type = ""
|
err_type = ""
|
||||||
|
summary = ctx["body_text"]
|
||||||
if err_type == "insufficient_quota":
|
if err_type == "insufficient_quota":
|
||||||
return SourceDeadError(
|
return SourceDeadError(
|
||||||
f"{source.name} 配额耗尽(insufficient_quota)",
|
compose_message(f"{source.name} 配额耗尽(insufficient_quota)", summary), **ctx
|
||||||
source_name=source.name,
|
|
||||||
status_code=429,
|
|
||||||
operation="chat",
|
|
||||||
)
|
)
|
||||||
return TransientError(
|
return TransientError(
|
||||||
f"{source.name} 限速: 429",
|
compose_message(f"{source.name} 限速: 429", summary),
|
||||||
retry_after_s=_parse_retry_after(headers.get("retry-after")),
|
retry_after_s=_parse_retry_after(headers.get("retry-after")),
|
||||||
source_name=source.name,
|
**ctx,
|
||||||
status_code=429,
|
|
||||||
operation="chat",
|
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _classify(status: int) -> tuple[type[PolyGatewayError], str]:
|
||||||
|
"""状态码 → (错误类, message 标签);映射与 ARCH §6.2 逐条相同,本次零变更。"""
|
||||||
|
if status in (401, 403):
|
||||||
|
return SourceDeadError, "凭据失效/欠费"
|
||||||
|
if status == 400:
|
||||||
|
return RequestRejectedError, "请求被拒"
|
||||||
|
if status >= 500:
|
||||||
|
return TransientError, "瞬时错误"
|
||||||
|
return RequestRejectedError, "客户端错误"
|
||||||
|
|
||||||
|
|
||||||
def _status_to_error(
|
def _status_to_error(
|
||||||
source: SourceConfig, status: int, body_text: str, headers: Mapping[str, str]
|
source: SourceConfig, status: int, body_text: str, headers: Mapping[str, str]
|
||||||
) -> Exception:
|
) -> Exception:
|
||||||
ctx: dict[str, Any] = {"source_name": source.name, "status_code": status, "operation": "chat"}
|
"""非 2xx → 领域错误,**全部分支**携带响应体摘要(issue #10)。
|
||||||
if status in (401, 403):
|
|
||||||
return SourceDeadError(f"{source.name} 凭据失效/欠费: {status}", **ctx)
|
摘要只算一次,message 与 `body_text` 共用同一份串: 两份不同长度会让"遥测里
|
||||||
if status == 400:
|
看到的"与"下游 catch 到的"对不上,排查时反而多一层困惑。
|
||||||
return RequestRejectedError(f"{source.name} 请求被拒: 400", **ctx)
|
"""
|
||||||
|
summary = summarize_body(body_text)
|
||||||
|
ctx: dict[str, Any] = {
|
||||||
|
"source_name": source.name,
|
||||||
|
"status_code": status,
|
||||||
|
"operation": "chat",
|
||||||
|
"body_text": summary,
|
||||||
|
}
|
||||||
if status == 429:
|
if status == 429:
|
||||||
return _translate_429(source, body_text, headers)
|
return _translate_429(source, body_text, headers, ctx)
|
||||||
if status >= 500:
|
cls, label = _classify(status)
|
||||||
return TransientError(f"{source.name} 瞬时错误: {status}", **ctx)
|
return cls(compose_message(f"{source.name} {label}: {status}", summary), **ctx)
|
||||||
return RequestRejectedError(f"{source.name} 客户端错误: {status}", **ctx)
|
|
||||||
|
|
||||||
|
|
||||||
def _strip_think(content: str) -> tuple[str, str]:
|
def _strip_think(content: str) -> tuple[str, str]:
|
||||||
@@ -177,6 +205,25 @@ def _coerce_cached_tokens(usage: Any) -> int | None:
|
|||||||
return cached
|
return cached
|
||||||
|
|
||||||
|
|
||||||
|
def _coerce_reasoning_tokens(usage: Any) -> int | None:
|
||||||
|
"""取 usage.completion_tokens_details.reasoning_tokens(issue #6);形态异常一律 None。
|
||||||
|
|
||||||
|
与 `_coerce_cached_tokens` 逐条同构(两者是 OpenAI 兼容 usage 里对称的一对):
|
||||||
|
`0` 如实保留、负数与非整数归 None、`bool` 显式排除。差别只在语义——本字段
|
||||||
|
的 None 是"**本次调用**未上报"而非"该源不上报": 中转在上游不返回 usage 时
|
||||||
|
会本地补算并整体替换 usage 对象,把 details 一并吃掉(findings §4c)。
|
||||||
|
"""
|
||||||
|
if not isinstance(usage, dict):
|
||||||
|
return None
|
||||||
|
details = usage.get("completion_tokens_details")
|
||||||
|
if not isinstance(details, dict):
|
||||||
|
return None
|
||||||
|
reasoning = details.get("reasoning_tokens")
|
||||||
|
if isinstance(reasoning, bool) or not isinstance(reasoning, int) or reasoning < 0:
|
||||||
|
return None
|
||||||
|
return reasoning
|
||||||
|
|
||||||
|
|
||||||
def _coerce_model_reported(value: Any) -> str | None:
|
def _coerce_model_reported(value: Any) -> str | None:
|
||||||
"""取响应体的 model 字段(issue #3);非 str 或空白串一律 None,收口时去空白。
|
"""取响应体的 model 字段(issue #3);非 str 或空白串一律 None,收口时去空白。
|
||||||
|
|
||||||
@@ -265,9 +312,14 @@ class OpenAICompatTransport:
|
|||||||
self,
|
self,
|
||||||
*,
|
*,
|
||||||
registry: Mapping[str, ProviderProfile] | None = None,
|
registry: Mapping[str, ProviderProfile] | None = None,
|
||||||
|
capabilities: Mapping[str, ThinkingCapability] | None = None,
|
||||||
client_factory: Callable[[SourceConfig], httpx.AsyncClient] | None = None,
|
client_factory: Callable[[SourceConfig], httpx.AsyncClient] | None = None,
|
||||||
) -> None:
|
) -> None:
|
||||||
self._registry = registry
|
self._registry = registry
|
||||||
|
self._capabilities = capabilities
|
||||||
|
# 未登记模型只喊一次: 装配期已喊过,逐次调用再喊是日志洪水。
|
||||||
|
# 实例级而非模块级 —— 模块级可变状态违反纯 asyncio 中立铁律
|
||||||
|
self._warned_models: set[str] = set()
|
||||||
self._client_factory = client_factory or _default_client_factory
|
self._client_factory = client_factory or _default_client_factory
|
||||||
self._clients: dict[str, httpx.AsyncClient] = {}
|
self._clients: dict[str, httpx.AsyncClient] = {}
|
||||||
|
|
||||||
@@ -290,10 +342,20 @@ class OpenAICompatTransport:
|
|||||||
payload: dict[str, Any] = {"model": source.model, "messages": messages, "stream": stream}
|
payload: dict[str, Any] = {"model": source.model, "messages": messages, "stream": stream}
|
||||||
if stream:
|
if stream:
|
||||||
payload["stream_options"] = {"include_usage": True} # 强制 usage 帧(三项目同款)
|
payload["stream_options"] = {"include_usage": True} # 强制 usage 帧(三项目同款)
|
||||||
if source.enable_thinking is True:
|
# 形态(provider 级)与能力(model 级)在此相遇;不可满足时 ValueError,
|
||||||
payload.update(profile.thinking_on)
|
# 由 complete() 翻译为四分类之一(issue #5)
|
||||||
elif source.enable_thinking is False:
|
capability = get_capability(source.model, table=self._capabilities)
|
||||||
payload.update(profile.thinking_off)
|
first_time = source.model not in self._warned_models
|
||||||
|
self._warned_models.add(source.model)
|
||||||
|
payload.update(
|
||||||
|
resolve_thinking(
|
||||||
|
profile,
|
||||||
|
capability,
|
||||||
|
source.enable_thinking,
|
||||||
|
model=source.model,
|
||||||
|
warn_unregistered=first_time,
|
||||||
|
)
|
||||||
|
)
|
||||||
# 顺序即优先级(issue #4 设计决策 A): 配置级 extra_body 在前,调用级
|
# 顺序即优先级(issue #4 设计决策 A): 配置级 extra_body 在前,调用级
|
||||||
# overlay(含结构化注入)在后覆盖之。两行不可调换
|
# overlay(含结构化注入)在后覆盖之。两行不可调换
|
||||||
payload.update(source.extra_body)
|
payload.update(source.extra_body)
|
||||||
@@ -311,9 +373,18 @@ class OpenAICompatTransport:
|
|||||||
) -> TransportResult:
|
) -> TransportResult:
|
||||||
"""一次原始调用;HTTP/线路/流式异常按 ARCH §6.2 翻译为领域错误。"""
|
"""一次原始调用;HTTP/线路/流式异常按 ARCH §6.2 翻译为领域错误。"""
|
||||||
profile = get_provider(source.provider, registry=self._registry)
|
profile = get_provider(source.provider, registry=self._registry)
|
||||||
|
try:
|
||||||
payload = self._build_payload(
|
payload = self._build_payload(
|
||||||
messages=messages, source=source, profile=profile, stream=stream, overlay=overlay
|
messages=messages, source=source, profile=profile, stream=stream, overlay=overlay
|
||||||
)
|
)
|
||||||
|
except ThinkingUnsupportedError as exc:
|
||||||
|
# 推理开关不可满足是**请求本身**的问题: 换源重试都救不了它。只捕这个
|
||||||
|
# 专用类型而非宽 catch ValueError —— 后者会把序列化等无关错误误贴标签
|
||||||
|
raise RequestRejectedError(
|
||||||
|
f"{source.name} 推理开关无法满足: {exc}",
|
||||||
|
source_name=source.name,
|
||||||
|
operation="chat",
|
||||||
|
) from exc
|
||||||
url = source.base_url.rstrip("/") + "/chat/completions"
|
url = source.base_url.rstrip("/") + "/chat/completions"
|
||||||
client = self._client_for(source)
|
client = self._client_for(source)
|
||||||
ctx: dict[str, Any] = {"source_name": source.name, "operation": "chat"}
|
ctx: dict[str, Any] = {"source_name": source.name, "operation": "chat"}
|
||||||
@@ -400,6 +471,7 @@ class OpenAICompatTransport:
|
|||||||
raw={"usage": sink.get("usage")},
|
raw={"usage": sink.get("usage")},
|
||||||
cached_prompt_tokens=_coerce_cached_tokens(sink.get("usage")),
|
cached_prompt_tokens=_coerce_cached_tokens(sink.get("usage")),
|
||||||
model_reported=_coerce_model_reported(sink.get("model")),
|
model_reported=_coerce_model_reported(sink.get("model")),
|
||||||
|
reasoning_tokens=_coerce_reasoning_tokens(sink.get("usage")),
|
||||||
)
|
)
|
||||||
|
|
||||||
def _check_done(
|
def _check_done(
|
||||||
@@ -484,6 +556,7 @@ class OpenAICompatTransport:
|
|||||||
raw={"usage": body.get("usage")},
|
raw={"usage": body.get("usage")},
|
||||||
cached_prompt_tokens=_coerce_cached_tokens(body.get("usage")),
|
cached_prompt_tokens=_coerce_cached_tokens(body.get("usage")),
|
||||||
model_reported=_coerce_model_reported(body.get("model")),
|
model_reported=_coerce_model_reported(body.get("model")),
|
||||||
|
reasoning_tokens=_coerce_reasoning_tokens(body.get("usage")),
|
||||||
)
|
)
|
||||||
|
|
||||||
async def aclose(self) -> None:
|
async def aclose(self) -> None:
|
||||||
|
|||||||
@@ -99,6 +99,15 @@ class LLMResponse:
|
|||||||
model_reported: str | None = None
|
model_reported: str | None = None
|
||||||
"""API 响应体里的 model 字段;None = 未上报。与 `model`(配置别名)可能
|
"""API 响应体里的 model 字段;None = 未上报。与 `model`(配置别名)可能
|
||||||
分叉——供应商把别名指向新权重时,实验复现必须认这个串。"""
|
分叉——供应商把别名指向新权重时,实验复现必须认这个串。"""
|
||||||
|
reasoning_tokens: int | None = None
|
||||||
|
"""推理消耗的输出 token 数(含在 `completion_tokens` 内,故不影响成本总额,
|
||||||
|
只补归因;issue #6)。
|
||||||
|
|
||||||
|
`None` = **本次调用**未上报,**不是**"该源不上报"——中转网关在上游不返回
|
||||||
|
usage 时会用本地 tokenizer 补算并整体替换 usage 对象,把
|
||||||
|
`completion_tokens_details` 一并吃掉(findings §4c 实测同一请求 10 轮呈
|
||||||
|
6:4 双峰)。实测三家供应商在未推理时都是整个 details 缺失、无人上报 `0`,
|
||||||
|
故下游判据须为 `in (None, 0)`,写 `== 0` 的条件永远不成立。"""
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
@@ -151,9 +160,10 @@ class TransportResult:
|
|||||||
ttft_ms: float | None
|
ttft_ms: float | None
|
||||||
max_inter_token_ms: float | None
|
max_inter_token_ms: float | None
|
||||||
raw: dict[str, Any]
|
raw: dict[str, Any]
|
||||||
# —— 可观测字段(issue #3;带默认值,非 OpenAI 兼容的 transport 可不填)——
|
# —— 可观测字段(issue #3/#6;带默认值,非 OpenAI 兼容的 transport 可不填)——
|
||||||
cached_prompt_tokens: int | None = None
|
cached_prompt_tokens: int | None = None
|
||||||
model_reported: str | None = None
|
model_reported: str | None = None
|
||||||
|
reasoning_tokens: int | None = None
|
||||||
|
|
||||||
|
|
||||||
@dataclass(frozen=True)
|
@dataclass(frozen=True)
|
||||||
|
|||||||
@@ -0,0 +1,444 @@
|
|||||||
|
"""真实 API 验证推理开关与 reasoning_tokens(issue #5 + #6)。
|
||||||
|
|
||||||
|
本组用例**必须真跑**: 改动的正确性与具体模型强相关,mock 只能验证代码路径,
|
||||||
|
验证不了"这个参数在这个模型上到底关没关掉推理"。
|
||||||
|
|
||||||
|
两条判据纪律(来自 findings §4c 的实测教训):
|
||||||
|
|
||||||
|
1. **判别量只能是 `reasoning_tokens`,不能是 `completion_tokens`。** 两档的输出
|
||||||
|
长度分布**是重叠的**: 实测关闭档最高 46 token(模型偶尔把解题过程写进正文),
|
||||||
|
开启档最低 13 token(medium 档想得少的那几轮),按长度阈值判两边都会误判。
|
||||||
|
而 `reasoning_tokens` 在同一批 30 轮里干净分开——关闭 15/15 为 None,
|
||||||
|
开启 15/15 大于 0。
|
||||||
|
2. **另配一个不含魔数的确定性锚点**(见 L2b): 同一模型上,关闭档的
|
||||||
|
`prompt_tokens` 严格小于开启档——供应商在开启时注入了推理指令,输入侧
|
||||||
|
token 数随之变大。这是相对比较,不硬编码任何具体数值。
|
||||||
|
3. **关闭方向要求每轮满足,开启方向只要求多数轮满足。** 中转在上游不返回
|
||||||
|
usage 时会本地补算并吃掉 `completion_tokens_details`(findings §4c),
|
||||||
|
开启方向因此可能偶尔观测不到;关闭方向不受影响。
|
||||||
|
|
||||||
|
源不可用一律 `skip` 并在报告中记为「未覆盖」,**绝不静默计入通过**。
|
||||||
|
"""
|
||||||
|
|
||||||
|
import dataclasses
|
||||||
|
import json
|
||||||
|
import os
|
||||||
|
from collections import Counter
|
||||||
|
from datetime import datetime
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
from dotenv import dotenv_values
|
||||||
|
|
||||||
|
from polygateway import GatewayClient, GatewaySettings
|
||||||
|
from polygateway.errors import (
|
||||||
|
AllSourcesExhausted,
|
||||||
|
RequestRejectedError,
|
||||||
|
SourceDeadError,
|
||||||
|
TransientError,
|
||||||
|
)
|
||||||
|
from polygateway.providers import DEFAULT_CAPABILITIES, get_capability
|
||||||
|
|
||||||
|
_ENV = {k: v for k, v in {**dotenv_values(".env"), **os.environ}.items() if v is not None}
|
||||||
|
_HAS_SOURCE = any(k.split("__")[0] == "LLM" and k.endswith("__API_KEY") for k in _ENV)
|
||||||
|
|
||||||
|
# slow: 本组 137 次真实调用、约 7 分钟,且判据是统计性的——网络抖动会让它偶发
|
||||||
|
# 失败(实测有一次 network_error 连续三次耗尽源)。让它阻断 `make ci` 会把测试
|
||||||
|
# 变成噪声源,故沿用项目既有的 slow 标记默认排除,合并前用 `-m slow` 显式真跑并
|
||||||
|
# 存档报告。"不自动门控"不等于"可跳过"。
|
||||||
|
pytestmark = [
|
||||||
|
pytest.mark.slow,
|
||||||
|
pytest.mark.skipif(
|
||||||
|
not _HAS_SOURCE, reason="需真实网关凭据: 在 .env 配置 LLM__{PROVIDER}__1__*(本组必须真跑)"
|
||||||
|
),
|
||||||
|
]
|
||||||
|
|
||||||
|
_OUT_DIR = Path("tests/outputs/e2e")
|
||||||
|
_ROUNDS = int(os.environ.get("PGW_E2E_THINKING_ROUNDS", "10"))
|
||||||
|
|
||||||
|
# 需要一点推理才能答对,但答案极短: 关掉推理时 completion 稳定在个位数,
|
||||||
|
# 开着时则是几百——两档之间隔着一个数量级,判据不必卡在噪声里
|
||||||
|
_PROMPT = "一个笼子里有若干鸡和兔,共 35 个头、94 只脚。鸡和兔各有多少只?只输出两个数字。"
|
||||||
|
|
||||||
|
_ON_MIN_COMPLETION = 100
|
||||||
|
"""仅用于 `reasoning_tokens` 被中转吃掉时的退路;关闭方向不设长度门(见 `_reasoning_off`)。"""
|
||||||
|
|
||||||
|
_ROWS: list[dict] = []
|
||||||
|
|
||||||
|
# 显式映射,不按模型名猜 provider —— 那正是 D11 要消灭的东西(providers.py 开篇)。
|
||||||
|
# 漏登记会被 test_every_capability_has_a_provider_mapping 当场抓住,而不是
|
||||||
|
# 在 L8 里被"源不可用"这个假理由吞掉
|
||||||
|
_MODEL_PROVIDER = {
|
||||||
|
"MiniMax-M3": "minimax",
|
||||||
|
"MiniMax-M2.7": "minimax",
|
||||||
|
"MiniMax-M2.5": "minimax",
|
||||||
|
"qwen3.7-plus": "qwen",
|
||||||
|
"deepseek-v4-pro": "deepseek",
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def _base_settings() -> GatewaySettings:
|
||||||
|
# 强制关缓存: 多轮测量要求每一轮都真的打到供应商,命中缓存会把后续轮次
|
||||||
|
# 变成对第一轮的回放,整组判据随之失效
|
||||||
|
return GatewaySettings.from_env("LLM", env={**_ENV, "PGW_CACHE_BACKEND": "none"})
|
||||||
|
|
||||||
|
|
||||||
|
def _settings(**source_overrides) -> GatewaySettings:
|
||||||
|
base = _base_settings()
|
||||||
|
source = dataclasses.replace(base.sources[0], **source_overrides)
|
||||||
|
return dataclasses.replace(base, sources=(source,))
|
||||||
|
|
||||||
|
|
||||||
|
async def _run_rounds(rounds: int, *, stream: bool = True, **source_overrides) -> list[dict]:
|
||||||
|
"""跑 N 轮真实调用,返回逐轮观测;任一轮抛错即向上冒泡由用例决定处置。"""
|
||||||
|
client = GatewayClient.from_settings(_settings(**source_overrides))
|
||||||
|
observations = []
|
||||||
|
try:
|
||||||
|
for i in range(rounds):
|
||||||
|
resp = await client.chat(
|
||||||
|
[{"role": "user", "content": _PROMPT}],
|
||||||
|
stream=stream,
|
||||||
|
# 每轮独立 salt: 即便某层缓存意外开着也不会回放
|
||||||
|
cache_salt=f"thinking-live-{i}",
|
||||||
|
)
|
||||||
|
observations.append(
|
||||||
|
{
|
||||||
|
"round": i + 1,
|
||||||
|
"prompt_tokens": resp.prompt_tokens,
|
||||||
|
"completion_tokens": resp.completion_tokens,
|
||||||
|
"reasoning_tokens": resp.reasoning_tokens,
|
||||||
|
"content": resp.content[:60],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
finally:
|
||||||
|
await client.aclose()
|
||||||
|
return observations
|
||||||
|
|
||||||
|
|
||||||
|
def _record(matrix_id: str, desc: str, status: str, detail, observations=None) -> None:
|
||||||
|
_ROWS.append(
|
||||||
|
{
|
||||||
|
"matrix": matrix_id,
|
||||||
|
"desc": desc,
|
||||||
|
"status": status,
|
||||||
|
"detail": detail,
|
||||||
|
"observations": observations or [],
|
||||||
|
}
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _reasoning_off(obs: dict) -> bool:
|
||||||
|
"""关闭方向: 只看 reasoning_tokens。
|
||||||
|
|
||||||
|
**刻意不设 completion_tokens 上限**: 实测关闭档偶尔会到 46 token(模型没照做
|
||||||
|
"只输出两个数字",把解题过程写进了正文),而那是正文不是推理。加长度门只会
|
||||||
|
把这种正常波动误判成"没关掉"。
|
||||||
|
"""
|
||||||
|
return obs["reasoning_tokens"] in (None, 0)
|
||||||
|
|
||||||
|
|
||||||
|
def _reasoning_on(obs: dict) -> bool:
|
||||||
|
"""开启方向: 有 reasoning_tokens 就以它为准,它是本次改动引入的直接判据。
|
||||||
|
|
||||||
|
不能拿 completion_tokens 当开启方向的主判据: medium 档的推理量方差极大
|
||||||
|
(实测 15 轮跨 7-170 token),按长度阈值判会把"推理了但想得少"误判成没推理。
|
||||||
|
仅当中转吃掉了 ctd(reasoning_tokens is None)才退回长度判据。
|
||||||
|
"""
|
||||||
|
reasoning = obs["reasoning_tokens"]
|
||||||
|
if reasoning is not None:
|
||||||
|
return reasoning > 0
|
||||||
|
return obs["completion_tokens"] > _ON_MIN_COMPLETION
|
||||||
|
|
||||||
|
|
||||||
|
def _skip_if_unreachable(exc: Exception, matrix_id: str, desc: str):
|
||||||
|
"""源不可用(渠道下线/模型未开通)→ 跳过并记为未覆盖,不伪装成通过。"""
|
||||||
|
_record(matrix_id, desc, "SKIP(源不可用)", str(exc)[:200])
|
||||||
|
pytest.skip(f"{matrix_id} 源不可用,已记为未覆盖: {str(exc)[:120]}")
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture(scope="module", autouse=True)
|
||||||
|
def _write_report():
|
||||||
|
yield
|
||||||
|
_OUT_DIR.mkdir(parents=True, exist_ok=True)
|
||||||
|
ts = datetime.now().strftime("%Y%m%d_%H%M%S")
|
||||||
|
path = _OUT_DIR / f"test_thinking_live_{ts}.md"
|
||||||
|
lines = [
|
||||||
|
"# 推理开关与 reasoning_tokens 真实 API 验证",
|
||||||
|
"",
|
||||||
|
f"- 时间: {ts}",
|
||||||
|
f"- 每档轮数: {_ROUNDS}",
|
||||||
|
"- 关闭判据: **每轮** reasoning_tokens in (None, 0);刻意不设输出长度上限"
|
||||||
|
"(两档的 completion 分布重叠: 实测关闭档最高 46、开启档最低 13)",
|
||||||
|
f"- 开启判据: **多数轮** reasoning_tokens > 0(被中转吃掉时退回 completion > {_ON_MIN_COMPLETION})",
|
||||||
|
"- 确定性锚点(L2b): 关闭档 prompt_tokens 最大值 < 开启档最小值,相对比较无魔数",
|
||||||
|
"",
|
||||||
|
"## 矩阵结论",
|
||||||
|
"",
|
||||||
|
"| 矩阵 | 场景 | 结论 | 说明 |",
|
||||||
|
"|---|---|---|---|",
|
||||||
|
]
|
||||||
|
total_calls = 0
|
||||||
|
for row in _ROWS:
|
||||||
|
detail = str(row["detail"]).replace("|", "\\|").replace("\n", " ")[:160]
|
||||||
|
lines.append(f"| {row['matrix']} | {row['desc']} | {row['status']} | {detail} |")
|
||||||
|
total_calls += len(row["observations"])
|
||||||
|
lines += ["", f"**总真实调用次数: {total_calls}**", "", "## 逐轮原始观测", ""]
|
||||||
|
for row in _ROWS:
|
||||||
|
if not row["observations"]:
|
||||||
|
continue
|
||||||
|
lines += [f"### {row['matrix']} — {row['desc']}", "", "```json"]
|
||||||
|
lines.append(json.dumps(row["observations"], ensure_ascii=False, indent=2))
|
||||||
|
lines += ["```", ""]
|
||||||
|
uncovered = [r["matrix"] for r in _ROWS if r["status"].startswith("SKIP")]
|
||||||
|
if uncovered:
|
||||||
|
lines += ["## 未覆盖", "", f"以下矩阵行未跑到: {', '.join(uncovered)}", ""]
|
||||||
|
path.write_text("\n".join(lines), encoding="utf-8")
|
||||||
|
print(f"\n[e2e 报告] {path}")
|
||||||
|
|
||||||
|
|
||||||
|
class TestMiniMaxM3:
|
||||||
|
"""M3 是唯一实测可关闭推理的 MiniMax 模型,修复的地基压在它身上。"""
|
||||||
|
|
||||||
|
async def test_l1_disable_actually_disables(self):
|
||||||
|
obs = await _run_rounds(_ROUNDS, model="MiniMax-M3", enable_thinking=False)
|
||||||
|
offs = [o for o in obs if _reasoning_off(o)]
|
||||||
|
_record(
|
||||||
|
"L1",
|
||||||
|
"enable_thinking=False(流式)",
|
||||||
|
"PASS" if len(offs) == len(obs) else "FAIL",
|
||||||
|
f"{len(offs)}/{len(obs)} 轮确认未推理",
|
||||||
|
obs,
|
||||||
|
)
|
||||||
|
assert len(offs) == len(obs), f"关闭方向要求每轮满足: {obs}"
|
||||||
|
|
||||||
|
async def test_l2_enable_actually_enables(self):
|
||||||
|
obs = await _run_rounds(_ROUNDS, model="MiniMax-M3", enable_thinking=True)
|
||||||
|
ons = [o for o in obs if _reasoning_on(o)]
|
||||||
|
_record(
|
||||||
|
"L2",
|
||||||
|
"enable_thinking=True(流式,注入 medium)",
|
||||||
|
"PASS" if len(ons) * 2 > len(obs) else "FAIL",
|
||||||
|
f"{len(ons)}/{len(obs)} 轮观察到推理",
|
||||||
|
obs,
|
||||||
|
)
|
||||||
|
assert len(ons) * 2 > len(obs), f"开启方向要求多数轮满足: {obs}"
|
||||||
|
|
||||||
|
async def test_l2b_off_and_on_are_distinguishable_without_magic_numbers(self):
|
||||||
|
"""确定性锚点: 开启档的 prompt_tokens 严格大于关闭档。
|
||||||
|
|
||||||
|
供应商在开启推理时会向模板注入推理指令,输入侧 token 数随之变大。这是
|
||||||
|
本组唯一不依赖输出侧噪声的证据,且是相对比较——不硬编码任何具体数值,
|
||||||
|
供应商改模板也不会让它假红。
|
||||||
|
"""
|
||||||
|
rounds = max(3, _ROUNDS // 3)
|
||||||
|
off = await _run_rounds(rounds, model="MiniMax-M3", enable_thinking=False)
|
||||||
|
on = await _run_rounds(rounds, model="MiniMax-M3", enable_thinking=True)
|
||||||
|
off_max = max(o["prompt_tokens"] for o in off)
|
||||||
|
on_min = min(o["prompt_tokens"] for o in on)
|
||||||
|
_record(
|
||||||
|
"L2b",
|
||||||
|
"关闭/开启的 prompt_tokens 可分",
|
||||||
|
"PASS" if off_max < on_min else "FAIL",
|
||||||
|
f"关闭档最大 {off_max} < 开启档最小 {on_min}",
|
||||||
|
off + on,
|
||||||
|
)
|
||||||
|
assert off_max < on_min, (
|
||||||
|
f"两档的 prompt_tokens 未分开(关闭最大 {off_max},开启最小 {on_min}): 注入可能没到达模型"
|
||||||
|
)
|
||||||
|
|
||||||
|
async def test_l3_no_opinion_is_the_model_default(self):
|
||||||
|
obs = await _run_rounds(_ROUNDS, model="MiniMax-M3", enable_thinking=None)
|
||||||
|
# M3 的默认档实测就是不推理(findings §2.1),所以不干预时也应观测不到推理。
|
||||||
|
# 注意这**不能**反过来证明关闭方向生效 —— L1 与本行同分布,区分二者的是
|
||||||
|
# L2b 的 prompt_tokens 与 L3b 的乱码值反证
|
||||||
|
quiet = [o for o in obs if _reasoning_off(o)]
|
||||||
|
_record(
|
||||||
|
"L3",
|
||||||
|
"enable_thinking=None(不干预,基线)",
|
||||||
|
"PASS" if len(quiet) == len(obs) else "FAIL",
|
||||||
|
f"{len(quiet)}/{len(obs)} 轮未推理(M3 默认档本就不推理)",
|
||||||
|
obs,
|
||||||
|
)
|
||||||
|
assert len(quiet) == len(obs), f"M3 默认档不应推理: {obs}"
|
||||||
|
|
||||||
|
async def test_l3b_none_is_recognised_not_silently_dropped(self):
|
||||||
|
"""反证: 关闭方向的观测必须排除"参数被静默丢弃"这一伪解释。
|
||||||
|
|
||||||
|
L1(关闭)与 L3(不干预)在 M3 上**同分布**——因为 M3 默认档本就不推理。
|
||||||
|
所以 L1 单独看不能区分"`none` 真的被消费"与"`none` 被中转吞了",而后者
|
||||||
|
正是 issue #5 的原始故障形态(`enable_thinking` 就是这么被吞的)。
|
||||||
|
|
||||||
|
判别方法: 发一个**非法值**。若未知值会被静默丢弃,它的表现应与"不注入"
|
||||||
|
一致(不推理);实测它反而开启了推理,说明网关认这个键、只是不认这个值。
|
||||||
|
既然非法值与 `none` 的表现不同,`none` 就必然是被识别的枚举值。
|
||||||
|
"""
|
||||||
|
rounds = max(3, _ROUNDS // 3)
|
||||||
|
bogus = await _run_rounds(
|
||||||
|
rounds,
|
||||||
|
model="MiniMax-M3",
|
||||||
|
enable_thinking=None,
|
||||||
|
extra_body={"reasoning_effort": "definitely-not-a-real-level"},
|
||||||
|
)
|
||||||
|
off = await _run_rounds(rounds, model="MiniMax-M3", enable_thinking=False)
|
||||||
|
bogus_on = [o for o in bogus if _reasoning_on(o)]
|
||||||
|
off_quiet = [o for o in off if _reasoning_off(o)]
|
||||||
|
ok = len(bogus_on) * 2 > len(bogus) and len(off_quiet) == len(off)
|
||||||
|
_record(
|
||||||
|
"L3b",
|
||||||
|
"非法值反证 none 被识别",
|
||||||
|
"PASS" if ok else "FAIL",
|
||||||
|
f"非法值 {len(bogus_on)}/{len(bogus)} 轮推理,none {len(off_quiet)}/{len(off)} 轮不推理"
|
||||||
|
"(两者表现不同 ⇒ none 非被丢弃)",
|
||||||
|
bogus + off,
|
||||||
|
)
|
||||||
|
assert len(bogus_on) * 2 > len(bogus), (
|
||||||
|
f"非法值未开启推理,无法排除'未知值被静默丢弃'这一伪解释: {bogus}"
|
||||||
|
)
|
||||||
|
assert len(off_quiet) == len(off), f"none 未关闭推理: {off}"
|
||||||
|
|
||||||
|
async def test_l4_extra_body_overrides_the_profile(self):
|
||||||
|
"""profile 注入 none,extra_body 要求 high —— 后者必须赢(优先级不可调换)。
|
||||||
|
|
||||||
|
判据是行为而非报文: 若 extra_body 没赢,拿到的就是 none 的结果(不推理)。
|
||||||
|
"""
|
||||||
|
rounds = max(3, _ROUNDS // 2)
|
||||||
|
obs = await _run_rounds(
|
||||||
|
rounds,
|
||||||
|
model="MiniMax-M3",
|
||||||
|
enable_thinking=False,
|
||||||
|
extra_body={"reasoning_effort": "high"},
|
||||||
|
)
|
||||||
|
ons = [o for o in obs if _reasoning_on(o)]
|
||||||
|
_record(
|
||||||
|
"L4",
|
||||||
|
"extra_body 覆盖 profile 注入",
|
||||||
|
"PASS" if len(ons) * 2 > len(obs) else "FAIL",
|
||||||
|
f"{len(ons)}/{len(obs)} 轮观察到推理(证明 high 生效而非 none)",
|
||||||
|
obs,
|
||||||
|
)
|
||||||
|
assert len(ons) * 2 > len(obs), f"extra_body 未能覆盖 profile: {obs}"
|
||||||
|
|
||||||
|
async def test_l5_non_stream_path_matches_stream(self):
|
||||||
|
"""非流式快路径独立于流式实现,采集与注入都要各自验一遍。"""
|
||||||
|
rounds = max(3, _ROUNDS // 2)
|
||||||
|
off = await _run_rounds(rounds, stream=False, model="MiniMax-M3", enable_thinking=False)
|
||||||
|
on = await _run_rounds(rounds, stream=False, model="MiniMax-M3", enable_thinking=True)
|
||||||
|
offs = [o for o in off if _reasoning_off(o)]
|
||||||
|
ons = [o for o in on if _reasoning_on(o)]
|
||||||
|
ok = len(offs) == len(off) and len(ons) * 2 > len(on)
|
||||||
|
_record(
|
||||||
|
"L5",
|
||||||
|
"非流式路径重跑 L1/L2",
|
||||||
|
"PASS" if ok else "FAIL",
|
||||||
|
f"关闭 {len(offs)}/{len(off)} 轮,开启 {len(ons)}/{len(on)} 轮",
|
||||||
|
off + on,
|
||||||
|
)
|
||||||
|
assert len(offs) == len(off), f"非流式关闭方向未满足: {off}"
|
||||||
|
assert len(ons) * 2 > len(on), f"非流式开启方向未满足: {on}"
|
||||||
|
|
||||||
|
|
||||||
|
class TestOtherProviders:
|
||||||
|
"""qwen / deepseek 的 profile 是既有实现,本组防的是"改 minimax 时误伤它们"。"""
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("matrix", "provider", "model"),
|
||||||
|
[("L6", "qwen", "qwen3.7-plus"), ("L7", "deepseek", "deepseek-v4-pro")],
|
||||||
|
)
|
||||||
|
async def test_existing_profiles_still_disable(self, matrix, provider, model):
|
||||||
|
desc = f"{provider} enable_thinking=False"
|
||||||
|
try:
|
||||||
|
obs = await _run_rounds(_ROUNDS, provider=provider, model=model, enable_thinking=False)
|
||||||
|
except (AllSourcesExhausted, SourceDeadError, TransientError) as exc:
|
||||||
|
# 只吞网关/网络类失败。**不吞 ValueError / RequestRejected** ——
|
||||||
|
# 那两类正是本次改动最可能的误伤方向,吞掉就成了纪律(c)要防的静默
|
||||||
|
_skip_if_unreachable(exc, matrix, desc)
|
||||||
|
offs = [o for o in obs if _reasoning_off(o)]
|
||||||
|
_record(
|
||||||
|
matrix,
|
||||||
|
desc,
|
||||||
|
"PASS" if len(offs) == len(obs) else "FAIL",
|
||||||
|
f"{len(offs)}/{len(obs)} 轮确认未推理",
|
||||||
|
obs,
|
||||||
|
)
|
||||||
|
assert len(offs) == len(obs), f"{provider} 关闭方向未满足: {obs}"
|
||||||
|
|
||||||
|
|
||||||
|
class TestCapabilityDrift:
|
||||||
|
"""L8 漂移哨兵: 能力表过期是必然事件,这里是它的过期告警。"""
|
||||||
|
|
||||||
|
def test_every_capability_has_a_provider_mapping(self):
|
||||||
|
"""能力表新增条目必须同步本测试的映射,否则该行会被静默跳过。"""
|
||||||
|
missing = sorted(set(DEFAULT_CAPABILITIES) - set(_MODEL_PROVIDER))
|
||||||
|
assert not missing, f"这些模型缺 provider 映射,L8 会漏测: {missing}"
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("model", sorted(DEFAULT_CAPABILITIES))
|
||||||
|
async def test_declared_capability_matches_reality(self, model):
|
||||||
|
cap = get_capability(model)
|
||||||
|
provider = _MODEL_PROVIDER[model]
|
||||||
|
rounds = max(3, _ROUNDS // 2)
|
||||||
|
desc = f"{model} 声明 can_disable={cap.can_disable}"
|
||||||
|
if not cap.can_disable:
|
||||||
|
# 声明关不掉: 装配期就该炸,炸了即与声明一致(不必真调用)
|
||||||
|
with pytest.raises(ValueError, match=model):
|
||||||
|
GatewayClient.from_settings(
|
||||||
|
_settings(provider=provider, model=model, enable_thinking=False)
|
||||||
|
)
|
||||||
|
_record("L8", desc, "PASS", "装配期按声明拒绝,与实测一致")
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
obs = await _run_rounds(rounds, provider=provider, model=model, enable_thinking=False)
|
||||||
|
except (AllSourcesExhausted, SourceDeadError, TransientError) as exc:
|
||||||
|
_skip_if_unreachable(exc, "L8", desc)
|
||||||
|
offs = [o for o in obs if _reasoning_off(o)]
|
||||||
|
verdict = Counter(_reasoning_off(o) for o in obs)
|
||||||
|
_record(
|
||||||
|
"L8",
|
||||||
|
desc,
|
||||||
|
"PASS" if len(offs) == len(obs) else "FAIL(能力表已漂移)",
|
||||||
|
f"实测 {dict(verdict)};声明 can_disable=True 要求每轮关闭",
|
||||||
|
obs,
|
||||||
|
)
|
||||||
|
assert len(offs) == len(obs), (
|
||||||
|
f"能力表漂移: {model} 声明可关闭推理,实测未关掉 —— 请复测后更新 DEFAULT_CAPABILITIES"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class TestAssemblyGuardAgainstRealConfig:
|
||||||
|
"""L9: 纯本地,但用的是 .env 里的真实配置形态,防"守卫只在合成配置上生效"。"""
|
||||||
|
|
||||||
|
def test_l9_m27_rejected_at_assembly(self):
|
||||||
|
with pytest.raises(ValueError, match="MiniMax-M2.7"):
|
||||||
|
GatewayClient.from_settings(
|
||||||
|
_settings(provider="minimax", model="MiniMax-M2.7", enable_thinking=False)
|
||||||
|
)
|
||||||
|
_record("L9", "M2.7 + enable_thinking=False", "PASS", "装配期报错,未发出任何请求")
|
||||||
|
|
||||||
|
def test_l9_unknown_shape_rejected_at_assembly(self):
|
||||||
|
with pytest.raises(ValueError, match="register_provider"):
|
||||||
|
GatewayClient.from_settings(
|
||||||
|
_settings(provider="openai", model="kimi-k3", enable_thinking=False)
|
||||||
|
)
|
||||||
|
_record("L9", "provider=openai 形态未知", "PASS", "装配期报错并指路")
|
||||||
|
|
||||||
|
async def test_transport_layer_rejects_when_guard_is_bypassed(self):
|
||||||
|
"""构造函数全量注入这条路绕过装配守卫,transport 必须兜住并归四分类。"""
|
||||||
|
settings = _settings(provider="minimax", model="MiniMax-M2.7", enable_thinking=False)
|
||||||
|
client = GatewayClient.from_settings(
|
||||||
|
dataclasses.replace(
|
||||||
|
settings, sources=(dataclasses.replace(settings.sources[0], enable_thinking=None),)
|
||||||
|
)
|
||||||
|
)
|
||||||
|
try:
|
||||||
|
# 装配用 None 绕过守卫,再把源换成 False 直接喂给 transport
|
||||||
|
bad = dataclasses.replace(settings.sources[0], enable_thinking=False)
|
||||||
|
with pytest.raises(RequestRejectedError, match="MiniMax-M2.7"):
|
||||||
|
await client._terminal._transport.complete(
|
||||||
|
messages=[{"role": "user", "content": _PROMPT}],
|
||||||
|
source=bad,
|
||||||
|
stream=True,
|
||||||
|
overlay={},
|
||||||
|
call_id="e2e-guard",
|
||||||
|
)
|
||||||
|
finally:
|
||||||
|
await client.aclose()
|
||||||
|
_record("L9", "绕过装配守卫时 transport 兜底", "PASS", "RequestRejectedError,属四分类")
|
||||||
@@ -11,7 +11,12 @@ import sqlite3
|
|||||||
import httpx
|
import httpx
|
||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
from polygateway import CircuitOpenError, GatewayClient, TransientError
|
from polygateway import (
|
||||||
|
CircuitOpenError,
|
||||||
|
GatewayClient,
|
||||||
|
RequestRejectedError,
|
||||||
|
TransientError,
|
||||||
|
)
|
||||||
from polygateway.backends.memory.breaker import InMemoryGate
|
from polygateway.backends.memory.breaker import InMemoryGate
|
||||||
from polygateway.backends.memory.cache import InMemoryCache
|
from polygateway.backends.memory.cache import InMemoryCache
|
||||||
from polygateway.backends.memory.limiter import InMemoryLimiter
|
from polygateway.backends.memory.limiter import InMemoryLimiter
|
||||||
@@ -173,6 +178,40 @@ class TestTelemetryAcrossPaths:
|
|||||||
assert len({cid for _, cid in rows}) == 2 # call_id 逐次独立
|
assert len({cid for _, cid in rows}) == 2 # call_id 逐次独立
|
||||||
|
|
||||||
|
|
||||||
|
class TestRejectionReasonIsQueryable:
|
||||||
|
"""issue #10 的验收主张: 400 之后,网关说的话必须能在遥测表里查到。
|
||||||
|
|
||||||
|
下游一轮 1050 张影像的批处理里,1 张在读表格时收到 400 被判确定性失败,
|
||||||
|
事后"这张图到底哪里不合规"无从查起——响应体在 transport 翻译层就没了。
|
||||||
|
"""
|
||||||
|
|
||||||
|
# issue #10 原文给出的真实响应体(一字不改)
|
||||||
|
_BODY = (
|
||||||
|
'{"error":{"message":"<400> ***.***.InvalidParameter: The image format is illegal '
|
||||||
|
'and cannot be opened","type":"invalid_request_error","param":"",'
|
||||||
|
'"code":"invalid_parameter_error"}}'
|
||||||
|
)
|
||||||
|
|
||||||
|
async def test_rejected_call_leaves_the_reason_in_telemetry(self, tmp_path):
|
||||||
|
recorder = SQLiteRecorder(tmp_path / "t.db")
|
||||||
|
client = _full_client(
|
||||||
|
lambda req: httpx.Response(400, content=self._BODY.encode()), telemetry=recorder
|
||||||
|
)
|
||||||
|
|
||||||
|
with pytest.raises(RequestRejectedError):
|
||||||
|
await client.chat([{"role": "user", "content": "hi"}])
|
||||||
|
|
||||||
|
recorder.close()
|
||||||
|
rows = sqlite3.connect(tmp_path / "t.db").execute("SELECT error FROM llm_calls").fetchall()
|
||||||
|
assert rows, "400 必须留下遥测行(遥测必录)"
|
||||||
|
errors = " ".join(r[0] or "" for r in rows)
|
||||||
|
# 修复前这里只有 "qwen_1 请求被拒: 400"——诊断信息一个字都不在
|
||||||
|
assert "InvalidParameter" in errors
|
||||||
|
assert "The image format is illegal" in errors
|
||||||
|
# 尾部的 code 才是向网关方追查的凭据,头部硬切正好会丢掉它
|
||||||
|
assert "invalid_parameter_error" in errors
|
||||||
|
|
||||||
|
|
||||||
class TestStructuredThroughStack:
|
class TestStructuredThroughStack:
|
||||||
async def test_feedback_reask_passes_through_governance(self):
|
async def test_feedback_reask_passes_through_governance(self):
|
||||||
"""重问经过内层治理: 第二次真实请求同样被限流/熔断记账。"""
|
"""重问经过内层治理: 第二次真实请求同样被限流/熔断记账。"""
|
||||||
|
|||||||
@@ -13,6 +13,7 @@ from __future__ import annotations
|
|||||||
import asyncio
|
import asyncio
|
||||||
import json
|
import json
|
||||||
import os
|
import os
|
||||||
|
import re
|
||||||
from uuid import uuid4
|
from uuid import uuid4
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
@@ -43,6 +44,7 @@ _EXPECTED_COLUMNS = [
|
|||||||
"cached_prompt_tokens",
|
"cached_prompt_tokens",
|
||||||
"model_reported",
|
"model_reported",
|
||||||
"sampling",
|
"sampling",
|
||||||
|
"reasoning_tokens",
|
||||||
]
|
]
|
||||||
|
|
||||||
# run 级前缀: 同库并存的其他运行(迁移批跑/另一开发机)互不可见
|
# run 级前缀: 同库并存的其他运行(迁移批跑/另一开发机)互不可见
|
||||||
@@ -107,6 +109,7 @@ async def _record_minimal(
|
|||||||
"cached_prompt_tokens": None,
|
"cached_prompt_tokens": None,
|
||||||
"model_reported": None,
|
"model_reported": None,
|
||||||
"sampling": None,
|
"sampling": None,
|
||||||
|
"reasoning_tokens": None,
|
||||||
}
|
}
|
||||||
fields.update(overrides)
|
fields.update(overrides)
|
||||||
await recorder.record_llm_call(**fields)
|
await recorder.record_llm_call(**fields)
|
||||||
@@ -297,3 +300,88 @@ class TestDegradation:
|
|||||||
await _record_minimal(recorder)
|
await _record_minimal(recorder)
|
||||||
await recorder.aclose()
|
await recorder.aclose()
|
||||||
await recorder.aclose()
|
await recorder.aclose()
|
||||||
|
|
||||||
|
|
||||||
|
_PROBE_PASSWORD = "pgw_issue9_probe" # 临时角色,teardown 删除;非任何真实凭据
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
async def least_privilege_dsn(dsn):
|
||||||
|
"""临时 schema + 临时角色: 只授表级 SELECT/INSERT,**不授 schema CREATE**。
|
||||||
|
|
||||||
|
这是 issue #9 的现场——最小权限部署的标准形态。fixture 建的一切
|
||||||
|
(schema、表、角色)都在 teardown 里删净,共享的 public.llm_calls 不受影响;
|
||||||
|
连不上或无权建角色(非超级用户)时 skip,不让 CI 假绿。
|
||||||
|
"""
|
||||||
|
import asyncpg
|
||||||
|
|
||||||
|
from polygateway.telemetry.postgres import _DDL
|
||||||
|
|
||||||
|
name = f"pgwtest_lp_{uuid4().hex[:8]}"
|
||||||
|
admin = await asyncpg.connect(dsn, timeout=10)
|
||||||
|
try:
|
||||||
|
if not await admin.fetchval(
|
||||||
|
"SELECT rolcreaterole OR rolsuper FROM pg_roles WHERE rolname = current_user"
|
||||||
|
):
|
||||||
|
pytest.skip("当前账号无权建临时角色,跳过最小权限用例")
|
||||||
|
await admin.execute(f"CREATE ROLE {name} LOGIN PASSWORD '{_PROBE_PASSWORD}'")
|
||||||
|
await admin.execute(f"CREATE SCHEMA {name}")
|
||||||
|
await admin.execute(f"SET search_path = {name}")
|
||||||
|
await admin.execute(_DDL) # 表由**别的账号**建好,与现场一致
|
||||||
|
await admin.execute(f"GRANT USAGE ON SCHEMA {name} TO {name}")
|
||||||
|
await admin.execute(f"GRANT SELECT, INSERT ON {name}.llm_calls TO {name}")
|
||||||
|
# 关键: 绝不 GRANT CREATE ON SCHEMA —— 缺的正是这一项
|
||||||
|
finally:
|
||||||
|
await admin.close()
|
||||||
|
low = re.sub(r"//[^@/]+@", f"//{name}:{_PROBE_PASSWORD}@", dsn, count=1)
|
||||||
|
sep = "&" if "?" in low else "?"
|
||||||
|
yield f"{low}{sep}options=-csearch_path%3D{name}", name
|
||||||
|
admin = await asyncpg.connect(dsn, timeout=10)
|
||||||
|
try:
|
||||||
|
await admin.execute(f"DROP SCHEMA IF EXISTS {name} CASCADE")
|
||||||
|
await admin.execute(f"DROP OWNED BY {name}")
|
||||||
|
await admin.execute(f"DROP ROLE IF EXISTS {name}")
|
||||||
|
finally:
|
||||||
|
await admin.close()
|
||||||
|
|
||||||
|
|
||||||
|
class TestLeastPrivilegeDeployment:
|
||||||
|
"""issue #9: 只有表级写权限的账号,遥测必须照常落库而不是整体判死。"""
|
||||||
|
|
||||||
|
async def test_create_table_if_not_exists_is_denied_for_this_role(self, least_privilege_dsn):
|
||||||
|
"""库外事实先钉死: 表存在、写得进去,DDL 仍被拒——PG 的权限检查早于 IF NOT EXISTS。
|
||||||
|
|
||||||
|
修复依赖的是这条 PG 语义;若某天它变了,这里先红,而不是让下面那条
|
||||||
|
用例悄悄变成"永远通过"的空断言。
|
||||||
|
"""
|
||||||
|
import asyncpg
|
||||||
|
|
||||||
|
low_dsn, _ = least_privilege_dsn
|
||||||
|
conn = await asyncpg.connect(low_dsn, timeout=10)
|
||||||
|
try:
|
||||||
|
assert await conn.fetchval("SELECT to_regclass('llm_calls')") is not None
|
||||||
|
with pytest.raises(asyncpg.exceptions.InsufficientPrivilegeError):
|
||||||
|
await conn.execute("CREATE TABLE IF NOT EXISTS llm_calls (call_id TEXT)")
|
||||||
|
finally:
|
||||||
|
await conn.close()
|
||||||
|
|
||||||
|
async def test_records_land_without_schema_create_privilege(self, least_privilege_dsn):
|
||||||
|
"""修复前: 建表被拒 → _failed → 整个进程一条不落(下游 150 次调用全丢)。"""
|
||||||
|
low_dsn, schema = least_privilege_dsn
|
||||||
|
recorder = PostgresRecorder(low_dsn)
|
||||||
|
try:
|
||||||
|
await _record_minimal(recorder, call_id=_cid("lp1"))
|
||||||
|
await _record_minimal(recorder, call_id=_cid("lp2"), cost=1.5)
|
||||||
|
assert recorder._failed is False # 判死开关不得被建表权限触发
|
||||||
|
rows = await _fetch(
|
||||||
|
low_dsn,
|
||||||
|
"SELECT call_id, cost FROM llm_calls WHERE call_id LIKE $1 ORDER BY call_id",
|
||||||
|
f"{_RUN_PREFIX}-lp%",
|
||||||
|
)
|
||||||
|
assert [(r["call_id"], r["cost"]) for r in rows] == [
|
||||||
|
(_cid("lp1"), None),
|
||||||
|
(_cid("lp2"), 1.5),
|
||||||
|
]
|
||||||
|
assert schema # teardown 会连表带角色删净
|
||||||
|
finally:
|
||||||
|
await recorder.aclose()
|
||||||
|
|||||||
@@ -16,7 +16,11 @@ import pytest
|
|||||||
from polygateway.backends.redis.breaker import RedisGate
|
from polygateway.backends.redis.breaker import RedisGate
|
||||||
from polygateway.backends.redis.limiter import RedisLimiter
|
from polygateway.backends.redis.limiter import RedisLimiter
|
||||||
from polygateway.client import GatewayClient
|
from polygateway.client import GatewayClient
|
||||||
from polygateway.errors import AllSourcesExhausted, GovernanceBackendError
|
from polygateway.errors import (
|
||||||
|
AllSourcesExhausted,
|
||||||
|
GatewayUnavailableError,
|
||||||
|
GovernanceBackendError,
|
||||||
|
)
|
||||||
from polygateway.sources import RoundRobinSelector
|
from polygateway.sources import RoundRobinSelector
|
||||||
from polygateway.types import (
|
from polygateway.types import (
|
||||||
BackpressurePolicy,
|
BackpressurePolicy,
|
||||||
@@ -225,7 +229,11 @@ async def test_cancel_in_flight_releases_lease(clients):
|
|||||||
|
|
||||||
|
|
||||||
async def test_redis_down_admission_fails_closed():
|
async def test_redis_down_admission_fails_closed():
|
||||||
"""Redis 不可达 → 准入侧抛 GovernanceBackendError,绝不放行(库铁律)。"""
|
"""Redis 不可达 → 准入侧报错绝不放行(库铁律),且以 scope 级形态到达调用方。
|
||||||
|
|
||||||
|
issue #7: 调用方只写 `except GatewayUnavailableError` 就该覆盖后端故障——
|
||||||
|
真实 Redis 掉线是这条链路唯一的端到端证据,故断言收紧到 scope 级语义。
|
||||||
|
"""
|
||||||
import redis.asyncio as aioredis
|
import redis.asyncio as aioredis
|
||||||
|
|
||||||
dead = aioredis.from_url(
|
dead = aioredis.from_url(
|
||||||
@@ -240,9 +248,12 @@ async def test_redis_down_admission_fails_closed():
|
|||||||
lease_ttl_s=30.0,
|
lease_ttl_s=30.0,
|
||||||
)
|
)
|
||||||
gate = RedisGate(config=_CFG, redis=dead, scope="t-dead")
|
gate = RedisGate(config=_CFG, redis=dead, scope="t-dead")
|
||||||
with pytest.raises(GovernanceBackendError):
|
for call in (limiter.try_acquire("s1", 0), gate.try_enter("s1", "w")):
|
||||||
await limiter.try_acquire("s1", 0)
|
with pytest.raises(GatewayUnavailableError) as ei:
|
||||||
with pytest.raises(GovernanceBackendError):
|
await call
|
||||||
await gate.try_enter("s1", "w")
|
assert isinstance(ei.value, GovernanceBackendError)
|
||||||
|
assert ei.value.reason == "governance_backend_down"
|
||||||
|
assert ei.value.scope == "t-dead"
|
||||||
|
assert ei.value.retry_after_s > 0
|
||||||
finally:
|
finally:
|
||||||
await dead.aclose()
|
await dead.aclose()
|
||||||
|
|||||||
+352
-10
@@ -11,7 +11,14 @@ import pytest
|
|||||||
|
|
||||||
from polygateway.backends.memory.breaker import InMemoryGate
|
from polygateway.backends.memory.breaker import InMemoryGate
|
||||||
from polygateway.backends.memory.limiter import InMemoryLimiter
|
from polygateway.backends.memory.limiter import InMemoryLimiter
|
||||||
from polygateway.errors import AllSourcesExhausted, GovernanceBackendError, TransientError
|
from polygateway.errors import (
|
||||||
|
AllSourcesExhausted,
|
||||||
|
GatewayUnavailableError,
|
||||||
|
GovernanceBackendError,
|
||||||
|
SourceNotConfiguredError,
|
||||||
|
TransientError,
|
||||||
|
)
|
||||||
|
from polygateway.middleware.ratelimit import QuotaGate
|
||||||
from polygateway.middleware.retry import RetryMW, backoff_delay
|
from polygateway.middleware.retry import RetryMW, backoff_delay
|
||||||
from polygateway.sources import RoundRobinSelector, SourceCooldownMemo
|
from polygateway.sources import RoundRobinSelector, SourceCooldownMemo
|
||||||
from polygateway.types import (
|
from polygateway.types import (
|
||||||
@@ -46,19 +53,31 @@ class BoundedSleep:
|
|||||||
await self._side_effect(len(self.delays))
|
await self._side_effect(len(self.delays))
|
||||||
|
|
||||||
|
|
||||||
def _mw(sources, limiter, script, *, clock, sleep, rng=lambda: 0.0, quota_full="wait", gate=None):
|
def _mw(
|
||||||
|
sources,
|
||||||
|
limiter,
|
||||||
|
script,
|
||||||
|
*,
|
||||||
|
clock,
|
||||||
|
sleep,
|
||||||
|
rng=lambda: 0.0,
|
||||||
|
quota_full="wait",
|
||||||
|
gate=None,
|
||||||
|
transport=None,
|
||||||
|
emitter=None,
|
||||||
|
):
|
||||||
return RetryMW(
|
return RetryMW(
|
||||||
scope="llm",
|
scope="llm",
|
||||||
sources=sources,
|
sources=sources,
|
||||||
selector=RoundRobinSelector(),
|
selector=RoundRobinSelector(),
|
||||||
limiter=limiter,
|
limiter=limiter,
|
||||||
gate=gate or InMemoryGate(config=_BREAKER, now=clock),
|
gate=gate or InMemoryGate(config=_BREAKER, now=clock),
|
||||||
transport=FakeTransport(script),
|
transport=transport or FakeTransport(script),
|
||||||
retry=RetryPolicy(max_attempts=3, backoff_base_s=2.0, backoff_max_s=30.0),
|
retry=RetryPolicy(max_attempts=3, backoff_base_s=2.0, backoff_max_s=30.0),
|
||||||
backpressure=BackpressurePolicy(stall_window_s=_STALL, poll_interval_s=0.01),
|
backpressure=BackpressurePolicy(stall_window_s=_STALL, poll_interval_s=0.01),
|
||||||
quota_full=quota_full,
|
quota_full=quota_full,
|
||||||
cooldown_memo=SourceCooldownMemo(now=clock),
|
cooldown_memo=SourceCooldownMemo(now=clock),
|
||||||
emitter=None,
|
emitter=emitter,
|
||||||
now=clock,
|
now=clock,
|
||||||
sleep=sleep,
|
sleep=sleep,
|
||||||
rng=rng,
|
rng=rng,
|
||||||
@@ -111,7 +130,11 @@ class TestStallQuadrants:
|
|||||||
assert resp.content == "ok"
|
assert resp.content == "ok"
|
||||||
|
|
||||||
async def test_global_stale_but_local_fresh_keeps_waiting(self):
|
async def test_global_stale_but_local_fresh_keeps_waiting(self):
|
||||||
"""仅全局超窗(从未出餐 age=inf): 本地才刚开始等 → 不判死。"""
|
"""仅全局超窗(从未出餐 age=inf): 本地才刚开始等 → 不判死。
|
||||||
|
|
||||||
|
`inf` 语义在 issue #8 后未变;变的是"本地"的口径——它现在度量的是
|
||||||
|
非生产性等待累计,不再是墙钟总耗时(见 TestStallBudget)。
|
||||||
|
"""
|
||||||
clock = FakeClock()
|
clock = FakeClock()
|
||||||
src, limiter = _blocked_limiter(clock)
|
src, limiter = _blocked_limiter(clock)
|
||||||
held = await limiter.try_acquire("s1", 0)
|
held = await limiter.try_acquire("s1", 0)
|
||||||
@@ -171,19 +194,235 @@ class TestStallQuadrants:
|
|||||||
await task
|
await task
|
||||||
|
|
||||||
|
|
||||||
|
class ClockAdvancingTransport:
|
||||||
|
"""按脚本 [(推进秒数, 动作), ...] 执行: 在一次尝试内部推进时钟, 模拟真实耗时。
|
||||||
|
|
||||||
|
动作语义同 `FakeTransport`(异常即抛、"hang" 即挂起、其余为返回值)。
|
||||||
|
stall 口径的关键区分在于"时间花在哪", 故必须能让时钟只在 transport 内前进。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, script, clock):
|
||||||
|
self.script = list(script)
|
||||||
|
self.clock = clock
|
||||||
|
self.calls = []
|
||||||
|
|
||||||
|
async def complete(self, *, messages, source, stream, overlay, call_id):
|
||||||
|
self.calls.append((source.name, call_id))
|
||||||
|
advance, action = self.script.pop(0)
|
||||||
|
self.clock.advance(advance)
|
||||||
|
if isinstance(action, Exception):
|
||||||
|
raise action
|
||||||
|
if action == "hang":
|
||||||
|
await asyncio.Event().wait()
|
||||||
|
return action
|
||||||
|
|
||||||
|
|
||||||
|
class _SlowEmitter:
|
||||||
|
"""遥测收尾中推进时钟: 钉住"遥测耗时属生产性"(设计 §3.1 边界声明)。"""
|
||||||
|
|
||||||
|
def __init__(self, clock, advance):
|
||||||
|
self._clock = clock
|
||||||
|
self._advance = advance
|
||||||
|
|
||||||
|
async def emit_attempt(self, *args, **kwargs):
|
||||||
|
self._clock.advance(self._advance)
|
||||||
|
|
||||||
|
|
||||||
|
class TestStallBudget:
|
||||||
|
"""stall 预算只计非生产性等待(issue #8 设计 §3.1)。
|
||||||
|
|
||||||
|
根因是两个预算重叠计费: 真实尝试的耗时同时烧重试预算与 stall 预算,
|
||||||
|
而 stall 预算更小必然先耗尽, 于是 max_attempts 在超时场景下永不生效。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def _free_limiter(self, clock):
|
||||||
|
src = make_source()
|
||||||
|
limiter = InMemoryLimiter(
|
||||||
|
scope="llm", sources={"s1": src}, global_limits=_NO_GLOBAL, now=clock
|
||||||
|
)
|
||||||
|
return src, limiter
|
||||||
|
|
||||||
|
async def test_single_timeout_does_not_exhaust_stall_budget(self):
|
||||||
|
"""timeout_s == stall_window_s 时, 一次超时不得判死——重试预算须真实可用。"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src, limiter = self._free_limiter(clock)
|
||||||
|
# 第一次尝试耗满 300s 超时后失败, 第二次立即成功
|
||||||
|
transport = ClockAdvancingTransport(
|
||||||
|
[(_STALL + 1, TransientError("timeout", status_code=504)), (0.0, _ok())], clock
|
||||||
|
)
|
||||||
|
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
|
||||||
|
resp = await mw(_REQ)
|
||||||
|
assert resp.content == "ok"
|
||||||
|
assert len(transport.calls) == 2 # 第二次尝试确实发出了
|
||||||
|
|
||||||
|
async def test_productive_time_excluded_from_stall(self):
|
||||||
|
"""连续多次长尝试也不烧 stall 预算: 它们烧的是重试预算。"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src, limiter = self._free_limiter(clock)
|
||||||
|
transport = ClockAdvancingTransport(
|
||||||
|
[
|
||||||
|
(_STALL + 100, TransientError("slow", status_code=500)),
|
||||||
|
(_STALL + 100, TransientError("slow", status_code=500)),
|
||||||
|
(0.0, _ok()),
|
||||||
|
],
|
||||||
|
clock,
|
||||||
|
)
|
||||||
|
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
|
||||||
|
resp = await mw(_REQ)
|
||||||
|
assert resp.content == "ok"
|
||||||
|
|
||||||
|
async def test_telemetry_time_counts_as_productive(self):
|
||||||
|
"""遥测收尾属 `_attempt` 边界内: 遥测抖动不得参与判死(设计 §3.1)。
|
||||||
|
|
||||||
|
必须走**失败**路径才有判别力: 成功后直接 return, 循环开头的 stall
|
||||||
|
判定根本不会再执行。此处让首次尝试快速失败、而遥测收尾慢得超窗,
|
||||||
|
下一轮循环开头即检验遥测耗时有没有被算进 stall 账。
|
||||||
|
"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src, limiter = self._free_limiter(clock)
|
||||||
|
transport = ClockAdvancingTransport(
|
||||||
|
[(0.1, TransientError("boom", status_code=500)), (0.0, _ok())], clock
|
||||||
|
)
|
||||||
|
mw = _mw(
|
||||||
|
[src],
|
||||||
|
limiter,
|
||||||
|
[],
|
||||||
|
clock=clock,
|
||||||
|
sleep=BoundedSleep(),
|
||||||
|
transport=transport,
|
||||||
|
emitter=_SlowEmitter(clock, _STALL + 100),
|
||||||
|
)
|
||||||
|
resp = await mw(_REQ)
|
||||||
|
assert resp.content == "ok"
|
||||||
|
|
||||||
|
async def test_nonproductive_wait_still_triggers_stall(self):
|
||||||
|
"""兜底未被削弱: 纯轮询等待累满窗口仍判死。"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src, limiter = _blocked_limiter(clock)
|
||||||
|
_held = await limiter.try_acquire("s1", 0)
|
||||||
|
|
||||||
|
async def advance(_n):
|
||||||
|
clock.advance(_STALL + 100)
|
||||||
|
|
||||||
|
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(advance))
|
||||||
|
with pytest.raises(AllSourcesExhausted) as ei:
|
||||||
|
await mw(_REQ)
|
||||||
|
assert ei.value.reason == "stalled"
|
||||||
|
|
||||||
|
async def test_saturation_429_still_stalls(self):
|
||||||
|
"""429 免预算不烧 fails, 主循环兜底须仍能判死而非无限循环(设计 §3.5)。"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src, limiter = self._free_limiter(clock)
|
||||||
|
# 429 往返本身极快(生产性可忽略), 退避 sleep 才是非生产性的大头
|
||||||
|
transport = ClockAdvancingTransport(
|
||||||
|
[(0.1, TransientError("429", status_code=429)) for _ in range(10)], clock
|
||||||
|
)
|
||||||
|
|
||||||
|
async def advance(_n):
|
||||||
|
clock.advance(_STALL)
|
||||||
|
|
||||||
|
mw = _mw(
|
||||||
|
[src], limiter, [], clock=clock, sleep=BoundedSleep(advance), transport=transport
|
||||||
|
)
|
||||||
|
with pytest.raises(AllSourcesExhausted) as ei:
|
||||||
|
await mw(_REQ)
|
||||||
|
assert ei.value.reason == "stalled" # 不是 retry_exhausted: 429 确实没烧重试预算
|
||||||
|
|
||||||
|
async def test_slow_429_does_not_escape_both_budgets(self):
|
||||||
|
"""排队型网关: 持满 timeout 才回 429。该耗时必须落进 stall 账。
|
||||||
|
|
||||||
|
429 免重试预算, 所以它的耗时若又算生产性就**两个预算都不烧**——调用
|
||||||
|
会挂满 stall_window/backoff_base 轮。修复前实测 301 次尝试、25.2 小时;
|
||||||
|
此处钉住"一轮 429 就把 stall 账推满"这个上界。
|
||||||
|
"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src, limiter = self._free_limiter(clock)
|
||||||
|
transport = ClockAdvancingTransport(
|
||||||
|
[(_STALL + 1, TransientError("429", status_code=429))] * 20, clock
|
||||||
|
)
|
||||||
|
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
|
||||||
|
with pytest.raises(AllSourcesExhausted) as ei:
|
||||||
|
await mw(_REQ)
|
||||||
|
assert ei.value.reason == "stalled"
|
||||||
|
# 一次持满超时的 429 即耗尽 stall 窗口, 不再无限排队
|
||||||
|
assert len(transport.calls) <= 2
|
||||||
|
|
||||||
|
async def test_cancel_inside_attempt_pierces(self):
|
||||||
|
"""取消发生在 `attempting()` 包裹内仍逐字穿透(库铁律)。"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src, limiter = self._free_limiter(clock)
|
||||||
|
transport = ClockAdvancingTransport([(0.0, "hang")], clock)
|
||||||
|
mw = _mw([src], limiter, [], clock=clock, sleep=asyncio.sleep, transport=transport)
|
||||||
|
task = asyncio.create_task(mw(_REQ))
|
||||||
|
while not transport.calls:
|
||||||
|
await asyncio.sleep(0.01)
|
||||||
|
task.cancel()
|
||||||
|
with pytest.raises(asyncio.CancelledError):
|
||||||
|
await task
|
||||||
|
assert (await limiter.source_stats("s1")).inflight == 0 # permit 在 finally 释放
|
||||||
|
|
||||||
|
async def test_clock_is_per_call_not_per_instance(self):
|
||||||
|
"""StallClock 必须是**调用级**局部状态,不得提升为 RetryMW 实例属性。
|
||||||
|
|
||||||
|
生产形态是一个长寿命 RetryMW 跑成千上万次调用。若 clock 成了实例属性,
|
||||||
|
`_entered_at` 会固定在进程启动时刻, 每次调用的 stalled_s() 随进程运行
|
||||||
|
时长单调增长, 最终所有调用被误判 stalled——这是本用例要拦的灾难。
|
||||||
|
|
||||||
|
判别力的关键是**复用同一个 mw**: 两个 mw 实例天然隔离, 抓不到实例共享。
|
||||||
|
"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src = make_source()
|
||||||
|
limiter = InMemoryLimiter(
|
||||||
|
scope="llm", sources={"s1": src}, global_limits=_NO_GLOBAL, now=clock
|
||||||
|
)
|
||||||
|
transport = ClockAdvancingTransport([(0.0, _ok("first")), (0.0, _ok("second"))], clock)
|
||||||
|
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
|
||||||
|
first = await mw(_REQ)
|
||||||
|
clock.advance(_STALL + 100) # 两次调用之间进程空转远超窗
|
||||||
|
second = await mw(_REQ)
|
||||||
|
assert (first.content, second.content) == ("first", "second")
|
||||||
|
|
||||||
|
async def test_concurrent_calls_do_not_share_clock(self):
|
||||||
|
"""并发两路共用同一个 mw: 一快一慢都能正常完成(形态冒烟)。
|
||||||
|
|
||||||
|
**这条不是回归防线**: 实测它在"clock 提为实例属性""去掉 refund""去掉
|
||||||
|
生产性扣减"三种变异下均保持绿色——共享 clock 时慢调用的耗时是作为
|
||||||
|
credit 记进共享账的,污染方向是让 stall 账**变小**(更宽松),而本用例
|
||||||
|
断言两路都成功。真正钉住调用级隔离的是上面那条
|
||||||
|
`test_clock_is_per_call_not_per_instance`。保留此条只为覆盖并发形态。
|
||||||
|
"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src = make_source(max_concurrency=2)
|
||||||
|
limiter = InMemoryLimiter(
|
||||||
|
scope="llm", sources={"s1": src}, global_limits=_NO_GLOBAL, now=clock
|
||||||
|
)
|
||||||
|
transport = ClockAdvancingTransport(
|
||||||
|
[(_STALL + 100, _ok("slow")), (0.0, _ok("fast"))], clock
|
||||||
|
)
|
||||||
|
mw = _mw([src], limiter, [], clock=clock, sleep=BoundedSleep(), transport=transport)
|
||||||
|
results = await asyncio.gather(mw(_REQ), mw(_REQ))
|
||||||
|
assert {r.content for r in results} == {"slow", "fast"}
|
||||||
|
|
||||||
|
|
||||||
class _GateSuccessBroken(InMemoryGate):
|
class _GateSuccessBroken(InMemoryGate):
|
||||||
async def record_success(self, entry):
|
async def record_success(self, entry):
|
||||||
raise GovernanceBackendError("redis 抖动")
|
raise GovernanceBackendError("redis 抖动", scope="llm")
|
||||||
|
|
||||||
|
|
||||||
class _GateFailureBroken(InMemoryGate):
|
class _GateFailureBroken(InMemoryGate):
|
||||||
async def record_failure(self, entry, reason, force_open):
|
async def record_failure(self, entry, reason, force_open):
|
||||||
raise GovernanceBackendError("redis 抖动")
|
raise GovernanceBackendError("redis 抖动", scope="llm")
|
||||||
|
|
||||||
|
|
||||||
class _LimiterProgressBroken(InMemoryLimiter):
|
class _LimiterProgressBroken(InMemoryLimiter):
|
||||||
async def mark_progress(self):
|
async def mark_progress(self):
|
||||||
raise GovernanceBackendError("redis 抖动")
|
raise GovernanceBackendError("redis 抖动", scope="llm")
|
||||||
|
|
||||||
|
|
||||||
|
class _GateSuccessMisconfigured(InMemoryGate):
|
||||||
|
# 签名须与端口一致(含 count_attempt),否则抛的是 TypeError 而非本类要测的异常
|
||||||
|
async def record_success(self, entry, *, count_attempt: bool = True):
|
||||||
|
raise SourceNotConfiguredError("未知源 's1'(scope=llm)")
|
||||||
|
|
||||||
|
|
||||||
class TestAccountingDegradation:
|
class TestAccountingDegradation:
|
||||||
@@ -200,6 +439,24 @@ class TestAccountingDegradation:
|
|||||||
resp = await mw(_REQ)
|
resp = await mw(_REQ)
|
||||||
assert resp.content == "ok" # 真实成功响应不因记账失败被丢弃
|
assert resp.content == "ok" # 真实成功响应不因记账失败被丢弃
|
||||||
|
|
||||||
|
async def test_assembly_defect_on_accounting_path_also_degrades(self):
|
||||||
|
"""记账侧降级按"路径性质"而非异常类型: 装配缺陷同样不得毁掉已完成的调用。
|
||||||
|
|
||||||
|
`SourceNotConfiguredError` 被放行穿透闸门包装器(issue #7 §T6)后,若
|
||||||
|
`_record_quietly` 只降级 `GovernanceBackendError`,它就会从记账侧冒泡、
|
||||||
|
销毁一个真实成功的响应——反转本类钉住的既有行为。当前无后端会从记账
|
||||||
|
方法抛它,此用例是为将来加了源名校验的后端守住这条不变式。
|
||||||
|
"""
|
||||||
|
clock = FakeClock()
|
||||||
|
src = make_source()
|
||||||
|
limiter = InMemoryLimiter(
|
||||||
|
scope="llm", sources={"s1": src}, global_limits=_NO_GLOBAL, now=clock
|
||||||
|
)
|
||||||
|
gate = _GateSuccessMisconfigured(config=_BREAKER, now=clock)
|
||||||
|
mw = _mw([src], limiter, [_ok()], clock=clock, sleep=BoundedSleep(), gate=gate)
|
||||||
|
resp = await mw(_REQ)
|
||||||
|
assert resp.content == "ok"
|
||||||
|
|
||||||
async def test_mark_progress_failure_does_not_lose_response(self):
|
async def test_mark_progress_failure_does_not_lose_response(self):
|
||||||
clock = FakeClock()
|
clock = FakeClock()
|
||||||
src = make_source()
|
src = make_source()
|
||||||
@@ -252,6 +509,91 @@ class TestQuotaGateProgressAge:
|
|||||||
async def progress_age_s(self):
|
async def progress_age_s(self):
|
||||||
raise OSError("down")
|
raise OSError("down")
|
||||||
|
|
||||||
assert await QuotaGate(_L()).progress_age_s() == 12.5
|
assert await QuotaGate(_L(), scope="llm").progress_age_s() == 12.5
|
||||||
with pytest.raises(GovernanceBackendError):
|
with pytest.raises(GovernanceBackendError):
|
||||||
await QuotaGate(_Broken()).progress_age_s()
|
await QuotaGate(_Broken(), scope="llm").progress_age_s()
|
||||||
|
|
||||||
|
|
||||||
|
class TestUnknownSourceIsAssemblyDefect:
|
||||||
|
"""未知源 = 限流后端的源名单与治理循环对不上,是装配缺陷不是后端故障。
|
||||||
|
|
||||||
|
两个后端行为必须一致(Redis 版对应用例在 `test_redis_key_layout.py::
|
||||||
|
TestConversions::test_unknown_source_rejected`);内存版此前无覆盖,
|
||||||
|
该分支从未被测过(issue #7 §3.4)。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def test_memory_limiter_rejects_unknown_source(self):
|
||||||
|
limiter = InMemoryLimiter(
|
||||||
|
scope="llm", sources={"s1": make_source("s1")}, global_limits=_NO_GLOBAL
|
||||||
|
)
|
||||||
|
with pytest.raises(SourceNotConfiguredError) as ei:
|
||||||
|
limiter._cfg("nope")
|
||||||
|
# 关键: 若归入 scope 级家族,配置写错的任务会永远延期重投、永不进死信
|
||||||
|
assert not isinstance(ei.value, GatewayUnavailableError)
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("method", ["try_acquire", "stats"])
|
||||||
|
async def test_survives_the_quota_gate_wrapper(self, method):
|
||||||
|
"""必须穿透 QuotaGate,否则整个拆分在生产路径上等于没做。
|
||||||
|
|
||||||
|
上面两条(以及 redis 版)打的都是私有 `_cfg`,绕过了包装器。而治理循环
|
||||||
|
只经 QuotaGate 访问后端,包装器的 `except Exception` 会把装配缺陷重新
|
||||||
|
包成 `GovernanceBackendError`——下游又拿到可重投异常,永远重投不告警。
|
||||||
|
"""
|
||||||
|
src = make_source("s1")
|
||||||
|
# 限流后端的源名单与治理循环拿到的源对不上 = 装配缺陷
|
||||||
|
limiter = InMemoryLimiter(
|
||||||
|
scope="llm", sources={"other": src}, global_limits=_NO_GLOBAL
|
||||||
|
)
|
||||||
|
gate = QuotaGate(limiter, scope="llm")
|
||||||
|
with pytest.raises(SourceNotConfiguredError) as ei:
|
||||||
|
await getattr(gate, method)(src)
|
||||||
|
assert not isinstance(ei.value, GatewayUnavailableError)
|
||||||
|
|
||||||
|
|
||||||
|
class TestGateFailuresReachCallersAsScopeLevel:
|
||||||
|
"""闸门泄漏路径必须以 scope 级不可用的形态到达调用方(issue #7)。
|
||||||
|
|
||||||
|
记账路径由 `_record_quietly` 降级为 warning,但闸门路径没有那层包裹,会一路
|
||||||
|
抛给调用方。只写 `except GatewayUnavailableError` 的调用方此前接不住,后果
|
||||||
|
是 Redis 抖一下就让积压任务烧掉业务失败预算进死信——而那是运维重启即可恢复
|
||||||
|
的故障。全部五条为: `QuotaGate` 的 try_acquire / stats / progress_age_s,
|
||||||
|
`BreakerGate` 的 try_enter / retry_after_s(判据是该调用点未被 `_record_quietly`
|
||||||
|
包裹)。此处钉住其中三条代表路径,余两条由同一注入机制覆盖。
|
||||||
|
"""
|
||||||
|
|
||||||
|
async def test_try_acquire_failure_is_scope_level(self):
|
||||||
|
from polygateway.middleware.ratelimit import QuotaGate
|
||||||
|
|
||||||
|
class _Broken:
|
||||||
|
async def try_acquire(self, name, est):
|
||||||
|
raise OSError("down")
|
||||||
|
|
||||||
|
with pytest.raises(GatewayUnavailableError) as ei:
|
||||||
|
await QuotaGate(_Broken(), scope="LLM").try_acquire(make_source("s1"))
|
||||||
|
assert ei.value.scope == "llm"
|
||||||
|
assert ei.value.reason == "governance_backend_down"
|
||||||
|
assert ei.value.retry_after_s > 0 # 0 会让积压任务零延迟冲击已挂的后端
|
||||||
|
|
||||||
|
async def test_try_enter_failure_is_scope_level(self):
|
||||||
|
from polygateway.middleware.breaker import BreakerGate
|
||||||
|
|
||||||
|
class _Broken:
|
||||||
|
async def try_enter(self, name, owner):
|
||||||
|
raise OSError("down")
|
||||||
|
|
||||||
|
with pytest.raises(GatewayUnavailableError) as ei:
|
||||||
|
await BreakerGate(_Broken(), scope="LLM").try_enter(make_source("s1"), "owner")
|
||||||
|
assert ei.value.scope == "llm"
|
||||||
|
assert ei.value.reason == "governance_backend_down"
|
||||||
|
|
||||||
|
async def test_progress_age_failure_is_scope_level(self):
|
||||||
|
from polygateway.middleware.ratelimit import QuotaGate
|
||||||
|
|
||||||
|
class _Broken:
|
||||||
|
async def progress_age_s(self):
|
||||||
|
raise OSError("down")
|
||||||
|
|
||||||
|
with pytest.raises(GatewayUnavailableError) as ei:
|
||||||
|
await QuotaGate(_Broken(), scope="LLM").progress_age_s()
|
||||||
|
assert ei.value.scope == "llm"
|
||||||
|
assert ei.value.reason == "governance_backend_down"
|
||||||
|
|||||||
@@ -202,6 +202,33 @@ class TestModelFingerprint:
|
|||||||
assert plain != tuned
|
assert plain != tuned
|
||||||
assert tuned.startswith("qwen-max|") # 旧字面量仍是前缀,便于人眼辨认
|
assert tuned.startswith("qwen-max|") # 旧字面量仍是前缀,便于人眼辨认
|
||||||
|
|
||||||
|
def test_enable_thinking_changes_fingerprint(self):
|
||||||
|
"""issue #5 配套: thinking 一旦真正改变请求体,就必须进缓存身份。
|
||||||
|
|
||||||
|
否则"关掉推理后重启"会读到开着推理时缓存的旧响应——issue #4 为
|
||||||
|
temperature 写过逐字相同的理由。
|
||||||
|
"""
|
||||||
|
from polygateway.client import build_model_fingerprint
|
||||||
|
|
||||||
|
plain = build_model_fingerprint([_source()])
|
||||||
|
off = build_model_fingerprint([_source(enable_thinking=False)])
|
||||||
|
on = build_model_fingerprint([_source(enable_thinking=True)])
|
||||||
|
assert len({plain, off, on}) == 3
|
||||||
|
|
||||||
|
def test_extra_body_only_fingerprint_is_byte_identical_to_before(self):
|
||||||
|
"""只配 extra_body、不表态 thinking 的存量源不得触发冷启动。
|
||||||
|
|
||||||
|
字面量在此硬编码: 这条断言的价值全在"逐字相同",改实现时必须先看见它红。
|
||||||
|
"""
|
||||||
|
import hashlib
|
||||||
|
import json
|
||||||
|
|
||||||
|
from polygateway.client import build_model_fingerprint
|
||||||
|
|
||||||
|
mark = json.dumps(["qwen-max", {"temperature": 0}], sort_keys=True, ensure_ascii=False)
|
||||||
|
expected = "qwen-max|" + hashlib.sha256(mark.encode("utf-8")).hexdigest()
|
||||||
|
assert build_model_fingerprint([_source(extra_body={"temperature": 0})]) == expected
|
||||||
|
|
||||||
def test_source_rename_does_not_change_fingerprint(self):
|
def test_source_rename_does_not_change_fingerprint(self):
|
||||||
"""指纹按 (model, extra_body) 而非源名: 改名不该误触全量冷启动。"""
|
"""指纹按 (model, extra_body) 而非源名: 改名不该误触全量冷启动。"""
|
||||||
from polygateway.client import build_model_fingerprint
|
from polygateway.client import build_model_fingerprint
|
||||||
|
|||||||
@@ -460,6 +460,36 @@ class TestCrossFieldInvariants:
|
|||||||
with pytest.raises(ValueError, match="lease_ttl_s"):
|
with pytest.raises(ValueError, match="lease_ttl_s"):
|
||||||
GatewayClient.from_settings(dataclasses.replace(base, lease_ttl_s=1.0))
|
GatewayClient.from_settings(dataclasses.replace(base, lease_ttl_s=1.0))
|
||||||
|
|
||||||
|
# —— 推理开关的装配守卫(issue #5)——
|
||||||
|
|
||||||
|
def _thinking_sources(self, provider, model, enable_thinking):
|
||||||
|
base = self._base()
|
||||||
|
src = dataclasses.replace(
|
||||||
|
base.sources[0], provider=provider, model=model, enable_thinking=enable_thinking
|
||||||
|
)
|
||||||
|
return dataclasses.replace(base, sources=(src,))
|
||||||
|
|
||||||
|
def test_model_that_cannot_disable_thinking_fails_at_assembly(self):
|
||||||
|
"""M2.x 关不掉推理: 配了 false 必须当场炸,而不是装出一个骗人的 client。"""
|
||||||
|
settings = self._thinking_sources("minimax", "MiniMax-M2.7", False)
|
||||||
|
with pytest.raises(ValueError, match="MiniMax-M2.7"):
|
||||||
|
GatewayClient.from_settings(settings)
|
||||||
|
|
||||||
|
def test_unknown_thinking_shape_fails_at_assembly(self):
|
||||||
|
"""provider=openai 是任意兼容厂商的兜底段名,形态未知即报错并指路。"""
|
||||||
|
settings = self._thinking_sources("openai", "kimi-k3", False)
|
||||||
|
with pytest.raises(ValueError, match="register_provider"):
|
||||||
|
GatewayClient.from_settings(settings)
|
||||||
|
|
||||||
|
def test_supported_combination_assembles(self):
|
||||||
|
settings = self._thinking_sources("minimax", "MiniMax-M3", False)
|
||||||
|
assert GatewayClient.from_settings(settings) is not None
|
||||||
|
|
||||||
|
def test_not_taking_a_position_never_trips_the_guard(self):
|
||||||
|
"""enable_thinking=None(不干预)对任何 provider 都不该被守卫拦下。"""
|
||||||
|
settings = self._thinking_sources("openai", "kimi-k3", None)
|
||||||
|
assert GatewayClient.from_settings(settings) is not None
|
||||||
|
|
||||||
def test_ocr_settings_cannot_wrap_invalid_gateway(self):
|
def test_ocr_settings_cannot_wrap_invalid_gateway(self):
|
||||||
"""OcrSettings/EmbeddingSettings 只是包一层 GatewaySettings,自动继承同一把关。"""
|
"""OcrSettings/EmbeddingSettings 只是包一层 GatewaySettings,自动继承同一把关。"""
|
||||||
base = self._base()
|
base = self._base()
|
||||||
|
|||||||
@@ -173,6 +173,8 @@ from polygateway.types import ( # noqa: E402
|
|||||||
GlobalLimits,
|
GlobalLimits,
|
||||||
RetryPolicy,
|
RetryPolicy,
|
||||||
)
|
)
|
||||||
|
from tests.contracts.conftest import FakeClock # noqa: E402
|
||||||
|
from tests.unit.test_backpressure import BoundedSleep # noqa: E402
|
||||||
|
|
||||||
_BREAKER = BreakerConfig(fail_threshold=3, cooldown_s=60.0, probe_ttl_s=120.0)
|
_BREAKER = BreakerConfig(fail_threshold=3, cooldown_s=60.0, probe_ttl_s=120.0)
|
||||||
_NO_GLOBAL = GlobalLimits(max_concurrency=0, rpm=0, tpm=0)
|
_NO_GLOBAL = GlobalLimits(max_concurrency=0, rpm=0, tpm=0)
|
||||||
@@ -208,6 +210,31 @@ class ScriptedEmbedTransport:
|
|||||||
return action
|
return action
|
||||||
|
|
||||||
|
|
||||||
|
class _ClockAdvancingEmbedTransport:
|
||||||
|
"""按脚本 [(推进秒数, 动作), ...] 执行: 在一次尝试内部推进时钟(issue #8)。
|
||||||
|
|
||||||
|
动作语义同 `ScriptedEmbedTransport`。stall 口径要区分"时间花在哪",
|
||||||
|
故必须能让时钟只在 transport 内前进。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, script, clock):
|
||||||
|
self.script = list(script)
|
||||||
|
self.clock = clock
|
||||||
|
self.calls = []
|
||||||
|
|
||||||
|
async def embed(self, *, texts, source, call_id):
|
||||||
|
self.calls.append((source.name, list(texts), call_id))
|
||||||
|
advance, action = self.script.pop(0)
|
||||||
|
self.clock.advance(advance)
|
||||||
|
if isinstance(action, Exception):
|
||||||
|
raise action
|
||||||
|
if action == "hang":
|
||||||
|
await asyncio.Event().wait()
|
||||||
|
if action == "ok":
|
||||||
|
return _vec_for(texts)
|
||||||
|
return action
|
||||||
|
|
||||||
|
|
||||||
class _MemoryRecorder:
|
class _MemoryRecorder:
|
||||||
def __init__(self):
|
def __init__(self):
|
||||||
self.rows = []
|
self.rows = []
|
||||||
@@ -338,6 +365,39 @@ class TestEmbedGovernance:
|
|||||||
await task
|
await task
|
||||||
assert (await limiter.source_stats("e1")).inflight == 0
|
assert (await limiter.source_stats("e1")).inflight == 0
|
||||||
|
|
||||||
|
async def test_single_timeout_does_not_exhaust_stall_budget(self):
|
||||||
|
"""issue #8: 一次耗满 timeout 的尝试不得吃掉 stall 预算而使重试失效。
|
||||||
|
|
||||||
|
embedding 只有 `_on_no_runnable` 一处 stall 判定, 故失效链条是
|
||||||
|
"先超时一次(墙钟耗尽) → 再遇到无可用源 → 判死"。此处正是这条路径。
|
||||||
|
"""
|
||||||
|
clock = FakeClock()
|
||||||
|
limiter = held = None # 闭包延迟求值: client 建好后才有 limiter
|
||||||
|
|
||||||
|
async def toggle_permit(n):
|
||||||
|
"""首次退避占满 permit, 迫使下一轮走 _on_no_runnable; 之后放行。"""
|
||||||
|
nonlocal held
|
||||||
|
if n == 1:
|
||||||
|
held = await limiter.try_acquire("e1", 0)
|
||||||
|
else:
|
||||||
|
await held.release()
|
||||||
|
|
||||||
|
# 第一次尝试耗满 300s 超时失败, 随后被迫走一轮 _on_no_runnable——
|
||||||
|
# stall 判定就在那里, 检验它有没有把这 300s 生产性时间算进 stall 账
|
||||||
|
transport = _ClockAdvancingEmbedTransport(
|
||||||
|
[(300.1, TransientError("timeout", status_code=504)), (0.0, "ok")], clock
|
||||||
|
)
|
||||||
|
client, limiter = _embed_client(
|
||||||
|
[_src(max_concurrency=1)],
|
||||||
|
[],
|
||||||
|
now=clock,
|
||||||
|
transport=transport,
|
||||||
|
sleep=BoundedSleep(toggle_permit),
|
||||||
|
)
|
||||||
|
resp = await client.embed(["a"])
|
||||||
|
assert resp.vectors == [[1.0]]
|
||||||
|
assert len(transport.calls) == 2 # 第二次尝试确实发出了
|
||||||
|
|
||||||
|
|
||||||
class TestEmbedTelemetry:
|
class TestEmbedTelemetry:
|
||||||
async def test_per_batch_rows_with_digest(self):
|
async def test_per_batch_rows_with_digest(self):
|
||||||
|
|||||||
@@ -3,6 +3,8 @@
|
|||||||
import pytest
|
import pytest
|
||||||
|
|
||||||
from polygateway.errors import (
|
from polygateway.errors import (
|
||||||
|
GOVERNANCE_BACKEND_RETRY_AFTER_S,
|
||||||
|
SCOPE_REASONS,
|
||||||
AllSourcesExhausted,
|
AllSourcesExhausted,
|
||||||
CircuitOpenError,
|
CircuitOpenError,
|
||||||
GatewayUnavailableError,
|
GatewayUnavailableError,
|
||||||
@@ -11,6 +13,7 @@ from polygateway.errors import (
|
|||||||
RequestRejectedError,
|
RequestRejectedError,
|
||||||
ResultInvalidError,
|
ResultInvalidError,
|
||||||
SourceDeadError,
|
SourceDeadError,
|
||||||
|
SourceNotConfiguredError,
|
||||||
TransientError,
|
TransientError,
|
||||||
)
|
)
|
||||||
|
|
||||||
@@ -31,6 +34,43 @@ class TestBaseShape:
|
|||||||
assert TransientError("429", retry_after_s=2.5).retry_after_s == 2.5
|
assert TransientError("429", retry_after_s=2.5).retry_after_s == 2.5
|
||||||
|
|
||||||
|
|
||||||
|
class TestBodyText:
|
||||||
|
"""issue #10: 非 2xx 的响应体摘要必须有承载处,否则拒绝理由事后不可查。"""
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"cls", (PolyGatewayError, TransientError, SourceDeadError, RequestRejectedError)
|
||||||
|
)
|
||||||
|
def test_defaults_empty_and_accepts_summary(self, cls):
|
||||||
|
assert cls("boom").body_text == ""
|
||||||
|
assert cls("boom", body_text='{"error":{"code":"bad"}}').body_text == (
|
||||||
|
'{"error":{"code":"bad"}}'
|
||||||
|
)
|
||||||
|
|
||||||
|
def test_result_invalid_keeps_both_fields_apart(self):
|
||||||
|
"""`body_text`(非 2xx 的拒绝理由)与 `raw_text`(2xx 的不可解析输出)不得混用。"""
|
||||||
|
exc = ResultInvalidError("bad json", raw_text="{oops", body_text="")
|
||||||
|
assert exc.raw_text == "{oops"
|
||||||
|
assert exc.body_text == ""
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
"exc",
|
||||||
|
(
|
||||||
|
AllSourcesExhausted(scope="llm", reason="stalled", retry_after_s=1.0),
|
||||||
|
CircuitOpenError(scope="llm", retry_after_s=1.0),
|
||||||
|
GovernanceBackendError("redis down", scope="llm"),
|
||||||
|
),
|
||||||
|
)
|
||||||
|
def test_scope_level_errors_carry_no_body(self, exc):
|
||||||
|
"""scope 级失败没有单一响应体可言,空串是如实表达而非噪音。"""
|
||||||
|
assert exc.body_text == ""
|
||||||
|
|
||||||
|
def test_body_text_does_not_leak_into_str(self):
|
||||||
|
"""字段是旁路数据: 加了它不得改变任何既有异常的 str() 输出。"""
|
||||||
|
assert str(RequestRejectedError("qwen_1 请求被拒: 400", body_text="whatever")) == (
|
||||||
|
"qwen_1 请求被拒: 400"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
class TestResultInvalid:
|
class TestResultInvalid:
|
||||||
def test_carries_diagnosis(self):
|
def test_carries_diagnosis(self):
|
||||||
exc = ResultInvalidError(
|
exc = ResultInvalidError(
|
||||||
@@ -86,6 +126,54 @@ class TestGatewayUnavailable:
|
|||||||
class TestBackendFailure:
|
class TestBackendFailure:
|
||||||
def test_governance_backend_error_is_not_transient(self):
|
def test_governance_backend_error_is_not_transient(self):
|
||||||
"""限流/熔断后端故障必须报错不放行,且不落入可重试分类。"""
|
"""限流/熔断后端故障必须报错不放行,且不落入可重试分类。"""
|
||||||
exc = GovernanceBackendError("redis down")
|
exc = GovernanceBackendError("redis down", scope="llm")
|
||||||
assert isinstance(exc, PolyGatewayError)
|
assert isinstance(exc, PolyGatewayError)
|
||||||
assert not isinstance(exc, TransientError)
|
assert not isinstance(exc, TransientError)
|
||||||
|
|
||||||
|
def test_is_scope_level_unavailability(self):
|
||||||
|
"""fail-closed 时整个 scope 一个请求都发不出去,调用方一条 except 应覆盖(issue #7)。"""
|
||||||
|
exc = GovernanceBackendError("限流后端 try_acquire 失败: boom", scope="LLM")
|
||||||
|
assert isinstance(exc, GatewayUnavailableError)
|
||||||
|
assert exc.reason == "governance_backend_down"
|
||||||
|
assert exc.scope == "llm" # 与既有 scope 级异常同款: 归一化小写
|
||||||
|
assert exc.retry_after_s == GOVERNANCE_BACKEND_RETRY_AFTER_S
|
||||||
|
|
||||||
|
def test_diagnostic_message_survives_reparenting(self):
|
||||||
|
"""父类把 message 覆写为模板串,而各构造点的诊断串是排障主线索(§3.5)。"""
|
||||||
|
exc = GovernanceBackendError("熔断后端 try_enter 失败: boom", scope="llm")
|
||||||
|
assert str(exc) == "熔断后端 try_enter 失败: boom"
|
||||||
|
|
||||||
|
def test_retry_after_overridable(self):
|
||||||
|
exc = GovernanceBackendError("redis down", scope="llm", retry_after_s=30.0)
|
||||||
|
assert exc.retry_after_s == 30.0
|
||||||
|
|
||||||
|
|
||||||
|
class TestSourceNotConfigured:
|
||||||
|
"""装配缺陷有意留在 scope 级家族之外(issue #7 §3.4,Q1 人类拍板)。"""
|
||||||
|
|
||||||
|
def test_is_domain_error_but_not_scope_level(self):
|
||||||
|
exc = SourceNotConfiguredError("未知源 'nope'(scope=llm)")
|
||||||
|
assert isinstance(exc, PolyGatewayError)
|
||||||
|
# 关键断言: 归入可重投家族会让配置写错的任务永远重投、永不进死信
|
||||||
|
assert not isinstance(exc, GatewayUnavailableError)
|
||||||
|
|
||||||
|
def test_exported_at_package_top_level(self):
|
||||||
|
import polygateway
|
||||||
|
|
||||||
|
assert polygateway.SourceNotConfiguredError is SourceNotConfiguredError
|
||||||
|
assert "SourceNotConfiguredError" in polygateway.__all__
|
||||||
|
|
||||||
|
|
||||||
|
class TestGovernanceBackendReason:
|
||||||
|
"""新 scope 级 reason 值域(issue #7 §3.1)。"""
|
||||||
|
|
||||||
|
def test_reason_admitted_to_scope_domain(self):
|
||||||
|
assert "governance_backend_down" in SCOPE_REASONS
|
||||||
|
|
||||||
|
def test_gateway_unavailable_accepts_the_new_reason(self):
|
||||||
|
exc = AllSourcesExhausted(scope="LLM", reason="governance_backend_down", retry_after_s=0.0)
|
||||||
|
assert exc.reason == "governance_backend_down"
|
||||||
|
|
||||||
|
def test_retry_after_default_is_non_zero(self):
|
||||||
|
"""取 0 会让积压任务零延迟冲击已挂掉的后端(§3.2)。"""
|
||||||
|
assert GOVERNANCE_BACKEND_RETRY_AFTER_S > 0
|
||||||
|
|||||||
@@ -0,0 +1,89 @@
|
|||||||
|
"""HTTP 错误响应体摘要口径(issue #10 设计 §3.2/§3.4)。
|
||||||
|
|
||||||
|
摘要是 message 与 `body_text` 共用的**同一份串**,故它的边界行为直接决定
|
||||||
|
遥测里看到的与下游 catch 到的是否一致——本组用例把规则钉成算术。
|
||||||
|
"""
|
||||||
|
|
||||||
|
import httpx
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from polygateway.transports._http_errors import (
|
||||||
|
_ERROR_BODY_CAP,
|
||||||
|
_HEAD_CHARS,
|
||||||
|
_TAIL_CHARS,
|
||||||
|
compose_message,
|
||||||
|
response_body,
|
||||||
|
summarize_body,
|
||||||
|
)
|
||||||
|
|
||||||
|
# issue #10 原文给出的真实响应体(一字不改),关键在于 code 收尾
|
||||||
|
_REAL_SAMPLE = (
|
||||||
|
'{"error":{"message":"<400> ***.***.InvalidParameter: The image format is illegal '
|
||||||
|
'and cannot be opened","type":"invalid_request_error","param":"",'
|
||||||
|
'"code":"invalid_parameter_error"}}'
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class TestSummarizeBody:
|
||||||
|
def test_short_body_passes_through(self):
|
||||||
|
assert summarize_body(_REAL_SAMPLE) == _REAL_SAMPLE
|
||||||
|
|
||||||
|
def test_whitespace_collapsed(self):
|
||||||
|
"""错误体常是缩进 JSON: 不折叠会把一行日志炸成多行、遥测列不可读。"""
|
||||||
|
assert (
|
||||||
|
summarize_body('{\n "error": {\n "code": "x"\n }\n}')
|
||||||
|
== '{ "error": { "code": "x" } }'
|
||||||
|
)
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("raw", ("", " ", "\n\t \n"))
|
||||||
|
def test_blank_yields_empty(self, raw):
|
||||||
|
assert summarize_body(raw) == ""
|
||||||
|
|
||||||
|
def test_exactly_at_cap_is_untouched(self):
|
||||||
|
body = "x" * _ERROR_BODY_CAP
|
||||||
|
assert summarize_body(body) == body
|
||||||
|
|
||||||
|
def test_one_over_cap_is_summarized(self):
|
||||||
|
summary = summarize_body("x" * (_ERROR_BODY_CAP + 1))
|
||||||
|
assert summary != "x" * (_ERROR_BODY_CAP + 1)
|
||||||
|
assert "略" in summary
|
||||||
|
|
||||||
|
def test_head_and_tail_both_survive(self):
|
||||||
|
"""头部硬切会丢掉尾部,而 JSON 错误体的 code/request_id 正在尾部。"""
|
||||||
|
body = "H" * 5000 + "T" * 5000
|
||||||
|
summary = summarize_body(body)
|
||||||
|
assert summary[:_HEAD_CHARS] == body[:_HEAD_CHARS]
|
||||||
|
assert summary[-_TAIL_CHARS:] == body[-_TAIL_CHARS:]
|
||||||
|
assert f"…(略 {10000 - _HEAD_CHARS - _TAIL_CHARS} 字)…" in summary
|
||||||
|
|
||||||
|
def test_real_sample_tail_visible_in_oversized_body(self):
|
||||||
|
"""设计 §7 用例 3c: 超长体里,追查网关方所需的 code 仍须可见。"""
|
||||||
|
summary = summarize_body("PADDING" * 1000 + _REAL_SAMPLE)
|
||||||
|
assert '"code":"invalid_parameter_error"}}' in summary
|
||||||
|
|
||||||
|
def test_idempotent(self):
|
||||||
|
"""再摘要一次不得嵌套标记,否则重复经手的串会层层套娃。"""
|
||||||
|
once = summarize_body("y" * 9999)
|
||||||
|
assert summarize_body(once) == once
|
||||||
|
|
||||||
|
|
||||||
|
class TestComposeMessage:
|
||||||
|
def test_empty_summary_leaves_message_intact(self):
|
||||||
|
assert compose_message("qwen_1 请求被拒: 400", "") == "qwen_1 请求被拒: 400"
|
||||||
|
|
||||||
|
def test_non_empty_summary_is_appended(self):
|
||||||
|
assert compose_message("qwen_1 请求被拒: 400", "{}") == "qwen_1 请求被拒: 400 | {}"
|
||||||
|
|
||||||
|
|
||||||
|
class TestResponseBody:
|
||||||
|
def test_reads_buffered_text(self):
|
||||||
|
assert response_body(httpx.Response(400, content=b'{"e":1}')) == '{"e":1}'
|
||||||
|
|
||||||
|
def test_unread_stream_degrades_to_empty(self):
|
||||||
|
"""取不到诊断信息绝不能升级为崩溃: 未读缓冲返回空串,且不触发网络读。"""
|
||||||
|
|
||||||
|
class _Unread(httpx.SyncByteStream):
|
||||||
|
def __iter__(self):
|
||||||
|
yield b"body"
|
||||||
|
|
||||||
|
assert response_body(httpx.Response(400, stream=_Unread())) == ""
|
||||||
@@ -19,7 +19,11 @@ from polygateway.errors import (
|
|||||||
SourceDeadError,
|
SourceDeadError,
|
||||||
TransientError,
|
TransientError,
|
||||||
)
|
)
|
||||||
from polygateway.transports.monkey_ocr import MonkeyOcrTransport, _parse_middle_json
|
from polygateway.transports.monkey_ocr import (
|
||||||
|
MonkeyOcrTransport,
|
||||||
|
_classify_status,
|
||||||
|
_parse_middle_json,
|
||||||
|
)
|
||||||
from polygateway.types import SourceConfig
|
from polygateway.types import SourceConfig
|
||||||
|
|
||||||
|
|
||||||
@@ -306,6 +310,45 @@ class TestErrorTranslation:
|
|||||||
await t.recognize_text(image=b"jpg", source=_source(), call_id="c1")
|
await t.recognize_text(image=b"jpg", source=_source(), call_id="c1")
|
||||||
assert ei.value.status_code == status
|
assert ei.value.status_code == status
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("status", "exc_type"),
|
||||||
|
[
|
||||||
|
(502, TransientError),
|
||||||
|
(429, TransientError), # 与 5xx 共用分支,但仍单列: 漏分支正是 issue #10 的成因
|
||||||
|
(401, SourceDeadError),
|
||||||
|
(404, RequestRejectedError),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
async def test_body_survives_every_branch(self, status, exc_type):
|
||||||
|
"""issue #10: OCR 侧 message 原本只有 HTTP 状态码,拒绝理由同样丢失。"""
|
||||||
|
body = '{"detail":"unsupported image mode CMYK"}'
|
||||||
|
t = _transport_for(_routes(text_resp=httpx.Response(status, content=body.encode())))
|
||||||
|
with pytest.raises(exc_type) as ei:
|
||||||
|
await t.recognize_text(image=b"jpg", source=_source(), call_id="c1")
|
||||||
|
assert ei.value.body_text == body
|
||||||
|
assert str(ei.value).endswith(f" | {body}")
|
||||||
|
|
||||||
|
def test_unread_body_degrades_without_changing_class(self):
|
||||||
|
"""取不到 body 时降级空串: 绝不能让 ResponseNotRead 逃出错误四分类。
|
||||||
|
|
||||||
|
直接测纯函数而非走 MockTransport——真实客户端对非 stream 请求总会读完
|
||||||
|
响应,未读态只可能在将来给 OCR 加 stream 时出现,而那正是要防的场景。
|
||||||
|
"""
|
||||||
|
|
||||||
|
class _Unread(httpx.SyncByteStream):
|
||||||
|
def __iter__(self):
|
||||||
|
yield b"body"
|
||||||
|
|
||||||
|
exc = httpx.HTTPStatusError(
|
||||||
|
"404",
|
||||||
|
request=httpx.Request("POST", "http://ocr.example/ocr/text"),
|
||||||
|
response=httpx.Response(404, stream=_Unread()),
|
||||||
|
)
|
||||||
|
err = _classify_status(exc, "monkey_1", "text")
|
||||||
|
assert isinstance(err, RequestRejectedError)
|
||||||
|
assert err.body_text == ""
|
||||||
|
assert str(err) == "monkey_1 OCR text HTTP 404"
|
||||||
|
|
||||||
async def test_connect_error_transient(self):
|
async def test_connect_error_transient(self):
|
||||||
def handler(request):
|
def handler(request):
|
||||||
raise httpx.ConnectError("refused", request=request)
|
raise httpx.ConnectError("refused", request=request)
|
||||||
|
|||||||
@@ -80,6 +80,22 @@ class ScriptedOcrTransport:
|
|||||||
raise NotImplementedError
|
raise NotImplementedError
|
||||||
|
|
||||||
|
|
||||||
|
class ClockAdvancingOcrTransport(ScriptedOcrTransport):
|
||||||
|
"""按脚本 [(推进秒数, 动作), ...] 在一次尝试内部推进时钟(issue #8)。
|
||||||
|
|
||||||
|
stall 口径要区分"时间花在哪",故必须能让时钟只在 transport 内前进。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, script, clock):
|
||||||
|
super().__init__([a for _, a in script])
|
||||||
|
self._advances = [d for d, _ in script]
|
||||||
|
self.clock = clock
|
||||||
|
|
||||||
|
async def _next(self, method, source, call_id):
|
||||||
|
self.clock.advance(self._advances.pop(0))
|
||||||
|
return await super()._next(method, source, call_id)
|
||||||
|
|
||||||
|
|
||||||
class StaticSelector:
|
class StaticSelector:
|
||||||
def order(self, sources, stats):
|
def order(self, sources, stats):
|
||||||
return list(sources)
|
return list(sources)
|
||||||
@@ -292,6 +308,38 @@ class TestBackpressure:
|
|||||||
await permit.settle(0)
|
await permit.settle(0)
|
||||||
await permit.release()
|
await permit.release()
|
||||||
|
|
||||||
|
async def test_single_timeout_does_not_exhaust_stall_budget(self):
|
||||||
|
"""issue #8: 一次耗满 timeout 的尝试不得吃掉 stall 预算而使重试失效。
|
||||||
|
|
||||||
|
OCR 只有 `_on_no_runnable` 一处 stall 判定,故失效链条是"先超时一次
|
||||||
|
(墙钟耗尽)→ 再遇到无可用源 → 判死"。此处正是这条路径。
|
||||||
|
"""
|
||||||
|
clock = FakeClock()
|
||||||
|
limiter = held = None
|
||||||
|
rounds = []
|
||||||
|
|
||||||
|
async def toggle_permit(_seconds):
|
||||||
|
"""首次退避占满 permit,迫使下一轮走 _on_no_runnable;之后放行。"""
|
||||||
|
nonlocal held
|
||||||
|
rounds.append(_seconds)
|
||||||
|
if len(rounds) > 10:
|
||||||
|
raise RuntimeError("超过 10 次轮询仍未判死/未获 permit")
|
||||||
|
if len(rounds) == 1:
|
||||||
|
held = await limiter.acquire("m1", 0)
|
||||||
|
else:
|
||||||
|
await held.settle(0)
|
||||||
|
await held.release()
|
||||||
|
|
||||||
|
transport = ClockAdvancingOcrTransport(
|
||||||
|
[(300.1, TransientError("timeout", status_code=504)), (0.0, "text")], clock
|
||||||
|
)
|
||||||
|
client, limiter, _ = _client(
|
||||||
|
[_src(max_concurrency=1)], [], now=clock, sleep=toggle_permit, transport=transport
|
||||||
|
)
|
||||||
|
r = await client.recognize_text(b"jpg")
|
||||||
|
assert r.text == "LINE-1"
|
||||||
|
assert len(transport.calls) == 2 # 第二次尝试确实发出了
|
||||||
|
|
||||||
|
|
||||||
class FakeClock:
|
class FakeClock:
|
||||||
def __init__(self, start=1000.0):
|
def __init__(self, start=1000.0):
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ import json
|
|||||||
|
|
||||||
import httpx
|
import httpx
|
||||||
import pytest
|
import pytest
|
||||||
|
from loguru import logger
|
||||||
|
|
||||||
from polygateway.errors import (
|
from polygateway.errors import (
|
||||||
RequestRejectedError,
|
RequestRejectedError,
|
||||||
@@ -15,6 +16,7 @@ from polygateway.errors import (
|
|||||||
)
|
)
|
||||||
from polygateway.middleware.telemetry import TelemetryEmitter
|
from polygateway.middleware.telemetry import TelemetryEmitter
|
||||||
from polygateway.pricing import ModelPrice, PricingTable
|
from polygateway.pricing import ModelPrice, PricingTable
|
||||||
|
from polygateway.transports._http_errors import summarize_body
|
||||||
from polygateway.transports.openai_compat import (
|
from polygateway.transports.openai_compat import (
|
||||||
OpenAICompatTransport,
|
OpenAICompatTransport,
|
||||||
_iter_sse_deltas,
|
_iter_sse_deltas,
|
||||||
@@ -394,6 +396,81 @@ class TestObservabilityFields:
|
|||||||
assert set(result.raw) == {"usage"}
|
assert set(result.raw) == {"usage"}
|
||||||
|
|
||||||
|
|
||||||
|
class TestReasoningTokens:
|
||||||
|
"""issue #6: 推理消耗的输出 token,与 issue #3 的 cached_tokens 对称。
|
||||||
|
|
||||||
|
实测三家供应商在"未推理"时是整个 completion_tokens_details 缺失,无人上报
|
||||||
|
0;且中转在上游不返回 usage 时会本地补算并吃掉该对象。故 None 的语义是
|
||||||
|
"本次调用未上报",不是"该源不上报"(findings §4c)。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def _reasoning_usage(self, reasoning):
|
||||||
|
return {**_USAGE, "completion_tokens_details": {"reasoning_tokens": reasoning}}
|
||||||
|
|
||||||
|
async def test_stream_reads_reasoning_tokens(self):
|
||||||
|
def handler(request):
|
||||||
|
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._reasoning_usage(7)))
|
||||||
|
|
||||||
|
result = await _complete(_transport_for(handler), _source())
|
||||||
|
assert result.reasoning_tokens == 7
|
||||||
|
|
||||||
|
async def test_non_stream_reads_reasoning_tokens(self):
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(
|
||||||
|
200,
|
||||||
|
json={
|
||||||
|
"choices": [{"message": {"content": "42"}}],
|
||||||
|
"usage": self._reasoning_usage(7),
|
||||||
|
},
|
||||||
|
)
|
||||||
|
|
||||||
|
result = await _complete(_transport_for(handler), _source(), stream=False)
|
||||||
|
assert result.reasoning_tokens == 7
|
||||||
|
|
||||||
|
async def test_zero_reasoning_tokens_is_a_real_zero(self):
|
||||||
|
"""0(上报了且确实没推理)与 None(本次未上报)必须可区分。"""
|
||||||
|
|
||||||
|
def handler(request):
|
||||||
|
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._reasoning_usage(0)))
|
||||||
|
|
||||||
|
result = await _complete(_transport_for(handler), _source())
|
||||||
|
assert result.reasoning_tokens == 0
|
||||||
|
|
||||||
|
async def test_usage_without_details_is_none(self):
|
||||||
|
def handler(request):
|
||||||
|
return _sse_stream(_chunk(content="ok"), _chunk(usage=_USAGE))
|
||||||
|
|
||||||
|
result = await _complete(_transport_for(handler), _source())
|
||||||
|
assert result.reasoning_tokens is None
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("bad", ["abc", -1, True, 1.5, None, [], {"x": 1}])
|
||||||
|
async def test_malformed_reasoning_tokens_degrade_to_none(self, bad):
|
||||||
|
"""`True` 必须排除: Python 里 isinstance(True, int) 为真。"""
|
||||||
|
|
||||||
|
def handler(request):
|
||||||
|
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._reasoning_usage(bad)))
|
||||||
|
|
||||||
|
result = await _complete(_transport_for(handler), _source())
|
||||||
|
assert result.reasoning_tokens is None
|
||||||
|
|
||||||
|
async def test_details_not_a_dict_is_none(self):
|
||||||
|
def handler(request):
|
||||||
|
usage = {**_USAGE, "completion_tokens_details": "oops"}
|
||||||
|
return _sse_stream(_chunk(content="ok"), _chunk(usage=usage))
|
||||||
|
|
||||||
|
result = await _complete(_transport_for(handler), _source())
|
||||||
|
assert result.reasoning_tokens is None
|
||||||
|
|
||||||
|
async def test_salvage_path_records_none_not_zero(self):
|
||||||
|
"""打捞路径拿不到 usage 帧: 记 None(未知)而非 0(确定没推理)。"""
|
||||||
|
|
||||||
|
def handler(request):
|
||||||
|
return _sse_stream(_chunk(content="ok"), done=False)
|
||||||
|
|
||||||
|
result = await _complete(_transport_for(handler), _source(missing_done="salvage"))
|
||||||
|
assert result.reasoning_tokens is None
|
||||||
|
|
||||||
|
|
||||||
class TestNonStreamFastPath:
|
class TestNonStreamFastPath:
|
||||||
async def test_non_stream_parses_message(self):
|
async def test_non_stream_parses_message(self):
|
||||||
def handler(request):
|
def handler(request):
|
||||||
@@ -431,6 +508,104 @@ class TestRequestShaping:
|
|||||||
assert "enable_thinking" not in seen
|
assert "enable_thinking" not in seen
|
||||||
assert seen["stream_options"] == {"include_usage": True}
|
assert seen["stream_options"] == {"include_usage": True}
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("enable_thinking", "expected"),
|
||||||
|
[(True, "medium"), (False, "none")],
|
||||||
|
)
|
||||||
|
async def test_minimax_injects_reasoning_effort(self, enable_thinking, expected):
|
||||||
|
"""issue #5: MiniMax 认的是 reasoning_effort,不是 enable_thinking。"""
|
||||||
|
seen = {}
|
||||||
|
|
||||||
|
def handler(request):
|
||||||
|
seen.update(json.loads(request.content))
|
||||||
|
return _sse_stream(_chunk(content="x"), _chunk(usage=_USAGE))
|
||||||
|
|
||||||
|
source = _source(
|
||||||
|
name="mm", provider="minimax", model="MiniMax-M3", enable_thinking=enable_thinking
|
||||||
|
)
|
||||||
|
await _complete(_transport_for(handler), source)
|
||||||
|
assert seen["reasoning_effort"] == expected
|
||||||
|
assert "enable_thinking" not in seen # 旧形态实测被静默丢弃,不再下发
|
||||||
|
|
||||||
|
async def test_extra_body_overrides_the_profile_slot(self):
|
||||||
|
"""注入顺序即优先级: profile → extra_body → overlay,两行不可调换。"""
|
||||||
|
seen = {}
|
||||||
|
|
||||||
|
def handler(request):
|
||||||
|
seen.update(json.loads(request.content))
|
||||||
|
return _sse_stream(_chunk(content="x"), _chunk(usage=_USAGE))
|
||||||
|
|
||||||
|
source = _source(
|
||||||
|
name="mm",
|
||||||
|
provider="minimax",
|
||||||
|
model="MiniMax-M3",
|
||||||
|
enable_thinking=True,
|
||||||
|
extra_body={"reasoning_effort": "high"},
|
||||||
|
)
|
||||||
|
await _complete(_transport_for(handler), source)
|
||||||
|
assert seen["reasoning_effort"] == "high"
|
||||||
|
|
||||||
|
async def test_model_that_cannot_disable_is_rejected_not_silently_ignored(self):
|
||||||
|
"""M2.x 关不掉推理: 必须是四分类之一的 RequestRejected,不是裸 ValueError。
|
||||||
|
|
||||||
|
裸异常会逃出 chat() —— 它不属错误四分类、TelemetryMW 也不捕,结果是一行
|
||||||
|
遥测都没有就崩了(设计 §5.1)。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def handler(request): # pragma: no cover - 不该走到发请求
|
||||||
|
raise AssertionError("请求不该发出")
|
||||||
|
|
||||||
|
source = _source(name="mm", provider="minimax", model="MiniMax-M2.7", enable_thinking=False)
|
||||||
|
with pytest.raises(RequestRejectedError, match="MiniMax-M2.7"):
|
||||||
|
await _complete(_transport_for(handler), source)
|
||||||
|
|
||||||
|
async def test_unregistered_model_warns_only_once_per_source(self):
|
||||||
|
"""未登记模型的告警不能打在请求热路径上: 装配期已喊过,逐次再喊是刷屏。"""
|
||||||
|
|
||||||
|
def handler(request):
|
||||||
|
return _sse_stream(_chunk(content="x"), _chunk(usage=_USAGE))
|
||||||
|
|
||||||
|
source = _source(name="mm", provider="minimax", model="MiniMax-M99", enable_thinking=False)
|
||||||
|
transport = _transport_for(handler)
|
||||||
|
messages: list[str] = []
|
||||||
|
sink_id = logger.add(messages.append, level="WARNING")
|
||||||
|
try:
|
||||||
|
await _complete(transport, source)
|
||||||
|
await _complete(transport, source)
|
||||||
|
await _complete(transport, source)
|
||||||
|
finally:
|
||||||
|
logger.remove(sink_id)
|
||||||
|
hits = [m for m in messages if "MiniMax-M99" in m]
|
||||||
|
assert len(hits) == 1, f"三次调用应只告警一次,实得 {len(hits)} 次"
|
||||||
|
|
||||||
|
async def test_unrelated_value_error_is_not_mislabelled(self, monkeypatch):
|
||||||
|
"""只捕 ThinkingUnsupportedError: 无关的 ValueError 不该被贴成推理开关的错。
|
||||||
|
|
||||||
|
今天 `_build_payload` 里只有 resolve_thinking 会抛 ValueError,所以这条
|
||||||
|
是防御未来 —— 但正因如此才要钉住: 将来谁在那里加一处校验,宽 catch 会
|
||||||
|
把它的错误信息盖掉,而这个用例会先红。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def handler(request): # pragma: no cover - 不该走到发请求
|
||||||
|
raise AssertionError("请求不该发出")
|
||||||
|
|
||||||
|
def _boom(*args, **kwargs):
|
||||||
|
raise ValueError("故意的无关错误")
|
||||||
|
|
||||||
|
monkeypatch.setattr("polygateway.transports.openai_compat.resolve_thinking", _boom)
|
||||||
|
with pytest.raises(ValueError, match="故意的无关错误") as exc:
|
||||||
|
await _complete(_transport_for(handler), _source(enable_thinking=False))
|
||||||
|
assert "推理开关" not in str(exc.value)
|
||||||
|
assert not isinstance(exc.value, RequestRejectedError)
|
||||||
|
|
||||||
|
async def test_unknown_shape_is_rejected(self):
|
||||||
|
def handler(request): # pragma: no cover - 不该走到发请求
|
||||||
|
raise AssertionError("请求不该发出")
|
||||||
|
|
||||||
|
source = _source(name="k3", provider="openai", model="kimi-k3", enable_thinking=False)
|
||||||
|
with pytest.raises(RequestRejectedError, match="register_provider"):
|
||||||
|
await _complete(_transport_for(handler), source)
|
||||||
|
|
||||||
async def test_overlay_merged_into_payload(self):
|
async def test_overlay_merged_into_payload(self):
|
||||||
seen = {}
|
seen = {}
|
||||||
|
|
||||||
@@ -517,6 +692,125 @@ class TestErrorTranslation:
|
|||||||
await _complete(_transport_for(handler), _source())
|
await _complete(_transport_for(handler), _source())
|
||||||
|
|
||||||
|
|
||||||
|
# issue #10 原文给出的真实响应体(一字不改): 关键在于 code 收尾
|
||||||
|
_REJECT_BODY = (
|
||||||
|
'{"error":{"message":"<400> ***.***.InvalidParameter: The image format is illegal '
|
||||||
|
'and cannot be opened","type":"invalid_request_error","param":"",'
|
||||||
|
'"code":"invalid_parameter_error"}}'
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class TestErrorBodyRetention:
|
||||||
|
"""issue #10: 网关说了什么必须活着离开翻译层——message 与字段各留一份。"""
|
||||||
|
|
||||||
|
@pytest.mark.parametrize(
|
||||||
|
("status", "exc"),
|
||||||
|
[
|
||||||
|
(400, RequestRejectedError),
|
||||||
|
(401, SourceDeadError),
|
||||||
|
(403, SourceDeadError),
|
||||||
|
(404, RequestRejectedError),
|
||||||
|
(500, TransientError),
|
||||||
|
(503, TransientError),
|
||||||
|
],
|
||||||
|
)
|
||||||
|
async def test_every_non_2xx_branch_keeps_the_body(self, status, exc):
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(status, content=_REJECT_BODY.encode())
|
||||||
|
|
||||||
|
with pytest.raises(exc) as ei:
|
||||||
|
await _complete(_transport_for(handler), _source())
|
||||||
|
assert ei.value.body_text == _REJECT_BODY
|
||||||
|
# message 与字段共用同一份串: 遥测里看到的与下游 catch 到的不得打架
|
||||||
|
assert str(ei.value).endswith(f" | {_REJECT_BODY}")
|
||||||
|
assert "invalid_parameter_error" in str(ei.value)
|
||||||
|
|
||||||
|
async def test_rate_limited_429_keeps_body_and_retry_after(self):
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(
|
||||||
|
429,
|
||||||
|
content=b'{"error":{"message":"per-minute cap 3"}}',
|
||||||
|
headers={"retry-after": "2.5"},
|
||||||
|
)
|
||||||
|
|
||||||
|
with pytest.raises(TransientError) as ei:
|
||||||
|
await _complete(_transport_for(handler), _source())
|
||||||
|
assert "per-minute cap 3" in str(ei.value)
|
||||||
|
assert ei.value.retry_after_s == 2.5 # 摘要不得干扰既有解析
|
||||||
|
|
||||||
|
async def test_insufficient_quota_429_keeps_body(self):
|
||||||
|
body = json.dumps(
|
||||||
|
{"error": {"type": "insufficient_quota", "message": "daily budget spent"}}
|
||||||
|
)
|
||||||
|
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(429, content=body.encode())
|
||||||
|
|
||||||
|
with pytest.raises(SourceDeadError) as ei:
|
||||||
|
await _complete(_transport_for(handler), _source())
|
||||||
|
assert "daily budget spent" in str(ei.value)
|
||||||
|
|
||||||
|
async def test_oversized_insufficient_quota_still_classified_dead(self):
|
||||||
|
"""实现红线: 类型判定必须读**原文**。
|
||||||
|
|
||||||
|
摘要会破坏 JSON 结构,若改用摘要解析,超长 body 的配额耗尽将退化成普通
|
||||||
|
限速——配额已耗尽的源不再 force_open,一个诊断改进就变成了治理 bug。
|
||||||
|
|
||||||
|
**填充必须是多个键**,不能是单个超长字符串值: 后者的截断点落在字符串
|
||||||
|
*内部*,省略标记成了合法的字符串内容,而头尾保留又让尾部的 error 对象
|
||||||
|
幸存——摘要照样解析得出 `insufficient_quota`,用例即告空转(2026-08-16
|
||||||
|
verifier 变异测试发现: 按错误写法实现,全套件 824 项依然全绿)。多键
|
||||||
|
填充让截断点落在结构记号之间,摘要才真正不可解析。
|
||||||
|
"""
|
||||||
|
body = json.dumps(
|
||||||
|
{**{f"k{i}": "v" * 10 for i in range(300)}, "error": {"type": "insufficient_quota"}}
|
||||||
|
)
|
||||||
|
assert len(body) > 2048
|
||||||
|
with pytest.raises(json.JSONDecodeError):
|
||||||
|
# 判别力的前提: 摘要确实不再是合法 JSON,读它必然拿不到 type
|
||||||
|
json.loads(summarize_body(body))
|
||||||
|
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(429, content=body.encode())
|
||||||
|
|
||||||
|
with pytest.raises(SourceDeadError):
|
||||||
|
await _complete(_transport_for(handler), _source())
|
||||||
|
|
||||||
|
async def test_empty_body_leaves_no_dangling_separator(self):
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(400, content=b"")
|
||||||
|
|
||||||
|
with pytest.raises(RequestRejectedError) as ei:
|
||||||
|
await _complete(_transport_for(handler), _source())
|
||||||
|
assert str(ei.value) == "qwen_1 请求被拒: 400"
|
||||||
|
assert ei.value.body_text == ""
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("body", [b"<html>gateway down</html>", b"\xff\xfe not utf-8"])
|
||||||
|
async def test_non_json_and_non_utf8_bodies_do_not_explode(self, body):
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(400, content=body)
|
||||||
|
|
||||||
|
with pytest.raises(RequestRejectedError) as ei:
|
||||||
|
await _complete(_transport_for(handler), _source())
|
||||||
|
assert ei.value.status_code == 400 # 分类不受 body 形态影响
|
||||||
|
|
||||||
|
async def test_non_stream_path_keeps_the_body(self):
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(400, content=_REJECT_BODY.encode())
|
||||||
|
|
||||||
|
with pytest.raises(RequestRejectedError) as ei:
|
||||||
|
await _complete(_transport_for(handler), _source(), stream=False)
|
||||||
|
assert ei.value.body_text == _REJECT_BODY
|
||||||
|
|
||||||
|
async def test_embedding_path_keeps_the_body(self):
|
||||||
|
def handler(request):
|
||||||
|
return httpx.Response(400, content=_REJECT_BODY.encode())
|
||||||
|
|
||||||
|
with pytest.raises(RequestRejectedError) as ei:
|
||||||
|
await _transport_for(handler).embed(texts=["hi"], source=_source(), call_id="cid-embed")
|
||||||
|
assert ei.value.body_text == _REJECT_BODY
|
||||||
|
|
||||||
|
|
||||||
class TestLifecycle:
|
class TestLifecycle:
|
||||||
async def test_aclose_idempotent(self):
|
async def test_aclose_idempotent(self):
|
||||||
def handler(request):
|
def handler(request):
|
||||||
|
|||||||
@@ -117,6 +117,7 @@ class _DummyRecorder:
|
|||||||
cached_prompt_tokens,
|
cached_prompt_tokens,
|
||||||
model_reported,
|
model_reported,
|
||||||
sampling,
|
sampling,
|
||||||
|
reasoning_tokens,
|
||||||
) -> None: ...
|
) -> None: ...
|
||||||
|
|
||||||
|
|
||||||
|
|||||||
@@ -1,12 +1,18 @@
|
|||||||
"""providers.py 注册表测试(M1 设计 §7;register_provider 为纯函数,无可变全局)。"""
|
"""providers.py 注册表测试(M1 设计 §7;register_provider 为纯函数,无可变全局)。"""
|
||||||
|
|
||||||
import pytest
|
import pytest
|
||||||
|
from loguru import logger
|
||||||
|
|
||||||
from polygateway.providers import (
|
from polygateway.providers import (
|
||||||
|
DEFAULT_CAPABILITIES,
|
||||||
DEFAULT_PROFILES,
|
DEFAULT_PROFILES,
|
||||||
ProviderProfile,
|
ProviderProfile,
|
||||||
|
ThinkingCapability,
|
||||||
|
get_capability,
|
||||||
get_provider,
|
get_provider,
|
||||||
|
register_capability,
|
||||||
register_provider,
|
register_provider,
|
||||||
|
resolve_thinking,
|
||||||
)
|
)
|
||||||
|
|
||||||
|
|
||||||
@@ -24,14 +30,21 @@ class TestDefaultProfiles:
|
|||||||
assert p.thinking_off == {"thinking": {"type": "disabled"}}
|
assert p.thinking_off == {"thinking": {"type": "disabled"}}
|
||||||
assert p.strip_think_tags is False
|
assert p.strip_think_tags is False
|
||||||
|
|
||||||
def test_openai_baseline_profile(self):
|
def test_openai_slots_are_unknown_not_empty(self):
|
||||||
|
"""issue #5: 该段名实践中被复用为任意兼容厂商的兜底(下游把 kimi 挂在此),
|
||||||
|
|
||||||
|
故不能下发任何厂商方言参数。None = 形态未知 → 配了 enable_thinking 即报错,
|
||||||
|
而不是空字典那种"注入了个寂寞"的静默失效。
|
||||||
|
"""
|
||||||
p = get_provider("openai")
|
p = get_provider("openai")
|
||||||
assert p.thinking_on == {} and p.thinking_off == {}
|
assert p.thinking_on is None and p.thinking_off is None
|
||||||
assert p.strip_think_tags is False
|
assert p.strip_think_tags is False
|
||||||
|
|
||||||
def test_minimax_baseline_profile(self):
|
def test_minimax_profile_uses_reasoning_effort(self):
|
||||||
|
"""2026-08-02 实测: reasoning_effort 才是 MiniMax 认的开关。"""
|
||||||
p = get_provider("minimax")
|
p = get_provider("minimax")
|
||||||
assert p.thinking_on == {} and p.thinking_off == {}
|
assert p.thinking_off == {"reasoning_effort": "none"}
|
||||||
|
assert p.thinking_on == {"reasoning_effort": "medium"}
|
||||||
assert p.strip_think_tags is False
|
assert p.strip_think_tags is False
|
||||||
|
|
||||||
def test_unknown_provider_fails_loudly(self):
|
def test_unknown_provider_fails_loudly(self):
|
||||||
@@ -62,3 +75,83 @@ class TestPureFunctionRegistration:
|
|||||||
def test_default_profiles_mapping_is_read_only(self):
|
def test_default_profiles_mapping_is_read_only(self):
|
||||||
with pytest.raises(TypeError):
|
with pytest.raises(TypeError):
|
||||||
DEFAULT_PROFILES["hack"] = None # type: ignore[index]
|
DEFAULT_PROFILES["hack"] = None # type: ignore[index]
|
||||||
|
|
||||||
|
|
||||||
|
def _warnings():
|
||||||
|
"""捕获库发出的 WARNING;loguru 不经标准 logging,pytest 的 caplog 抓不到。"""
|
||||||
|
messages: list[str] = []
|
||||||
|
sink_id = logger.add(messages.append, level="WARNING")
|
||||||
|
return messages, sink_id
|
||||||
|
|
||||||
|
|
||||||
|
class TestThinkingCapability:
|
||||||
|
"""issue #5: 能力按 model 登记——同一 provider 内部代际差异是决定性的。"""
|
||||||
|
|
||||||
|
def test_registered_models_carry_evidence(self):
|
||||||
|
"""登记必须附实测证据: 表会过期,没有出处就无从判断该不该信。"""
|
||||||
|
for model in ("MiniMax-M3", "MiniMax-M2.7", "MiniMax-M2.5"):
|
||||||
|
cap = get_capability(model)
|
||||||
|
assert cap is not None and cap.evidence.strip()
|
||||||
|
|
||||||
|
def test_m3_can_disable_but_m2x_cannot(self):
|
||||||
|
assert get_capability("MiniMax-M3").can_disable is True
|
||||||
|
assert get_capability("MiniMax-M2.7").can_disable is False
|
||||||
|
assert get_capability("MiniMax-M2.5").can_disable is False
|
||||||
|
|
||||||
|
def test_unregistered_model_is_unknown(self):
|
||||||
|
assert get_capability("some-brand-new-model") is None
|
||||||
|
|
||||||
|
def test_register_capability_is_pure(self):
|
||||||
|
table = register_capability("x-1", ThinkingCapability(True, "实测"))
|
||||||
|
assert get_capability("x-1", table=table) is not None
|
||||||
|
assert get_capability("x-1") is None # 默认表未被污染
|
||||||
|
|
||||||
|
def test_default_capabilities_mapping_is_read_only(self):
|
||||||
|
with pytest.raises(TypeError):
|
||||||
|
DEFAULT_CAPABILITIES["hack"] = None # type: ignore[index]
|
||||||
|
|
||||||
|
|
||||||
|
class TestResolveThinking:
|
||||||
|
"""五条判定规则(顺序即语义);设计 §5 真值表。"""
|
||||||
|
|
||||||
|
def test_rule1_none_injects_nothing(self):
|
||||||
|
got = resolve_thinking(get_provider("minimax"), None, None, model="MiniMax-M3")
|
||||||
|
assert got == {}
|
||||||
|
|
||||||
|
@pytest.mark.parametrize("enable", [True, False])
|
||||||
|
def test_rule2_unknown_shape_raises_and_points_the_way(self, enable):
|
||||||
|
with pytest.raises(ValueError, match="register_provider") as exc:
|
||||||
|
resolve_thinking(get_provider("openai"), None, enable, model="kimi-k3")
|
||||||
|
assert "extra_body" in str(exc.value)
|
||||||
|
|
||||||
|
def test_rule3_unregistered_model_warns_but_passes(self):
|
||||||
|
messages, sink_id = _warnings()
|
||||||
|
try:
|
||||||
|
got = resolve_thinking(get_provider("minimax"), None, False, model="MiniMax-M9")
|
||||||
|
finally:
|
||||||
|
logger.remove(sink_id)
|
||||||
|
assert got == {"reasoning_effort": "none"}
|
||||||
|
assert any("MiniMax-M9" in m for m in messages)
|
||||||
|
|
||||||
|
def test_rule4_cannot_disable_raises_with_the_model_name(self):
|
||||||
|
cap = get_capability("MiniMax-M2.7")
|
||||||
|
with pytest.raises(ValueError, match="MiniMax-M2.7"):
|
||||||
|
resolve_thinking(get_provider("minimax"), cap, False, model="MiniMax-M2.7")
|
||||||
|
|
||||||
|
def test_rule4_only_blocks_the_off_direction(self):
|
||||||
|
"""关不掉 ≠ 开不了: M2.x 默认就在推理,开的方向不该被拦。"""
|
||||||
|
cap = get_capability("MiniMax-M2.7")
|
||||||
|
got = resolve_thinking(get_provider("minimax"), cap, True, model="MiniMax-M2.7")
|
||||||
|
assert got == {"reasoning_effort": "medium"}
|
||||||
|
|
||||||
|
def test_rule5_normal_path(self):
|
||||||
|
cap = get_capability("MiniMax-M3")
|
||||||
|
assert resolve_thinking(get_provider("minimax"), cap, False, model="MiniMax-M3") == {
|
||||||
|
"reasoning_effort": "none"
|
||||||
|
}
|
||||||
|
|
||||||
|
def test_unknown_shape_beats_capability_check(self):
|
||||||
|
"""第 2 步先于第 4 步: 形态未知时无从注入,能力如何无关紧要。"""
|
||||||
|
cap = ThinkingCapability(can_disable=False, evidence="构造")
|
||||||
|
with pytest.raises(ValueError, match="register_provider"):
|
||||||
|
resolve_thinking(get_provider("openai"), cap, False, model="whatever")
|
||||||
|
|||||||
@@ -68,10 +68,13 @@ class TestConversions:
|
|||||||
_limiter(lease_ttl_s=0)
|
_limiter(lease_ttl_s=0)
|
||||||
|
|
||||||
def test_unknown_source_rejected(self):
|
def test_unknown_source_rejected(self):
|
||||||
from polygateway.errors import GovernanceBackendError
|
"""未知源是装配缺陷,不是后端故障(issue #7 §3.4)。"""
|
||||||
|
from polygateway.errors import GatewayUnavailableError, SourceNotConfiguredError
|
||||||
|
|
||||||
with pytest.raises(GovernanceBackendError):
|
with pytest.raises(SourceNotConfiguredError) as ei:
|
||||||
_limiter()._cfg("nope")
|
_limiter()._cfg("nope")
|
||||||
|
# 关键: 若归入 scope 级家族,配置写错的任务会永远延期重投、永不进死信
|
||||||
|
assert not isinstance(ei.value, GatewayUnavailableError)
|
||||||
|
|
||||||
|
|
||||||
class TestLuaFidelity:
|
class TestLuaFidelity:
|
||||||
|
|||||||
@@ -209,11 +209,13 @@ class TestObservabilityPassthrough:
|
|||||||
raw={},
|
raw={},
|
||||||
cached_prompt_tokens=64,
|
cached_prompt_tokens=64,
|
||||||
model_reported="MiniMax-Text-01-250321",
|
model_reported="MiniMax-Text-01-250321",
|
||||||
|
reasoning_tokens=7,
|
||||||
)
|
)
|
||||||
mw, *_ = _harness([_src("a")], [result])
|
mw, *_ = _harness([_src("a")], [result])
|
||||||
resp = await mw(_REQ)
|
resp = await mw(_REQ)
|
||||||
assert resp.cached_prompt_tokens == 64
|
assert resp.cached_prompt_tokens == 64
|
||||||
assert resp.model_reported == "MiniMax-Text-01-250321"
|
assert resp.model_reported == "MiniMax-Text-01-250321"
|
||||||
|
assert resp.reasoning_tokens == 7
|
||||||
# model 仍是配置别名: 真实版本是旁证,不顶替溯源主字段
|
# model 仍是配置别名: 真实版本是旁证,不顶替溯源主字段
|
||||||
assert resp.model == "m"
|
assert resp.model == "m"
|
||||||
|
|
||||||
@@ -221,6 +223,7 @@ class TestObservabilityPassthrough:
|
|||||||
mw, *_ = _harness([_src("a")], [_ok()])
|
mw, *_ = _harness([_src("a")], [_ok()])
|
||||||
resp = await mw(_REQ)
|
resp = await mw(_REQ)
|
||||||
assert resp.cached_prompt_tokens is None and resp.model_reported is None
|
assert resp.cached_prompt_tokens is None and resp.model_reported is None
|
||||||
|
assert resp.reasoning_tokens is None
|
||||||
|
|
||||||
|
|
||||||
class TestRetryAndFailover:
|
class TestRetryAndFailover:
|
||||||
|
|||||||
@@ -1,4 +1,4 @@
|
|||||||
"""遥测子系统测试: SQLiteRecorder(21 列)+ TelemetryEmitter 单一 helper + TelemetryMW。"""
|
"""遥测子系统测试: SQLiteRecorder(22 列)+ TelemetryEmitter 单一 helper + TelemetryMW。"""
|
||||||
|
|
||||||
import asyncio
|
import asyncio
|
||||||
import json
|
import json
|
||||||
@@ -39,6 +39,7 @@ _EXPECTED_COLUMNS = [
|
|||||||
"cached_prompt_tokens",
|
"cached_prompt_tokens",
|
||||||
"model_reported",
|
"model_reported",
|
||||||
"sampling",
|
"sampling",
|
||||||
|
"reasoning_tokens",
|
||||||
]
|
]
|
||||||
|
|
||||||
|
|
||||||
@@ -102,6 +103,7 @@ async def _record_minimal(recorder, call_id="c1", **overrides):
|
|||||||
"cached_prompt_tokens": None,
|
"cached_prompt_tokens": None,
|
||||||
"model_reported": None,
|
"model_reported": None,
|
||||||
"sampling": None,
|
"sampling": None,
|
||||||
|
"reasoning_tokens": None,
|
||||||
}
|
}
|
||||||
fields.update(overrides)
|
fields.update(overrides)
|
||||||
await recorder.record_llm_call(**fields)
|
await recorder.record_llm_call(**fields)
|
||||||
@@ -158,6 +160,22 @@ class TestSQLiteRecorder:
|
|||||||
assert rows["c-zero"] == 0 # 真实零命中,读回仍是 0 而非 NULL
|
assert rows["c-zero"] == 0 # 真实零命中,读回仍是 0 而非 NULL
|
||||||
assert rows["c-none"] is None
|
assert rows["c-none"] is None
|
||||||
|
|
||||||
|
async def test_reasoning_tokens_column_round_trip(self, tmp_path):
|
||||||
|
"""issue #6: 7 / 0 / None 三种值各自如实落库,0 与 NULL 不得混同。"""
|
||||||
|
recorder = SQLiteRecorder(tmp_path / "t.db")
|
||||||
|
await _record_minimal(recorder, call_id="r-some", reasoning_tokens=7)
|
||||||
|
await _record_minimal(recorder, call_id="r-zero", reasoning_tokens=0)
|
||||||
|
await _record_minimal(recorder, call_id="r-none", reasoning_tokens=None)
|
||||||
|
recorder.close()
|
||||||
|
rows = dict(
|
||||||
|
sqlite3.connect(tmp_path / "t.db")
|
||||||
|
.execute("SELECT call_id, reasoning_tokens FROM llm_calls")
|
||||||
|
.fetchall()
|
||||||
|
)
|
||||||
|
assert rows["r-some"] == 7
|
||||||
|
assert rows["r-zero"] == 0 # 上报了且确实没推理
|
||||||
|
assert rows["r-none"] is None # 本次调用未上报
|
||||||
|
|
||||||
async def test_sampling_column_round_trips(self, tmp_path):
|
async def test_sampling_column_round_trips(self, tmp_path):
|
||||||
"""issue #4: 采样参数落库,否则事后无法证明某批数据跑在什么温度下。"""
|
"""issue #4: 采样参数落库,否则事后无法证明某批数据跑在什么温度下。"""
|
||||||
recorder = SQLiteRecorder(tmp_path / "t.db")
|
recorder = SQLiteRecorder(tmp_path / "t.db")
|
||||||
@@ -239,17 +257,41 @@ class TestSQLiteColumnBackfill:
|
|||||||
|
|
||||||
|
|
||||||
class _FakePgConn:
|
class _FakePgConn:
|
||||||
"""记录执行过的语句;可让 ALTER 抛错以模拟权限不足。"""
|
"""记录执行过的语句;可让 ALTER/CREATE/探测抛错以模拟权限不足与抖动。
|
||||||
|
|
||||||
def __init__(self, existing: list[str], *, fail_alter: bool = False):
|
`existing` 为空列表即表示**表不存在**(与真实 PG 一致: `to_regclass` 为 NULL
|
||||||
|
时列探测必然零行),故 `fetchval` 与 `fetch` 共用同一份事实。
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
existing: list[str],
|
||||||
|
*,
|
||||||
|
fail_alter: bool = False,
|
||||||
|
fail_create: bool = False,
|
||||||
|
probe_errors: int = 0,
|
||||||
|
):
|
||||||
self.existing = existing
|
self.existing = existing
|
||||||
self.fail_alter = fail_alter
|
self.fail_alter = fail_alter
|
||||||
|
self.fail_create = fail_create
|
||||||
|
self.probe_errors = probe_errors
|
||||||
self.statements: list[str] = []
|
self.statements: list[str] = []
|
||||||
|
|
||||||
async def execute(self, sql, *args):
|
async def execute(self, sql, *args):
|
||||||
self.statements.append(sql)
|
self.statements.append(sql)
|
||||||
if sql.startswith("ALTER TABLE") and self.fail_alter:
|
if sql.startswith("ALTER TABLE") and self.fail_alter:
|
||||||
raise RuntimeError("must be owner of table llm_calls")
|
raise RuntimeError("must be owner of table llm_calls")
|
||||||
|
if sql.lstrip().startswith("CREATE TABLE"):
|
||||||
|
if self.fail_create:
|
||||||
|
raise RuntimeError("permission denied for schema public")
|
||||||
|
self.existing = list(_EXPECTED_COLUMNS)
|
||||||
|
|
||||||
|
async def fetchval(self, sql, *args):
|
||||||
|
self.statements.append(sql)
|
||||||
|
if self.probe_errors > 0:
|
||||||
|
self.probe_errors -= 1
|
||||||
|
raise RuntimeError("connection was closed in the middle of operation")
|
||||||
|
return "llm_calls" if self.existing else None
|
||||||
|
|
||||||
async def fetch(self, sql, *args):
|
async def fetch(self, sql, *args):
|
||||||
self.statements.append(sql)
|
self.statements.append(sql)
|
||||||
@@ -284,6 +326,7 @@ class TestPostgresBackfillDiscipline:
|
|||||||
"cached_prompt_tokens",
|
"cached_prompt_tokens",
|
||||||
"model_reported",
|
"model_reported",
|
||||||
"sampling",
|
"sampling",
|
||||||
|
"reasoning_tokens",
|
||||||
]
|
]
|
||||||
|
|
||||||
def _recorder(self, conn):
|
def _recorder(self, conn):
|
||||||
@@ -319,6 +362,76 @@ class TestPostgresBackfillDiscipline:
|
|||||||
assert all("IF NOT EXISTS" not in s for s in altered) # 探测已确认缺列,无需再判
|
assert all("IF NOT EXISTS" not in s for s in altered) # 探测已确认缺列,无需再判
|
||||||
|
|
||||||
|
|
||||||
|
class TestPostgresTableProbe:
|
||||||
|
"""建表必须先探测,且"判死"只认"确定写不进去"(issue #9)。
|
||||||
|
|
||||||
|
实测(PostgreSQL 16.14,只有表级 SELECT/INSERT 的角色): `CREATE TABLE IF NOT
|
||||||
|
EXISTS` 被拒 permission denied for schema,而同一连接的 `INSERT` 通过——
|
||||||
|
PG 对 schema 的 CREATE 权限检查早于 `IF NOT EXISTS` 的存在性判断。无条件发
|
||||||
|
DDL 会让这类最小权限部署的整个进程静默失遥测。
|
||||||
|
"""
|
||||||
|
|
||||||
|
_CURRENT = [
|
||||||
|
"call_id",
|
||||||
|
"cost",
|
||||||
|
"created_at",
|
||||||
|
"cached_prompt_tokens",
|
||||||
|
"model_reported",
|
||||||
|
"sampling",
|
||||||
|
"reasoning_tokens",
|
||||||
|
]
|
||||||
|
|
||||||
|
def _recorder(self, conn):
|
||||||
|
from polygateway.telemetry.postgres import PostgresRecorder
|
||||||
|
|
||||||
|
return PostgresRecorder("postgresql://u:p@h:5432/polygateway", pool=_FakePgPool(conn))
|
||||||
|
|
||||||
|
def _created(self, conn):
|
||||||
|
return [s for s in conn.statements if s.lstrip().startswith("CREATE TABLE")]
|
||||||
|
|
||||||
|
async def test_existing_table_is_never_recreated(self):
|
||||||
|
"""表已存在就一条 DDL 都不发——这是权限被拒的唯一根治办法。"""
|
||||||
|
conn = _FakePgConn(self._CURRENT)
|
||||||
|
await _record_minimal(self._recorder(conn))
|
||||||
|
assert not self._created(conn)
|
||||||
|
|
||||||
|
async def test_create_denied_on_existing_table_keeps_recording(self):
|
||||||
|
"""就算 DDL 仍被发出并被拒,表存在时也不得判死整个 recorder。"""
|
||||||
|
conn = _FakePgConn(self._CURRENT, fail_create=True)
|
||||||
|
recorder = self._recorder(conn)
|
||||||
|
await _record_minimal(recorder) # 不得抛
|
||||||
|
assert recorder._failed is False
|
||||||
|
assert any(s.startswith("INSERT INTO llm_calls") for s in conn.statements)
|
||||||
|
|
||||||
|
async def test_missing_table_is_created_and_not_backfilled(self):
|
||||||
|
"""表不存在→建表;新建表列已齐全,不得再发补列 ALTER。"""
|
||||||
|
conn = _FakePgConn([])
|
||||||
|
recorder = self._recorder(conn)
|
||||||
|
await _record_minimal(recorder)
|
||||||
|
assert len(self._created(conn)) == 1
|
||||||
|
assert not [s for s in conn.statements if s.startswith("ALTER TABLE")]
|
||||||
|
assert recorder._failed is False
|
||||||
|
assert any(s.startswith("INSERT INTO llm_calls") for s in conn.statements)
|
||||||
|
|
||||||
|
async def test_create_failure_on_missing_table_degrades_to_noop(self):
|
||||||
|
"""表确定不存在且建不出来 = 确定写不进去: 此时才允许永久 no-op。"""
|
||||||
|
conn = _FakePgConn([], fail_create=True)
|
||||||
|
recorder = self._recorder(conn)
|
||||||
|
await _record_minimal(recorder) # 不得抛
|
||||||
|
assert recorder._failed is True
|
||||||
|
assert not [s for s in conn.statements if s.startswith("INSERT INTO llm_calls")]
|
||||||
|
|
||||||
|
async def test_probe_failure_is_transient_not_terminal(self):
|
||||||
|
"""探测失败多为连接抖动: 跳过本次,下次调用必须重试,绝不永久判死。"""
|
||||||
|
conn = _FakePgConn(self._CURRENT, probe_errors=1)
|
||||||
|
recorder = self._recorder(conn)
|
||||||
|
await _record_minimal(recorder, call_id="first") # 不得抛
|
||||||
|
assert recorder._failed is False
|
||||||
|
assert not [s for s in conn.statements if s.startswith("INSERT INTO llm_calls")]
|
||||||
|
await _record_minimal(recorder, call_id="second")
|
||||||
|
assert [s for s in conn.statements if s.startswith("INSERT INTO llm_calls")]
|
||||||
|
|
||||||
|
|
||||||
class _MemoryRecorder:
|
class _MemoryRecorder:
|
||||||
def __init__(self):
|
def __init__(self):
|
||||||
self.rows = []
|
self.rows = []
|
||||||
@@ -384,11 +497,12 @@ class TestEmitterObservabilityFields:
|
|||||||
source=_source(),
|
source=_source(),
|
||||||
call_id="cid-1",
|
call_id="cid-1",
|
||||||
latency_ms=42,
|
latency_ms=42,
|
||||||
response=_resp(cached_prompt_tokens=64, model_reported="m-real"),
|
response=_resp(cached_prompt_tokens=64, model_reported="m-real", reasoning_tokens=7),
|
||||||
error=None,
|
error=None,
|
||||||
)
|
)
|
||||||
assert rec.rows[0]["cached_prompt_tokens"] == 64
|
assert rec.rows[0]["cached_prompt_tokens"] == 64
|
||||||
assert rec.rows[0]["model_reported"] == "m-real"
|
assert rec.rows[0]["model_reported"] == "m-real"
|
||||||
|
assert rec.rows[0]["reasoning_tokens"] == 7
|
||||||
|
|
||||||
async def test_failed_attempt_has_no_provider_facts(self):
|
async def test_failed_attempt_has_no_provider_facts(self):
|
||||||
rec = _MemoryRecorder()
|
rec = _MemoryRecorder()
|
||||||
@@ -402,16 +516,19 @@ class TestEmitterObservabilityFields:
|
|||||||
)
|
)
|
||||||
assert rec.rows[0]["cached_prompt_tokens"] is None
|
assert rec.rows[0]["cached_prompt_tokens"] is None
|
||||||
assert rec.rows[0]["model_reported"] is None
|
assert rec.rows[0]["model_reported"] is None
|
||||||
|
assert rec.rows[0]["reasoning_tokens"] is None
|
||||||
|
|
||||||
async def test_cache_hit_replays_the_recorded_values(self):
|
async def test_cache_hit_replays_the_recorded_values(self):
|
||||||
"""决策 B1: 命中行原样回放,故命中率统计必须带 WHERE cache_hit = false。"""
|
"""决策 B1: 命中行原样回放,故命中率统计必须带 WHERE cache_hit = false。"""
|
||||||
rec = _MemoryRecorder()
|
rec = _MemoryRecorder()
|
||||||
await TelemetryEmitter(rec).emit_cache_hit(
|
await TelemetryEmitter(rec).emit_cache_hit(
|
||||||
request=_REQ, response=_resp(cached_prompt_tokens=64, model_reported="m-real")
|
request=_REQ,
|
||||||
|
response=_resp(cached_prompt_tokens=64, model_reported="m-real", reasoning_tokens=7),
|
||||||
)
|
)
|
||||||
row = rec.rows[0]
|
row = rec.rows[0]
|
||||||
assert row["cache_hit"] is True
|
assert row["cache_hit"] is True
|
||||||
assert row["cached_prompt_tokens"] == 64 and row["model_reported"] == "m-real"
|
assert row["cached_prompt_tokens"] == 64 and row["model_reported"] == "m-real"
|
||||||
|
assert row["reasoning_tokens"] == 7 # 与 cached 同口径原样回放
|
||||||
|
|
||||||
async def test_terminal_failure_records_none(self):
|
async def test_terminal_failure_records_none(self):
|
||||||
rec = _MemoryRecorder()
|
rec = _MemoryRecorder()
|
||||||
@@ -420,6 +537,7 @@ class TestEmitterObservabilityFields:
|
|||||||
)
|
)
|
||||||
assert rec.rows[0]["cached_prompt_tokens"] is None
|
assert rec.rows[0]["cached_prompt_tokens"] is None
|
||||||
assert rec.rows[0]["model_reported"] is None
|
assert rec.rows[0]["model_reported"] is None
|
||||||
|
assert rec.rows[0]["reasoning_tokens"] is None
|
||||||
|
|
||||||
|
|
||||||
class TestEmitterSamplingColumn:
|
class TestEmitterSamplingColumn:
|
||||||
|
|||||||
@@ -54,6 +54,7 @@ class TestLLMResponse:
|
|||||||
resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid")
|
resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid")
|
||||||
assert resp.cached_prompt_tokens is None
|
assert resp.cached_prompt_tokens is None
|
||||||
assert resp.model_reported is None
|
assert resp.model_reported is None
|
||||||
|
assert resp.reasoning_tokens is None # issue #6: 本次调用未上报
|
||||||
filled = LLMResponse(
|
filled = LLMResponse(
|
||||||
"c",
|
"c",
|
||||||
"t",
|
"t",
|
||||||
@@ -68,9 +69,11 @@ class TestLLMResponse:
|
|||||||
"cid",
|
"cid",
|
||||||
cached_prompt_tokens=0,
|
cached_prompt_tokens=0,
|
||||||
model_reported="MiniMax-Text-01-250321",
|
model_reported="MiniMax-Text-01-250321",
|
||||||
|
reasoning_tokens=0,
|
||||||
)
|
)
|
||||||
assert filled.cached_prompt_tokens == 0 # 真实零命中,不得与 None 混同
|
assert filled.cached_prompt_tokens == 0 # 真实零命中,不得与 None 混同
|
||||||
assert filled.model_reported == "MiniMax-Text-01-250321"
|
assert filled.model_reported == "MiniMax-Text-01-250321"
|
||||||
|
assert filled.reasoning_tokens == 0 # 上报了且确实没推理,不得与 None 混同
|
||||||
|
|
||||||
def test_frozen(self):
|
def test_frozen(self):
|
||||||
resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid")
|
resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid")
|
||||||
@@ -247,6 +250,7 @@ class TestAuxTypes:
|
|||||||
assert s.raw["id"] == "x"
|
assert s.raw["id"] == "x"
|
||||||
# issue #3: 新字段带默认值,不填也能构造(OCR 等其他 transport 零改动)
|
# issue #3: 新字段带默认值,不填也能构造(OCR 等其他 transport 零改动)
|
||||||
assert s.cached_prompt_tokens is None and s.model_reported is None
|
assert s.cached_prompt_tokens is None and s.model_reported is None
|
||||||
|
assert s.reasoning_tokens is None
|
||||||
|
|
||||||
|
|
||||||
class TestOcrTypes:
|
class TestOcrTypes:
|
||||||
|
|||||||
@@ -0,0 +1,220 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""机械校验 Gitea Wiki 与源码的可比对事实(签名/导出/字段序/列清单/env 键)。
|
||||||
|
|
||||||
|
**只查机械可比对的部分**——机制语义、行为口径这类需要读懂代码才能判断的断言
|
||||||
|
不在此列(那部分靠 `解释-治理行为` 的适用性总表做单一事实源 + 人工审查)。
|
||||||
|
|
||||||
|
设计动机: 2026-08 对 wiki 做了七轮人工审查,88 条发现里有相当一部分属于
|
||||||
|
"机械可校验却写错"——`gather_bounded(coros, limit)`(实为 keyword-only 的
|
||||||
|
`concurrency`)、`LLMResponse` 字段表把 `source_name` 排进前 11 位、遥测列数
|
||||||
|
写成 20/21(实为 22 列)、`__version__` 在自称"全集"的页面缺席。这类偏差不该
|
||||||
|
靠人一轮轮追,故收敛为脚本。
|
||||||
|
|
||||||
|
用法(wiki 是独立仓库,须显式给路径;**不做 skip 静默降级**):
|
||||||
|
python3 tools/check_wiki_alignment.py --wiki /path/to/PolyGateway.wiki
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import dataclasses
|
||||||
|
import inspect
|
||||||
|
import re
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import polygateway
|
||||||
|
from polygateway import EmbeddingClient, GatewayClient, LLMResponse
|
||||||
|
from polygateway.config import _SOURCE_FIELDS
|
||||||
|
from polygateway.ocr import OcrClient
|
||||||
|
from polygateway.providers import register_provider
|
||||||
|
from polygateway.telemetry.sqlite import _COLUMNS as TELEMETRY_COLUMNS
|
||||||
|
|
||||||
|
# 参数名允许在 wiki 里以别名出现的白名单(仅限确无歧义的自解释形参)
|
||||||
|
_PARAM_ALIASES: dict[str, set[str]] = {"env": {"env"}}
|
||||||
|
|
||||||
|
# (符号, 可调用对象) —— 这些的签名必须在 wiki 里逐参数出现
|
||||||
|
_SIGNATURE_TARGETS = [
|
||||||
|
("GatewayClient.chat", GatewayClient.chat),
|
||||||
|
("GatewayClient.from_env", GatewayClient.from_env),
|
||||||
|
("EmbeddingClient.from_env", EmbeddingClient.from_env),
|
||||||
|
("EmbeddingClient.embed", EmbeddingClient.embed),
|
||||||
|
("OcrClient.from_env", OcrClient.from_env),
|
||||||
|
("OcrClient.recognize_text", OcrClient.recognize_text),
|
||||||
|
("gather_bounded", polygateway.gather_bounded),
|
||||||
|
("register_provider", register_provider),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def _wiki_text(wiki: Path) -> dict[str, str]:
|
||||||
|
"""读全部 .md;文件名(不含后缀)→ 正文。"""
|
||||||
|
pages = {p.stem: p.read_text(encoding="utf-8") for p in sorted(wiki.glob("*.md"))}
|
||||||
|
if not pages:
|
||||||
|
raise SystemExit(f"错误: {wiki} 下没有 .md 文件,路径是否指向 wiki 克隆?")
|
||||||
|
return pages
|
||||||
|
|
||||||
|
|
||||||
|
def check_exports_documented(pages: dict[str, str]) -> list[str]:
|
||||||
|
"""`__all__` 每一项都得在某页出现过(R1 漏 gather_bounded、R6 漏 __version__)。"""
|
||||||
|
blob = "\n".join(pages.values())
|
||||||
|
missing = [name for name in polygateway.__all__ if name not in blob]
|
||||||
|
return [f"__all__ 的 {name!r} 在全部 wiki 页面中零命中(页首自称『顶层导出全集』)"
|
||||||
|
for name in missing]
|
||||||
|
|
||||||
|
|
||||||
|
def check_signatures(pages: dict[str, str]) -> list[str]:
|
||||||
|
"""提到某个公共可调用的那一行,必须列全它的参数名。
|
||||||
|
|
||||||
|
只查「参数名是否出现」,不查顺序与类型——后者用自然语言表述合法。
|
||||||
|
历史命中: gather_bounded 的 concurrency 被写成 limit;EmbeddingClient/
|
||||||
|
OcrClient 的 from_env 用省略号承接 chat 的关键字集合,掩盖了没有 cache=。
|
||||||
|
"""
|
||||||
|
problems = []
|
||||||
|
for label, func in _SIGNATURE_TARGETS:
|
||||||
|
symbol = label.split(".")[-1]
|
||||||
|
params = [
|
||||||
|
p.name
|
||||||
|
for p in inspect.signature(func).parameters.values()
|
||||||
|
if p.name not in ("self", "cls")
|
||||||
|
]
|
||||||
|
# 找出提到该符号的所有行,任一行列全即算通过
|
||||||
|
lines = [
|
||||||
|
line
|
||||||
|
for text in pages.values()
|
||||||
|
for line in text.splitlines()
|
||||||
|
if f"`{symbol}`" in line or f"{symbol}(" in line
|
||||||
|
]
|
||||||
|
if not lines:
|
||||||
|
problems.append(f"{label}: wiki 里找不到任何提及")
|
||||||
|
continue
|
||||||
|
best_missing: list[str] | None = None
|
||||||
|
for line in lines:
|
||||||
|
missing = [
|
||||||
|
p for p in params
|
||||||
|
if p not in line and not (_PARAM_ALIASES.get(p, set()) & set(line.split()))
|
||||||
|
]
|
||||||
|
if not missing:
|
||||||
|
best_missing = []
|
||||||
|
break
|
||||||
|
if best_missing is None or len(missing) < len(best_missing):
|
||||||
|
best_missing = missing
|
||||||
|
if best_missing:
|
||||||
|
problems.append(
|
||||||
|
f"{label}: 没有任何一行列全参数,最接近的一行仍缺 {best_missing}"
|
||||||
|
f"(实际签名 {inspect.signature(func)})"
|
||||||
|
)
|
||||||
|
return problems
|
||||||
|
|
||||||
|
|
||||||
|
def check_llmresponse_field_order(pages: dict[str, str]) -> list[str]:
|
||||||
|
"""字段表出现顺序须与 dataclass 声明顺序一致。
|
||||||
|
|
||||||
|
R2 命中: wiki 把 source_name 与 model/provider 并成一行(位置 5),而它实为
|
||||||
|
第 12 个字段。迁移中的三项目按位置构造 fake,照 wiki 写会静默错位。
|
||||||
|
"""
|
||||||
|
page = pages.get("参考-公共API")
|
||||||
|
if page is None:
|
||||||
|
return ["缺少 参考-公共API.md"]
|
||||||
|
# 只在 LLMResponse 小节内找: 别的类型(EmbeddingResponse 等)也有同名字段,
|
||||||
|
# 全页搜索会命中它们、把顺序判断带偏
|
||||||
|
start = page.find("## LLMResponse")
|
||||||
|
if start < 0:
|
||||||
|
return ["参考-公共API.md 缺少 `## LLMResponse` 小节"]
|
||||||
|
end = page.find("\n## ", start + 1)
|
||||||
|
section = page[start : end if end > 0 else len(page)]
|
||||||
|
declared = [f.name for f in dataclasses.fields(LLMResponse)]
|
||||||
|
positions = []
|
||||||
|
for name in declared:
|
||||||
|
# 取该字段在小节内最早的出现位置(表格首列可能写成 `a / b` 合并形式)
|
||||||
|
cands = [
|
||||||
|
section.find(pat)
|
||||||
|
for pat in (f"| {name} ", f"{name} /", f"/ {name} ", f"| {name}\n")
|
||||||
|
]
|
||||||
|
hits = [i for i in cands if i >= 0]
|
||||||
|
positions.append((name, min(hits) if hits else -1))
|
||||||
|
documented = [n for n, i in positions if i >= 0]
|
||||||
|
missing = [n for n, i in positions if i < 0]
|
||||||
|
problems = [f"LLMResponse 字段 {missing} 未在 参考-公共API 的字段表出现"] if missing else []
|
||||||
|
ordered = sorted((i, n) for n, i in positions if i >= 0)
|
||||||
|
actual = [n for _, n in ordered]
|
||||||
|
expected = [n for n in declared if n in documented]
|
||||||
|
if actual != expected:
|
||||||
|
problems.append(
|
||||||
|
f"LLMResponse 字段表顺序与声明顺序不符\n"
|
||||||
|
f" wiki 顺序: {actual}\n"
|
||||||
|
f" 声明顺序: {expected}"
|
||||||
|
)
|
||||||
|
return problems
|
||||||
|
|
||||||
|
|
||||||
|
def check_telemetry_columns(pages: dict[str, str]) -> list[str]:
|
||||||
|
"""遥测列清单与 sqlite 后端的 _COLUMNS 对齐(created_at 由 DDL 生成,单列)。"""
|
||||||
|
page = pages.get("指南-遥测与成本")
|
||||||
|
if page is None:
|
||||||
|
return ["缺少 指南-遥测与成本.md"]
|
||||||
|
expected = [*TELEMETRY_COLUMNS, "created_at"]
|
||||||
|
missing = [c for c in expected if c not in page]
|
||||||
|
problems = [f"遥测列 {missing} 未在 指南-遥测与成本 出现"] if missing else []
|
||||||
|
# 列数声明: 表 = _COLUMNS + created_at;端口 = _COLUMNS。
|
||||||
|
# 页内**所有** "N 列" 声明都必须等于真实列数——只查"正确值是否出现"会被
|
||||||
|
# 漏改的旧数字骗过(它们同时存在时检查照样通过)
|
||||||
|
table_n, port_n = len(expected), len(TELEMETRY_COLUMNS)
|
||||||
|
declared_counts = {int(m) for m in re.findall(r"(\d+)\s*列", page)}
|
||||||
|
if not declared_counts:
|
||||||
|
problems.append(f"指南-遥测与成本 未声明表列数(应为 {table_n} 列)")
|
||||||
|
elif declared_counts != {table_n}:
|
||||||
|
wrong = sorted(declared_counts - {table_n})
|
||||||
|
problems.append(
|
||||||
|
f"指南-遥测与成本 的列数声明 {wrong} 与实际 {table_n} 列不符"
|
||||||
|
f"(端口是 {port_n} 参数,两者差 created_at)"
|
||||||
|
)
|
||||||
|
return problems
|
||||||
|
|
||||||
|
|
||||||
|
def check_source_env_fields(pages: dict[str, str]) -> list[str]:
|
||||||
|
"""`_SOURCE_FIELDS` 的每个 FIELD 段都得在 参考-配置键 出现(R1 命中 EXTRA_BODY)。"""
|
||||||
|
page = pages.get("参考-配置键")
|
||||||
|
if page is None:
|
||||||
|
return ["缺少 参考-配置键.md"]
|
||||||
|
# 用词边界匹配: `f in page` 会让 EXTRA_BODY 被 EXTRA_BODYY 蒙混过关
|
||||||
|
missing = [f for f in _SOURCE_FIELDS if not re.search(rf"\b{re.escape(f)}\b", page)]
|
||||||
|
return [f"源键 FIELD 段 {missing} 未在 参考-配置键 文档化"] if missing else []
|
||||||
|
|
||||||
|
|
||||||
|
_CHECKS = [
|
||||||
|
("顶层导出覆盖", check_exports_documented),
|
||||||
|
("公共签名参数", check_signatures),
|
||||||
|
("LLMResponse 字段序", check_llmresponse_field_order),
|
||||||
|
("遥测列清单", check_telemetry_columns),
|
||||||
|
("源 env 键覆盖", check_source_env_fields),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
parser = argparse.ArgumentParser(description=__doc__)
|
||||||
|
parser.add_argument("--wiki", required=True, type=Path, help="PolyGateway.wiki 克隆目录")
|
||||||
|
args = parser.parse_args()
|
||||||
|
if not args.wiki.is_dir():
|
||||||
|
raise SystemExit(f"错误: {args.wiki} 不是目录")
|
||||||
|
|
||||||
|
pages = _wiki_text(args.wiki)
|
||||||
|
failed = 0
|
||||||
|
for label, check in _CHECKS:
|
||||||
|
problems = check(pages)
|
||||||
|
if problems:
|
||||||
|
failed += len(problems)
|
||||||
|
print(f"✗ {label}")
|
||||||
|
for p in problems:
|
||||||
|
print(f" {p}")
|
||||||
|
else:
|
||||||
|
print(f"✓ {label}")
|
||||||
|
print()
|
||||||
|
if failed:
|
||||||
|
print(f"{failed} 处机械偏差 —— wiki 与源码不一致")
|
||||||
|
return 1
|
||||||
|
print(f"{len(pages)} 页机械校验通过")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
sys.exit(main())
|
||||||
Reference in New Issue
Block a user