35 Commits

Author SHA1 Message Date
iomgaa ce630a37ef feat: pass sampling parameters through to the provider (issue #4) 2026-08-01 00:00:44 -04:00
iomgaa dfda59fec2 chore: release 1.0.5 with sampling parameter passthrough 2026-07-31 23:52:49 -04:00
iomgaa 15b9b02e96 fix: make the sampling invariant test actually enforce the constraint
The test passed overlay and sampling as separate objects while production
aliases them, so an in-place mutation slipped through it. Also syncs the
telemetry schema page and adds the missing postgres round-trip assertion.
2026-07-31 22:01:51 -04:00
iomgaa cce7562d07 test: verify sampling parameters through the full governance stack 2026-07-31 21:46:05 -04:00
iomgaa 2958dc8231 docs: document sampling passthrough and the empty thinking profiles 2026-07-31 21:41:34 -04:00
iomgaa a5ebf72f17 feat: strip extra_body on the embedding and OCR paths with a warning
Stripping is load-bearing, not tidying: those transports never send the
value, so leaving it would make telemetry record a parameter never sent.
2026-07-31 21:36:50 -04:00
iomgaa 4516761dbe feat: record sampling parameters in telemetry (port 20 to 21 fields)
Each of the three emitter entry points has a pinned meaning: only the
attempt path has an effective source, so only it merges extra_body.
2026-07-31 21:30:45 -04:00
iomgaa b6e4cc3f3b test: lock the sampling snapshot invariant across the onion
Verified by breaking structured.py so the reask drops sampling: the test
goes red for the right reason, not merely because something errored.
2026-07-31 21:23:50 -04:00
iomgaa c31cc1adad feat: fold sampling parameters into the cache key
Without this, five seeds over identical messages all hit the first cached
response and the reported standard deviation is silently always zero.
2026-07-31 21:20:56 -04:00
iomgaa 6bb64ca938 feat: accept sampling overlay on chat() and per-source extra_body
Priority is structured injection > per-call overlay > per-source config,
and extra_body now takes part in the cache fingerprint.
2026-07-31 21:17:49 -04:00
iomgaa 152fa264ed feat: parse EXTRA_BODY as a JSON object per source 2026-07-31 21:12:06 -04:00
iomgaa 6023d11bfb feat: add sampling overlay validation and source extra_body
Three pure helpers in the innermost layer plus ChatRequest.sampling as a
cross-layer snapshot, so cache keys and telemetry read one stable value.
2026-07-31 21:09:03 -04:00
iomgaa b12bf6ce79 docs: plan the sampling parameter implementation (issue #4)
Eleven verifiable tasks covering decisions A-G and the 14 test items,
with the reviewer-found execution traps written into the tasks.
2026-07-31 13:01:53 -04:00
iomgaa b24e224beb docs: soften the non-chat extra_body gate to strip-and-warn
Stripping is load-bearing: without it telemetry would record a sampling
parameter that was never sent on the OCR and embedding paths.
2026-07-31 12:31:31 -04:00
iomgaa 09e77f11f8 docs: register the sampling design in the research wiki 2026-07-31 12:00:20 -04:00
iomgaa 0cc89fb03c docs: harden the sampling design against the reviewer findings
Pin the telemetry column semantics across all three emitter entry points,
add the cross-layer sampling snapshot, and reject extra_body on the
embedding and OCR paths instead of accepting it silently.
2026-07-31 11:57:39 -04:00
iomgaa 20fd899d93 docs: design sampling parameter passthrough (issue #4)
Two-layer entry: per-call overlay on chat() and per-source extra_body.
Covers the cache-key and telemetry interactions the issue omitted.
2026-07-31 11:40:14 -04:00
iomgaa 486809b08b feat: expose provider cache tokens and reported model (issue #3) 2026-07-31 11:15:18 -04:00
iomgaa 58cb55b869 chore: release 1.0.4 instead of a minor bump 2026-07-31 11:11:55 -04:00
iomgaa 86fb4d5536 fix: keep the postgres backfill from disabling telemetry or locking the table 2026-07-31 10:42:55 -04:00
iomgaa 32d7869043 fix: harden the observability fields against the verifier findings 2026-07-31 08:28:41 -04:00
iomgaa 966d548245 chore: release 1.1.0 with the response observability fields 2026-07-31 08:08:37 -04:00
iomgaa c2fcd5b1f8 feat: record the observability fields end to end through telemetry 2026-07-31 08:03:43 -04:00
iomgaa 0ed9dc107c feat: support a cached input price tier in the pricing table 2026-07-31 07:58:37 -04:00
iomgaa c4eda119ac feat: carry the new observability fields through retry and cache 2026-07-31 07:56:04 -04:00
iomgaa cd1a9520ff docs: spell out that a reported zero is not a missing value 2026-07-31 07:53:42 -04:00
iomgaa 0aa7202c87 feat: collect provider cache tokens and reported model in transport 2026-07-31 07:51:47 -04:00
iomgaa 4841d901af feat: add cached prompt tokens and reported model to response types 2026-07-31 07:48:35 -04:00
iomgaa 30d7ffd94a docs: register the implementation plan in the research wiki 2026-07-31 07:11:52 -04:00
iomgaa 037e7a011e docs: fold the plan review findings into the plan 2026-07-31 07:09:52 -04:00
iomgaa 0e2f734b0b docs: plan the implementation of the observability fields 2026-07-31 06:59:42 -04:00
iomgaa 7ccb25e8f1 docs: record the human approval of the design decisions 2026-07-31 06:54:55 -04:00
iomgaa 42d16919fc docs: register the design in the research wiki 2026-07-31 04:37:52 -04:00
iomgaa 005a90ca19 docs: fold the independent review findings into the design 2026-07-31 04:35:51 -04:00
iomgaa 8a824e2000 docs: design the response observability fields for issue 3 2026-07-31 04:22:51 -04:00
47 changed files with 3051 additions and 53 deletions
+6
View File
@@ -18,6 +18,10 @@ LLM__QWEN__1__TIMEOUT_S=120
# LLM__QWEN__1__ENABLE_THINKING=true # 三态: 缺省=不注入 / true=注入开启 / false=注入关闭 # LLM__QWEN__1__ENABLE_THINKING=true # 三态: 缺省=不注入 / true=注入开启 / false=注入关闭
# LLM__QWEN__1__MISSING_DONE=retry # SSE 缺 [DONE]: retry(默认) | salvage # LLM__QWEN__1__MISSING_DONE=retry # SSE 缺 [DONE]: retry(默认) | salvage
# LLM__QWEN__1__TRUST_ENV=true # false = 绕过本地代理(LAN 直连) # LLM__QWEN__1__TRUST_ENV=true # false = 绕过本地代理(LAN 直连)
# LLM__QWEN__1__EXTRA_BODY={"temperature":0} # 本源恒定的采样参数(JSON 对象串)
# 并入请求体,优先级低于 chat(overlay=...);受控实验固定解码用它,免得漏传
# 禁用键 model/messages/stream/stream_options(会击穿治理),配了直接报错
# OCR/EMBED scope 不消费该键: 配了会被忽略并 warning(见 issue #4 决策 G)
# ══ scope 级全局闸(跨源合计;0/缺省 = 不启用)══ # ══ scope 级全局闸(跨源合计;0/缺省 = 不启用)══
# LLM__GLOBAL__MAX_CONCURRENCY=8 # LLM__GLOBAL__MAX_CONCURRENCY=8
@@ -54,6 +58,8 @@ PGW_TELEMETRY_BACKEND=none # sqlite | postgres | none(必填)
# PGW_TELEMETRY_SQLITE_PATH=logs/telemetry.db # sqlite 时必填 # PGW_TELEMETRY_SQLITE_PATH=logs/telemetry.db # sqlite 时必填
# PGW_TELEMETRY_PG_DSN=postgresql://user:pass@host:5432/polygateway # postgres 时必填;严禁指向在用业务库(实验室约定: 专用库 polygateway) # PGW_TELEMETRY_PG_DSN=postgresql://user:pass@host:5432/polygateway # postgres 时必填;严禁指向在用业务库(实验室约定: 专用库 polygateway)
# PGW_PRICING_PATH=config/prices.json # 可选: {"<model>": {"input_per_1m": x, "output_per_1m": y}};缺省 cost 恒 None # PGW_PRICING_PATH=config/prices.json # 可选: {"<model>": {"input_per_1m": x, "output_per_1m": y}};缺省 cost 恒 None
# # 可选第三档 "cached_input_per_1m": z —— 供应商 prompt cache 命中部分的单价;
# # 不填即命中部分也按 input 全额计(库不猜折扣率),cost 会偏高
# PGW_CACHE_NAMESPACE=<项目名或租户前缀> # 缓存启用时必填(防跨项目毒化) # PGW_CACHE_NAMESPACE=<项目名或租户前缀> # 缓存启用时必填(防跨项目毒化)
# PGW_CACHE_TTL_S=604800 # 缓存启用时必填,须 > 0 # PGW_CACHE_TTL_S=604800 # 缓存启用时必填,须 > 0
# PGW_STRUCTURED_MAX_RETRIES=2 # 缺省 2(M2.5);0 = 解析失败不重问(CHS 策略) # PGW_STRUCTURED_MAX_RETRIES=2 # 缺省 2(M2.5);0 = 解析失败不重问(CHS 策略)
+38
View File
@@ -1,5 +1,43 @@
# Changelog # Changelog
## 1.0.5(2026-07-31)
采样参数透传(issue #4)。`chat()` 此前没有任何途径设置 `temperature` / `seed` / `max_tokens`——全库检索 `temperature` 零命中,`ChatRequest.overlay` 虽会被并进请求体却只由结构化中间件填充,调用方够不着。对受控实验而言这是阻塞性的:解码温度未知且可能随供应商默认值变化,每格配置跑 5 个 seed 报出的标准差无从解释。
### 新增(纯增,不破坏任何现有调用方)
- **`chat()` 新增 keyword-only 参数 `overlay: Mapping[str, Any] | None = None`**,承载逐次变化的采样参数(每个 rollout 不同的 `seed`)。带默认值的 keyword-only 参数不改变既有调用点。
- **`SourceConfig` 新增 `extra_body` 字段**,对应环境键 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`(JSON **对象**串),承载全局恒定的参数(`temperature=0`)——免得每个调用点都要记得传,而漏传一次不会报错、只会让数字悄悄不可比。
- **优先级为 结构化注入 > 调用级 `overlay` > 源级 `extra_body`。** 由现有层序天然给出,未引入新机制。
- **遥测表 `llm_calls` 新增 `sampling` 列**,`TelemetryRecorder` 端口由 20 字段扩为 21;补列走 1.0.4 已建立的"先探测缺列再 ALTER、失败只逐行降级"套路。列语义是「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,**不含**结构化输出注入的 `response_format`(列名是采样参数,而数 KB 的 schema 逐行落库只会让审计表膨胀)。
### 下游请读
- **采样参数进缓存 key,所以逐次变化的 `seed` 天然全部 miss。** 这是正确语义而非缺陷:不进 key 的话,同 messages 跑 5 个 seed 会全部命中第一次的响应,标准差恒为 0 且不报错。代价是缓存对这条路径不再省钱。**不传采样参数时 key 逐字不变**,存量缓存不受影响。
- **`model_fingerprint` 是集合级指纹,不是本次选中源的指纹。** 同 scope 下各源 `extra_body` 不同时,缓存仍可能返回另一源、另一组解码参数下产生的响应(这是既有取舍的延续,`model` 一直如此)。要求逐源可复现的实验应让每个源独享 scope 或 namespace。
- **`{model, messages, stream, stream_options}` 是保护键,配了直接报 `ValueError`。** 它们由治理层拥有:`model` 被覆盖会让成本按错单价算,`stream`/`stream_options` 会绕过流式看门狗、丢掉 usage 帧。不可 JSON 序列化的值(如 numpy 标量)同样在进洋葱之前报错——否则会在缓存层的降级保护之外抛裸 `TypeError`,连一行遥测都留不下。
- **`SourceConfig` 不再 hashable**,`dataclasses.asdict()` / `copy.deepcopy()` 也不再适用(加任何 mapping 字段的固有代价,裸 dict 亦然)。要可变副本用 `dict(source.extra_body)`,要改字段用 `dataclasses.replace(source, ...)`
- **OCR / embedding 路径不消费 `extra_body`**:配了会被**剥离并 warning**,装配照常成功。这两条路径的 transport 根本不发这个值(embed payload 硬编码 `{model, input}`、MonkeyOCR 只发 multipart 表单),剥离是为了让遥测不至于记录一个从未发出的参数。需要 `dimensions` 等 embedding 参数请提 issue。
- **`enable_thinking``openai` / `minimax` 两个 provider 不产生任何效果**(它们的 thinking profile 两档皆空)。此前没有任何地方说明这一点,调用方可能以为自己关掉了推理。需要下发自定义参数请用 `extra_body`
## 1.0.4(2026-07-31)
响应可观测字段扩展(issue #3)。下游 dissect 要把每次调用落成一行审计记录,其中两列拿不到值:供应商侧 prompt cache 命中了多少 token、这次调用实际跑的是哪个模型版本。前者关系到能否把「缓存命中率差异带来的成本」与「实验条件本身带来的成本」分开,后者关系到实验快照的可复现性。本次把两者暴露到公共类型与遥测表,并让成本换算认识缓存单价。
### 新增(纯增字段,不破坏任何现有调用方)
- **`LLMResponse` 新增 `cached_prompt_tokens: int | None``model_reported: str | None`。** 前者是供应商 prompt cache 命中的输入 token 数(OpenAI 兼容格式的 `usage.prompt_tokens_details.cached_tokens`),后者是 API 响应体里的 `model` 字段(与 `.env` 配的别名可能分叉——供应商把别名指向新权重时,只有它认得出真正跑的那个版本)。两者均带默认值 `None`,逐字段传参的 fake 构造零改动。
- **`None``0` 是两回事,不可混同。** `None` = 该源不上报这个数(下游据此声明「本源不可做缓存成本校正」);`0` = 该源上报了一次真实零命中。网关报文一律不可信:形态异常(负数、字符串、`bool``prompt_tokens_details` 非 dict)一律归 `None` 且绝不抛异常——可观测字段缺失不得打断调用。
- **遥测表 `llm_calls` 新增 `cached_prompt_tokens``model_reported` 两列**,`TelemetryRecorder` 端口由 18 字段扩为 20。两个后端在初始化期对**已存在的旧表幂等补列**——`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入都被逐行 warning 丢弃、遥测静默全失。两侧都是**先探测缺列、只在真缺列时才 ALTER**(SQLite 查 `PRAGMA table_info`,Postgres 查 `pg_attribute`):`ADD COLUMN IF NOT EXISTS` 即使列已存在也会先取 ACCESS EXCLUSIVE 锁,而遥测是内联 await,让每个进程的首次写入都去锁共享审计表会拖垮业务调用;稳态下一条 ALTER 都不会发。**补列失败只降级为逐行丢弃,绝不会让 recorder 整体失能**(应用账号只有 INSERT 权限时,`ALTER TABLE` 的 ownership 检查早于存在性判断,列齐全也会失败)。
- **`PricingTable` 支持可选的缓存读取单价 `cached_input_per_1m`。** 配了该档且本次有命中时按 `(prompt - cached) × input + cached × cached_input` 分段计价,消除 cost 的系统性高估;**未配则不猜折扣率**,退化为现状全额输入价(P5 严禁默认值掩盖)。旧价格表文件与 embedding 侧的三参 `cost()` 调用零改动。命中数超过输入总数时按总数夹取并 warning,不产生负成本。
### 下游请读
- **`cache_hit` 与新字段是两个不同的东西。** `cache_hit` 指的始终是 **PolyGateway 自身的响应缓存**(未产生网关调用),而 `cached_prompt_tokens` 指的是**供应商服务器**复用了提示词前缀、那部分按更低单价计费——真实调用里天天发生,`cache_hit` 永远看不见它。字段名保持不变(改名会破坏迁移兼容),语义已在 docstring 中消歧。
- **统计供应商缓存命中率必须写 `WHERE cache_hit = false`。** 缓存命中行的这两个字段是**原样回放**的历史值(与 `model``prompt_tokens` 同一口径:`CacheMW` 只覆写与本次调用相关的时序字段),计入会重复计数。这与 1.0.3 里 `cost` 缺口口径的坑是同一类。
- 缓存命中行的 `cost` 仍恒为 `0.0`(未产生新调用),该短路排在任何单价换算之前,不受缓存单价档影响。
- 旧格式的缓存条目(缺这两个键)照常可重建为 `None`,不会回源;历史遥测行的新列为 NULL。
## 1.0.3(2026-07-30) ## 1.0.3(2026-07-30)
`est_tokens` 解耦(issue #2):一个常量此前被派了两份对"保守"定义相反的差事——TPM 入场预扣(押多了只是慢,安全)与 usage 缺失时的用量兜底(按上界记账只会账单虚高)。本次把两者拆开。 `est_tokens` 解耦(issue #2):一个常量此前被派了两份对"保守"定义相反的差事——TPM 入场预扣(押多了只是慢,安全)与 usage 缺失时的用量兜底(按上界记账只会账单虚高)。本次把两者拆开。
+1 -1
View File
@@ -80,7 +80,7 @@ async def main() -> None:
await client.aclose() # 归还连接与治理后端资源 await client.aclose() # 归还连接与治理后端资源
``` ```
`chat()` 原生接受 OpenAI 多模态 content 数组(`image_url` data URL),VLM 调用无需专门客户端;`session_id` / `parent_call_id` / `cache_salt` 关键字参数用于链路追踪与缓存控制。 `chat()` 原生接受 OpenAI 多模态 content 数组(`image_url` data URL),VLM 调用无需专门客户端;`session_id` / `parent_call_id` / `cache_salt` 关键字参数用于链路追踪与缓存控制;`overlay` 传采样参数(`temperature` / `seed` / `max_tokens` 等,恒定值宜配在源的 `EXTRA_BODY` 上)——它会进缓存 key,故逐次变化的 `seed` 天然不命中缓存
### 3. OCR 与 Embedding ### 3. OCR 与 Embedding
+6
View File
@@ -0,0 +1,6 @@
{
"MiniMax-M3": {
"input_per_1m": 2.1,
"output_per_1m": 8.4
}
}
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
[project] [project]
name = "polygateway" name = "polygateway"
version = "1.0.3" version = "1.0.5"
description = "PolyGateway:实验室统一的大语言模型(LLM/VLM/OCR)调度与中转库——多源、限流、重试、熔断、缓存、遥测" description = "PolyGateway:实验室统一的大语言模型(LLM/VLM/OCR)调度与中转库——多源、限流、重试、熔断、缓存、遥测"
requires-python = ">=3.11" requires-python = ">=3.11"
dependencies = [ dependencies = [
+25 -5
View File
@@ -328,7 +328,18 @@ flowchart TB
| `cache_hit` | bool | 是否缓存命中 | | `cache_hit` | bool | 是否缓存命中 |
| `call_id` | str | UUID,每次**尝试**独立 | | `call_id` | str | UUID,每次**尝试**独立 |
新增字段(库扩展,全部带默认值): `source_name`(多源溯源)、`cost`(pricing 换算,可为 None)、`usage_source`(三态,见下)、`structured_data`(D14 阶梯通过后的解析产物;不参与缓存序列化,命中时由 CacheMW 复用 strategy 零网络重建)。 新增字段(库扩展,全部带默认值): `source_name`(多源溯源)、`cost`(pricing 换算,可为 None)、`usage_source`(三态,见下)、`structured_data`(D14 阶梯通过后的解析产物;不参与缓存序列化,命中时由 CacheMW 复用 strategy 零网络重建)`cached_prompt_tokens``model_reported`(2026-07-31,issue #3,见下)
**可观测字段(2026-07-31,issue #3;下游 dissect 的调用审计需求)**:
| 字段 | 含义 | 生产者 |
|---|---|---|
| `cached_prompt_tokens` | **供应商侧** prompt cache 命中的输入 token 数(OpenAI 兼容格式的 `usage.prompt_tokens_details.cached_tokens`)。`None` = 该源未上报;`0` = 上报了一次真实零命中——两者对下游处置不同(前者不可做缓存成本校正),故不可混同 | `openai_compat` 两条路径解析后经 `TransportResult` 上浮 |
| `model_reported` | API 响应体里的 `model` 字段;`None` = 未上报。与 `model`(`.env` 配置别名)可能分叉——供应商把别名指向新权重时,实验复现必须认这个串 | 流式取首个含 `model` 的 chunk(首次写入即固定),非流式取 body 顶层 |
`cache_hit` 指的始终是 **PolyGateway 自身响应缓存**,与供应商 prompt cache 无关;两者语义不同但名字相近,docstring 已消歧(改名会破坏迁移兼容,故只注释)。
**缓存命中行的口径(决策 B1)**: 与 `model`/`prompt_tokens` 同一规则——`CacheMW._rehydrate` 只覆写与本次调用相关的时序字段,这两个新字段**原样回放**历史值。故**统计供应商缓存命中率必须写 `WHERE cache_hit = false`**,否则回放行会被重复计数(与 §5.1 `cost` 缺口口径同款教训)。
**`usage_source` 三态值域(2026-07-30,est_tokens 解耦设计;此前为 measured/estimated 两态)**: **`usage_source` 三态值域(2026-07-30,est_tokens 解耦设计;此前为 measured/estimated 两态)**:
@@ -350,6 +361,8 @@ flowchart TB
**`chat()` 公共签名定稿(2026-07-20,GovDoc 迁移缺口 G1/G2)**: `chat(messages, *, session_id=None, parent_call_id=None, cache_salt=None, cache_namespace=None, structured=None, stream=True)`。要点: ① `session_id`/`parent_call_id` 与三项目现有 `LLMProvider.chat` Protocol 逐字兼容——这是"调用点零改动"承诺的前提;② **per-call `cache_namespace`**: GovDoc 是单 client 服务多租户、tenant 每请求变化,装配级 namespace 只是默认值,per-call 传入时覆盖并进入缓存 key(§7.5);③ `cache_salt` per-call 可传(Video-Tree 跨 epoch 重采样);④ `structured` 三档语义(D14),类型定稿 `type[BaseModel] | Literal["json"] | None`(M1 设计): 不传 = 原始文本,`"json"` = 仅修复,pydantic 模型 = 完整阶梯(修复+形态校验+有界带反馈重问)。 **`chat()` 公共签名定稿(2026-07-20,GovDoc 迁移缺口 G1/G2)**: `chat(messages, *, session_id=None, parent_call_id=None, cache_salt=None, cache_namespace=None, structured=None, stream=True)`。要点: ① `session_id`/`parent_call_id` 与三项目现有 `LLMProvider.chat` Protocol 逐字兼容——这是"调用点零改动"承诺的前提;② **per-call `cache_namespace`**: GovDoc 是单 client 服务多租户、tenant 每请求变化,装配级 namespace 只是默认值,per-call 传入时覆盖并进入缓存 key(§7.5);③ `cache_salt` per-call 可传(Video-Tree 跨 epoch 重采样);④ `structured` 三档语义(D14),类型定稿 `type[BaseModel] | Literal["json"] | None`(M1 设计): 不传 = 原始文本,`"json"` = 仅修复,pydantic 模型 = 完整阶梯(修复+形态校验+有界带反馈重问)。
**`overlay` 追加(2026-07-31,issue #4)**: 签名末尾增 `overlay: Mapping[str, Any] | None = None`,承载采样参数(`temperature`/`seed`/`max_tokens` 等)。带默认值的 keyword-only 参数不改变既有调用点,"签名冻结"承诺不破。要点: ① 优先级 **结构化注入 > 调用级 overlay > 源级 `extra_body`**,由 `StructuredMW``{**request.overlay, **strategy_overlay}` 与 transport `_build_payload` 的 update 顺序天然给出,无新机制;② 保护键 `{model, messages, stream, stream_options}` 与不可 JSON 序列化的值在**进洋葱之前**报 `ValueError`(前者被覆盖会击穿成本换算/缓存口径/流式看门狗/usage 帧,后者会在 `CacheMW` 的降级 try 之外抛裸 `TypeError` 且一行遥测都没有);③ 同时填 `ChatRequest.sampling` 快照字段——`overlay` 在洋葱不同深度取值不同(内层含 `response_format`),缓存 key 与遥测需要一个跨层恒定的读取点,否则同一列在不同行口径分叉。
--- ---
## 6. 错误模型 ## 6. 错误模型
@@ -424,11 +437,13 @@ flowchart TB
### 7.5 响应缓存 ### 7.5 响应缓存
**key 公式**: `sha256(canonical_json({model, messages_digest, namespace, salt}))`,前缀 `pgw:cache:` **key 公式**: `sha256(canonical_json({model, messages_digest, namespace, salt, sampling}))`,前缀 `pgw:cache:`
- `messages_digest`: 文本部分原文参与;多模态 content part(base64 图像等)先各自 sha256 摘要再参与——修正 Video-Tree 把整段 base64 进 hash 的开销问题,且 key 稳定性不变。 - `messages_digest`: 文本部分原文参与;多模态 content part(base64 图像等)先各自 sha256 摘要再参与——修正 Video-Tree 把整段 base64 进 hash 的开销问题,且 key 稳定性不变。
- `namespace`: 必填(项目名/租户 id),修正 GovDoc 缓存 key 缺租户隔离与多项目共用 Redis 时的互相毒化风险。 - `namespace`: 必填(项目名/租户 id),修正 GovDoc 缓存 key 缺租户隔离与多项目共用 Redis 时的互相毒化风险。
- `salt`: 可选,跨 epoch 强制重采样(Video-Tree 需求)。 - `salt`: 可选,跨 epoch 强制重采样(Video-Tree 需求)。
- `sampling`(2026-07-31,issue #4): 调用级采样参数,**仅非空时参与**(注意与 `salt` 的"仅非 None"不同——空串是有意义的 salt,而空采样参数与不传无差别),故空 overlay 时旧键逐字不变、存量缓存不冷启动。读 `request.sampling` 而非 `request.overlay`,不依赖"CacheMW 恰在 StructuredMW 外侧"的层序巧合。**不进 key 的后果**: 同 messages 跑 5 个 seed 会全部命中第一次的响应,标准差恒为 0 且不报错——受控实验静默作废。源级 `extra_body` 同理并入 `model_fingerprint`(全源皆空时字面量不变,否则追加 `|sha256(...)`,摘要对象是各源 `(model, extra_body)` 的 canonical JSON 排序去重——按模型而非源名,改源名不误触冷启动)。
- **两条已知副作用**: ① 逐 rollout 变化的 `seed` 进 key 后该路径天然全部 miss(正确语义,但缓存对它不再省钱);② `model_fingerprint` 是**集合级**指纹而非本次选中源的指纹,同 scope 各源 `extra_body` 不同时仍可能返回另一源的响应(既有取舍的延续,与 `model` 同),要求逐源可复现应让每源独享 scope 或 namespace。
- value = `LLMResponse` 的 JSON;TTL 必填且 > 0(禁止永不过期,继承 Video-Tree 校验);Redis 不可用 → get 返回 None、set 吞异常记 warning(静默降级)。**只缓存成功响应**;`ResultInvalidError` 的原始响应不缓存(避免固化坏结果)。 - value = `LLMResponse` 的 JSON;TTL 必填且 > 0(禁止永不过期,继承 Video-Tree 校验);Redis 不可用 → get 返回 None、set 吞异常记 warning(静默降级)。**只缓存成功响应**;`ResultInvalidError` 的原始响应不缓存(避免固化坏结果)。
### 7.6 流式活性看门狗 ### 7.6 流式活性看门狗
@@ -437,7 +452,7 @@ flowchart TB
### 7.7 多源与选源 ### 7.7 多源与选源
`SourceConfig`: name/provider/base_url/api_key/model/超时组/限额组(单源并发/RPM/TPM)/`est_tokens`(TPM 预扣量的**可选调优覆盖**,移植 CHS `config.py:55`;2026-07-20 缺口 G2 补,2026-07-30 由必填降为可选)/enable_thinking。聚合自环境变量 `{SCOPE}__{PROVIDER}__{N}__{FIELD}`(§9)。 `SourceConfig`: name/provider/base_url/api_key/model/超时组/限额组(单源并发/RPM/TPM)/`est_tokens`(TPM 预扣量的**可选调优覆盖**,移植 CHS `config.py:55`;2026-07-20 缺口 G2 补,2026-07-30 由必填降为可选)/enable_thinking/`extra_body`(2026-07-31 issue #4: 本源恒定的采样参数,构造期校验保护键后转 `MappingProxyType`;**该字段令 SourceConfig 不再 hashable**——加任何 mapping 字段的固有代价,库内无以源作 dict key/set 元素的写法,要可变副本用 `dict(...)`、要改字段用 `dataclasses.replace`)。聚合自环境变量 `{SCOPE}__{PROVIDER}__{N}__{FIELD}`(§9)。
**TPM 有效预扣量(2026-07-30,est_tokens 解耦设计,G2 闭环)**: `try_acquire`(§7.3)传入的 est 来自 `SourceConfig.effective_est_tokens()` 这一份纯方法,五个调用点(`QuotaGate` 入场 + chat/embedding 各自的成功侧与失败侧结算)共用,保证预扣与结算恒取同一值(`delta == 0`,否则押金会被整笔退回、TPM 闸退化成进门即放行)。规则:显式 `est_tokens > 0` 则原样用;否则 `tpm > 0` 时派生 `max(1, tpm // 60)`;`tpm == 0`(该闸不启用)时为 0。 **TPM 有效预扣量(2026-07-30,est_tokens 解耦设计,G2 闭环)**: `try_acquire`(§7.3)传入的 est 来自 `SourceConfig.effective_est_tokens()` 这一份纯方法,五个调用点(`QuotaGate` 入场 + chat/embedding 各自的成功侧与失败侧结算)共用,保证预扣与结算恒取同一值(`delta == 0`,否则押金会被整笔退回、TPM 闸退化成进门即放行)。规则:显式 `est_tokens > 0` 则原样用;否则 `tpm > 0` 时派生 `max(1, tpm // 60)`;`tpm == 0`(该闸不启用)时为 0。
@@ -449,11 +464,15 @@ flowchart TB
### 7.8 遥测与成本 ### 7.8 遥测与成本
**必录字段**(继承三项目 15 字段规范): call_id、parent_call_id、session_id、model、provider、source_name、messages(JSON)、response、thinking、prompt_tokens、completion_tokens、usage_source、latency_ms、ttft_ms、max_inter_token_ms、cache_hit、error、**cost**。链路: `session_id`/`parent_call_id` 由调用方传入贯穿(agent step → LLM call)。`messages` 落库前对多模态 part 先摘要(与缓存 key 共用同一摘要函数,§7.5)——Video-Tree 现状 base64 整段进 SQLite 导致 db 膨胀(`llm.py:330`),库内修复(2026-07-20,VT 迁移缺口 R12) **必录字段**(继承三项目 15 字段规范): call_id、parent_call_id、session_id、model、provider、source_name、messages(JSON)、response、thinking、prompt_tokens、completion_tokens、usage_source、latency_ms、ttft_ms、max_inter_token_ms、cache_hit、error、**cost**、**cached_prompt_tokens**、**model_reported**、**sampling**
**`sampling` 列(2026-07-31,issue #4,端口 20 → 21)**: 列语义 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。三个 emit 入口口径必须各自定死,否则同一列在不同行含义不同: `emit_attempt`(RetryMW 调用,**唯一**有生效源者)并上 `source.extra_body`;`emit_cache_hit` / `emit_terminal_failure`(TelemetryMW 最外层调用)无 source 可言,只记调用级——与 `model`/`source_name` 在终态行置空是同一先例,且缓存命中行无损(`sampling` 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同)。三者统一读 `request.sampling` 而非 `request.overlay`(后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处未被污染,直接用必然三行分叉)。OCR/embedding 路径因决策 G 剥离 `extra_body`,该列恒 NULL。
(`cached_prompt_tokens`/`model_reported` 为 2026-07-31 issue #3 新增,端口由 18 字段扩为 20;两个后端在初始化期对已存在的旧表幂等补列——`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入都被逐行 warning 丢弃。补列一律**先探测缺列再 ALTER**(`ADD COLUMN IF NOT EXISTS` 即使列已存在也先取 ACCESS EXCLUSIVE 锁,而遥测内联 await,锁共享审计表会拖垮业务调用),且**失败只逐行降级、绝不置结构性失能标志**。新列在 DDL 里必须排在 `created_at` **之后**,与 `ALTER TABLE ADD COLUMN` 的追加位置一致,否则新建库与升级库的物理列序分叉)。链路: `session_id`/`parent_call_id` 由调用方传入贯穿(agent step → LLM call)。`messages` 落库前对多模态 part 先摘要(与缓存 key 共用同一摘要函数,§7.5)——Video-Tree 现状 base64 整段进 SQLite 导致 db 膨胀(`llm.py:330`),库内修复(2026-07-20,VT 迁移缺口 R12)。
- 后端: `SQLiteRecorder`(默认;WAL + busy_timeout、`INSERT OR IGNORE` 幂等、`asyncio.to_thread` 桥接、初始化/写入失败全降级不冒泡)与 `PostgresRecorder` - 后端: `SQLiteRecorder`(默认;WAL + busy_timeout、`INSERT OR IGNORE` 幂等、`asyncio.to_thread` 桥接、初始化/写入失败全降级不冒泡)与 `PostgresRecorder`
- **单一 helper 铁律**: 遥测调用点收敛为一个内部函数/上下文管理器;Video-Tree 与 GovDoc 各有 4-5 处逐字复制的 `record_llm_call(15 个参数)` 是本条的直接教训。 - **单一 helper 铁律**: 遥测调用点收敛为一个内部函数/上下文管理器;Video-Tree 与 GovDoc 各有 4-5 处逐字复制的 `record_llm_call(15 个参数)` 是本条的直接教训。
- 成本: `pricing.py` 维护 model → (input 单价, output 单价) 表,遥测时换算 `cost` 字段;查不到价格记 None 并 warning,**不阻塞调用**。 - 成本: `pricing.py` 维护 model → (input 单价, output 单价, **可选** cached_input 单价) 表,遥测时换算 `cost` 字段;查不到价格记 None 并 warning,**不阻塞调用**。缓存读取单价(2026-07-31,issue #3)只在配置了该档且本次有命中时启用,按 `(prompt - cached) × input + cached × cached_input` 分段计价;**未配该档绝不按经验折扣率猜**,退化为全额输入价(P5)。命中数超过输入总数时按总数夹取并 warning,不产生负成本。
### 7.9 结构化输出阶梯(D14) ### 7.9 结构化输出阶梯(D14)
@@ -509,6 +528,7 @@ src/polygateway/
- **载体**: `.env` + 环境变量(工程配置);缺失关键配置直接报错,严禁硬编码默认值兜底(三项目共同铁律)。**实现勘误(2026-07-20 M1,人类确认)**: 多源 `{SCOPE}__{PROVIDER}__{N}__{FIELD}` 是动态键族,pydantic-settings 的静态字段模型无法表达,故 `GatewaySettings` 为 frozen dataclass + python-dotenv(显式核心依赖)读取,fail-loud 校验语义与 pydantic-settings 一致。 - **载体**: `.env` + 环境变量(工程配置);缺失关键配置直接报错,严禁硬编码默认值兜底(三项目共同铁律)。**实现勘误(2026-07-20 M1,人类确认)**: 多源 `{SCOPE}__{PROVIDER}__{N}__{FIELD}` 是动态键族,pydantic-settings 的静态字段模型无法表达,故 `GatewaySettings` 为 frozen dataclass + python-dotenv(显式核心依赖)读取,fail-loud 校验语义与 pydantic-settings 一致。
- **多源命名**: `{SCOPE}__{PROVIDER}__{N}__{FIELD}`(如 `LLM__QWEN__1__API_KEY``OCR__MONKEY__1__BASE_URL`),聚合为 `list[SourceConfig]`;SCOPE 支持逻辑角色前缀(§7.7)。 - **多源命名**: `{SCOPE}__{PROVIDER}__{N}__{FIELD}`(如 `LLM__QWEN__1__API_KEY``OCR__MONKEY__1__BASE_URL`),聚合为 `list[SourceConfig]`;SCOPE 支持逻辑角色前缀(§7.7)。
- **`{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`(2026-07-31,issue #4)**: 值为 JSON **对象**串(数组/标量报错),解析为源级恒定采样参数。`_SOURCE_FIELDS` 是跨 scope 共用的一张表,故该键在 `OCR__`/`EMBED__` 下也语法合法,但那两条路径不消费它(embed payload 硬编码 `{model, input}`、MonkeyOCR 只发 multipart)——处置为**构造期剥离 + warning 放行**而非报错(2026-07-31 人类拍板: 这两条路径本无采样语义,配错后果远轻于 chat,不值得让下游装配起不来)。剥离本身是承重的: 不剥离则遥测 `sampling` 列会记录一个从未发出的参数(§7.8),那是数据造假而非参数失效。
- **韧性参数键名**沿用三项目习惯(`LLM_TIMEOUT` / `LLM_MAX_RETRIES` / `LLM_RETRY_BASE_DELAY` / `LLM_RETRY_MAX_DELAY` / `LLM_CIRCUIT_BREAKER_THRESHOLD` / `LLM_CIRCUIT_BREAKER_COOLDOWN` / `LLM_TTFT_TIMEOUT` / `LLM_INTER_TOKEN_TIMEOUT`),降低三项目迁移改名成本。 - **韧性参数键名**沿用三项目习惯(`LLM_TIMEOUT` / `LLM_MAX_RETRIES` / `LLM_RETRY_BASE_DELAY` / `LLM_RETRY_MAX_DELAY` / `LLM_CIRCUIT_BREAKER_THRESHOLD` / `LLM_CIRCUIT_BREAKER_COOLDOWN` / `LLM_TTFT_TIMEOUT` / `LLM_INTER_TOKEN_TIMEOUT`),降低三项目迁移改名成本。
- **per-scope 韧性配置(2026-07-20,CHS 迁移缺口 G4)**: 韧性参数支持按 scope 覆盖——`{SCOPE}__RETRY__MAX_ATTEMPTS` / `{SCOPE}__BREAKER__FAIL_THRESHOLD` / `{SCOPE}__BREAKER__COOLDOWN_S` / `{SCOPE}__BACKPRESSURE__STALL_WINDOW_S` / `{SCOPE}__SELECTOR` / `{SCOPE}__GLOBAL__MAX_CONCURRENCY|RPM|TPM`(CHS 现状: VLM 与 OCR 两 scope 参数各异)。平铺键(`LLM_*`)是单 scope 场景的简写;两者并存时 scope 键优先。 - **per-scope 韧性配置(2026-07-20,CHS 迁移缺口 G4)**: 韧性参数支持按 scope 覆盖——`{SCOPE}__RETRY__MAX_ATTEMPTS` / `{SCOPE}__BREAKER__FAIL_THRESHOLD` / `{SCOPE}__BREAKER__COOLDOWN_S` / `{SCOPE}__BACKPRESSURE__STALL_WINDOW_S` / `{SCOPE}__SELECTOR` / `{SCOPE}__GLOBAL__MAX_CONCURRENCY|RPM|TPM`(CHS 现状: VLM 与 OCR 两 scope 参数各异)。平铺键(`LLM_*`)是单 scope 场景的简写;两者并存时 scope 键优先。
- **装配只有两条路**: `GatewayClient.from_env()`/`from_settings(settings)`(工厂,覆盖 90% 用户;补上三项目每次手写、GovDoc 缺失的"配置→client"一段)或构造函数全量依赖注入(测试/高级用户)。库内部任何组件**不得自读环境变量**(显式优于隐式)。 - **装配只有两条路**: `GatewayClient.from_env()`/`from_settings(settings)`(工厂,覆盖 90% 用户;补上三项目每次手写、GovDoc 缺失的"配置→client"一段)或构造函数全量依赖注入(测试/高级用户)。库内部任何组件**不得自读环境变量**(显式优于隐式)。
@@ -0,0 +1,164 @@
# 响应可观测字段扩展设计(Issue #3)
- **日期**: 2026-07-31
- **来源**: Gitea Issue #3(下游 dissect 审计需求)
- **状态**: 已批准(2026-07-31,人类逐条确认 A2 / B1 / C1 / D1)
- **触发档位**: 强制(变更 `types.py` 公共类型 + `ports.py` 端口签名 + 遥测持久化 schema)
## 1. 目标与非目标
| 项 | 内容 |
|---|---|
| 目标 1 | `LLMResponse` 暴露供应商侧 prompt cache 命中的输入 token 数 |
| 目标 2 | `LLMResponse` 暴露 API 响应体实际返回的模型版本串 |
| 目标 3 | 两字段同步落 `llm_calls` 遥测表(端口 18 → 20 字段) |
| 目标 4 | `PricingTable` 支持可选的缓存读取单价,消除 cost 高估 |
| 非目标 1 | 不改 `EmbeddingResponse` / OCR 响应——embedding 与 OCR 无 prompt cache 语义,且 issue 未提;`cost()` 新增参数带默认值,embedding 调用点(`embedding.py:419`)零改动 |
| 非目标 2 | 不改 `cache_hit` 字段名/类型(破兼容),只在 docstring 消歧 |
| 非目标 3 | 不为 `reasoning_tokens` 等其他 usage 细项开口(YAGNI,无下游需求) |
### 1.1 Issue 前提的一处修正
Issue 称「两者的数据都已经存在于 `TransportResult.raw` 里」。核查结果:
| 数据 | 实际所在 | 结论 |
|---|---|---|
| `usage.prompt_tokens_details.cached_tokens` | `raw={"usage": ...}`(流式 `openai_compat.py:362`、非流式 `:444`) | ✅ 已在 raw 内 |
| 响应体顶层 `model` | **不在**。非流式 raw 只放 `body["usage"]`;流式 sink 只吸收 `usage``done` 两键(`_sse_delta`,`:44-47`),chunk 的 `model` 从未收集 | ❌ 需改 transport 采集 |
故本变更**不是纯字段暴露**,必须同时改 `transports/`。这决定了下面决策 A 的必要性。
## 2. 决策 A:字段的采集与传递路径
| 方案 | 做法 | 权衡 |
|---|---|---|
| A1 raw 约定键 | transport 往 `raw` 里塞 `{"model": ...}`;RetryMW 读 `raw.get("model")``raw["usage"]["prompt_tokens_details"]["cached_tokens"]` | 改动最小;但 `raw: dict[str, Any]` 变成隐式契约,键名靠约定;且 middleware 要懂 OpenAI 报文嵌套结构 |
| A2 TransportResult 强类型字段(**推荐**) | `TransportResult` 追加 `cached_prompt_tokens: int \| None = None``model_reported: str \| None = None`;解析逻辑留在 `openai_compat.py`;RetryMW 直接搬运 | 报文格式知识不出 `transports/`,middleware 只做搬运,符合 P7(middleware 只依赖端口、不懂具体报文);两字段带默认值,`monkey_ocr` 的 OCR 结果类型不受影响 |
| A3 middleware 解析 raw | RetryMW 内写 OpenAI 嵌套路径解析 | 把 provider 报文格式知识放进 middleware 层,新增非 OpenAI 兼容 transport 时会分叉;违反分层,否决 |
**选 A2**`TransportResult` 是库内部流转类型(非三项目消费面),但仍按「新增必带默认值」处理,使 `openai_compat` 之外的构造点零改动;全库该类型仅 2 处构造(`openai_compat.py:354/436`)。
解析纪律(P5 一切外部输入校验后使用):`cached_tokens``model` 均来自网关响应,类型不可信。取值走防御 helper,不抛异常(可观测字段缺失绝不能打断主路径):
| 输入 | 结果 |
|---|---|
| `cached_tokens` 为非负 `int`(**含 `0`**) | 如实保留——`0` 是「该源上报了一次真实零命中」,与「未上报」的 `None` 语义不同,这正是本 issue 的核心诉求 |
| `cached_tokens` 为负数 / 非 `int` / `bool` | `None`(`bool` 必须显式排除:`isinstance(True, int)` 在 Python 里为真) |
| `usage``prompt_tokens_details` 非 dict | `None` |
| `model` 为非空 `str` | 保留 |
| `model` 为非 `str` / 空白串 | `None` |
## 3. 决策 B:缓存命中回放时两字段取什么值
| 方案 | LLMResponse 层 | 遥测层 | 权衡 |
|---|---|---|---|
| B1 原样回放(**推荐**) | 随缓存 JSON 回放原值 | 照记回放值 | 与既有口径一致——`CacheMW._rehydrate`(`cache.py:113-119`)只覆写与本次调用相关的时序字段(`latency_ms`/`ttft_ms`/`max_inter_token_ms`/`call_id`/`cache_hit`),`model`/`provider`/`prompt_tokens` 全部回放。新字段与它们同类(溯源 + 用量),按同一规则处理 |
| B2 命中时置 None | 覆写为 None | NULL | 语义上「本次未打供应商,无供应商侧事实」也成立,但与同层的 `prompt_tokens` 回放行为不一致,下游要记两套规则 |
| B3 混合 | `model_reported` 回放、`cached_prompt_tokens` 置 None | 同左 | 最难解释,否决 |
**选 B1**,并写入文档一条度量口径约束(与 `cost` 缺口口径同款教训,ARCHITECTURE §5.1):
> 统计供应商缓存命中率必须写 `WHERE cache_hit = false`——缓存命中行的 `cached_prompt_tokens` 是历史回放值,计入会重复计数。
`cost` 不受影响:遥测层 `cache_hit=True` 分支仍短路为 `0.0`,早于任何单价换算。
## 4. 决策 C:缓存读取单价(人类已选「增加可选档」)
| 方案 | 做法 | 权衡 |
|---|---|---|
| C1 ModelPrice 可选第三档(**推荐**) | `cached_input_per_1m: float \| None = None`;`cost()` 增可选参 `cached_prompt_tokens: int \| None = None` | 价格表旧文件零改动仍可加载;`embedding.py:419` 的三参调用零改动 |
| C2 cost() 收 LLMResponse | 换算函数直接吃响应对象 | `pricing.py` 会反向依赖 `types.py` 且难以单测纯函数,否决 |
换算规则与退化路径:
| 条件 | 计价方式 |
|---|---|
| 配了 `cached_input_per_1m` 且本次 `cached_prompt_tokens` 为正 | `(prompt - cached) × input + cached × cached_input` |
| 未配该档,或本次 `cached_prompt_tokens` 为 None/0 | 全额按 `input` 计(现状行为,不变) |
| `cached > prompt`(网关口径异常) | 按 `cached = prompt` 夹取并记一次 warning;不抛异常、不产生负成本 |
**不猜折扣率**:未配置缓存档时绝不按「五分之一」之类经验值折算(P5 严禁默认值掩盖)。`from_file` 的 fail-loud 校验对新档同样适用:出现该键但非数或为负 → `ValueError`
## 5. 决策 D:遥测表扩列的落地方式
人类确认「现在不存在必须保留的生产库」。但两个后端的 DDL 都是 `CREATE TABLE IF NOT EXISTS`,**已存在的开发库/下游库不会自动获得新列**,INSERT 会失败。两侧的失败形态都是**逐行 warning 丢弃**(SQLite `sqlite.py:93`;PG `postgres.py:127`——`_failed` 结构性标志只在 `_ensure_ready` 建池/建表失败时置位,与写入路径无关),即每一次调用的遥测行都丢,却不会有任何一次硬失败提示,与「遥测必录」相悖。
| 方案 | 做法 | 权衡 |
|---|---|---|
| D1 初始化期幂等补列(**推荐**) | DDL 加新列;初始化时按需 `ALTER TABLE ADD COLUMN`——PG 用原生 `IF NOT EXISTS`,SQLite 先查 `PRAGMA table_info` 再按需 ALTER | 旧库自动升列,新库无副作用;两处各约 5 行;补列失败沿用现有降级策略(warning,不冒泡) |
| D2 只改 DDL,文档写「删表重建」 | 零代码 | 已建表的开发机/下游踩坑后只看到降级 warning,排查成本高;违反防御性 |
| D3 引入迁移框架(alembic) | 正规版本化迁移 | 新增依赖,与「依赖极简」铁律冲突,规模严重不匹配,否决 |
**选 D1**。列类型:SQLite `cached_prompt_tokens INTEGER` / `model_reported TEXT`;PG `INTEGER` / `TEXT`。两列均可空(NULL = 该源未上报),不设 NOT NULL 与默认值——0 与 NULL 的区分正是本 issue 的核心诉求。
**D1 的实现纪律(必须钉进计划,否则补列会把降级放大成永久失能)**:
| 约束 | 原因 |
|---|---|
| SQLite 的 ALTER 必须用**独立 try**,且置于 `self._conn = conn` **之后** | `__init__` 现有 try 的最后一句才是 `self._conn = conn`(`sqlite.py:74-84`);ALTER 抛异常会让 `_conn` 停在 `None`,`record_llm_call` 首行即 return —— 整个 recorder 永久 no-op,比逐行丢弃严重得多 |
| `duplicate column name` 视为成功吞掉 | `PRAGMA table_info` 探测 + ALTER 是 TOCTOU:多 worker 共用同一 db 文件时后到者必然撞上 |
| 不得为补列加宽 `except` | `sqlite.py:93` 只捕 `(OSError, sqlite3.Error)`,取消是天然穿透的;PG 侧的 `except asyncio.CancelledError: raise` 必须留在最前 |
| PG 用原生 `ALTER TABLE ... ADD COLUMN IF NOT EXISTS` | 无 TOCTOU;落在既有 `_init_lock` 保护的 `_ensure_ready` 内 |
端口 `TelemetryRecorder.record_llm_call` 由 18 字段扩为 20 字段(关键字参数),`ports.py:248` 的「18 字段冻结」注释与 ARCHITECTURE 相应表述同步更新。新增参数在 Protocol 上**不设默认值**——依据不是「漏改会报错」(本仓无 mypy,`make lint` 只有 ruff + import-linter,8 个测试 fake 全是 `**fields`,漏改根本不会自动红),而是**库外不存在第三方实现者**:三项目迁移文档明确删除各自的 TelemetryRecorder Protocol 与实现(`migrations/govdoc-saas.md:36``video-tree-trm5.md:36/51`),端口的唯一实现者就是库内两个后端,完整签名的成本为零。漏改的兜底靠 §8 的键集合断言测试,不靠类型检查。
## 6. 行为审计(既有行为逐条标注)
| 既有行为 | 处置 |
|---|---|
| `LLMResponse` 前 11 字段顺序即公共承诺 | **保留**,新字段追加到尾部(`structured_data` 之后) |
| 缓存序列化 `_serialize``asdict` 后 pop 掉 `structured_data``_rehydrate``_RESPONSE_FIELDS` 过滤 | **保留**。新字段自动进出;旧缓存条目缺这两键时,`LLMResponse(**fields)` 靠默认值构造成功(向后兼容已验证) |
| `cache_hit` 语义 = PolyGateway 自身响应缓存 | **保留**,仅补 docstring 消歧 |
| `TelemetryEmitter` 单一 `_record` helper(遥测必录铁律:禁止复制参数列表) | **保留**,新字段只在 `_record` 增两个参数,三个 `emit_*` 入口各传一次 |
| 失败尝试 / 终态失败行记 `usage_source="unavailable"` | **保留**,两个新字段在这些路径记 `None` |
| `pricing.cost()` 是唯一换算点(注释语)| **修正**:实际有 `TelemetryEmitter``embedding.py:419` 两个调用点,顺带订正该 docstring(限于一行注释,不做结构重构) |
| OCR / embedding 各自构造 `LLMResponse` | **保留**,两字段取默认 `None`(该路径无供应商 cache 概念) |
## 7. 非功能维度
| 维度 | 结论 |
|---|---|
| 并发与取消 | 纯数据字段,无新增 await 点、无共享状态。PG 补列在既有 `_init_lock` 保护的 `_ensure_ready` 内,并发首调用不会重复 ALTER;SQLite 的 `__init__` **不持** `_lock`(它只保护 `_write`/`close`),跨进程共库靠上面 D1 纪律里的「duplicate column 视为成功」兜底。取消穿透不变:PG 两处 `except asyncio.CancelledError: raise` 保持在最前,SQLite 侧只捕 `(OSError, sqlite3.Error)` 故天然穿透 |
| 降级方向 | 遥测属「静默降级」侧:补列失败 → warning 并沿用既有逐行丢弃,绝不冒泡到调用方,也绝不让 recorder 整体失能(见 D1 纪律)。解析失败 → 字段记 `None`,不影响响应返回 |
| 幂等与重复 | 补列幂等(PG `IF NOT EXISTS`;SQLite 先探测)。写入幂等性不变(`INSERT OR IGNORE` / `ON CONFLICT DO NOTHING``call_id`) |
| 持久化与原子性 | 单行 INSERT 原子性不变;新增两列不参与主键与冲突判定。缓存 JSON 是整值覆写,无部分写入 |
| 向后兼容 | 下游三项目 + dissect:纯增字段带默认值,逐字段传参的 fake 构造零改动;旧价格表文件、旧缓存条目、旧遥测表均可继续工作 |
## 8. 错误处理与测试策略
错误分类:本变更**不新增任何错误路径**。网关报文里这两项缺失或类型异常 → 记 `None`,不归入四分类(它们不是失败,是「该源没给」)。价格表配置错误仍走装配期 `ValueError`(fail-loud,不属运行时四分类)。
| 层 | 测试(先失败后通过) |
|---|---|
| types(unit) | 新字段默认值为 `None`;字段顺序不变(前 11 位置构造仍成立) |
| transports(unit) | 用真实网关响应二次构造样本:① 流式含 `prompt_tokens_details.cached_tokens` → 解析出正整数;② 非流式同上;③ 无该键 → `None`;④ 值为 `"abc"`/负数 → `None` 不抛;⑤ 流式 chunk 的 `model` 被 sink 采集;⑥ 顶层无 `model``None` |
| retry(unit) | `_build_response` 透传两字段;失败尝试路径不受影响 |
| cache(unit) | ① 新字段随序列化往返;② **旧格式**缓存条目(缺这两键)仍能 rehydrate;③ 命中回放值符合 B1 |
| pricing(unit) | ① 配缓存档 + 命中 → 成本低于全额;② 未配该档 → 与现状逐位相等;③ `cached > prompt` → 夹取且不为负;④ 三参旧调用签名仍可用(embedding 调用形态);⑤ 价格表含负缓存单价 → `ValueError` |
| telemetry(integration) | ① 20 字段写入 SQLite/PG 成功并可读回;② **旧表**(18 列)在初始化后自动补列并写入成功;③ 补列失败时降级为 warning 且 recorder 仍能工作(SQLite `_conn` 不得因此为 None) |
| 契约(**新增,不可省**) | 断言 `TelemetryEmitter` 传给 recorder 的实参键集合 == 两个后端的 `_COLUMNS`。理由:`row = tuple(fields[col] for col in _COLUMNS)` 位于两个后端的 try **之外**(`sqlite.py:90` / `postgres.py:121`),emitter 漏传新字段会抛 `KeyError`,被 `_record``except Exception` 吞成 warning → **静默丢遥测**。这是本变更最危险的失败形态,而现有 8 个 `**fields` 形态的 fake 一个都拦不住 |
> integration 层的 Redis/PG 测试遵守既有纪律:共享后端严禁并跑,`conda run -n PolyGateway --no-capture-output`。
## 9. 影响面清单
| 文件 | 改动 |
|---|---|
| `src/polygateway/types.py` | `LLMResponse` +2 字段;`TransportResult` +2 字段;`cache_hit` docstring 消歧 |
| `src/polygateway/transports/openai_compat.py` | sink 采集 `model`;两处 `TransportResult` 构造填新字段;新增防御解析 helper |
| `src/polygateway/middleware/retry.py` | `_build_response` 透传 2 字段 |
| `src/polygateway/middleware/telemetry.py` | `_record` + 三个 `emit_*` 各透传 2 字段;cost 换算传入 `cached_prompt_tokens` |
| `src/polygateway/pricing.py` | `ModelPrice` +可选档;`cost()` +可选参;`from_file` 校验;订正唯一换算点注释 |
| `src/polygateway/ports.py` | `TelemetryRecorder` 18 → 20 字段 |
| `src/polygateway/telemetry/{sqlite,postgres}.py` | DDL +2 列;`_COLUMNS` +2;初始化期幂等补列 |
| `tests/` | 四处天然拦截点必须同步(漏改即红): 两个 `_record_minimal` 手写 18 键 dict(`unit/test_telemetry.py:76` 起、`integration/test_postgres_telemetry.py:81-105`)与两个 `_EXPECTED_COLUMNS` 列序断言(`unit/test_telemetry.py:18-40``integration/test_postgres_telemetry.py:22-41`);`unit/test_ports.py:96` 的全签名 fake 同步(它**不会**红,Protocol 的 isinstance 不校验签名);新增契约测试 |
| `research-wiki/ARCHITECTURE.md` | §5.1 字段表 + 遥测表定义 + 「18 字段冻结」表述 |
| 「18 字段冻结」的其余措辞点 | `ports.py:248``pricing.py:6`(币种说明里引用了该数字)、`telemetry/sqlite.py:87``tests/unit/test_telemetry.py:1` |
| Wiki 站 + `CHANGELOG.md` | 按 `docs-convention.md` §2 清单同步(公共行为变更,版本 bump 不得裸发) |
| `.env.example:56` | 该行内联注释是仓内**唯一**的价格表格式说明(无独立模板文件,`config/prices.json` 是未入库的本地文件),补 `cached_input_per_1m` 可选档 |
## 10. 审批记录
2026-07-31 人类逐条确认: **A2**(TransportResult 强类型字段)、**B1**(缓存命中原样回放 + 度量口径带 `cache_hit = false`)、**C1**(ModelPrice 可选缓存单价档)、**D1**(DDL 加列 + 初始化期幂等补列)。设计获批,进入 `writing-plans`
版本号按 `1.1.0` 推进(纯增字段不破坏下游,但触及端口签名与表结构,minor 位比 patch 位更能提示下游);发版前若人类另有指示以指示为准。
@@ -0,0 +1,234 @@
# 采样参数透传设计(issue #4)
- **日期**: 2026-07-31
- **状态**: 待人类审批
- **触发**: issue #4 —— `chat()` 无法设置 `temperature`/`seed`/`max_tokens`,下游受控实验无法固定解码
- **影响面**: `chat()` 公共签名、`SourceConfig` 公共类型、缓存 key 公式(ARCH §7.5)、遥测端口(20 → 21 字段)
---
## 1. 诉求与现状审计
下游 dissect 是一组受控实验:解码固定 `temperature=0`,每格配置跑 5 个 seed 报标准差。标准差必须只反映被研究的变量,不能混进解码随机性。
代码事实(本会话核实):
| 事实 | 位置 | 后果 |
|---|---|---|
| 全库 `temperature` 零命中 | `grep -rn temperature src/` | 解码跑在供应商默认值上,不可复现 |
| `chat()` 签名无 overlay 入口 | `client.py:143-153` | 调用方够不着 `ChatRequest.overlay` |
| `overlay` 唯一写入点是结构化中间件 | `middleware/structured.py:98` | 字段存在但只服务库内 |
| `payload.update(overlay)` 是最后一步 | `transports/openai_compat.py:297` | overlay 可覆盖 `model`/`messages`/`stream`/`stream_options` |
| 缓存 key 公式不含 overlay | `middleware/cache.py:52-64` | **见 §2 决策 C** |
| `model_fingerprint` 只由源 `model` 名算 | `client.py:117` | 配置级采样参数变更不改 key |
| minimax / openai profile 均 `thinking_off={}` | `providers.py:49,56` | `enable_thinking=False` 对两源均无效果 |
**issue 未提及但必须一并处理的**: 缓存与遥测的交互。不处理的话,failure mode 恰是 issue 自己最担心的那种——数字悄悄不可比,且不报错。
---
## 2. 设计决策
### 决策 A: 两层入口,合并优先级由现有层序天然给出
| 层 | 载体 | 用途 | 生效点 |
|---|---|---|---|
| 调用级 | `chat(..., overlay: Mapping[str, Any] \| None = None)` | 逐次变化(每 rollout 不同的 `seed`) | 填入 `ChatRequest` |
| 配置级 | `SourceConfig.extra_body: Mapping[str, Any]` | 全局恒定(`temperature=0`) | transport `_build_payload` |
优先级 **结构化注入 > 调用级 > 配置级**,无需任何新机制:
```text
_build_payload: payload{model,messages,stream} → thinking_profile
→ source.extra_body ← 配置级(新增一行)
→ overlay ← 调用级 ⊎ 结构化注入
StructuredMW: {**request.overlay, **strategy_overlay} ← 结构化已在最右,天然最高
```
配置级放在 transport 而非装配层合并,是因为 `extra_body` 是 per-source 的,选源在 RetryMW 之后才确定;放 transport 无需改动任何端口签名。
**`ChatRequest` 增第二个字段 `sampling: Mapping[str, Any] = field(default_factory=dict)`**(调用方原始采样意图的快照,库内中间件**永不修改**),与 `overlay`(请求体覆盖层,会被结构化注入)分开。`chat()` 同时填两者。理由是 `overlay` 在洋葱不同深度取值不同——`StructuredMW` 内侧含 `response_format`、外侧不含——缓存 key 与遥测若各自依赖"在哪一层读"就会口径分叉(见决策 C/D)。`sampling` 提供一个跨层恒定的读取点。
类型定死为 `Mapping` 而非 `dict[str, Any] | None`:空 dict 与 `None` 在此无语义差别(都是"没传采样参数"),多一种表示只会让 key 公式与 `merge()` 签名各选各的。因此决策 C 的 key 公式一律按**仅非空**参与(注意与同处的 `salt` 不同——`salt` 是"仅非 None",空串是有意义的 salt)。
### 决策 B: 保护键黑名单,构造期显式报错
`{model, messages, stream, stream_options}` 禁止出现在 overlay/extra_body 中。理由逐条:
| 键 | 被覆盖的后果 |
|---|---|
| `model` | 遥测记录的 model 与实际请求分叉 → 成本按错单价算 |
| `messages` | 缓存 key 与遥测口径同时失真 |
| `stream` | 绕过流式看门狗(TTFT/inter-token 三层超时全失效) |
| `stream_options` | 丢 usage 帧 → 成本遥测归零、TPM 闸按预扣量结算失准 |
同一校验函数还必须验**值可 JSON 序列化**。理由:`CacheMW.__call__` 第 95 行的 `build_cache_key` 内部 `json.dumps`,**不在 `_safe_get`/`_safe_set` 的降级 try 内**;`TelemetryMW` 只捕 `GatewayUnavailableError`/`GovernanceBackendError`/`CancelledError`。调用方传 `{"temperature": np.float32(0)}`(温度扫描用 numpy 生成极自然)会抛裸 `TypeError`:不属四分类、一行遥测都没有、RetryMW 从未执行。构造期一次校验即可保住"overlay 错误全部发生在进洋葱之前"这条不变式。
校验函数落在 `types.py`(最内层,无依赖),两个入口各调一次:`chat()` 参数在进洋葱**之前**校验(与既有 `structured` 的 ImportError 同款先例),`SourceConfig.__post_init__` 在装配期校验(符合 §4.5「缺失/非法关键配置直接报错」)。抛裸 `ValueError`——这是调用方编程错误,不属 §6 四分类,不应被 RetryMW 当作可重试失败。
transport 不重复校验:三个 overlay 来源(chat 参数、SourceConfig 字段、库内策略)已全部在构造期收口,库内策略只注入 `response_format`(`json_repair.py:41` 恒空,`native_schema.py:21-29` 只产该键)。
### 决策 C: 调用级 overlay 进缓存 key —— 本设计的关键点
不做的话:同 messages 跑 5 个 seed,后 4 次命中第一次的缓存,返回同一 response,**标准差恒为 0**,实验静默作废。这正是「无缓存毒化」铁律的场景。
key 公式扩展(ARCH §7.5 需同步修订),读 `request.sampling` 而非 `request.overlay`——语义明确、不依赖"CacheMW 恰在 StructuredMW 外侧"这一层序巧合:
```text
key_obj = {model, messages_digest, namespace, [salt], [sampling]}
仅非 None 仅非空
```
沿用 `salt` 的「仅非空时参与」写法,保证**空采样参数时旧键逐字不变**,不触发存量缓存全量冷启动。
配置级同理:`model_fingerprint``",".join(sorted(models))` 扩展为——所有源 `extra_body` 皆空时字面不变;否则追加 `"|" + sha256(...)`,摘要对象是「每个源的 `(model, extra_body)` 先各自 canonical-JSON 化成字符串,再排序去重」(dict 本身既不可排序也不可哈希,必须先序列化;`extra_body` 若存为 `MappingProxyType``dict(...)` 后再 `json.dumps`)。取 `(model, extra_body)` 而非 `(name, ...)`,语义是「本 scope 会用哪些(模型,解码参数)组合」,改源名不会误触冷启动。该计算在 `client.py:117` 且不在任何降级 try 内,写错即装配期崩——实施时须有直接单测。
**两条已知副作用(须写进 wiki)**:
1. 逐 rollout 变化的 `seed` 进 key 后,该路径**天然全部 miss**。这是正确语义而非缺陷,但下游要知道缓存对这条路径不再省钱。
2. `model_fingerprint` 是**集合级**指纹,不是本次实际选中源的指纹。同 scope 下各源 `extra_body` 不同时,缓存仍可能返回另一源、另一组解码参数下产生的响应。这是既有取舍的延续(`cache.py:68-72``model` 已如此),不是本设计引入的新缺口,但"配置级采样参数进 key"容易被读成更强的保证,须写明边界。受控实验若要求逐源可复现,应让每个源独享 scope 或 namespace。
### 决策 D: 采样参数入遥测(端口 20 → 21 字段)
「实验可复现」的另一半是参数落库。不记的话,同 messages 不同输出在审计表里无法解释。与 issue #3 新增 `model_reported` 同类动机(供应商把别名指向新权重时,复现必须认真实串)。
**列语义定死**:`sampling: str | None` = 「调用方采样意图 ⊎ 生效源的 `extra_body`」的 canonical JSON,**不含库内结构化注入的 `response_format`**。两个理由:该列名叫采样参数,`response_format` 不是;schema 可达数 KB,逐行记会让审计表无谓膨胀。
`TelemetryEmitter` 有三个入口且都汇入同一个 `_record`(显式关键字参数,加列必须三处都传),必须逐个定死,否则同一列在不同行口径分叉——这正是 1.0.4 里 `cached_prompt_tokens` 不得不写"下游请读"警告的同类坑:
| 入口 | 调用者 | 有 `source`? | `sampling` 记什么 |
|---|---|---|---|
| `emit_attempt` | RetryMW(最内) | 有 | `merge(source.extra_body, request.sampling)` |
| `emit_cache_hit` | TelemetryMW(最外) | **无** | 仅 `request.sampling` |
| `emit_terminal_failure` | TelemetryMW | **无** | 仅 `request.sampling` |
后两行缺 `extra_body` 是**客观事实而非口径瑕疵**:它们没有"生效源"可言——与 `model`/`provider`/`source_name` 在终态行置空是同一先例。缓存命中行尤其无损:`sampling` 已进缓存 key,能命中就意味着历史那次的调用级采样参数与本次逐字相同;`extra_body` 亦已进 `model_fingerprint`,命中意味着源集合的配置指纹相同。
三个入口统一读 `request.sampling`(决策 A 的新字段)而非 `request.overlay`,是因为后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处则未被污染,直接用会让三行天然分叉。
**共用范围写清楚**:`types.py` 提供「2 参 dict 合并 + canonical 序列化」这一个原语,transport 与 emitter 共用它。**不追求统一到两者之上**——transport 是往更大的 payload 上依次 `update(thinking_profile) → update(extra_body) → update(overlay)`,emitter 算的是 `merge(extra_body, sampling)`,参与方与顺序本就不同,强行统一是错的。这不影响正确性:该列语义已定义为「调用方意图 ⊎ 生效源 `extra_body`」,而非 payload 的逐字回显。共用原语的目的只是让"合并语义与序列化口径"这一件事不出现两份实现。
两个后端按 issue #3 已建立的套路幂等补列:**先探测缺列再 ALTER**、失败只逐行降级不置结构性失能标志、新列排在 `created_at` 之后。
一次做完而非分两步:「能传参数但没记」的中间状态最危险——数据已产生且事后无法追溯,且分步要做两遍 DDL 迁移。
**OCR/Embedding 的 emit 调用点零改动**:`ocr.py:418``embedding.py:372` 也调 `emit_attempt` 且都传 `source`,只要 `sampling` 由 emitter 内部推导(而非作为新必填参数由调用者传入),这两处调用不动一行——反之立刻 TypeError,实施时必须走推导路线。两个文件本身仍有改动,即决策 G 的构造期剥离(它正是让这里的推导对 OCR/embedding 恒得 NULL 的前提)。
### 决策 E: 入参拷贝语义与两条只读约束
`chat()` 对传入 overlay 做**一次** `dict(overlay)` 浅拷贝,同一份快照对象同时填 `overlay``sampling` 两个字段(不做两份独立拷贝——它们在进入 `StructuredMW` 之前本就应当逐字相同,两份拷贝反而给"两者可以分叉"留了口子)。
issue 场景就是逐次改 `seed`——调用方复用同一 dict 对象改值是极可能的模式,不拷贝会出现「请求已发出、key 用了新 seed」的竞态。`ChatRequest` 虽 frozen 但 dict 是浅冻结,拦不住。`SourceConfig.extra_body``__post_init__``MappingProxyType` 同理(成本近零)。
拷贝之外的第二条约束:**任何中间件不得就地修改这两个 dict**,只能经 `dataclasses.replace` 派生新请求。现状已满足(`StructuredMW``{**a, **b}` 生成新 dict,`_build_payload` 只往 payload 上 `update`,全库无就地改写),本设计只是把它写成明文约束——决策 C 与 D 都建立在 `sampling` 跨层恒定之上,这条被破坏则两者同时失效(测试 #14 为此加机械执法)。
### 决策 F: 空 thinking profile 的诚实性缺口(issue 附带项)
`minimax``openai``thinking_on/thinking_off` 均为空字典。`providers.py:52` 那条「OpenAI 兼容基线,无已知注入差异」的注释在词法上属于紧随其后的 **minimax** 条目,`openai` 条目没有任何注释。所以现状是:已有的注释解释了"为何为空",但两个 provider 都没点明**后果**——`enable_thinking=False` 对它们不产生任何效果,调用方以为关掉了实际没关。
补的是这一句后果说明(覆盖两个 provider),不是重复已有的"为何为空"。不改行为:真需要关时经 `extra_body` 绕过。
### 决策 G: 非 chat 路径的 `extra_body` —— 剥离并 warning,不中断装配
`_SOURCE_FIELDS`(`config.py:33-47`)是**跨 scope 共用**的一张表,加了 `EXTRA_BODY` 之后 `OCR__MONKEY__1__EXTRA_BODY` / `EMBED__QWEN__1__EXTRA_BODY` 会被合法接受、进 `SourceConfig`、进遥测 `sampling` 列,但两条路径都不消费它:`monkey_ocr.py:225,247` 只发 multipart `files=`(**根本没有 JSON body**),`OpenAICompatTransport.embed`(`openai_compat.py:343`)payload 硬编码 `{"model", "input"}`。放任即**静默无效**,正是 §4.5 要禁的形态。
**处置(2026-07-31 人类拍板改此档)**:`EmbeddingClient` / `OcrClient` 构造期发现源带非空 `extra_body` → 记 warning 并 `dataclasses.replace(source, extra_body={})` **剥离后放行**,不抛异常。
剥离是这一档的**必要组成部分,不是顺手清理**。`ocr.py:390``embedding.py:350` 构造 `ChatRequest` 时不带 `sampling`,但传给 `emit_attempt``source` 是真实配置对象;若不剥离,决策 D 的 `merge(source.extra_body, request.sampling)` 会让遥测**记录一个从未发出的参数**——审计表显示该次 OCR 调用带了 `temperature=0`,实际请求体里没有。那不是"参数不生效",是遥测造假,污染的恰是事后复现的唯一依据。替代方案是在 emitter 里特判调用方身份,直接违背「遥测调用点收敛为单一 helper」铁律,否决。
剥离后该列在 OCR/embedding 行恒为 NULL,语义干净,emitter 零特判。
**被否决的原方案**: 装配期 `ValueError` 直接拒绝。理由是这两条路径本无采样语义,配错的后果远轻于 chat 路径,不值得让下游整个装配起不来。**残余风险须写进 wiki**: loguru warning 在生产中容易被淹没,运维可能仍以为参数生效——这是"不中断装配"换来的代价,故 warning 文案必须**指路**:`dimensions` 是 OpenAI embeddings 的正式参数,下游想调向量维度时会第一个撞上,文案应写明"embedding 路径暂不支持 `extra_body`,该配置已被忽略;需要 `dimensions` 等参数请提 issue"。
不顺手给 embed 加透传:embedding 没有采样一说,issue 也未提出诉求(YAGNI);真有需求时单独设计。
---
## 3. 关键岔路与否决记录
| 岔路 | 否决方 | 理由 |
|---|---|---|
| `chat()` 展开为 `temperature=`/`seed=`/`max_tokens=` 具名参数 | 否决 | 供应商私有参数无穷尽(`top_k`/`repetition_penalty`/`thinking_budget`),具名等于永久追加签名;且违背「深模块窄接口」(ARCH §132) |
| 配置级放装配层全局字典而非 `SourceConfig` | 否决 | 采样参数与源强相关(不同供应商键名不同),全局字典会把无效键发给不认识它的源 |
| overlay 不进缓存 key,靠调用方传 `cache_salt` 区分 | 否决 | 把毒化防护的责任推给调用方,漏传不报错——正是 issue 抱怨的失败形态 |
| 采样参数不入遥测,由下游 run 快照自记 | 否决 | 见决策 D |
| 缓存 key 与遥测都直接读 `request.overlay`,不加 `sampling` 字段 | 否决 | `overlay` 在洋葱不同深度取值不同(结构化注入),三个 emit 入口与 CacheMW 会各记各的,同一列口径分叉 |
| `sampling` 列记「实际发出的完整合并结果」(含 `response_format`) | 否决 | 该列名为采样参数,schema 不是;且数 KB schema 逐行落库无谓膨胀 |
| 给 embedding 路径也加 `extra_body` 透传 | 否决 | embedding 无采样一说,issue 未提诉求(决策 G) |
| 非 chat 路径带 `extra_body` 时装配期 `ValueError` | 否决(人类拍板) | 这两条路径无采样语义,配错后果远轻于 chat,不值得让下游装配起不来;改为剥离 + warning |
| 允许放行但**不剥离** `extra_body` | 否决 | 遥测会记录一个从未发出的参数(决策 D 的 merge 读 `source.extra_body`),是数据造假而非参数失效 |
| 放行不剥离,改在 emitter 内特判 OCR/embedding 不记 | 否决 | emitter 是「遥测调用点收敛单一 helper」的产物,让它识别调用方身份是开倒车 |
| transport 层再兜一次保护键校验 | 否决 | 三个入口已构造期收口,重复校验属 gold-plating |
---
## 4. 非功能维度
| 维度 | 回答 |
|---|---|
| **并发** | 无新增共享状态。`extra_body` 装配后只读(MappingProxyType);调用级 overlay 每调用独立拷贝,并发调用互不可见 |
| **取消** | 无新增 await 点与等待循环,`CancelledError` 穿透路径完全不变 |
| **降级方向** | 不涉及新后端。遥测新列写失败沿用既有逐行 warning 降级;缓存 key 变更不影响 Redis 掉线的静默降级方向。决策 G 的剥离 + warning 是**配置面**降级(装配期一次性、可复现、部署即暴露),与铁律里"限流/熔断后端不可用须报错"的**运行时**降级方向是两回事,不冲突 |
| **幂等与重复** | 保护键校验是纯函数,重复调用安全;遥测补列先探测后 ALTER,重启幂等 |
| **持久化与原子性** | 遥测单行写入,无部分写入风险。缓存 value 结构不变(`sampling` 只进遥测不进 `LLMResponse`,避免动已被三项目消费的公共类型) |
| **重试交互** | overlay 在 RetryMW 循环外确定,换源重试时同一 overlay 应用到新源的 `extra_body` 之上——语义正确(调用级意图跨源保持) |
| **限流交互** | overlay 里的 `max_tokens` 不影响入场预扣(取 `effective_est_tokens()`)。调用方把 `max_tokens` 抬到远超预扣量时 TPM 入场保护会短暂失真,结算侧(`retry.py:338-343`)按实测用量回填自愈。已知且可接受,不为此加机制 |
---
## 5. 错误处理与测试策略
**错误分类**: 保护键违规与 `EXTRA_BODY` JSON 解析失败均为裸 `ValueError`,发生在进入洋葱之前/装配期,不入四分类、不触发重试或熔断。运行时若供应商拒绝某个采样参数(如不支持 `seed`),网关返回 4xx,由既有 `RequestRejectedError` 路径处置——无需新增分类。
**测试清单**(每条须先失败后通过):
| # | 用例 | 层 |
|---|---|---|
| 1 | 同 messages 不同 `seed` → 两次 miss、两个不同 key(issue 场景直接回归) | unit |
| 2 | 空 overlay 时 key 与旧实现逐字相同(防存量冷启动) | unit |
| 3 | 全源 `extra_body` 为空时 fingerprint 与旧实现逐字相同 | unit |
| 4 | 保护键:`chat(overlay={"stream": False})``SourceConfig(extra_body={"model": "x"})``ValueError` | unit |
| 5 | 优先级:配置 `temperature=0` + 调用级 `temperature=1` → payload 为 1;结构化 `response_format` 覆盖调用级同名键 | unit |
| 6 | 调用方在 `chat()` 返回前修改自己的 dict,不影响已发请求与已算 key(拷贝语义) | unit |
| 7 | env 解析:`EXTRA_BODY` 合法 JSON 对象 → dict;非法 JSON / 非对象 → `ValueError` | unit |
| 8 | 不可 JSON 序列化的值(如 `np.float32`)在 `chat()` 入口即 `ValueError`,不进洋葱 | unit |
| 9 | 三个 emit 入口的 `sampling` 口径:attempt 含 `extra_body`、cache_hit 与 terminal 只含调用级、结构化注入的 `response_format` **三行都不出现** | unit |
| 10 | `EmbeddingClient`/`OcrClient` 装配时源带 `extra_body` → 记 warning、装配成功、源上 `extra_body` 已被剥空,且该路径遥测 `sampling` 为 NULL(决策 G;后半段是防遥测造假的真正断言) | unit |
| 11 | `_EXPECTED_COLUMNS` 断言更新后仍逐字匹配实际列序(见 §6,两处会直接红) | unit + integration |
| 12 | 遥测 `sampling` 落库正确;两后端对既有旧表幂等补列 | integration |
| 13 | 采样参数经全链路(chat → 选源 → transport payload)到达请求体 | integration |
| 14 | **地基不变式**:走结构化重问阶梯(至少重问一次)后,RetryMW 每次尝试看到的 `request.sampling``chat()` 传入值逐字相同,且同一时刻 `request.overlay``response_format` | unit |
第 14 条是决策 C/D 共同的承重前提。它现在只靠"`dataclasses.replace` 恰好保留未提及字段"这一约定成立,无任何机械执法;缺这条测试则决策 E 的只读约束被破坏时不会有人发现。
第 2 条(空采样参数时旧键逐字不变)需自行先固化旧 key 值再比对——现有 `tests/unit/test_cache.py:39-54` 只有相等/不等与前缀断言,没有 golden hash 可依。
---
## 6. 配置与文档同步
env 键名沿用既有约定:`{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`,值为 JSON 对象串;`_SOURCE_FIELDS` 增一项、`_cast``json` 分支(解析失败与非 dict 均报错)。
同步清单(docs-convention §2):
| 目标 | 改什么 |
|---|---|
| ARCH §5.2 | `chat()` 签名定稿段追加 `overlay` 要点 |
| ARCH §7.5 | key 公式补 `sampling` 项 + 两条已知副作用 |
| ARCH §7.7 | 该节逐字段枚举 `SourceConfig` 构成(`ARCHITECTURE.md:452`),补 `extra_body` |
| ARCH §7.8 | 必录字段 20 → 21 |
| ARCH §9 | 配置面键族事实源(`:519-527`),登记 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY` |
| `.env.example` | `client.py:247` docstring 声明它是键名清单的事实源,新键不写进去等于无处可查 |
| `README.md:83` | 该行逐一列举 `chat()` 关键字参数,补 `overlay` |
| wiki how-to | 增「固定解码参数」条目,写明 seed 进 key 导致缓存必 miss、以及 OCR/embedding 路径的 `extra_body` 会被忽略(仅 warning) |
| CHANGELOG | 公共 API 新增 + 遥测端口扩列 |
## 7. 实施范围
`types.py`(保护键与 JSON 可序列化校验、合并纯函数、`ChatRequest.sampling``SourceConfig.extra_body`)、`client.py`(`chat()` 参数 + fingerprint)、`middleware/cache.py`(key 公式)、`transports/openai_compat.py`(`_build_payload` 一行)、`config.py`(env 解析)、`ports.py` + `middleware/telemetry.py` + `telemetry/{sqlite,postgres}.py`(第 21 字段与补列)、`ocr.py` + `embedding.py`(仅决策 G 的构造期剥离 + warning)、`providers.py`(注释)。
**测试侧必改**(否则直接红):`tests/unit/test_telemetry.py:18,113``tests/integration/test_postgres_telemetry.py:22,210,231``_EXPECTED_COLUMNS` 断言完整列表与列序。
无需改动:import-linter 契约(校验函数落最内层 `types.py`,分层关系不变)。
不做:给 embedding/OCR 加采样参数透传(决策 G)、任何任务外重构。
@@ -0,0 +1,40 @@
---
type: design
node_id: design:response-observability-fields
title: "响应可观测字段扩展(Issue #3)"
date: 2026-07-31
---
# 响应可观测字段扩展(Issue #3)
全文见 `2026-07-31-response-observability-fields-design.md`。来源: Gitea Issue #3(下游 dissect 的调用审计需求)。
## 选定方案
| 决策 | 选定 | 关键理由 |
|---|---|---|
| A 采集路径 | `TransportResult` 追加 `cached_prompt_tokens` / `model_reported` 强类型字段,解析留在 `openai_compat.py` | OpenAI 报文格式知识不出 `transports/`,middleware 只做搬运(P7) |
| B 缓存命中语义 | 原样回放;度量口径必须带 `cache_hit = false` | 与 `_rehydrate` 既有口径一致——它只覆写时序字段,`model`/`prompt_tokens` 全回放 |
| C 缓存单价 | `ModelPrice` 加可选 `cached_input_per_1m`,`cost()` 加可选参 | 旧价格表与 `embedding.py:419` 三参调用零改动;未配置该档时**不猜折扣率**,退化为全额计价 |
| D 遥测扩列 | 端口 18 → 20 字段;DDL 加列 + 初始化期幂等补列 | `CREATE TABLE IF NOT EXISTS` 不会给旧库补列,INSERT 会**逐行 warning 丢弃**——遥测全失却无硬失败提示 |
## 被否决的备选
| 备选 | 否决原因 |
|---|---|
| A1 往 `raw` 里塞约定键 | `dict[str, Any]` 沦为隐式契约,且 middleware 要懂 OpenAI 嵌套结构 |
| A3 middleware 内解析 raw | 报文格式知识进 middleware,新增非 OpenAI 兼容 transport 时会分叉,违反分层 |
| B2 命中时置 None / B3 混合 | 与同层 `prompt_tokens` 的回放行为不一致,下游要记两套规则 |
| C2 `cost()` 直接收 `LLMResponse` | `pricing.py` 会反向依赖 `types.py`,且纯函数难单测 |
| D2 只改 DDL、文档写「删表重建」 | 已建表的开发机/下游只会看到降级 warning,排查成本高 |
| D3 引入 alembic 迁移框架 | 新增依赖违反「依赖极简」铁律,规模严重不匹配 |
## 独立审查修正(2026-07-31)
Codex CLI 安装损坏(vendor 二进制缺失),改由全新上下文的 Claude subagent 审。三条问题全部核实属实并已折回设计:
1. PG 缺列时**不是**结构性短路,而是逐行 warning(`_failed` 仅在 `_ensure_ready` 置位)。
2. SQLite 补列若塞进 `__init__` 现有 try,异常会让 `_conn` 停在 `None` → recorder 永久 no-op。已定纪律: 独立 try、置于 `self._conn = conn` 之后、duplicate column 视为成功。
3. 「端口无默认值 → 漏改即报错」不成立(无 mypy,8 个 fake 全是 `**fields`)。改为新增「emitter 实参键集合 == `_COLUMNS`」契约测试兜底——否则 `KeyError` 会被 `_record``except Exception` 吞成 warning,静默丢遥测。
相关: [[m1-core-design]]、[[est-tokens-decoupling]]
+17
View File
@@ -0,0 +1,17 @@
---
type: design
node_id: design:sampling-params
title: "采样参数透传设计(issue #4)"
date: 2026-07-31
---
# 采样参数透传设计(issue #4)
正文: `2026-07-31-sampling-params-design.md`。状态: 待人类审批。
- **选定方案**: 两层入口——调用级 `chat(..., overlay=)` 供逐 rollout 变化的 `seed`,配置级 `SourceConfig.extra_body`(env 键 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`,JSON 串)供恒定的 `temperature=0`。优先级 **结构化注入 > 调用级 > 配置级** 由现有层序天然给出,不加新机制。
- **issue 未提但必须一并处理的四件事**: ① 采样参数进缓存 key(否则 5 个 seed 全命中同一缓存、标准差恒为 0,实验静默作废——「无缓存毒化」铁律);② 保护键黑名单 `{model, messages, stream, stream_options}` 与值可 JSON 序列化,均在构造期报错(覆盖它们会击穿流式看门狗、成本遥测与 TPM 结算;不可序列化的值会在 `CacheMW` 降级 try 之外抛裸 `TypeError`,一行遥测都没有);③ 采样参数入遥测(端口 20 → 21 字段,列名 `sampling`);④ 入参拷贝语义。
- **关键结构决策**: `ChatRequest``sampling` 快照字段作为**跨洋葱层恒定的读取点**。`request.overlay``StructuredMW` 内侧含 `response_format`、外侧不含,缓存 key 与三个遥测 emit 入口若各读各的层就会口径分叉。`sampling` 列语义定死为「调用方意图 ⊎ 生效源 `extra_body`」,**不含**结构化注入。
- **被否决备选及理由**: `chat()` 展开为 `temperature=`/`seed=` 具名参数(供应商私有参数无穷尽,等于永久追加签名,违「深模块窄接口」);配置级放装配层全局字典(采样参数与源强相关,会把无效键发给不认识它的源);overlay 不进 key 靠调用方传 `cache_salt`(把毒化防护责任推给调用方,漏传不报错——正是 issue 抱怨的失败形态);采样参数不入遥测由下游 run 快照自记(中间态数据不可追溯,且分两步要做两遍 DDL 迁移);缓存与遥测直接读 `request.overlay` 不加 `sampling` 字段(口径必分叉);`sampling` 记含 `response_format` 的完整合并结果(列名为采样参数,且数 KB schema 逐行落库无谓膨胀);给 embedding 加 `extra_body` 透传(embedding 无采样一说,装配期报错比静默无效更能指路);transport 层重复校验保护键(三入口已构造期收口,属 gold-plating)。
- **附带修正**: `providers.py``minimax`/`openai` 空 thinking profile 补后果说明(`enable_thinking=False` 对两者不产生效果,调用方以为关掉了实际没关);`_SOURCE_FIELDS` 跨 scope 共用导致 `EXTRA_BODY` 在 OCR/EMBED scope 静默无效,改为构造期**剥离 + warning**(2026-07-31 人类拍板由原「装配期 `ValueError`」改此档: 这两条路径无采样语义,不值得让下游装配起不来)。**剥离不可省**——不剥离则遥测会记录一个从未发出的参数(决策 D 的 merge 读 `source.extra_body`,而 `monkey_ocr` 只发 multipart、`embed` payload 硬编码),那是数据造假而非参数失效;在 emitter 内特判调用方身份则违「遥测调用点收敛单一 helper」铁律。
- **审查留痕**: Codex CLI 不可用(vendor 二进制缺失),改派全新上下文 subagent 两轮只读审查。首轮报 5 项必修(三个 emit 入口口径分叉、OCR/embedding 耦合、JSON 序列化缺口、注释归属写反、同步清单漏 4 处),逐条核实后全部采纳;次轮结论通过,其 5 条建议(承重不变式测试、`sampling` 类型定死、拷贝语义跟进、共用范围收窄、报错文案指路)亦已就地收进。
+34
View File
@@ -110,6 +110,26 @@
"id": "plan:est-tokens-decoupling", "id": "plan:est-tokens-decoupling",
"label": "est_tokens 解耦实施计划", "label": "est_tokens 解耦实施计划",
"type": "plan" "type": "plan"
},
{
"id": "design:response-observability-fields",
"label": "响应可观测字段扩展(Issue #3)",
"type": "design"
},
{
"id": "plan:response-observability-fields",
"label": "响应可观测字段扩展实现计划",
"type": "plan"
},
{
"id": "design:sampling-params",
"label": "采样参数透传设计(issue #4)",
"type": "design"
},
{
"id": "plan:sampling-params-plan",
"label": "采样参数透传实现计划(issue #4)",
"type": "plan"
} }
], ],
"links": [ "links": [
@@ -189,6 +209,20 @@
"relation": "implements", "relation": "implements",
"evidence": "5 任务实现设计 §3.2 的 11 条改动项;任务排序经中间态破窗分析(先加能力→切调用点→三态生效→解绑约束)", "evidence": "5 任务实现设计 §3.2 的 11 条改动项;任务排序经中间态破窗分析(先加能力→切调用点→三态生效→解绑约束)",
"added": "2026-07-30T09:39:26.442986+00:00" "added": "2026-07-30T09:39:26.442986+00:00"
},
{
"source": "plan:response-observability-fields",
"target": "design:response-observability-fields",
"relation": "implements",
"evidence": "计划 T1-T7 逐条实现设计的 A2/B1/C1/D1 四个决策",
"added": "2026-07-31T11:10:03.872049+00:00"
},
{
"source": "plan:sampling-params-plan",
"target": "design:sampling-params",
"relation": "implements",
"evidence": "11 个任务逐条覆盖设计的决策 A-G 与 §5 的 14 条测试清单",
"added": "2026-07-31T16:59:35.657367+00:00"
} }
] ]
} }
+12 -4
View File
@@ -1,8 +1,8 @@
# Research Wiki 索引 # Research Wiki 索引
> 自动生成,更新时间:2026-07-30 09:39 UTC > 自动生成,更新时间:2026-08-01 01:58 UTC
## design (16) ## design (20)
- [2026-07-20-m1-core-design](designs/2026-07-20-m1-core-design.md) `design:2026-07-20-m1-core-design` - [2026-07-20-m1-core-design](designs/2026-07-20-m1-core-design.md) `design:2026-07-20-m1-core-design`
- [2026-07-20-m2-distributed-design](designs/2026-07-20-m2-distributed-design.md) `design:2026-07-20-m2-distributed-design` - [2026-07-20-m2-distributed-design](designs/2026-07-20-m2-distributed-design.md) `design:2026-07-20-m2-distributed-design`
- [2026-07-21-m25-resilience-design](designs/2026-07-21-m25-resilience-design.md) `design:2026-07-21-m25-resilience-design` - [2026-07-21-m25-resilience-design](designs/2026-07-21-m25-resilience-design.md) `design:2026-07-21-m25-resilience-design`
@@ -11,6 +11,8 @@
- [2026-07-29-settings-invariant-guards-design](designs/2026-07-29-settings-invariant-guards-design.md) `design:2026-07-29-settings-invariant-guards-design` - [2026-07-29-settings-invariant-guards-design](designs/2026-07-29-settings-invariant-guards-design.md) `design:2026-07-29-settings-invariant-guards-design`
- [2026-07-30-est-tokens-decoupling-design](designs/2026-07-30-est-tokens-decoupling-design.md) `design:2026-07-30-est-tokens-decoupling-design` - [2026-07-30-est-tokens-decoupling-design](designs/2026-07-30-est-tokens-decoupling-design.md) `design:2026-07-30-est-tokens-decoupling-design`
- [2026-07-30-settings-invariants-round-2-design](designs/2026-07-30-settings-invariants-round-2-design.md) `design:2026-07-30-settings-invariants-round-2-design` - [2026-07-30-settings-invariants-round-2-design](designs/2026-07-30-settings-invariants-round-2-design.md) `design:2026-07-30-settings-invariants-round-2-design`
- [2026-07-31-response-observability-fields-design](designs/2026-07-31-response-observability-fields-design.md) `design:2026-07-31-response-observability-fields-design`
- [2026-07-31-sampling-params-design](designs/2026-07-31-sampling-params-design.md) `design:2026-07-31-sampling-params-design`
- [est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)](designs/est-tokens-decoupling.md) `design:est-tokens-decoupling` - [est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)](designs/est-tokens-decoupling.md) `design:est-tokens-decoupling`
- [GatewaySettings 装配校验补齐(第二轮)](designs/settings-invariants-round-2.md) `design:settings-invariants-round-2` - [GatewaySettings 装配校验补齐(第二轮)](designs/settings-invariants-round-2.md) `design:settings-invariants-round-2`
- [GatewaySettings 跨字段不变量守卫的生效范围](designs/settings-invariant-guards.md) `design:settings-invariant-guards` - [GatewaySettings 跨字段不变量守卫的生效范围](designs/settings-invariant-guards.md) `design:settings-invariant-guards`
@@ -19,6 +21,8 @@
- [M2.5 治理韧性: 半死源隔离与健康感知调度](designs/m25-resilience.md) `design:m25-resilience` - [M2.5 治理韧性: 半死源隔离与健康感知调度](designs/m25-resilience.md) `design:m25-resilience`
- [M3 OCR 端口族设计](designs/m3-ocr.md) `design:m3-ocr` - [M3 OCR 端口族设计](designs/m3-ocr.md) `design:m3-ocr`
- [M4 迁移验证设计(GovDoc→CHS,发 v1.0)](designs/m4-migration.md) `design:m4-migration` - [M4 迁移验证设计(GovDoc→CHS,发 v1.0)](designs/m4-migration.md) `design:m4-migration`
- [响应可观测字段扩展(Issue #3)](designs/response-observability-fields.md) `design:response-observability-fields`
- [采样参数透传设计(issue #4)](designs/sampling-params.md) `design:sampling-params`
## finding (11) ## finding (11)
- [2026-07-20-m2-soak-workload](findings/2026-07-20-m2-soak-workload.md) `finding:2026-07-20-m2-soak-workload` - [2026-07-20-m2-soak-workload](findings/2026-07-20-m2-soak-workload.md) `finding:2026-07-20-m2-soak-workload`
@@ -33,22 +37,26 @@
- [P6 混合浸泡首跑基线与记分板三重伪击穿修复](findings/p6-soak-baseline.md) `finding:p6-soak-baseline` - [P6 混合浸泡首跑基线与记分板三重伪击穿修复](findings/p6-soak-baseline.md) `finding:p6-soak-baseline`
- [P7 OCR soak 验收: 99.73% 与 13 不变量全 PASS](findings/p7-ocr-soak.md) `finding:p7-ocr-soak` - [P7 OCR soak 验收: 99.73% 与 13 不变量全 PASS](findings/p7-ocr-soak.md) `finding:p7-ocr-soak`
## plan (12) ## plan (16)
- [2026-07-20-m1-core-plan](plans/2026-07-20-m1-core-plan.md) `plan:2026-07-20-m1-core-plan` - [2026-07-20-m1-core-plan](plans/2026-07-20-m1-core-plan.md) `plan:2026-07-20-m1-core-plan`
- [2026-07-20-m2-distributed-plan](plans/2026-07-20-m2-distributed-plan.md) `plan:2026-07-20-m2-distributed-plan` - [2026-07-20-m2-distributed-plan](plans/2026-07-20-m2-distributed-plan.md) `plan:2026-07-20-m2-distributed-plan`
- [2026-07-21-m25-resilience-plan](plans/2026-07-21-m25-resilience-plan.md) `plan:2026-07-21-m25-resilience-plan` - [2026-07-21-m25-resilience-plan](plans/2026-07-21-m25-resilience-plan.md) `plan:2026-07-21-m25-resilience-plan`
- [2026-07-21-m3-ocr-plan](plans/2026-07-21-m3-ocr-plan.md) `plan:2026-07-21-m3-ocr-plan` - [2026-07-21-m3-ocr-plan](plans/2026-07-21-m3-ocr-plan.md) `plan:2026-07-21-m3-ocr-plan`
- [2026-07-22-m4-migration-plan](plans/2026-07-22-m4-migration-plan.md) `plan:2026-07-22-m4-migration-plan` - [2026-07-22-m4-migration-plan](plans/2026-07-22-m4-migration-plan.md) `plan:2026-07-22-m4-migration-plan`
- [2026-07-30-est-tokens-decoupling-plan](plans/2026-07-30-est-tokens-decoupling-plan.md) `plan:2026-07-30-est-tokens-decoupling-plan` - [2026-07-30-est-tokens-decoupling-plan](plans/2026-07-30-est-tokens-decoupling-plan.md) `plan:2026-07-30-est-tokens-decoupling-plan`
- [2026-07-31-response-observability-fields](plans/2026-07-31-response-observability-fields.md) `plan:2026-07-31-response-observability-fields`
- [2026-07-31-sampling-params](plans/2026-07-31-sampling-params.md) `plan:2026-07-31-sampling-params`
- [est_tokens 解耦实施计划](plans/est-tokens-decoupling.md) `plan:est-tokens-decoupling` - [est_tokens 解耦实施计划](plans/est-tokens-decoupling.md) `plan:est-tokens-decoupling`
- [M1 核心里程碑实现计划](plans/m1-core-plan.md) `plan:m1-core-plan` - [M1 核心里程碑实现计划](plans/m1-core-plan.md) `plan:m1-core-plan`
- [M2 分布式实现计划](plans/m2-distributed.md) `plan:m2-distributed` - [M2 分布式实现计划](plans/m2-distributed.md) `plan:m2-distributed`
- [M2.5 治理韧性实现计划](plans/m25-resilience.md) `plan:m25-resilience` - [M2.5 治理韧性实现计划](plans/m25-resilience.md) `plan:m25-resilience`
- [M3 OCR 实现计划](plans/m3-ocr.md) `plan:m3-ocr` - [M3 OCR 实现计划](plans/m3-ocr.md) `plan:m3-ocr`
- [M4 迁移实现计划(T0-T14)](plans/m4-migration.md) `plan:m4-migration` - [M4 迁移实现计划(T0-T14)](plans/m4-migration.md) `plan:m4-migration`
- [响应可观测字段扩展实现计划](plans/response-observability-fields.md) `plan:response-observability-fields`
- [采样参数透传实现计划(issue #4)](plans/sampling-params-plan.md) `plan:sampling-params-plan`
## schema (1) ## schema (1)
- [表结构: llm_calls(遥测 18 字段)](schemas/llm-calls.md) `schema:llm-calls` - [表结构: llm_calls(遥测 21 字段)](schemas/llm-calls.md) `schema:llm-calls`
## metric (2) ## metric (2)
- [OCR 治理调用成功率与错误分类分布](metrics/ocr-call-success.md) `metric:ocr-call-success` - [OCR 治理调用成功率与错误分类分布](metrics/ocr-call-success.md) `metric:ocr-call-success`
+16
View File
@@ -57,3 +57,19 @@
- [2026-07-30 09:39 UTC] 新增 plan: est_tokens 解耦实施计划 (plan:est-tokens-decoupling) - [2026-07-30 09:39 UTC] 新增 plan: est_tokens 解耦实施计划 (plan:est-tokens-decoupling)
- [2026-07-30 09:39 UTC] 新增边: plan:est-tokens-decoupling --implements--> design:est-tokens-decoupling - [2026-07-30 09:39 UTC] 新增边: plan:est-tokens-decoupling --implements--> design:est-tokens-decoupling
- [2026-07-30 09:39 UTC] 重建索引: 42 篇页面 - [2026-07-30 09:39 UTC] 重建索引: 42 篇页面
- [2026-07-31 08:35 UTC] 新增 design: 响应可观测字段扩展(Issue #3) (design:response-observability-fields)
- [2026-07-31 08:37 UTC] 重建索引: 44 篇页面
- [2026-07-31 11:10 UTC] 新增 plan: 响应可观测字段扩展实现计划 (plan:response-observability-fields)
- [2026-07-31 11:10 UTC] 新增边: plan:response-observability-fields --implements--> design:response-observability-fields
- [2026-07-31 11:10 UTC] 重建索引: 46 篇页面
- [2026-07-31 11:11 UTC] 重建索引: 46 篇页面
- [2026-07-31 12:25 UTC] 重建索引: 46 篇页面
- [2026-07-31 15:57 UTC] 新增 design: 采样参数透传设计(issue #4) (design:sampling-params)
- [2026-07-31 15:57 UTC] 重建索引: 48 篇页面
- [2026-07-31 16:00 UTC] 重建索引: 48 篇页面
- [2026-07-31 16:31 UTC] 重建索引: 48 篇页面
- [2026-07-31 16:59 UTC] 新增 plan: 采样参数透传实现计划(issue #4) (plan:sampling-params-plan)
- [2026-07-31 16:59 UTC] 新增边: plan:sampling-params-plan --implements--> design:sampling-params
- [2026-07-31 16:59 UTC] 重建索引: 50 篇页面
- [2026-07-31 17:01 UTC] 重建索引: 50 篇页面
- [2026-08-01 01:58 UTC] 重建索引: 50 篇页面
@@ -0,0 +1,224 @@
# 实现计划: 响应可观测字段扩展(Issue #3)
- **目标**: 让 `LLMResponse` 与遥测表如实暴露「供应商 prompt cache 命中的输入 token 数」与「API 实际返回的模型版本串」。
- **方案概述**: 报文解析留在 `transports/`(新增两个强类型字段随 `TransportResult` 上浮),`RetryMW` 只搬运;遥测端口由 18 字段扩到 20 并给两个后端加幂等补列;`PricingTable` 增加可选缓存单价档消除 cost 高估。缓存命中行按既有口径原样回放。
- **依据设计**: `research-wiki/designs/2026-07-31-response-observability-fields-design.md`(2026-07-31 已获人类批准,决策 A2/B1/C1/D1)。
- **涉及技术**: Python 3.11 frozen dataclass、httpx SSE 解析、sqlite3、asyncpg、pytest。
- **保真校验**: 本计划**不涉及** `reference/` 参考实现迁移,保真校验不适用。但遥测后端属 ARCHITECTURE §1.4 资产,T5 明确约束「不得改变既有降级语义」。
## 文件结构
| 文件 | 职责 | 本次改动 |
|---|---|---|
| `src/polygateway/types.py` | 冻结公共类型 | `LLMResponse` / `TransportResult` 各 +2 字段;`cache_hit` docstring 消歧 |
| `src/polygateway/transports/openai_compat.py` | OpenAI 兼容报文解析 | 防御解析 helper;SSE sink 采集 `model`;两处 `TransportResult` 构造填新字段 |
| `src/polygateway/middleware/retry.py` | 尝试循环 | `_build_response` 搬运两字段 |
| `src/polygateway/middleware/cache.py` | 响应缓存 | **零代码改动**(自动透传),仅补测试固化行为 |
| `src/polygateway/pricing.py` | 单价换算 | `ModelPrice` +可选档;`cost()` +可选参;`from_file` 校验 |
| `src/polygateway/ports.py` | 端口契约 | `TelemetryRecorder` 18 → 20 字段 |
| `src/polygateway/telemetry/{sqlite,postgres}.py` | 遥测后端 | DDL +2 列;`_COLUMNS` +2;初始化期幂等补列 |
| `src/polygateway/middleware/telemetry.py` | 遥测唯一调用点 | `_record` 与三个 `emit_*` 搬运两字段;cost 换算传入缓存 token |
字段定义(全库唯一权威,后续任务一律引用此处):
```python
# LLMResponse 与 TransportResult 尾部,同名同类型同默认值
cached_prompt_tokens: int | None = None # 供应商 prompt cache 命中的输入 token;None = 该源未上报
model_reported: str | None = None # API 响应体的 model 字段;None = 未上报
```
---
## T1. 类型层加字段
- [ ] **改**: `src/polygateway/types.py`
**行为**: 在 `LLMResponse` 尾部(`structured_data` 之后)与 `TransportResult` 尾部(`raw` 之后)各追加上面两个字段。`cache_hit` 的语义在 `LLMResponse` docstring 中写明是「**PolyGateway 自身响应缓存**命中,与供应商 prompt cache 无关,后者见 `cached_prompt_tokens`」。
**验收**: 前 11 个字段的顺序与名字一字不动;新字段有默认值,`LLMResponse(...)` 按前 11 位置参数构造仍成立;`TransportResult` 现有两处构造(`openai_compat.py:354/436`)不传新字段也能构造。
**测试**(`tests/unit/test_types.py`): ① 不传新字段时两个类型的新字段均为 `None`;② 按位置构造 `LLMResponse` 的前 11 字段仍可用(迁移兼容承诺)。
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_types.py -v` → PASS。
**提交**: `feat: add cached prompt tokens and reported model to response types`
## T2. transport 采集与防御解析
- [ ] **改**: `src/polygateway/transports/openai_compat.py`
**行为**分三处:
1. 新增两个模块级防御 helper(网关返回一律不可信,解析失败**返回 None,不抛异常**):
```python
def _coerce_cached_tokens(usage: Any) -> int | None:
"""从 usage.prompt_tokens_details.cached_tokens 取非负整数;任何形态异常 → None。"""
def _coerce_model_reported(value: Any) -> str | None:
"""响应体 model 字段: 非空 str 才收,其余(含空串/非 str)→ None。"""
```
`_coerce_cached_tokens` 需容忍:`usage` 为 None、`prompt_tokens_details` 缺失或非 dict、`cached_tokens``bool`/`str`/负数/浮点。`bool` 必须排除(Python 中 `isinstance(True, int)` 为真)。**`0` 必须如实保留而非归 None**——真实零命中与未上报是两回事,这是 issue 的核心诉求。
2. 流式路径:`_sse_delta`(`:44-47`)当前只把 `usage` 旁路进 sink。补一条——chunk 里出现 `model` 时写 `usage_sink["model"]`(**首次写入即固定**,后续 chunk 不覆盖,避免末帧异常值污染)。`_stream_once``TransportResult` 构造(`:354`)填 `cached_prompt_tokens=_coerce_cached_tokens(sink.get("usage"))``model_reported=_coerce_model_reported(sink.get("model"))`
3. 非流式路径:`TransportResult` 构造(`:436`)填 `_coerce_cached_tokens(body.get("usage"))``_coerce_model_reported(body.get("model"))`
**验收**: `raw` 的内容保持原样不动(新字段是独立格子,不是杂物袋的扩充);OCR 与 embedding 的解析路径一行不改。
**测试**(`tests/unit/test_openai_compat.py`,用现有 fake 响应二次构造):
| 用例 | 期望 |
|---|---|
| 非流式 usage 含 `prompt_tokens_details.cached_tokens: 128` | `cached_prompt_tokens == 128` |
| 流式 usage 帧同上 | 同上 |
| 无 `prompt_tokens_details` / usage 帧缺失 | `None` |
| `cached_tokens``"abc"` / `-1` / `True` / `1.5` / `[]` / dict | `None`,且**不抛异常** |
| `cached_tokens``0` | `0`(真实零命中,**不得**归 None) |
| `prompt_tokens_details` 非 dict | `None` |
| 非流式 body 含 `model: "MiniMax-Text-01-250321"` | `model_reported` 为该串 |
| 流式首个含 model 的 chunk 后又出现不同 model | 取**首个** |
| body 无 `model` / `model``""` | `None` |
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_openai_compat.py -v` → PASS(新增用例先失败后通过)。
**提交**: `feat: collect provider cache tokens and reported model in transport`
## T3. RetryMW 搬运与缓存回放固化
- [ ] **改**: `src/polygateway/middleware/retry.py`
**行为**: `_build_response`(`:419-438`)追加 `cached_prompt_tokens=result.cached_prompt_tokens``model_reported=result.model_reported``model=source.model` **保持不变**——别名仍是主字段,真实版本是旁证(设计非目标 2)。
`middleware/cache.py` **不改一行**:`_RESPONSE_FIELDS``dataclasses.fields(LLMResponse)` 动态生成(`:26`)、`_serialize``asdict`(`:139`),新字段自动进出;决策 B1 要求命中时原样回放,而 `_rehydrate` 的覆写清单(`:113-119`)本就不含新字段,零改动即是正确行为。本任务用测试把它钉死。
**测试**:
- `tests/unit/test_retry.py`: transport 返回带两字段的 `TransportResult``chat()` 返回的 `LLMResponse` 上两字段一致;transport 未上报时为 `None`
- `tests/unit/test_cache.py`: ① 带两字段的响应写入缓存再命中,回放值与原值相等且 `cache_hit=True`;② **旧格式兼容**——手工构造缺这两个键的缓存 JSON 塞进后端,命中后能正常 rehydrate 且两字段为 `None`(不得抛异常回源)。
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_retry.py tests/unit/test_cache.py -v` → PASS。
**提交**: `feat: carry the new observability fields through retry and cache`
## T4. 缓存读取单价
- [ ] **改**: `src/polygateway/pricing.py`
**行为**:
```python
class ModelPrice: # 追加第三档,可选
cached_input_per_1m: float | None = None
def cost(self, model: str, prompt_tokens: int, completion_tokens: int,
cached_prompt_tokens: int | None = None) -> float | None:
```
换算规则(设计 §4):配了缓存档**且** `cached_prompt_tokens` 为正 → `(prompt - cached) × input + cached × cached_input`;否则全额按 `input`(现状,逐位不变)。`cached > prompt` 时按 `prompt` 夹取并 `logger.warning` 一次(沿用 `_warned` 的去重思路,按 model 去重,防日志风暴),**绝不产生负成本**。
`from_file` 的 fail-loud 扩展:条目出现 `cached_input_per_1m` 键时必须可转 float 且非负,否则 `ValueError`;不出现该键 = 合法(旧价格表零改动)。`__post_init__` 同步校验非负。
顺带订正 `PricingTable` docstring(`pricing.py:36`)那句「cost() 是全库唯一换算点(经 TelemetryEmitter)」——实际有 `TelemetryEmitter`(`middleware/telemetry.py:137`)与 `embedding.py:419` 两个调用点(设计 §6 行为审计已声明)。**只改这一行注释,不做任何结构重构**。
**验收**: `embedding.py:419` 的三参调用形态**一行不改**仍可用;未配缓存档时,任意输入下 `cost()` 结果与改前逐位相等。
**测试**(`tests/unit/test_pricing.py`): ① 配缓存档 + 命中 → 成本严格低于全额且等于手算值;② 未配缓存档 + 命中 → 与不传该参数结果相等;③ `cached > prompt` → 结果等于全部按缓存价、非负、有 warning;④ `cached_prompt_tokens=None/0` → 全额;⑤ 三参旧调用签名可用;⑥ 价格表含 `cached_input_per_1m: -1``"x"``from_file``ValueError`;⑦ 无该键的旧价格表照常加载。
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_pricing.py -v` → PASS。
**提交**: `feat: support a cached input price tier in the pricing table`
## T5. 端口扩字段与后端补列
- [ ] **改**: `src/polygateway/ports.py``src/polygateway/telemetry/sqlite.py``src/polygateway/telemetry/postgres.py`
**行为**:
1. `TelemetryRecorder.record_llm_call`(`ports.py:250-271`)在 `cost` 之后追加 `cached_prompt_tokens: int | None``model_reported: str | None`,**不设默认值**(设计 §5:库外无第三方实现者)。同步该 Protocol 的「18 字段冻结」docstring。
2. 两个后端的 `_DDL` 加列(SQLite `INTEGER`/`TEXT`;PG `INTEGER`/`TEXT`,均可空、无默认值)。**新列在 DDL 里必须放在 `created_at` 之后(即表的最末尾),不得插在 `cost` 之后**——旧表走 `ALTER TABLE ADD COLUMN` 只能追加到末尾,若新建库把新列插在 `created_at` 前面,两条路径的物理列序就会分叉,而 `tests/integration/test_postgres_telemetry.py:117-128``test_schema_has_frozen_columns_in_order``ordinal_position` 逐位断言,且该表是与真实批跑共享的表、**严禁 DROP/TRUNCATE**(文件头隔离纪律),分叉后没有合规修法。
`_COLUMNS``"cost"` 之后追加同名两项即可——`_INSERT` 是显式列名拼装(`sqlite.py:62-65`/`postgres.py:67-71`),`_COLUMNS` 只需与自身的 `row = tuple(...)` 自洽,**与 DDL 物理列序无关**。
3. **幂等补列**,按设计 D1 纪律执行:
- **SQLite**(`sqlite.py:74-84`):补列代码必须放在 `self._conn = conn` **之后**、用**独立 try**,且**首行必须守卫 `if self._conn is None: return`**——初始化 try 吞掉失败时 `self._conn` 仍是 `None`(局部 `conn` 甚至未绑定),无守卫的补列块会抛 `AttributeError`/`NameError`,这两者不被 `sqlite3.Error` 捕获,会直接逃出 `__init__`,打破「初始化失败静默降级」的对外契约(既有测试 `tests/unit/test_telemetry.py:132-135` `test_unwritable_path_degrades_silently` 会红)。守卫之后:`PRAGMA table_info(llm_calls)` 取现有列名集合,缺哪列补哪列;捕获 `sqlite3.Error` 时消息含 `duplicate column` 视为成功(多进程共库的 TOCTOU),其余记 warning。**绝不允许**因补列失败把 `self._conn` 置回 `None`——那会让整个 recorder 永久 no-op。
- **Postgres**(`postgres.py:_ensure_ready` 内、`_DDL` 执行之后):两条 `ALTER TABLE llm_calls ADD COLUMN IF NOT EXISTS ...`,共享既有 `_init_lock``except asyncio.CancelledError: raise` 结构。
**验收(降级语义不得改变)**: SQLite 侧 `except` 不得加宽(取消天然穿透);PG 侧 `CancelledError` 分支保持在最前;写入失败仍是逐行 warning 丢弃,不冒泡。
**必须同步改的测试(共 5 处)**:
| 位置 | 内容 | 漏改会怎样 |
|---|---|---|
| `tests/unit/test_telemetry.py:76``_record_minimal` | 手写 18 键 dict | **红**(`KeyError`) |
| `tests/integration/test_postgres_telemetry.py:81-105` `_record_minimal` | 同上 | **红** |
| `tests/unit/test_telemetry.py:18-40` `_EXPECTED_COLUMNS` | 19 项列序断言(含 `created_at`) | **红**;新列追加到 `created_at` **之后** |
| `tests/integration/test_postgres_telemetry.py:22-41` `_EXPECTED_COLUMNS` | 同上 | **红**;同上 |
| `tests/unit/test_ports.py:96` `_DummyRecorder` | 唯一写全签名的 fake | **不会红**(它只被 `:131``isinstance` 使用,`runtime_checkable` Protocol 只校验方法名不校验签名),但仍应同步以免误导后来者 |
前四处是本次仅有的天然拦截点;端口加参数**不会**带来编译期保护(本仓无 mypy,其余 8 个 fake 全是 `**fields`)。
**测试**:
- `tests/unit/test_telemetry.py`(SQLite):① 20 字段写入后可读回两个新列的值(含 `None`);② **旧表升级**——先用 18 列 DDL 手工建表,再实例化 `SQLiteRecorder`,写入成功且新列有值;③ **补列失败路径**(设计 §8 第 ③ 条,最危险的分支,不可用成功路径顶替)——构造一个 ALTER 必然失败的场景(把 `llm_calls` 建成同名 view,或注入在 ALTER 上抛 `sqlite3.OperationalError` 的连接),断言构造**不抛异常**、`recorder._conn` 仍非 `None`、后续 `record_llm_call` 不抛(降级为逐行 warning);④ 初始化路径不可写时仍静默降级(`test_unwritable_path_degrades_silently` 保持绿)。
- `tests/integration/test_postgres_telemetry.py`:① 20 字段写入 PG 并 `SELECT` 回读;② 18 列旧表经初始化后自动补列并写入成功。
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_telemetry.py tests/unit/test_ports.py -v` → PASS;PG 部分 `conda run -n PolyGateway --no-capture-output pytest tests/integration/test_postgres_telemetry.py -v` → PASS。**PG/Redis 属共享后端,严禁与其他会话或钩子测试并跑**,起跑前确认无并发占用。
**提交**: **与 T6 合并为一次提交**,不得单独落地。理由:T5 落地后 emitter 仍只传 18 键,后端的 `row = tuple(fields[col] for col in _COLUMNS)` 会抛 `KeyError`,被 `_record``except Exception`(`middleware/telemetry.py:164-165`)吞成 warning → **该 commit 处于全量遥测静默丢失的状态**,且现有测试无一能捕获。提交信息见 T6。
## T6. Emitter 搬运与契约测试(与 T5 同一次提交)
- [ ] **改**: `src/polygateway/middleware/telemetry.py`
**行为**: `_record`(`:108-125`)新增两个参数并透传给 `record_llm_call`;三个入口各自提供取值——
| 入口 | `cached_prompt_tokens` | `model_reported` |
|---|---|---|
| `emit_attempt` | `response.cached_prompt_tokens if response else None` | 同左 |
| `emit_cache_hit` | `response.cached_prompt_tokens`(B1 原样回放) | 同左 |
| `emit_terminal_failure` | `None` | `None` |
cost 换算(`:137`)改为把 `cached_prompt_tokens` 传进 `self._pricing.cost(...)``cache_hit → 0.0``usage_source == "unavailable" → None` 两条短路的**先后顺序一字不动**(ARCHITECTURE §5.1 cost 口径不变式)。
**测试**(`tests/unit/test_telemetry.py`):
- **契约测试(不可省)**: 用记录 kwargs 的 fake recorder 跑一次 `emit_attempt`,断言 `set(kwargs) == set(sqlite._COLUMNS) == set(postgres._COLUMNS)`。理由:`row = tuple(fields[col] for col in _COLUMNS)` 位于两个后端 try 之外(`sqlite.py:90`/`postgres.py:121`),emitter 漏传字段会抛 `KeyError` 并被 `_record``except Exception` 吞成 warning → 静默丢遥测;现有 8 个 `**fields` 形态的 fake 一个都拦不住。
- 三个入口各记一行,断言新字段取值符合上表。
- cost 回归:配了缓存档且响应带 `cached_prompt_tokens` → 落库 cost 低于全额;缓存命中行 cost 仍为 `0.0`;`unavailable` 行仍为 `None`
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_telemetry.py -v` → PASS;随后 `make ci` 全绿(含 ruff 与 import-linter)。
**提交**(含 T5 全部改动): `feat: record the observability fields end to end through telemetry`
## T7. 文档同步与发版
- [ ] **改**: `research-wiki/ARCHITECTURE.md``CHANGELOG.md``.env.example``pyproject.toml``src/polygateway/__init__.py`、四处「18 字段冻结」措辞点、Gitea Wiki 站
**行为**:
1. `ARCHITECTURE.md`:§5.1 新增字段表补两行;**§7.8「必录字段」的行内清单**(`:452`)补两项——该文件不含字面「18 字段冻结」,§7.8 与 D8(`:202`)才是遥测字段的落点;§7.8 末条「`pricing.py` 维护 model →(input 单价, output 单价)表」同步第三档。补一条度量口径警示(与 cost 缺口同款):**统计供应商缓存命中率必须带 `WHERE cache_hit = false`**,否则缓存回放行会被重复计入。
2. 代码里的「18 字段/18 列」措辞共 **6 处**,全部订正(`grep -rn "18 字段\|18 列" src/ tests/` 可复核):`ports.py:248``middleware/telemetry.py:31``pricing.py:6``telemetry/sqlite.py:87``telemetry/postgres.py:9`(「18 列 schema 与 SQLite 版同名同序」)、`tests/unit/test_telemetry.py:1`
3. `.env.example:56` 是仓内**唯一**的价格表格式说明(无独立模板文件),补 `cached_input_per_1m` 可选档与「不填即全额计价、库不猜折扣率」的说明。
4. 版本 bump `1.0.3``1.1.0`,**两处必须同步**(`pyproject.toml:7``src/polygateway/__init__.py:34`;`tests/unit/test_package.py:11` 会断言二者相等)。
5. `CHANGELOG.md` 顶部新增 `## 1.1.0` 段,沿用既有写法(先讲问题、再讲变更、点明下游要读什么):两个新字段的语义与 `None`/`0` 之别、`cache_hit` 与供应商 prompt cache 的区分、遥测表新增两列与自动补列、价格表可选缓存档、度量口径的 `cache_hit = false` 约束。
6. Gitea Wiki 站(需单独 `git clone https://gitea.iomgaa.online/iomgaa/PolyGateway.wiki.git`)按 `docs-convention.md` §2 清单同步:`参考-公共API`(LLMResponse 字段表)、`参考-配置键`(价格表格式)、`指南-遥测与成本`(新列与成本校正口径)、`Home.md` 版本号与安装命令、`_Sidebar.md` 如有结构变化。
**验收**: 版本 bump 的提交**不允许单独存在**(docs-convention §2 门),必须与 wiki/CHANGELOG 同步在同一次交付内。
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_package.py -v` → PASS;`make ci` 全绿。
**提交**: `chore: release 1.1.0 with the response observability fields`
---
## 合并前门(逐条对应 CLAUDE.md §3)
- [ ] 每个行为变更都有「先失败后通过」的测试证据(T1-T6 各自的新增用例)。
- [ ] `make ci` 全绿(ruff + import-linter + pytest + 覆盖率)。
- [ ] 派**全新上下文**的 verifier subagent 独立验证(跨多文件,`verification-before-completion` 强制档)。
- [ ] 合并前整分支代码审查(`requesting-code-review`)。
- [ ] Gitea Issue #3 的关闭说明:两个字段的最终名字与语义、缓存命中行的回放口径、遥测新列与补列行为。
@@ -0,0 +1,396 @@
# 实现计划: 采样参数透传(issue #4)
- **设计**: `research-wiki/designs/2026-07-31-sampling-params-design.md`(2026-07-31 人类批准)
- **分支**: `feat/issue-4-sampling-params`
- **目标**: 让下游能固定解码参数(`temperature`/`seed`/`max_tokens`),且不破坏缓存隔离与遥测诚实性。
- **方案概述**: `chat()` 增 keyword-only `overlay` 参数(调用级),`SourceConfig``extra_body` 字段(配置级)。`ChatRequest``sampling` 快照字段作为跨洋葱层恒定读取点,供缓存 key 与遥测消费。遥测端口 20 → 21 字段。
- **技术**: Python 3.11+,frozen dataclass,`MappingProxyType`,sqlite3 / asyncpg DDL 幂等补列。
**保真校验**: 本计划不涉及 `reference/` 参考实现迁移,保真校验不适用。
---
## 1. 文件结构
| 文件 | 职责变更 |
|---|---|
| `src/polygateway/types.py` | 新增 `validate_request_overlay()``merge_sampling()` 两个纯函数;`ChatRequest.sampling` 字段;`SourceConfig.extra_body` 字段与构造期校验 |
| `src/polygateway/client.py` | `chat()``overlay` 参数;`model_fingerprint` 计算纳入 `extra_body` |
| `src/polygateway/middleware/cache.py` | `build_cache_key()``sampling` 入参并纳入 key |
| `src/polygateway/transports/openai_compat.py` | `_build_payload` 在 thinking profile 之后、overlay 之前应用 `source.extra_body` |
| `src/polygateway/config.py` | `_SOURCE_FIELDS``EXTRA_BODY`;`_cast``json` 分支 |
| `src/polygateway/ports.py` | `TelemetryRecorder.record_llm_call` 增第 21 参 `sampling` |
| `src/polygateway/middleware/telemetry.py` | 三个 emit 入口按设计表格产出 `sampling`;`_record` 透传 |
| `src/polygateway/telemetry/sqlite.py` | DDL / `_BACKFILL_COLUMNS` / `_COLUMNS``sampling` |
| `src/polygateway/telemetry/postgres.py` | DDL / `_BACKFILL` / `_COLUMNS``sampling` |
| `src/polygateway/ocr.py` / `embedding.py` | 构造期剥离 `extra_body` + warning(决策 G) |
| `src/polygateway/providers.py` | minimax/openai 空 thinking profile 补后果注释(决策 F) |
| `.env.example` / `README.md` / `CHANGELOG.md` / `research-wiki/ARCHITECTURE.md` | 文档同步(设计 §6) |
**各任务需新增的 import**(现状核实,不加即 NameError):
| 文件 | 需新增 |
|---|---|
| `types.py` | `from collections.abc import Mapping``from types import MappingProxyType``import json`。**该文件无 `from __future__ import annotations`**,注解在类体求值,`Mapping` 必须真导入 |
| `client.py` | `import json``import hashlib` |
| `config.py` | `import json` |
| `ocr.py` / `embedding.py` | `import dataclasses`(现只有 `from dataclasses import dataclass`)、`from loguru import logger`(若未导入) |
| `middleware/telemetry.py` | `merge_sampling`/`canonical_sampling_json` 需**运行时**导入(现对 `polygateway.types` 只在 `TYPE_CHECKING` 下导入) |
**关键接口**(跨任务消费,此处定死):
```python
# types.py —— 两个纯函数 + 两个字段
_PROTECTED_OVERLAY_KEYS = frozenset({"model", "messages", "stream", "stream_options"})
def validate_request_overlay(overlay: Mapping[str, Any], *, origin: str) -> dict[str, Any]:
"""校验采样参数覆盖层并返回浅拷贝;origin 用于错误信息定位来源。
保护键会击穿治理(model→成本算错、messages→缓存与遥测口径失真、
stream/stream_options→绕过看门狗与 usage 帧);值必须 JSON 可序列化,
否则会在 CacheMW 的降级 try 之外抛裸 TypeError(设计 §决策 B)。
"""
def merge_sampling(extra_body: Mapping[str, Any], sampling: Mapping[str, Any]) -> dict[str, Any]:
"""合并配置级与调用级采样参数(调用级优先);两者皆空返回空 dict。"""
def canonical_sampling_json(merged: Mapping[str, Any]) -> str | None:
"""遥测列与缓存 key 共用的序列化口径;空 mapping → None。"""
@dataclass(frozen=True)
class ChatRequest:
...
overlay: dict[str, Any] = field(default_factory=dict)
sampling: Mapping[str, Any] = field(default_factory=dict) # 新增
@dataclass(frozen=True)
class SourceConfig:
...
extra_body: Mapping[str, Any] = field(default_factory=dict) # 新增,__post_init__ 转 MappingProxyType
```
```python
# middleware/cache.py —— 签名扩展(sampling 为 keyword-only)
# 默认值用 None 而非 {}: dict 字面量作默认参数会被 ruff B006 拦下
def build_cache_key(
model_fingerprint: str,
messages: list[dict[str, Any]],
namespace: str,
salt: str | None,
*,
sampling: Mapping[str, Any] | None = None,
) -> str: ...
```
```python
# client.py —— chat() 新签名
async def chat(
self, messages: list[dict[str, Any]], *,
session_id: str | None = None, parent_call_id: str | None = None,
cache_salt: str | None = None, cache_namespace: str | None = None,
structured: type[BaseModel] | Literal["json"] | None = None,
stream: bool = True,
overlay: Mapping[str, Any] | None = None, # 新增
) -> LLMResponse: ...
```
---
## 2. 任务清单
任务按依赖排序;每个任务一次提交、独立可验证。每个任务合并前必须出示**先失败后通过**的测试证据(先写测试跑红,再实现跑绿)。
统一验证命令前缀:`conda run -n PolyGateway --no-capture-output pytest`
> **共享后端纪律**: 涉及 Redis/Postgres 的 integration 测试严禁与其他会话并跑(含 git 钩子触发的测试)。Task 7、Task 11 受此约束。
---
### - [ ] Task 1: `types.py` 内核 —— 校验与合并纯函数 + 两个新字段
**文件**: 改 `src/polygateway/types.py`;测试 `tests/unit/test_types.py`
**实现行为**:
1. `validate_request_overlay(overlay, *, origin)`,**校验顺序即下列顺序**:
- 键必须是 `str`,否则 `ValueError`(canonical JSON 要求)。**必须排在序列化试探之前**——`{1: "a", "b": 2}``sort_keys=True` 下抛的是 `TypeError: '<' not supported between 'str' and 'int'`,若先试序列化会被误报成"值不可 JSON 序列化",指错方向;
- 命中 `_PROTECTED_OVERLAY_KEYS` 任一键 → `ValueError`,信息含 origin、违规键名、以及**为什么**(如 `stream` 会绕过流式看门狗);
- 对整个 mapping 做 `json.dumps(..., sort_keys=True)` 试序列化,`TypeError` → 转 `ValueError` 并指出该值不可 JSON 序列化(信息提示改用 `float(x)` 等原生类型);
- 返回 `dict(overlay)` 浅拷贝。
2. `merge_sampling(extra_body, sampling)``{**extra_body, **sampling}`(调用级优先)。
3. `canonical_sampling_json(merged)` → 空则 `None`,否则 `json.dumps(merged, sort_keys=True, ensure_ascii=False)`
4. `ChatRequest``sampling` 字段(见 §1 关键接口)。
5. `SourceConfig``extra_body` 字段;`__post_init__` 新增 `_validate_extra_body()`:调 `validate_request_overlay(self.extra_body, origin=f"SourceConfig({self.name}).extra_body")`,再 `object.__setattr__(self, "extra_body", MappingProxyType(dict(...)))`(frozen dataclass 需用 `object.__setattr__`)。
**已知后果(必须显式接受,不是疏漏)**: `SourceConfig` 加 mapping 字段后**不再 hashable**(`hash()``TypeError`),且因 `MappingProxyType` 不可 pickle,`dataclasses.asdict()` / `copy.deepcopy()` 也会失败。
- 不可 hash 是**加任何 mapping 字段的固有代价**,与是否用 `MappingProxyType` 无关(裸 `dict` 同样不可 hash),无法规避;
- 库内当前无调用点会踩:`asdict` 只用于 `LLMResponse`/`EmbeddingResponse`(`cache.py:140`),全库无 `set(sources)` 或以源作 dict key 的写法;
- 保留 `MappingProxyType` 而非裸 dict,是因为决策 E 的只读约束值得这个代价;下游要可变副本用 `dict(source.extra_body)`,要改字段用 `dataclasses.replace(source, ...)`(已验证可行,会重跑 `__post_init__` 重新包 proxy,不递归)。
**验收标准**: 四个保护键各自触发 `ValueError` 且信息含原因;非 str 键报的是"键必须是 str"而非"不可序列化";`{"temperature": object()}` 类不可序列化值报 `ValueError` 而非 `TypeError`;合法 `{"temperature": 0, "seed": 42}` 通过并返回独立副本(改原 dict 不影响返回值);`SourceConfig.extra_body` 构造后为 `MappingProxyType` 且不可改。
**测试要求**: 新增 `tests/unit/test_types.py::TestSamplingValidation`,覆盖上述每条。不可序列化值用 `object()` 实例即可,不引入 numpy 依赖。**另加一条锁定测试**:`pytest.raises(TypeError): hash(source_config)`,把"不再 hashable"钉成有意行为——否则将来有人踩到时会以为是 bug 并"修"回去。
**验证**: `pytest tests/unit/test_types.py -v` → 全 PASS
---
### - [ ] Task 2: `config.py` —— `EXTRA_BODY` env 解析
**文件**: 改 `src/polygateway/config.py`;测试 `tests/unit/test_config.py`
**实现行为**:
- `_SOURCE_FIELDS``"EXTRA_BODY": ("extra_body", "json")`;
- `_cast``json` 分支:`json.loads` 失败 → `ValueError`(沿用既有 `配置 {key} 解析失败: {exc}` 包装);解析结果**非 dict** → `ValueError`,信息说明必须是 JSON 对象(而非数组/标量)。
**验收标准**: `LLM__QWEN__1__EXTRA_BODY={"temperature":0}``SourceConfig.extra_body == {"temperature": 0}`;`{invalid``ValueError`;`[1,2]``ValueError`;`{"model":"x"}``ValueError`(经 Task 1 的 `SourceConfig.__post_init__` 保护键校验)。
**测试要求**: 新增 4 个 case 覆盖上述。**注意**: 这里同时验证了 Task 1 的校验确实挂在装配路径上。
**验证**: `pytest tests/unit/test_config.py -v` → 全 PASS
---
### - [ ] Task 3: `chat()` 入口 + transport 应用 + fingerprint
**文件**: 改 `src/polygateway/client.py``src/polygateway/transports/openai_compat.py`;测试 `tests/unit/test_client.py``tests/unit/test_openai_compat.py`
**实现行为**:
1. `chat()``overlay` 参数(见 §1 签名)。进洋葱**之前**:
```python
validated = validate_request_overlay(overlay or {}, origin="chat(overlay=...)")
```
同一份 `validated` 对象同时填 `ChatRequest.overlay` 与 `.sampling`(设计决策 E:一次拷贝、两个字段指向同一快照,不做两份独立拷贝)。
2. `_build_payload`:在 thinking profile 之后、`payload.update(overlay)` 之前插入 `payload.update(source.extra_body)`。**顺序即优先级,不可调换**。
3. `model_fingerprint`(`client.py:117`)改为:
```python
fingerprint = ",".join(sorted({s.model for s in sources}))
marks = sorted({json.dumps([s.model, dict(s.extra_body)], sort_keys=True, ensure_ascii=False)
for s in sources if s.extra_body})
if marks:
fingerprint += "|" + hashlib.sha256("".join(marks).encode()).hexdigest()
```
全源 `extra_body` 皆空时字面量与旧实现**逐字相同**。`dict(...)` 是因为 `MappingProxyType` 不能直接进 `json.dumps`。
**验收标准**: 配置 `temperature=0` + 调用级 `temperature=1` → payload 中为 1;结构化注入的 `response_format` 覆盖调用级同名键;保护键在 `chat()` 入口即 `ValueError`(未进洋葱,可用 mock handler 断言未被调用);全源无 `extra_body` 时 fingerprint 与旧值逐字相同;有 `extra_body` 时不同;改源 `name` 不改变 fingerprint。
**测试要求**: 覆盖设计 §5 测试 #3、#4(chat 侧)、#5、#6(拷贝语义:调用方在 `chat()` 返回后修改自己的 dict,不影响已构造的 request)、#8(不可 JSON 序列化的值在 `chat()` 入口即 `ValueError`,断言洋葱 handler 未被调用)。
**验证**: `pytest tests/unit/test_client.py tests/unit/test_openai_compat.py -v` → 全 PASS
---
### - [ ] Task 4: 缓存 key 纳入 `sampling`
**文件**: 改 `src/polygateway/middleware/cache.py`;测试 `tests/unit/test_cache.py`
**实现行为**:
- `build_cache_key` 增 keyword-only `sampling` 参数(见 §1 签名),非空时以 `"sampling"` 键并入 `key_obj`(**仅非空参与**,与 `salt` 的"仅非 None"不同——见设计决策 A 末段);
- `CacheMW.__call__` 传 `sampling=request.sampling`(**不是 `request.overlay`**——后者在此层虽尚未被结构化注入污染,但读 `sampling` 才是语义正确且不依赖层序巧合的写法)。
**验收标准**:
- 同 messages、不同 `seed` → 两个不同 key,第二次 miss(**issue 场景的直接回归**);
- 空 `sampling` 时 key 与旧实现**逐字相同**——测试须先把旧实现的 key 值固化为常量再比对(现有 `tests/unit/test_cache.py:39-54` 只有相等/不等断言,无 golden hash 可依);
- 同 `sampling` 不同键序 → 同一 key(canonical 序列化)。
**测试要求**: 覆盖设计 §5 测试 #1、#2。golden hash 的取法:在改动前先运行一次现有 `build_cache_key` 打印结果,写死进测试。
**验证**: `pytest tests/unit/test_cache.py -v` → 全 PASS
---
### - [ ] Task 5: 地基不变式回归(承重)
**文件**: 测试 `tests/unit/test_structured.py`(或就近的洋葱集成测试文件)
**实现行为**: 纯测试任务,不改产品代码。
落点:`tests/unit/test_structured.py` 里既有的 `ScriptedTerminal` 恰好站在 RetryMW 的位置(`client.py:91` 的 `terminal = RetryMW(...)`,StructuredMW 是最内中间件),扩写它即可,**无需搭全洋葱**。
断言:走结构化重问阶梯(强制至少重问一次,用先返回坏 JSON 再返回好 JSON 的 scripted terminal)后——
1. terminal 每次收到的 `request.sampling` 与**构造 `ChatRequest` 时传入的 `sampling`** 逐字相同;
2. 同一时刻 `request.overlay` **含** `response_format`(证明两者确实分叉,`sampling` 不是冗余字段)。
**为什么单列一个任务**: 决策 C 与 D 都建立在"`sampling` 跨层恒定"之上,而这条目前只靠"`dataclasses.replace` 恰好保留未提及字段"的约定成立,无任何机械执法。这条测试同时钉死决策 A 的"库内中间件永不修改"与决策 E 的只读约束。缺它则约束被破坏时无人发现。
**验收标准**: 该测试在故意把 `structured.py` 的 `replace` 改成重建 `ChatRequest`(丢掉 `sampling`)时**必须变红**——实施时须实际验证这一点,否则测试是空的。
**验证**: `pytest tests/unit/test_structured.py -v` → 全 PASS,且上述"故意破坏"实验红过一次
---
### - [ ] Task 6: 遥测端口扩至 21 字段 + 三入口口径
**文件**: 改 `src/polygateway/ports.py`、`src/polygateway/middleware/telemetry.py`;测试 `tests/unit/test_telemetry.py`
**实现行为**:
1. `ports.TelemetryRecorder.record_llm_call` 增第 21 参 `sampling: str | None`(排在 `model_reported` 之后)。
2. `TelemetryEmitter._record` 增同名参数并透传给 recorder。
3. 三个入口按设计决策 D 的表格产出(**不含**结构化注入的 `response_format`):
| 入口 | `sampling` 取值 |
|---|---|
| `emit_attempt` | `canonical_sampling_json(merge_sampling(source.extra_body, request.sampling))` |
| `emit_cache_hit` | `canonical_sampling_json(request.sampling)` |
| `emit_terminal_failure` | `canonical_sampling_json(request.sampling)` |
后两者无 `source` 可言(由最外层 TelemetryMW 调用),与 `model`/`provider`/`source_name` 在终态行置空是同一先例。
**关键约束**: `sampling` 必须由 emitter **内部推导**,**不得**作为新必填参数由调用者传入——否则 `ocr.py:418` 与 `embedding.py:372` 立刻 TypeError。
**验收标准**: 三个入口各自的 `sampling` 值符合上表;`response_format` **三行都不出现**;`request.sampling` 与 `source.extra_body` 皆空时为 `None`。
**测试要求**: 覆盖设计 §5 测试 #9。用 fake recorder 捕获 kwargs 断言。
**验证**: `pytest tests/unit/test_telemetry.py -v` → 全 PASS
---
### - [ ] Task 7: 两个遥测后端落列 + 幂等补列
**文件**: 改 `src/polygateway/telemetry/sqlite.py`、`src/polygateway/telemetry/postgres.py`;测试 `tests/unit/test_telemetry.py`、`tests/integration/test_postgres_telemetry.py`
**实现行为**(逐字沿用 issue #3 建立的套路):
- **sqlite.py**: DDL 在 `model_reported` **之后**加 `sampling TEXT`;`_BACKFILL_COLUMNS` 追加 `("sampling", "TEXT")`;`_COLUMNS` 末尾追加 `"sampling"`。
- **postgres.py**: DDL 同位置加 `sampling TEXT`;`_BACKFILL` 追加 `("sampling", "ALTER TABLE llm_calls ADD COLUMN sampling TEXT")`;`_COLUMNS` 末尾追加。
- 两处 `record_llm_call(**fields)` 按 `_COLUMNS` 取值,**无需改动**。
**硬约束**: 新列必须排在 `created_at` **之后**(两文件既有注释已说明理由:旧表只能 ALTER 追加到末尾,新建库若插在前面,两条路径物理列序分叉)。补列一律**先探测缺列再 ALTER**;失败只逐行降级,**绝不置结构性失能标志**(postgres 的 `_failed`)。
**连带必改**(不改则直接红):
| 位置 | 改什么 | 不改的后果 |
|---|---|---|
| `tests/unit/test_telemetry.py:78-102` 的 `_record_minimal()` | `fields` dict 加 `"sampling": None` | **两侧所有落库测试全红**:`sqlite.py:126` / `postgres.py:161` 的 `row = tuple(fields[col] for col in _COLUMNS)` **在 try 之外**,`_COLUMNS` 加列后抛裸 `KeyError: 'sampling'` 冒泡出 `record_llm_call` |
| `tests/integration/test_postgres_telemetry.py:88-109` 的 `_record_minimal()` | 同上 | 同上 |
| `tests/unit/test_telemetry.py:18` 的 `_EXPECTED_COLUMNS` | 追加 `"sampling"` | 列序断言红 |
| `tests/integration/test_postgres_telemetry.py:22` 的 `_EXPECTED_COLUMNS` | 追加(另见 `:210,231` 引用点) | 列序断言红 |
| `tests/unit/test_ports.py:95-119` 的 `_DummyRecorder.record_llm_call` | 显式 20 参签名同步为 21 | **不会红**(`runtime_checkable` 的 isinstance 只查方法存在不查签名),但会与端口脱节,顺带同步 |
| `sqlite.py:123` docstring、`test_telemetry.py:1` 文案 | "20 字段" → "21 字段" | 无功能影响,文案与事实脱节 |
**验收标准**: 新建库列序正确;对**已存在的 20 列旧表**能幂等补列且补后列序与新建库一致;重复初始化不报错;补列失败(模拟只有 INSERT 权限)时仅 warning、后续写入不被禁用。
**测试要求**: 覆盖设计 §5 测试 #11、#12。Postgres 部分是 integration,**须独占 PG `polygateway` 库时序,严禁并跑**。
**验证**:
- `pytest tests/unit/test_telemetry.py -v` → 全 PASS
- `pytest tests/integration/test_postgres_telemetry.py -v` → 全 PASS(确认无其他会话在用 PG)
---
### - [ ] Task 8: 决策 G —— OCR/embedding 构造期剥离 + warning
**文件**: 改 `src/polygateway/ocr.py`、`src/polygateway/embedding.py`;测试 `tests/unit/test_ocr_client.py`、`tests/unit/test_embedding.py`
**实现行为**: 两个 `__init__` 在既有校验块(`quota_full` 域校验附近)之后、`self._sources = list(sources)` 之前:
```python
stripped = []
for src in sources:
if src.extra_body:
logger.warning(
"{} 路径暂不支持 extra_body,源 {} 的该配置已被忽略"
"(需要 dimensions 等参数请提 issue): {}",
<"embedding"|"OCR">, src.name, dict(src.extra_body),
)
src = dataclasses.replace(src, extra_body={})
stripped.append(src)
self._sources = stripped
```
**剥离不是顺手清理,是承重的**: 不剥离则 Task 6 的 `merge_sampling(source.extra_body, ...)` 会让遥测**记录一个从未发出的参数**——`monkey_ocr.py:225,247` 只发 multipart `files=`(根本没有 JSON body),`openai_compat.py:343` 的 embed payload 硬编码 `{"model","input"}`。那是数据造假而非参数失效。替代方案(emitter 内特判调用方身份)违「遥测调用点收敛单一 helper」铁律,已否决。
**验收标准**: 带 `extra_body` 的源 → 装配**成功**(不抛异常)、记一条 warning、`client._sources` 上 `extra_body` 为空;该路径遥测 `sampling` 列为 `None`;不带 `extra_body` 时无 warning。
**测试要求**: 覆盖设计 §5 测试 #10。**后半段(遥测 `sampling` 为 None)是防遥测造假的真正断言,不可省**——只断言"装配成功 + 有 warning"是不够的。用 `caplog`/loguru 捕获断言 warning 存在。
**验证**: `pytest tests/unit/test_ocr_client.py tests/unit/test_embedding.py -v` → 全 PASS
---
### - [ ] Task 9: 决策 F —— 空 thinking profile 的后果注释
**文件**: 改 `src/polygateway/providers.py`
**实现行为**: 给 `openai`(`:46-51`)与 `minimax`(`:53-58`)两个 profile 各补一句**后果**说明:`enable_thinking=False` 对本 provider 不产生任何效果,需要关闭推理请用 `SourceConfig.extra_body`。
**注意**: `:52` 那条既有注释(「OpenAI 兼容基线,无已知注入差异」)在词法上属于紧随其后的 **minimax** 条目,`openai` 条目**没有**任何注释。补的是"后果"而非重复"为何为空"——不要写出与既有注释重复或矛盾的内容。
**验收标准**: 两个 profile 都能让读者明白 `enable_thinking=False` 对它们无效。纯注释变更,无行为变化。
**测试要求**: 无(纯注释)。此任务不单独提交,与 Task 10 合并提交。
**验证**: `make lint` → PASS
---
### - [ ] Task 10: 文档同步(设计 §6 清单)
**文件**: 改 `.env.example`、`README.md`、`CHANGELOG.md`、`research-wiki/ARCHITECTURE.md`
| 目标 | 具体改动 |
|---|---|
| `.env.example` | 在 `LLM__QWEN__1__TRUST_ENV` 注释行(`:20`)后加 `# LLM__QWEN__1__EXTRA_BODY={"temperature":0}` 及说明(JSON 对象串;保护键会报错;OCR/EMBED scope 会被忽略并 warning)。`client.py:251` docstring 声明本文件是键名清单事实源,漏写等于新键无处可查 |
| `README.md:83` | 该行逐一列举 `chat()` 关键字参数,补 `overlay` 及一句用途 |
| ARCH §5.2 | `chat()` 签名定稿段追加 `overlay` 要点(带默认值的 keyword-only,不破坏"调用点零改动"承诺) |
| ARCH §7.5 | key 公式补 `sampling` 项 + 两条已知副作用(seed 进 key 导致该路径必 miss;`model_fingerprint` 是集合级指纹,同 scope 各源 `extra_body` 不同时仍可能跨源命中) |
| ARCH §7.7(`:451`) | 该节逐字段枚举 `SourceConfig` 构成(`name/provider/.../enable_thinking`),补 `extra_body` |
| ARCH §7.8(`:463`) | 必录字段 20 → 21,补 `sampling` 及其列语义(不含 `response_format`) |
| ARCH §9(`:519-527`) | 配置面键族事实源,登记 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY` |
| `CHANGELOG.md` | 公共 API 新增(`chat(overlay=)`、`SourceConfig.extra_body`)+ 遥测端口扩列 |
**Gitea Wiki 同步**(`docs-convention.md` §2,CLAUDE.md §6 标为硬门)。本次同时命中该表两行:
| 命中行 | 必同步页 |
|---|---|
| 新公共 API / 新能力 | 对应指南页(新增「固定解码参数」内容,落在 `指南-遥测与成本` 或新页)+ `参考-公共API`(`chat()` 签名、`SourceConfig.extra_body`)+ `_Sidebar.md` + CHANGELOG |
| 新增配置键 | `参考-配置键`(登记 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`)+ 相关指南页的配置片段 + `.env.example` |
指南页必须写明三条坑:① `seed` 逐次变化时该路径缓存**必 miss**;② `model_fingerprint` 是集合级指纹,同 scope 各源 `extra_body` 不同时仍可能跨源命中(要逐源可复现需每源独享 scope 或 namespace);③ OCR/EMBED scope 的 `EXTRA_BODY` 会被忽略并 warning。
**验收标准**: 每条都能在文件中指到具体位置;ARCH 的改动与设计文档不矛盾;wiki 两行清单逐页落实。
**测试要求**: 无(纯文档)。与 Task 9 合并提交。
**验证**: `make lint` → PASS
---
### - [ ] Task 11: 全链路集成验证与合并前检查
**文件**: 测试 `tests/integration/`(就近文件或新增)
**实现行为**: 端到端断言采样参数经 `chat()` → 选源 → transport payload 到达请求体(设计 §5 测试 #13),用 fake HTTP 层捕获实际 payload。
**合并前门(逐条出示证据)**:
1. `make ci` → 全绿(`make lint` + `make test` + 覆盖率)
2. import-linter 契约无新违规(校验函数落最内层 `types.py`,分层关系不变)
3. 设计 §5 的 14 条测试全部有对应实现,逐条对应到具体测试函数名
4. 派**全新上下文** verifier subagent 独立验证(CLAUDE.md §3.2 里程碑级/跨多文件硬门)
**验证**:
- `make ci` → 全 PASS(**不要**在外面套 `conda run`:`Makefile` 每条 target 内部已是 `conda run -n PolyGateway ...`,嵌套会让内层输出被缓冲)
- verifier 报告无 blocking 问题
---
## 3. 提交节奏
| 提交 | 内容 |
|---|---|
| 1 | Task 1(types 内核) |
| 2 | Task 2(env 解析) |
| 3 | Task 3(chat 入口 + transport + fingerprint) |
| 4 | Task 4(缓存 key) |
| 5 | Task 5(地基不变式测试) |
| 6 | Task 6(遥测三入口) |
| 7 | Task 7(两后端落列) |
| 8 | Task 8(决策 G) |
| 9 | Task 9 + 10(注释与文档) |
| 10 | Task 11(集成验证,如有修补) |
每次提交调 `commit` skill。Task 1-4 是 issue 诉求的最小闭环;Task 5-8 是设计中"issue 未提但必须处理"的部分,**不可跳过**。
@@ -0,0 +1,30 @@
---
type: plan
node_id: plan:response-observability-fields
title: 响应可观测字段扩展实现计划
date: 2026-07-31
---
# 响应可观测字段扩展实现计划
全文见 `2026-07-31-response-observability-fields.md`。实现 [[response-observability-fields]] 设计(A2/B1/C1/D1)。
## 任务序列
| 任务 | 内容 | 提交 |
|---|---|---|
| T1 | `types.py` 两个类型各 +2 字段;`cache_hit` docstring 消歧 | 独立 |
| T2 | `openai_compat.py` 防御解析 + SSE sink 采集 `model` + 两处构造填值 | 独立 |
| T3 | `retry.py` 搬运;`cache.py` 零改动但用测试固化 B1 回放语义 | 独立 |
| T4 | `pricing.py` 可选缓存单价档 + 夹取防负 | 独立 |
| T5+T6 | 端口 18→20、两后端 DDL 加列与幂等补列、emitter 搬运、契约测试 | **必须合一次提交** |
| T7 | ARCHITECTURE §7.8 / CHANGELOG / `.env.example:56` / 6 处「18 字段」措辞 / 版本 1.1.0 / Gitea Wiki 站 | 独立 |
## 独立审查抓出的四个坑(已折回计划)
1. **DDL 新列必须放在 `created_at` 之后**(表末尾)。旧表走 `ALTER ADD COLUMN` 只能追加到末尾,若新建库把新列插在 `created_at` 前,两条路径列序分叉 —— 而 `test_schema_has_frozen_columns_in_order``ordinal_position` 逐位断言,且该 PG 表与真实批跑共享、严禁 DROP,分叉后无合规修法。
2. **SQLite 补列块首行必须守卫 `if self._conn is None: return`**。否则初始化失败时补列块抛 `AttributeError`/`NameError`(不被 `sqlite3.Error` 捕获)逃出 `__init__`,打破「初始化失败静默降级」契约。
3. **T5 与 T6 不得分开提交**。中间状态下 emitter 只传 18 键,后端抛 `KeyError` 被吞成 warning,该 commit 全量遥测静默丢失。
4. **天然拦截点是四处而非三处**:两个 `_record_minimal` + 两个 `_EXPECTED_COLUMNS`;`test_ports.py:96` 的全签名 fake **不会**红(Protocol 的 isinstance 不校验签名),不能当作覆盖保证。
相关: [[response-observability-fields]]、[[est-tokens-decoupling]]
@@ -0,0 +1,17 @@
---
type: plan
node_id: plan:sampling-params-plan
title: "采样参数透传实现计划(issue #4)"
date: 2026-07-31
---
# 采样参数透传实现计划(issue #4)
正文: `2026-07-31-sampling-params.md`。实现 `design:sampling-params`
- **任务数**: 11 个,每个一次提交、独立可验证。Task 1-4 是 issue 诉求的最小闭环;Task 5-8 是设计中「issue 未提但必须处理」的部分(地基不变式、遥测三入口、两后端落列、决策 G 剥离),不可跳过。
- **关键接口已在计划 §1 定死**: `validate_request_overlay()` / `merge_sampling()` / `canonical_sampling_json()` 三个纯函数落 `types.py`(最内层),`ChatRequest.sampling``SourceConfig.extra_body` 两个新字段,`build_cache_key()``chat()` 的新签名。
- **审查暴露的执行陷阱(已写进计划)**: ① 两个 `_record_minimal()` 的硬编码 20 键 fields dict 必须同步,否则 `row = tuple(fields[col] for col in _COLUMNS)`(在 try 之外)抛裸 `KeyError` 让两侧落库测试全红;② `SourceConfig` 加 mapping 字段后不再 hashable、`asdict`/`deepcopy` 失效——已核实库内无调用点会踩,作为已知后果显式接受并加锁定测试;③ 各文件需新增的 import 逐一列出(`types.py``from __future__ import annotations`,注解在类体求值);④ `validate_request_overlay` 的校验顺序必须先查 str 键再试序列化,否则非 str 键会被误报成"值不可序列化"。
- **测试证据门**: 每个任务合并前须出示先失败后通过的证据。Task 5 单列一条地基不变式回归——决策 C/D 都建立在「`sampling` 跨层恒定」之上,而这条目前只靠 `dataclasses.replace` 的约定,无机械执法;该测试须在故意破坏 `structured.py` 时验证过确实变红。
- **共享后端纪律**: Task 7、Task 11 涉及 PG `polygateway` 库,严禁与其他会话并跑(含 git 钩子触发的测试)。
- **审查留痕**: Codex CLI 不可用(vendor 二进制缺失),派全新上下文 subagent 只读审查。报 5 项必修(两处测试文件路径不存在、`_record_minimal` 漏项、Gitea Wiki 同步漏整块、`SourceConfig` 可哈希性后果未声明),逐条核实后全部采纳;5 条建议(import 清单、校验顺序、Task 5 落点表述、`make ci` 勿嵌套 `conda run``test_ports.py``_DummyRecorder` 同步)亦已收进。
+41 -2
View File
@@ -1,11 +1,11 @@
--- ---
type: schema type: schema
node_id: schema:llm-calls node_id: schema:llm-calls
title: "表结构: llm_calls(遥测 18 字段)" title: "表结构: llm_calls(遥测 21 字段)"
date: 2026-07-20 date: 2026-07-20
--- ---
# 表结构: llm_calls(遥测 18 字段) # 表结构: llm_calls(遥测 21 字段)
## 列定义(冻结,M1 设计 §4.4 / ARCH §7.8) ## 列定义(冻结,M1 设计 §4.4 / ARCH §7.8)
@@ -24,6 +24,9 @@ date: 2026-07-20
| error | TEXT | 异常信息;取消记 "cancelled" | | error | TEXT | 异常信息;取消记 "cancelled" |
| cost | REAL | M2 起 pricing 换算;`usage_source='unavailable'` 的真实调用行为 NULL(缓存命中行例外,仍为 0.0) | | cost | REAL | M2 起 pricing 换算;`usage_source='unavailable'` 的真实调用行为 NULL(缓存命中行例外,仍为 0.0) |
| created_at | TEXT NOT NULL DEFAULT (datetime('now')) | 落库时刻 | | created_at | TEXT NOT NULL DEFAULT (datetime('now')) | 落库时刻 |
| cached_prompt_tokens | INTEGER | 供应商 prompt cache 命中的输入 token(2026-07-31,issue #3);NULL = 该源未上报,`0` = 上报了真实零命中,两者不可混同 |
| model_reported | TEXT | API 响应体实际返回的 model;NULL = 未上报。与 `model`(配置别名)可能分叉 |
| sampling | TEXT | 本次调用的采样参数 canonical JSON(2026-07-31,issue #4);NULL = 未传。见下方口径 |
## usage/成本口径(2026-07-30,est_tokens 解耦) ## usage/成本口径(2026-07-30,est_tokens 解耦)
@@ -35,6 +38,42 @@ date: 2026-07-20
`SUM(cost)` 天然跳过 NULL,故账单汇总不再被虚构的估值污染;账目缺口的度量口径固定为 `WHERE usage_source = 'unavailable' AND cache_hit = false`。**`cache_hit` 限定不可省**:缓存命中行未产生新调用,cost 是事实上的 `0.0` 而非未知,本无账目缺口,漏掉该条件会让缺口度量偏高。 `SUM(cost)` 天然跳过 NULL,故账单汇总不再被虚构的估值污染;账目缺口的度量口径固定为 `WHERE usage_source = 'unavailable' AND cache_hit = false`。**`cache_hit` 限定不可省**:缓存命中行未产生新调用,cost 是事实上的 `0.0` 而非未知,本无账目缺口,漏掉该条件会让缺口度量偏高。
## 供应商 prompt cache 口径(2026-07-31,issue #3)
新增两列排在 `created_at` **之后**——旧表只能经 `ALTER TABLE ADD COLUMN` 追加到末尾,DDL 里若插在前面,新建库与升级库的物理列序会分叉(列序断言无合规修法)。两个后端在初始化期幂等补列:`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入被逐行 warning 丢弃、遥测静默全失。两侧都**先探测缺列再 ALTER**(`ADD COLUMN IF NOT EXISTS` 即使列已存在也先取 ACCESS EXCLUSIVE 锁,遥测是内联 await,锁共享审计表会拖垮业务调用),且**补列失败只降级为逐行丢弃,绝不让 recorder 整体失能**——两侧纪律必须对称。
`cache_hit`**PolyGateway 自身响应缓存**,与供应商 prompt cache 是两回事。缓存命中行的这两列是**原样回放**的历史值(与 `model`/`prompt_tokens` 同一口径),故命中率度量口径固定为:
```sql
SELECT SUM(cached_prompt_tokens)::float / NULLIF(SUM(prompt_tokens), 0)
FROM llm_calls WHERE cache_hit = false AND cached_prompt_tokens IS NOT NULL;
```
`WHERE cache_hit = false` 不可省,理由与上面 cost 缺口口径同源:回放行计入即重复计数。
## 采样参数口径(2026-07-31,issue #4)
`sampling` 列 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化输出注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。补列纪律与 issue #3 两列逐字相同(排在末尾、先探测再 ALTER、失败只逐行降级)。
三个 emit 入口的取值必须各自定死,否则同一列在不同行含义不同:
| 入口 | 调用者 | 有生效源? | 记什么 |
|---|---|---|---|
| `emit_attempt` | RetryMW(最内) | 有 | `merge(source.extra_body, request.sampling)` |
| `emit_cache_hit` | TelemetryMW(最外) | 无 | 仅 `request.sampling` |
| `emit_terminal_failure` | TelemetryMW | 无 | 仅 `request.sampling` |
后两行缺 `extra_body` 是客观事实而非口径瑕疵——它们没有"生效源"可言,与 `model`/`source_name` 在终态行置空是同一先例;缓存命中行亦无损:`sampling` 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同。三者统一读 `request.sampling` 而非 `request.overlay`(后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处未被污染,直接用必然三行分叉)。
OCR / embedding 路径的该列**恒为 NULL**:两条路径的 transport 不发 `extra_body`(embed payload 硬编码 `{model, input}`、MonkeyOCR 只发 multipart),故其源在构造期就被剥离——不剥离则该列会记录一个从未发出的参数,那是数据造假而非参数失效。
复现某批实验的解码条件:
```sql
SELECT DISTINCT sampling FROM llm_calls
WHERE session_id = $1 AND cache_hit = false AND error IS NULL;
```
## 埋点位置(单一 helper 铁律) ## 埋点位置(单一 helper 铁律)
- `middleware/telemetry.py::TelemetryEmitter` 是全库**唯一** `record_llm_call` 调用点; - `middleware/telemetry.py::TelemetryEmitter` 是全库**唯一** `record_llm_call` 调用点;
+1 -1
View File
@@ -31,7 +31,7 @@ from polygateway.types import (
SourceConfig, SourceConfig,
) )
__version__ = "1.0.3" __version__ = "1.0.5"
__all__ = [ __all__ = [
"DEFAULT_PROFILES", "DEFAULT_PROFILES",
+42 -5
View File
@@ -9,6 +9,8 @@
from __future__ import annotations from __future__ import annotations
import asyncio import asyncio
import hashlib
import json
import random import random
import time import time
from typing import TYPE_CHECKING, Any, Literal, TypeVar from typing import TYPE_CHECKING, Any, Literal, TypeVar
@@ -32,7 +34,7 @@ from polygateway.sources import (
SourceCooldownMemo, SourceCooldownMemo,
) )
from polygateway.transports.openai_compat import OpenAICompatTransport from polygateway.transports.openai_compat import OpenAICompatTransport
from polygateway.types import ChatRequest, LLMResponse from polygateway.types import ChatRequest, LLMResponse, validate_request_overlay
if TYPE_CHECKING: if TYPE_CHECKING:
from collections.abc import Awaitable, Iterable, Mapping from collections.abc import Awaitable, Iterable, Mapping
@@ -59,6 +61,29 @@ if TYPE_CHECKING:
_T = TypeVar("_T") _T = TypeVar("_T")
def build_model_fingerprint(sources: Iterable[SourceConfig]) -> str:
"""缓存 key 的模型身份: 多源 scope = 排序去重的 model 合集。
配置级采样参数(`extra_body`)必须参与,否则把 temperature 从 0 改成 1
后重启仍会读到旧缓存(issue #4 设计决策 C)。全源 `extra_body` 皆空时
字面量与历史实现逐字相同,不触发存量缓存冷启动。
"""
fingerprint = ",".join(sorted({s.model for s in sources}))
# 按 (model, extra_body) 而非源名摘要: 语义是"本 scope 会用哪些
# (模型, 解码参数)组合",改源名不该误触全量冷启动
marks = sorted(
{
json.dumps([s.model, dict(s.extra_body)], sort_keys=True, ensure_ascii=False)
for s in sources
if s.extra_body
}
)
if marks:
digest = hashlib.sha256("".join(marks).encode("utf-8")).hexdigest()
fingerprint = f"{fingerprint}|{digest}"
return fingerprint
class GatewayClient: class GatewayClient:
"""统一治理入口;构造函数全量注入(测试/高级),工厂覆盖 90% 场景。""" """统一治理入口;构造函数全量注入(测试/高级),工厂覆盖 90% 场景。"""
@@ -113,12 +138,11 @@ class GatewayClient:
if cache is not None: if cache is not None:
if cache_namespace is None or cache_ttl_s is None: if cache_namespace is None or cache_ttl_s is None:
raise ValueError("启用缓存必须提供 cache_namespace 与 cache_ttl_s") raise ValueError("启用缓存必须提供 cache_namespace 与 cache_ttl_s")
# 多源 scope 的 key 身份 = 排序去重的 model 合集;源集合变化 → 一次性冷启动 # 多源 scope 的 key 身份;源集合或其 extra_body 变化 → 一次性冷启动
fingerprint = ",".join(sorted({s.model for s in sources}))
middlewares.append( middlewares.append(
CacheMW( CacheMW(
backend=cache, backend=cache,
model_fingerprint=fingerprint, model_fingerprint=build_model_fingerprint(sources),
default_namespace=cache_namespace, default_namespace=cache_namespace,
ttl_s=cache_ttl_s, ttl_s=cache_ttl_s,
strategy=structured_strategy, strategy=structured_strategy,
@@ -150,12 +174,23 @@ class GatewayClient:
cache_namespace: str | None = None, cache_namespace: str | None = None,
structured: type[BaseModel] | Literal["json"] | None = None, structured: type[BaseModel] | Literal["json"] | None = None,
stream: bool = True, stream: bool = True,
overlay: Mapping[str, Any] | None = None,
) -> LLMResponse: ) -> LLMResponse:
"""一次治理调用(签名冻结,ARCH §5.2;与三项目 LLMProvider 协议兼容)。""" """一次治理调用(签名冻结,ARCH §5.2;与三项目 LLMProvider 协议兼容)。
`overlay` 是采样参数覆盖层(`temperature`/`seed`/`max_tokens` 等),优先级
高于源级 `extra_body`、低于结构化输出的注入。带默认值的 keyword-only
参数不影响既有调用点(issue #4)。
"""
if structured is not None and not self._structured_available: if structured is not None and not self._structured_available:
raise ImportError( raise ImportError(
"结构化输出未启用: 安装 pip install 'polygateway[structured]' 后重新装配" "结构化输出未启用: 安装 pip install 'polygateway[structured]' 后重新装配"
) )
# 进洋葱之前校验并拷贝: 保护键/不可序列化值在此收口(否则会在 CacheMW
# 的降级 try 之外抛裸 TypeError);拷贝防调用方复用同一 dict 逐次改 seed
# 造成的竞态。同一份快照填 overlay 与 sampling——前者会被结构化注入,
# 后者跨层恒定,供缓存 key 与遥测读取(设计决策 A/B/E)
sampling = validate_request_overlay(overlay or {}, origin="chat(overlay=...)")
request = ChatRequest( request = ChatRequest(
messages=messages, messages=messages,
session_id=session_id, session_id=session_id,
@@ -164,6 +199,8 @@ class GatewayClient:
cache_namespace=cache_namespace, cache_namespace=cache_namespace,
structured=structured, structured=structured,
stream=stream, stream=stream,
overlay=sampling,
sampling=sampling,
) )
return await self._handler(request) return await self._handler(request)
+8
View File
@@ -11,6 +11,7 @@ fail-loud 校验语义与 pydantic-settings 一致。
from __future__ import annotations from __future__ import annotations
import json
import os import os
from dataclasses import dataclass from dataclasses import dataclass
from typing import TYPE_CHECKING from typing import TYPE_CHECKING
@@ -44,6 +45,7 @@ _SOURCE_FIELDS: dict[str, tuple[str, str]] = {
"ENABLE_THINKING": ("enable_thinking", "bool"), "ENABLE_THINKING": ("enable_thinking", "bool"),
"MISSING_DONE": ("missing_done", "str"), "MISSING_DONE": ("missing_done", "str"),
"TRUST_ENV": ("trust_env", "bool"), "TRUST_ENV": ("trust_env", "bool"),
"EXTRA_BODY": ("extra_body", "json"),
} }
_RESERVED_SEGMENTS = frozenset({"GLOBAL", "RETRY", "BREAKER", "BACKPRESSURE"}) _RESERVED_SEGMENTS = frozenset({"GLOBAL", "RETRY", "BREAKER", "BACKPRESSURE"})
_SELECTORS = frozenset({"round_robin", "least_inflight", "health_aware"}) _SELECTORS = frozenset({"round_robin", "least_inflight", "health_aware"})
@@ -74,6 +76,12 @@ def _cast(raw: str, kind: str, key: str) -> object:
if lowered in ("0", "false", "no", "off"): if lowered in ("0", "false", "no", "off"):
return False return False
raise ValueError(f"非法布尔值: {raw!r}") raise ValueError(f"非法布尔值: {raw!r}")
if kind == "json":
# JSONDecodeError 是 ValueError 子类,复用下方的统一包装
parsed = json.loads(raw)
if not isinstance(parsed, dict):
raise ValueError(f"必须是 JSON 对象(而非数组/标量): {raw!r}")
return parsed
return raw return raw
except ValueError as exc: except ValueError as exc:
raise ValueError(f"配置 {key} 解析失败: {exc}") from exc raise ValueError(f"配置 {key} 解析失败: {exc}") from exc
+9 -2
View File
@@ -41,7 +41,12 @@ from polygateway.middleware.ratelimit import QuotaGate
from polygateway.middleware.retry import _failure_reason, backoff_delay from polygateway.middleware.retry import _failure_reason, backoff_delay
from polygateway.middleware.telemetry import TelemetryEmitter from polygateway.middleware.telemetry import TelemetryEmitter
from polygateway.sources import SourceCooldownMemo from polygateway.sources import SourceCooldownMemo
from polygateway.types import ChatRequest, EmbeddingResponse, LLMResponse from polygateway.types import (
ChatRequest,
EmbeddingResponse,
LLMResponse,
strip_unsupported_extra_body,
)
if TYPE_CHECKING: if TYPE_CHECKING:
from collections.abc import Awaitable, Callable, Mapping from collections.abc import Awaitable, Callable, Mapping
@@ -111,7 +116,9 @@ class EmbeddingClient:
if expected_dim is not None and expected_dim < 1: if expected_dim is not None and expected_dim < 1:
raise ValueError("expected_dim 必须 ≥ 1") raise ValueError("expected_dim 必须 ≥ 1")
self._scope = scope self._scope = scope
self._sources = list(sources) # embed payload 硬编码 {model, input},带 extra_body 的源必须先剥离,
# 否则遥测会记录一个从未发出的采样参数(issue #4 决策 G)
self._sources = strip_unsupported_extra_body(list(sources), path="embedding")
self._selector = selector self._selector = selector
self._quota = QuotaGate(limiter) self._quota = QuotaGate(limiter)
self._breaker = BreakerGate(breaker) self._breaker = BreakerGate(breaker)
+25 -3
View File
@@ -20,6 +20,8 @@ from loguru import logger
from polygateway.types import ChatRequest, LLMResponse from polygateway.types import ChatRequest, LLMResponse
if TYPE_CHECKING: if TYPE_CHECKING:
from collections.abc import Mapping
from polygateway.ports import CacheBackend, CallNext, StructuredOutputStrategy from polygateway.ports import CacheBackend, CallNext, StructuredOutputStrategy
_KEY_PREFIX = "pgw:cache:" _KEY_PREFIX = "pgw:cache:"
@@ -50,9 +52,19 @@ def _digest_part(part: Any) -> Any:
def build_cache_key( def build_cache_key(
model_fingerprint: str, messages: list[dict[str, Any]], namespace: str, salt: str | None model_fingerprint: str,
messages: list[dict[str, Any]],
namespace: str,
salt: str | None,
*,
sampling: Mapping[str, Any] | None = None,
) -> str: ) -> str:
"""缓存 key 公式;salt 仅非 None 时参与(VT 旧键语义: 不传 salt 键形不变)。""" """缓存 key 公式;salt 仅非 None 时参与(VT 旧键语义: 不传 salt 键形不变)。
`sampling` 仅**非空**时参与(与 salt 的"仅非 None"不同——空串是有意义的
salt,而空采样参数与不传无语义差别)。它必须进 key: 否则同 messages 跑 5 个
seed 会全部命中第一次的响应,标准差恒为 0 且不报错(issue #4 决策 C)。
"""
key_obj: dict[str, Any] = { key_obj: dict[str, Any] = {
"model": model_fingerprint, "model": model_fingerprint,
"messages": digest_messages(messages), "messages": digest_messages(messages),
@@ -60,6 +72,8 @@ def build_cache_key(
} }
if salt is not None: if salt is not None:
key_obj["salt"] = salt key_obj["salt"] = salt
if sampling:
key_obj["sampling"] = dict(sampling)
payload = json.dumps(key_obj, sort_keys=True, ensure_ascii=False) payload = json.dumps(key_obj, sort_keys=True, ensure_ascii=False)
return _KEY_PREFIX + hashlib.sha256(payload.encode("utf-8")).hexdigest() return _KEY_PREFIX + hashlib.sha256(payload.encode("utf-8")).hexdigest()
@@ -92,7 +106,15 @@ class CacheMW:
async def __call__(self, request: ChatRequest, call_next: CallNext) -> LLMResponse: async def __call__(self, request: ChatRequest, call_next: CallNext) -> LLMResponse:
namespace = request.cache_namespace or self._namespace namespace = request.cache_namespace or self._namespace
key = build_cache_key(self._fingerprint, request.messages, namespace, request.cache_salt) # 读 sampling 而非 overlay: 语义明确,且不依赖"CacheMW 恰在 StructuredMW
# 外侧"这一层序巧合——结构化注入不该改变缓存身份(设计决策 C)
key = build_cache_key(
self._fingerprint,
request.messages,
namespace,
request.cache_salt,
sampling=request.sampling,
)
cached = await self._safe_get(key) cached = await self._safe_get(key)
if cached is not None: if cached is not None:
hit = self._rehydrate(cached, request) hit = self._rehydrate(cached, request)
+2
View File
@@ -436,6 +436,8 @@ class RetryMW:
source_name=source.name, source_name=source.name,
cost=None, cost=None,
usage_source=result.usage_source, usage_source=result.usage_source,
cached_prompt_tokens=result.cached_prompt_tokens,
model_reported=result.model_reported,
) )
async def _settle_and_release(self, permit: Permit, actual: int) -> None: async def _settle_and_release(self, permit: Permit, actual: int) -> None:
+26 -2
View File
@@ -18,6 +18,7 @@ from loguru import logger
from polygateway.errors import GatewayUnavailableError, GovernanceBackendError from polygateway.errors import GatewayUnavailableError, GovernanceBackendError
from polygateway.middleware.cache import digest_messages from polygateway.middleware.cache import digest_messages
from polygateway.types import canonical_sampling_json, merge_sampling
if TYPE_CHECKING: if TYPE_CHECKING:
from collections.abc import Callable from collections.abc import Callable
@@ -28,7 +29,7 @@ if TYPE_CHECKING:
class TelemetryEmitter: class TelemetryEmitter:
"""从请求与结果组装 18 字段并写入 recorder;一切写失败降级 warning。""" """从请求与结果组装 21 字段并写入 recorder;一切写失败降级 warning。"""
def __init__(self, recorder: TelemetryRecorder, *, pricing: PricingTable | None = None) -> None: def __init__(self, recorder: TelemetryRecorder, *, pricing: PricingTable | None = None) -> None:
self._recorder = recorder self._recorder = recorder
@@ -61,6 +62,10 @@ class TelemetryEmitter:
max_inter_token_ms=response.max_inter_token_ms if response else None, max_inter_token_ms=response.max_inter_token_ms if response else None,
cache_hit=False, cache_hit=False,
error=error, error=error,
cached_prompt_tokens=response.cached_prompt_tokens if response else None,
model_reported=response.model_reported if response else None,
# 唯一有"生效源"的入口,故是唯一能并上 extra_body 的(设计决策 D)
sampling=canonical_sampling_json(merge_sampling(source.extra_body, request.sampling)),
) )
async def emit_cache_hit(self, *, request: ChatRequest, response: LLMResponse) -> None: async def emit_cache_hit(self, *, request: ChatRequest, response: LLMResponse) -> None:
@@ -81,6 +86,13 @@ class TelemetryEmitter:
max_inter_token_ms=None, max_inter_token_ms=None,
cache_hit=True, cache_hit=True,
error=None, error=None,
# 决策 B1: 与 model/prompt_tokens 同一口径,原样回放历史值。
# 统计供应商缓存命中率必须带 WHERE cache_hit = false,否则重复计数。
cached_prompt_tokens=response.cached_prompt_tokens,
model_reported=response.model_reported,
# 由最外层 TelemetryMW 调用,手上没有 source。缓存命中行无损:
# sampling 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同
sampling=canonical_sampling_json(request.sampling),
) )
async def emit_terminal_failure( async def emit_terminal_failure(
@@ -103,6 +115,10 @@ class TelemetryEmitter:
max_inter_token_ms=None, max_inter_token_ms=None,
cache_hit=False, cache_hit=False,
error=error, error=error,
cached_prompt_tokens=None,
model_reported=None,
# 无具体源,与 model/provider/source_name 置空同一先例(设计决策 D)
sampling=canonical_sampling_json(request.sampling),
) )
async def _record( async def _record(
@@ -123,6 +139,9 @@ class TelemetryEmitter:
max_inter_token_ms: float | None, max_inter_token_ms: float | None,
cache_hit: bool, cache_hit: bool,
error: str | None, error: str | None,
cached_prompt_tokens: int | None,
model_reported: str | None,
sampling: str | None,
) -> None: ) -> None:
try: try:
# 成本换算(M2 §6): 成功行按单价换算;缓存命中 0.0(未产生新调用); # 成本换算(M2 §6): 成功行按单价换算;缓存命中 0.0(未产生新调用);
@@ -134,7 +153,9 @@ class TelemetryEmitter:
# 必须排在 cache_hit 之后——缓存命中未产生新调用,0.0 是事实而非未知 # 必须排在 cache_hit 之后——缓存命中未产生新调用,0.0 是事实而非未知
cost = None cost = None
elif error is None and model and self._pricing is not None: elif error is None and model and self._pricing is not None:
cost = self._pricing.cost(model, prompt_tokens, completion_tokens) cost = self._pricing.cost(
model, prompt_tokens, completion_tokens, cached_prompt_tokens
)
else: else:
cost = None cost = None
# messages 落库前多模态摘要,与缓存 key 共用同一函数(VT R12) # messages 落库前多模态摘要,与缓存 key 共用同一函数(VT R12)
@@ -158,6 +179,9 @@ class TelemetryEmitter:
cache_hit=cache_hit, cache_hit=cache_hit,
error=error, error=error,
cost=cost, cost=cost,
cached_prompt_tokens=cached_prompt_tokens,
model_reported=model_reported,
sampling=sampling,
) )
except asyncio.CancelledError: except asyncio.CancelledError:
raise raise
+4 -1
View File
@@ -44,6 +44,7 @@ from polygateway.types import (
OcrLayoutResult, OcrLayoutResult,
OcrTextResult, OcrTextResult,
Usage, Usage,
strip_unsupported_extra_body,
) )
if TYPE_CHECKING: if TYPE_CHECKING:
@@ -113,7 +114,9 @@ class OcrClient:
if quota_full not in ("wait", "fail_fast"): if quota_full not in ("wait", "fail_fast"):
raise ValueError(f"quota_full 必须是 wait|fail_fast: {quota_full!r}") raise ValueError(f"quota_full 必须是 wait|fail_fast: {quota_full!r}")
self._scope = scope self._scope = scope
self._sources = list(sources) # MonkeyOCR 只发 multipart 表单,带 extra_body 的源必须先剥离,否则
# 遥测会记录一个从未发出的采样参数(issue #4 决策 G)
self._sources = strip_unsupported_extra_body(list(sources), path="OCR")
self._selector = selector self._selector = selector
self._feed_health = isinstance(selector, OutcomeAwareSelector) self._feed_health = isinstance(selector, OutcomeAwareSelector)
self._quota = QuotaGate(limiter) self._quota = QuotaGate(limiter)
+8 -1
View File
@@ -245,7 +245,11 @@ class StructuredOutputStrategy(Protocol):
@runtime_checkable @runtime_checkable
class TelemetryRecorder(Protocol): class TelemetryRecorder(Protocol):
"""遥测后端;18 字段冻结(M1 设计 §4.4),唯一调用点是 TelemetryEmitter。""" """遥测后端;20 字段冻结(M1 设计 §4.4 + issue #3),唯一调用点是 TelemetryEmitter。
新增参数不设默认值: 库外无第三方实现者(三项目迁移时删除了各自的同名
Protocol),完整签名的成本为零,而少写一列会被 emitter 的降级吞成 warning。
"""
async def record_llm_call( async def record_llm_call(
self, self,
@@ -268,4 +272,7 @@ class TelemetryRecorder(Protocol):
cache_hit: bool, cache_hit: bool,
error: str | None, error: str | None,
cost: float | None, cost: float | None,
cached_prompt_tokens: int | None,
model_reported: str | None,
sampling: str | None,
) -> None: ... ) -> None: ...
+60 -9
View File
@@ -3,7 +3,7 @@
**零内置单价**: 实验室走中转网关,计费非官方牌价;库内硬编码单价表 **零内置单价**: 实验室走中转网关,计费非官方牌价;库内硬编码单价表
必然过时并掩盖真实成本(P5 严禁默认值掩盖错误)。价格一律由使用方 必然过时并掩盖真实成本(P5 严禁默认值掩盖错误)。价格一律由使用方
提供——JSON 文件(`PGW_PRICING_PATH`)或 dict 注入;币种由使用方全表 提供——JSON 文件(`PGW_PRICING_PATH`)或 dict 注入;币种由使用方全表
统一口径,库不设币种字段(18 字段冻结)。查不到的 model → cost=None 统一口径,库不设币种字段(20 字段冻结)。查不到的 model → cost=None
且每 model 仅首次 warning(防日志风暴),不阻塞调用。 且每 model 仅首次 warning(防日志风暴),不阻塞调用。
""" """
@@ -22,22 +22,33 @@ if TYPE_CHECKING:
@dataclass(frozen=True) @dataclass(frozen=True)
class ModelPrice: class ModelPrice:
"""每百万 token 的输入/输出单价(币种由使用方口径统一)。""" """每百万 token 的输入/输出单价(币种由使用方口径统一)。
`cached_input_per_1m` 是可选的**缓存读取单价**(issue #3): 供应商 prompt
cache 命中的那部分输入按更低单价计费。不填即不启用——库绝不按经验折扣率
猜一个数(P5 严禁默认值掩盖),未填时全额按 `input_per_1m` 计。
"""
input_per_1m: float input_per_1m: float
output_per_1m: float output_per_1m: float
cached_input_per_1m: float | None = None
def __post_init__(self) -> None: def __post_init__(self) -> None:
if self.input_per_1m < 0 or self.output_per_1m < 0: if self.input_per_1m < 0 or self.output_per_1m < 0:
raise ValueError("单价不能为负") raise ValueError("单价不能为负")
if self.cached_input_per_1m is not None and self.cached_input_per_1m < 0:
raise ValueError("缓存读取单价不能为负")
class PricingTable: class PricingTable:
"""model → 单价 的只读表;cost() 是全库唯一换算点(经 TelemetryEmitter)。""" """model → 单价 的只读表;cost() 有两个调用点: `TelemetryEmitter`(chat 主路径)
与 `embedding.py` 的批量换算。"""
def __init__(self, prices: Mapping[str, ModelPrice]) -> None: def __init__(self, prices: Mapping[str, ModelPrice]) -> None:
self._prices = dict(prices) self._prices = dict(prices)
self._warned: set[str] = set() self._warned: set[str] = set()
# 独立集合: 与"未知 model"的告警去重键分开,避免 model 名恰好撞上时互相抑制
self._warned_clamp: set[str] = set()
@classmethod @classmethod
def from_file(cls, path: Path | str) -> PricingTable: def from_file(cls, path: Path | str) -> PricingTable:
@@ -53,21 +64,61 @@ class PricingTable:
for model, entry in data.items(): for model, entry in data.items():
if not isinstance(entry, dict) or not {"input_per_1m", "output_per_1m"} <= set(entry): if not isinstance(entry, dict) or not {"input_per_1m", "output_per_1m"} <= set(entry):
raise ValueError(f"价格表 {p} 条目 {model!r} 须含 input_per_1m 与 output_per_1m") raise ValueError(f"价格表 {p} 条目 {model!r} 须含 input_per_1m 与 output_per_1m")
cached_raw = entry.get("cached_input_per_1m")
try:
cached = None if cached_raw is None else float(cached_raw)
except (TypeError, ValueError) as exc:
raise ValueError(
f"价格表 {p} 条目 {model!r} 的 cached_input_per_1m 必须是数字: {cached_raw!r}"
) from exc
if cached is not None and cached < 0:
raise ValueError(f"价格表 {p} 条目 {model!r} 的 cached_input_per_1m 不能为负")
prices[model] = ModelPrice( prices[model] = ModelPrice(
input_per_1m=float(entry["input_per_1m"]), input_per_1m=float(entry["input_per_1m"]),
output_per_1m=float(entry["output_per_1m"]), output_per_1m=float(entry["output_per_1m"]),
cached_input_per_1m=cached,
) )
return cls(prices) return cls(prices)
def cost(self, model: str, prompt_tokens: int, completion_tokens: int) -> float | None: def cost(
"""换算一次调用成本;未知 model 记 None 并仅首次 warning。""" self,
model: str,
prompt_tokens: int,
completion_tokens: int,
cached_prompt_tokens: int | None = None,
) -> float | None:
"""换算一次调用成本;未知 model 记 None 并仅首次 warning。
`cached_prompt_tokens` 是供应商 prompt cache 命中的输入 token 数
(issue #3);仅当该 model 配了 `cached_input_per_1m` 时才分段计价,
否则全额按输入价——不猜折扣率。参数带默认值: embedding 侧的三参调用
形态不受影响。
"""
price = self._prices.get(model) price = self._prices.get(model)
if price is None: if price is None:
if model not in self._warned: if model not in self._warned:
self._warned.add(model) self._warned.add(model)
logger.warning("pricing 表无 model {!r} 的单价,cost 记 None", model) logger.warning("pricing 表无 model {!r} 的单价,cost 记 None", model)
return None return None
return ( billed_input = prompt_tokens / 1_000_000 * price.input_per_1m
prompt_tokens / 1_000_000 * price.input_per_1m # 负数按"无命中"处理: cost() 是公共方法,不能假定调用方已过 transport 的校验
+ completion_tokens / 1_000_000 * price.output_per_1m if price.cached_input_per_1m is not None and (cached_prompt_tokens or 0) > 0:
) cached = self._clamp_cached(model, prompt_tokens, cached_prompt_tokens)
billed_input = (prompt_tokens - cached) / 1_000_000 * price.input_per_1m + (
cached / 1_000_000 * price.cached_input_per_1m
)
return billed_input + completion_tokens / 1_000_000 * price.output_per_1m
def _clamp_cached(self, model: str, prompt_tokens: int, cached: int) -> int:
"""命中数按输入总数夹取: 网关口径异常不得算出负成本(每 model 只警告一次)。"""
if cached <= prompt_tokens:
return cached
if model not in self._warned_clamp:
self._warned_clamp.add(model)
logger.warning(
"model {!r} 上报的缓存命中 {} 超过输入总数 {},按总数夹取计价",
model,
cached,
prompt_tokens,
)
return prompt_tokens
+8 -1
View File
@@ -19,6 +19,10 @@ class ProviderProfile:
True/False 时并入请求体的参数片段(None 时二者都不注入,用模型默认); True/False 时并入请求体的参数片段(None 时二者都不注入,用模型默认);
strip_think_tags 声明响应 content 需剥离 ``<think>`` 标签(qwen 系); strip_think_tags 声明响应 content 需剥离 ``<think>`` 标签(qwen 系);
supports_native_schema 供 D14 阶梯选择原生 response_format 策略。 supports_native_schema 供 D14 阶梯选择原生 response_format 策略。
注: 某个 provider 的两档若皆为空字典(如 openai/minimax),说明该 provider
无已知的推理开关参数——此时 `enable_thinking` 对它**不产生任何效果**,
而非静默生效。需要下发自定义参数时用 `SourceConfig.extra_body`。
""" """
name: str name: str
@@ -43,13 +47,16 @@ DEFAULT_PROFILES: Mapping[str, ProviderProfile] = MappingProxyType(
thinking_off={"thinking": {"type": "disabled"}}, thinking_off={"thinking": {"type": "disabled"}},
strip_think_tags=False, strip_think_tags=False,
), ),
# 两档皆空 ⇒ `enable_thinking` 对本 provider **不产生任何效果**(调用方
# 以为关掉了实际没关)。真需要控制推理时经 `SourceConfig.extra_body` 下发
"openai": ProviderProfile( "openai": ProviderProfile(
name="openai", name="openai",
thinking_on={}, thinking_on={},
thinking_off={}, thinking_off={},
strip_think_tags=False, strip_think_tags=False,
), ),
# OpenAI 兼容基线,无已知注入差异;reasoning_content 由 transport 通用处理 # OpenAI 兼容基线,无已知注入差异;reasoning_content 由 transport 通用处理
# 同上: 两档皆空 ⇒ `enable_thinking` 对 MiniMax 源不产生任何效果
"minimax": ProviderProfile( "minimax": ProviderProfile(
name="minimax", name="minimax",
thinking_on={}, thinking_on={},
+45 -2
View File
@@ -6,7 +6,7 @@
① 结构性失败(建池/建表)→ warning 一次后永久降级(池置 None 短路); ① 结构性失败(建池/建表)→ warning 一次后永久降级(池置 None 短路);
② 运行时单条写失败 → 逐条 warning 丢弃,不降级不重试(连接抖动由 ② 运行时单条写失败 → 逐条 warning 丢弃,不降级不重试(连接抖动由
asyncpg 池自恢复;避免浸泡开头一次抖动导致后续全程失遥测)。 asyncpg 池自恢复;避免浸泡开头一次抖动导致后续全程失遥测)。
构造不连库(lazy),18 列 schema 与 SQLite 版同名同序。 构造不连库(lazy),20 列 schema 与 SQLite 版同名同序。
""" """
from __future__ import annotations from __future__ import annotations
@@ -39,10 +39,26 @@ CREATE TABLE IF NOT EXISTS llm_calls (
cache_hit BOOLEAN NOT NULL DEFAULT FALSE, cache_hit BOOLEAN NOT NULL DEFAULT FALSE,
error TEXT, error TEXT,
cost DOUBLE PRECISION, cost DOUBLE PRECISION,
created_at TIMESTAMPTZ NOT NULL DEFAULT now() created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
cached_prompt_tokens INTEGER,
model_reported TEXT,
sampling TEXT
); );
""" """
# 新列排在 created_at 之后: 与旧表 ALTER 追加的位置一致(见 sqlite.py 同款注释)
_BACKFILL = (
("cached_prompt_tokens", "ALTER TABLE llm_calls ADD COLUMN cached_prompt_tokens INTEGER"),
("model_reported", "ALTER TABLE llm_calls ADD COLUMN model_reported TEXT"),
("sampling", "ALTER TABLE llm_calls ADD COLUMN sampling TEXT"),
)
# 探测现有列;尊重 search_path(to_regclass 按当前 search_path 解析)
_EXISTING_COLUMNS = (
"SELECT attname FROM pg_attribute "
"WHERE attrelid = to_regclass('llm_calls') AND attnum > 0 AND NOT attisdropped"
)
_COLUMNS = ( _COLUMNS = (
"call_id", "call_id",
"parent_call_id", "parent_call_id",
@@ -62,6 +78,9 @@ _COLUMNS = (
"cache_hit", "cache_hit",
"error", "error",
"cost", "cost",
"cached_prompt_tokens",
"model_reported",
"sampling",
) )
_INSERT = ( _INSERT = (
@@ -104,6 +123,7 @@ class PostgresRecorder:
self._pool = await asyncpg.create_pool(self._dsn, timeout=10) self._pool = await asyncpg.create_pool(self._dsn, timeout=10)
async with self._pool.acquire() as conn: async with self._pool.acquire() as conn:
await conn.execute(_DDL) await conn.execute(_DDL)
await self._backfill_columns(conn)
self._schema_ready = True self._schema_ready = True
return self._pool return self._pool
except asyncio.CancelledError: except asyncio.CancelledError:
@@ -113,6 +133,29 @@ class PostgresRecorder:
logger.warning("Postgres 遥测初始化失败,后续记录降级为 no-op: {}", exc) logger.warning("Postgres 遥测初始化失败,后续记录降级为 no-op: {}", exc)
return None return None
async def _backfill_columns(self, conn: object) -> None:
"""给已存在的旧表补新列(issue #3);**先探测再 ALTER,失败绝不置 `_failed`**。
两条纪律各有实测理由:
① 不置 `_failed`: 应用账号只有 INSERT 权限时,`ALTER TABLE` 的 ownership
检查早于 `IF NOT EXISTS` 的存在性判断——列明明齐全也会失败。置位会让
整个 recorder 永久 no-op,与「补列失败只降级为逐行丢弃」的承诺相悖
(SQLite 侧同款守卫,两侧必须对称)。
② 先探测: `ADD COLUMN IF NOT EXISTS` 即便列已存在,也会**先取 ACCESS
EXCLUSIVE 锁**再判存在性(实测会被一个开着的读事务阻塞)。遥测是内联
await,让每个进程的首次写入都去抢共享审计表的排他锁,等于用记录基础设施
拖垮业务调用。探测走 ACCESS SHARE,稳态下一条 ALTER 都不会发。
"""
try:
existing = {row["attname"] for row in await conn.fetch(_EXISTING_COLUMNS)} # type: ignore[attr-defined]
for column, statement in _BACKFILL:
if column not in existing:
await conn.execute(statement) # type: ignore[attr-defined]
except asyncio.CancelledError:
raise
except Exception as exc:
logger.warning("Postgres 遥测补列失败(写入将逐行降级): {}", exc)
async def record_llm_call(self, **fields: object) -> None: async def record_llm_call(self, **fields: object) -> None:
"""写一行遥测;单条失败逐条 warning 丢弃(两级降级之二),绝不冒泡。""" """写一行遥测;单条失败逐条 warning 丢弃(两级降级之二),绝不冒泡。"""
pool = await self._ensure_ready() pool = await self._ensure_ready()
+44 -2
View File
@@ -34,10 +34,21 @@ CREATE TABLE IF NOT EXISTS llm_calls (
cache_hit INTEGER NOT NULL DEFAULT 0, cache_hit INTEGER NOT NULL DEFAULT 0,
error TEXT, error TEXT,
cost REAL, cost REAL,
created_at TEXT NOT NULL DEFAULT (datetime('now')) created_at TEXT NOT NULL DEFAULT (datetime('now')),
cached_prompt_tokens INTEGER,
model_reported TEXT,
sampling TEXT
); );
""" """
# 新列必须排在 created_at 之后: 旧表只能经 ALTER 追加到末尾,新建库若把它们
# 插在前面,两条路径的物理列序会分叉(列序断言测试无合规修法)。
_BACKFILL_COLUMNS = (
("cached_prompt_tokens", "INTEGER"),
("model_reported", "TEXT"),
("sampling", "TEXT"),
)
_COLUMNS = ( _COLUMNS = (
"call_id", "call_id",
"parent_call_id", "parent_call_id",
@@ -57,6 +68,9 @@ _COLUMNS = (
"cache_hit", "cache_hit",
"error", "error",
"cost", "cost",
"cached_prompt_tokens",
"model_reported",
"sampling",
) )
_INSERT = ( _INSERT = (
@@ -82,9 +96,37 @@ class SQLiteRecorder:
self._conn = conn self._conn = conn
except (OSError, sqlite3.Error) as exc: except (OSError, sqlite3.Error) as exc:
logger.warning("SQLite 遥测初始化失败,后续记录降级为 no-op: {}", exc) logger.warning("SQLite 遥测初始化失败,后续记录降级为 no-op: {}", exc)
self._backfill_columns()
def _backfill_columns(self) -> None:
"""给已存在的旧表补新列(issue #3);独立 try,失败只降级为逐行丢弃。
必须放在 `self._conn` 赋值**之后**并先判空: 初始化失败时连接为 None,
无守卫的补列会抛 AttributeError 逃出 `__init__`,把"静默降级"变成崩溃。
补列失败也绝不清空 `self._conn`——那会让整个 recorder 永久 no-op,
比逐行丢弃严重得多。
"""
if self._conn is None:
return
try:
existing = {row[1] for row in self._conn.execute("PRAGMA table_info(llm_calls)")}
except sqlite3.Error as exc:
logger.warning("SQLite 遥测列探测失败(写入将逐行降级): {}", exc)
return
for column, decl in _BACKFILL_COLUMNS:
if column in existing:
continue
# 逐列独立 try: 一列撞上 duplicate 不得让后面的列漏补
try:
self._conn.execute(f"ALTER TABLE llm_calls ADD COLUMN {column} {decl}")
self._conn.commit()
except sqlite3.Error as exc:
# duplicate column: 多进程共库时后到者必然撞上,属预期竞态,视为成功
if "duplicate column" not in str(exc).lower():
logger.warning("SQLite 遥测补列失败(写入将逐行降级): {}", exc)
async def record_llm_call(self, **fields: object) -> None: async def record_llm_call(self, **fields: object) -> None:
"""写一行遥测;字段集合即 18 字段冻结签名(ports.TelemetryRecorder)。""" """写一行遥测;字段集合即 21 字段冻结签名(ports.TelemetryRecorder)。"""
if self._conn is None: if self._conn is None:
return return
row = tuple(fields[col] for col in _COLUMNS) row = tuple(fields[col] for col in _COLUMNS)
@@ -45,6 +45,12 @@ def _sse_delta(chunk: dict[str, Any], usage_sink: dict[str, Any]) -> tuple[bool,
"""从 chunk 提取增量: (True, content) 或 (False, reasoning);usage 帧旁路进 sink。""" """从 chunk 提取增量: (True, content) 或 (False, reasoning);usage 帧旁路进 sink。"""
if chunk.get("usage"): if chunk.get("usage"):
usage_sink["usage"] = chunk["usage"] usage_sink["usage"] = chunk["usage"]
if "model" not in usage_sink:
# 首个**有效**值即固定: 末帧的异常值不得覆盖它;但首帧报空串也不能锁死
# sink——否则后续真实版本会丢(issue #3)
reported = _coerce_model_reported(chunk.get("model"))
if reported is not None:
usage_sink["model"] = reported
choices = chunk.get("choices") or [] choices = chunk.get("choices") or []
if not choices: if not choices:
return None return None
@@ -152,6 +158,35 @@ def _resolve_usage(usage: dict[str, Any]) -> tuple[int, int, str]:
return 0, 0, "unavailable" return 0, 0, "unavailable"
def _coerce_cached_tokens(usage: Any) -> int | None:
"""取 usage.prompt_tokens_details.cached_tokens(issue #3);形态异常一律 None。
`0` 与 `None` 必须可区分: 前者是"该源上报了一次真实零命中",后者是"该源
不报这个数",下游对两者的处置不同(后者不可做缓存成本校正)。故只把
**负数与非整数**归 None,`0` 如实保留。`bool` 显式排除——isinstance(True, int)
在 Python 里为真,放行会把 `True` 记成 1 个命中 token。
"""
if not isinstance(usage, dict):
return None
details = usage.get("prompt_tokens_details")
if not isinstance(details, dict):
return None
cached = details.get("cached_tokens")
if isinstance(cached, bool) or not isinstance(cached, int) or cached < 0:
return None
return cached
def _coerce_model_reported(value: Any) -> str | None:
"""取响应体的 model 字段(issue #3);非 str 或空白串一律 None,收口时去空白。
去空白不是洁癖: 下游拿这个串做实验快照的 key,`" m "` 与 `"m"` 会造成假分叉。
"""
if not isinstance(value, str) or not value.strip():
return None
return value.strip()
def _resolve_stream_usage(sink: dict[str, Any], salvaged: bool) -> tuple[int, int, str]: def _resolve_stream_usage(sink: dict[str, Any], salvaged: bool) -> tuple[int, int, str]:
"""流式用量口径: 打捞路径把 measured 降级为 estimated,unavailable 原样保留。 """流式用量口径: 打捞路径把 measured 降级为 estimated,unavailable 原样保留。
@@ -259,6 +294,9 @@ class OpenAICompatTransport:
payload.update(profile.thinking_on) payload.update(profile.thinking_on)
elif source.enable_thinking is False: elif source.enable_thinking is False:
payload.update(profile.thinking_off) payload.update(profile.thinking_off)
# 顺序即优先级(issue #4 设计决策 A): 配置级 extra_body 在前,调用级
# overlay(含结构化注入)在后覆盖之。两行不可调换
payload.update(source.extra_body)
payload.update(overlay) payload.update(overlay)
return payload return payload
@@ -360,6 +398,8 @@ class OpenAICompatTransport:
ttft_ms=ttft_ms, ttft_ms=ttft_ms,
max_inter_token_ms=(max_gap if ttft_ms is not None else None), max_inter_token_ms=(max_gap if ttft_ms is not None else None),
raw={"usage": sink.get("usage")}, raw={"usage": sink.get("usage")},
cached_prompt_tokens=_coerce_cached_tokens(sink.get("usage")),
model_reported=_coerce_model_reported(sink.get("model")),
) )
def _check_done( def _check_done(
@@ -442,6 +482,8 @@ class OpenAICompatTransport:
ttft_ms=None, ttft_ms=None,
max_inter_token_ms=None, max_inter_token_ms=None,
raw={"usage": body.get("usage")}, raw={"usage": body.get("usage")},
cached_prompt_tokens=_coerce_cached_tokens(body.get("usage")),
model_reported=_coerce_model_reported(body.get("model")),
) )
async def aclose(self) -> None: async def aclose(self) -> None:
+112
View File
@@ -4,11 +4,27 @@
fake,字段顺序即公共承诺;新增字段只增不删且必带默认值。 fake,字段顺序即公共承诺;新增字段只增不删且必带默认值。
""" """
import dataclasses
import json
from collections.abc import Mapping
from dataclasses import dataclass, field from dataclasses import dataclass, field
from types import MappingProxyType
from typing import Any from typing import Any
from loguru import logger
_MISSING_DONE_DOMAIN = frozenset({"retry", "salvage"}) _MISSING_DONE_DOMAIN = frozenset({"retry", "salvage"})
_PROTECTED_OVERLAY_KEYS: Mapping[str, str] = MappingProxyType(
{
"model": "会让遥测记录的 model 与实际请求分叉,成本按错单价换算",
"messages": "会同时破坏缓存 key 与遥测的 messages 口径",
"stream": "会绕过流式活性看门狗(TTFT/inter-token 超时全部失效)",
"stream_options": "会丢 usage 帧,导致成本遥测归零、TPM 闸按预扣量结算失准",
}
)
"""禁止出现在采样参数覆盖层里的键: 它们由治理层拥有,被覆盖即击穿治理。"""
USAGE_SOURCES = frozenset({"measured", "estimated", "unavailable"}) USAGE_SOURCES = frozenset({"measured", "estimated", "unavailable"})
"""usage_source 值域;仅约束库内生产侧取值,不在 frozen dataclass 上做运行时校验。""" """usage_source 值域;仅约束库内生产侧取值,不在 frozen dataclass 上做运行时校验。"""
@@ -16,6 +32,45 @@ _EST_TOKENS_QUOTA_DIVISOR = 60
"""未显式配置时的预扣量除数: 假定一次调用约占一秒钟的 TPM 配额份额。""" """未显式配置时的预扣量除数: 假定一次调用约占一秒钟的 TPM 配额份额。"""
def validate_request_overlay(overlay: Mapping[str, Any], *, origin: str) -> dict[str, Any]:
"""校验采样参数覆盖层并返回浅拷贝;origin 用于把错误指回配置/调用点。
两类校验缺一不可(issue #4 设计决策 B):保护键会击穿治理;不可 JSON
序列化的值会在 `CacheMW` 的降级 try **之外**抛裸 `TypeError`——那条路径
不属错误四分类、`TelemetryMW` 也不捕,结果是一行遥测都没有就崩了。
两者都在进洋葱之前收口,故抛裸 `ValueError`(调用方编程错误,不可重试)。
"""
# Phase 1: 键形态——必须先于序列化试探,否则非 str 键会因 sort_keys 的
# 比较失败被误报成"值不可序列化",把人指向错误的方向
for key in overlay:
if not isinstance(key, str):
raise ValueError(f"{origin} 的键必须是 str: {key!r}(canonical JSON 要求)")
# Phase 2: 保护键
for key, reason in _PROTECTED_OVERLAY_KEYS.items():
if key in overlay:
raise ValueError(f"{origin} 不得覆盖 {key!r}: {reason}")
# Phase 3: 值可序列化(缓存 key 与遥测列都要 json.dumps)
try:
json.dumps(dict(overlay), sort_keys=True, ensure_ascii=False)
except (TypeError, ValueError) as exc:
raise ValueError(
f"{origin} 的值必须可 JSON 序列化(如 numpy 标量请先转 float/int): {exc}"
) from exc
return dict(overlay)
def merge_sampling(extra_body: Mapping[str, Any], sampling: Mapping[str, Any]) -> dict[str, Any]:
"""合并配置级与调用级采样参数;调用级优先(issue #4 设计决策 A)。"""
return {**extra_body, **sampling}
def canonical_sampling_json(merged: Mapping[str, Any]) -> str | None:
"""缓存 key 与遥测 sampling 列共用的序列化口径;空 mapping → None。"""
if not merged:
return None
return json.dumps(dict(merged), sort_keys=True, ensure_ascii=False)
@dataclass(frozen=True) @dataclass(frozen=True)
class LLMResponse: class LLMResponse:
"""一次治理调用的统一响应(与三项目超集兼容,ARCH §5.1)。""" """一次治理调用的统一响应(与三项目超集兼容,ARCH §5.1)。"""
@@ -30,12 +85,20 @@ class LLMResponse:
ttft_ms: float | None ttft_ms: float | None
max_inter_token_ms: float | None max_inter_token_ms: float | None
cache_hit: bool cache_hit: bool
"""**PolyGateway 自身响应缓存**命中(未产生网关调用);与供应商侧 prompt
cache 无关,后者见 `cached_prompt_tokens`。"""
call_id: str call_id: str
# —— 库新增(只增不删,必带默认值;迁移兼容硬约束)—— # —— 库新增(只增不删,必带默认值;迁移兼容硬约束)——
source_name: str = "" source_name: str = ""
cost: float | None = None cost: float | None = None
usage_source: str = "measured" usage_source: str = "measured"
structured_data: Any | None = None structured_data: Any | None = None
cached_prompt_tokens: int | None = None
"""供应商 prompt cache 命中的输入 token 数(issue #3);None = 该源未上报,
"上报了但是 0"(真实零命中)区分——两者对下游的处置不同。"""
model_reported: str | None = None
"""API 响应体里的 model 字段;None = 未上报。与 `model`(配置别名)可能
分叉——供应商把别名指向新权重时,实验复现必须认这个串。"""
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -50,6 +113,12 @@ class ChatRequest:
structured: Any | None = None structured: Any | None = None
stream: bool = True stream: bool = True
overlay: dict[str, Any] = field(default_factory=dict) overlay: dict[str, Any] = field(default_factory=dict)
sampling: Mapping[str, Any] = field(default_factory=dict)
"""调用方采样意图的快照,库内中间件**永不修改**(issue #4 设计决策 A)。
与 `overlay` 分开是因为后者会被结构化中间件注入 `response_format`,在洋葱
不同深度取值不同;缓存 key 与三个遥测入口需要一个跨层恒定的读取点,否则
同一列在不同行口径分叉。"""
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -82,6 +151,9 @@ class TransportResult:
ttft_ms: float | None ttft_ms: float | None
max_inter_token_ms: float | None max_inter_token_ms: float | None
raw: dict[str, Any] raw: dict[str, Any]
# —— 可观测字段(issue #3;带默认值,非 OpenAI 兼容的 transport 可不填)——
cached_prompt_tokens: int | None = None
model_reported: str | None = None
@dataclass(frozen=True) @dataclass(frozen=True)
@@ -107,11 +179,18 @@ class SourceConfig:
enable_thinking: bool | None = None enable_thinking: bool | None = None
missing_done: str = "retry" missing_done: str = "retry"
trust_env: bool = True trust_env: bool = True
extra_body: Mapping[str, Any] = field(default_factory=dict)
"""本源恒定的采样参数(如 `temperature=0`),并入请求体(issue #4)。
优先级低于调用级 overlay。注: 本字段令 `SourceConfig` 不再 hashable
(加任何 mapping 字段的固有代价,裸 dict 亦然),库内无以源作 key 的写法;
要可变副本用 `dict(source.extra_body)`,要改字段用 `dataclasses.replace`。"""
def __post_init__(self) -> None: def __post_init__(self) -> None:
self._validate_identity() self._validate_identity()
self._validate_gates() self._validate_gates()
self._validate_watchdog() self._validate_watchdog()
self._freeze_extra_body()
def effective_est_tokens(self) -> int: def effective_est_tokens(self) -> int:
"""TPM 入场预扣量: 显式配置优先,否则按 tpm 派生(设计 §2.2)。""" """TPM 入场预扣量: 显式配置优先,否则按 tpm 派生(设计 §2.2)。"""
@@ -148,6 +227,39 @@ class SourceConfig:
): ):
raise ValueError("看门狗不变式要求 0 < inter_token < ttft < timeout_s") raise ValueError("看门狗不变式要求 0 < inter_token < ttft < timeout_s")
def _freeze_extra_body(self) -> None:
"""校验后转只读视图: 装配完成的源不应再被就地改采样参数(设计决策 E)。"""
validated = validate_request_overlay(
self.extra_body, origin=f"SourceConfig({self.name}).extra_body"
)
object.__setattr__(self, "extra_body", MappingProxyType(validated))
def strip_unsupported_extra_body(sources: list[SourceConfig], *, path: str) -> list[SourceConfig]:
"""剥离非 chat 路径不消费的 `extra_body` 并 warning(issue #4 决策 G)。
剥离是必需的而非顺手清理: embedding 的 payload 硬编码 `{model, input}`、
MonkeyOCR 只发 multipart 表单,两者都不会把 `extra_body` 发出去;但遥测的
`sampling` 列会并上 `source.extra_body`,不剥离就等于**记录一个从未发出的
参数**——那是数据造假,污染的恰是事后复现的唯一依据。
选择 warning 放行而非报错: 这两条路径本无采样语义,配错的后果远轻于 chat
路径,不值得让下游整个装配起不来(2026-07-31 人类拍板)。
"""
stripped = []
for source in sources:
if source.extra_body:
logger.warning(
"{} 路径暂不支持 extra_body,源 {} 的该配置已被忽略"
"(需要 dimensions 等参数请提 issue): {}",
path,
source.name,
dict(source.extra_body),
)
source = dataclasses.replace(source, extra_body={})
stripped.append(source)
return stripped
@dataclass(frozen=True) @dataclass(frozen=True)
class RetryPolicy: class RetryPolicy:
@@ -4,6 +4,7 @@
""" """
import asyncio import asyncio
import dataclasses
import json import json
import sqlite3 import sqlite3
@@ -204,3 +205,74 @@ class TestTransientErrorExport:
assert isinstance(ei.value, polygateway.AllSourcesExhausted) assert isinstance(ei.value, polygateway.AllSourcesExhausted)
assert isinstance(ei.value.__cause__, TransientError) assert isinstance(ei.value.__cause__, TransientError)
class TestSamplingThroughStack:
"""issue #4: 采样参数经完整洋葱到达请求体,且缓存/遥测口径一致。"""
async def test_reaches_wire_and_lands_in_telemetry(self, tmp_path):
seen = []
def handler(request):
seen.append(json.loads(request.content))
return _sse()
db = tmp_path / "t.db"
recorder = SQLiteRecorder(db)
client = _full_client(handler, telemetry=recorder)
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 42})
recorder.close()
assert seen[0]["seed"] == 42 # 穿过全栈到达线上
rows = sqlite3.connect(db).execute("SELECT sampling FROM llm_calls").fetchall()
assert json.loads(rows[0][0]) == {"seed": 42}
async def test_config_level_merges_and_records(self, tmp_path):
"""源级 extra_body 只有 emit_attempt 记得到(唯一有生效源的入口)。"""
seen = []
def handler(request):
seen.append(json.loads(request.content))
return _sse()
src = dataclasses.replace(_source(), extra_body={"temperature": 0})
db = tmp_path / "t.db"
recorder = SQLiteRecorder(db)
client = GatewayClient(
scope="llm",
sources=[src],
selector=RoundRobinSelector(),
limiter=InMemoryLimiter(
scope="llm", sources={src.name: src}, global_limits=GlobalLimits(0, 0, 0)
),
breaker=InMemoryGate(config=_BREAKER),
transport=OpenAICompatTransport(
client_factory=lambda s: httpx.AsyncClient(transport=httpx.MockTransport(handler))
),
retry=RetryPolicy(2, 2.0, 30.0),
backpressure=BackpressurePolicy(300.0, 0.01),
telemetry=recorder,
structured_strategy=JsonRepairStrategy(),
sleep=_noop_sleep,
)
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 1})
recorder.close()
assert seen[0]["temperature"] == 0 and seen[0]["seed"] == 1
rows = sqlite3.connect(db).execute("SELECT sampling FROM llm_calls").fetchall()
assert json.loads(rows[0][0]) == {"seed": 1, "temperature": 0}
async def test_differing_seed_bypasses_cache_end_to_end(self):
"""issue 场景全栈回归: 逐 rollout 变 seed 必须真的回源。"""
calls = []
def handler(request):
calls.append(json.loads(request.content)["seed"])
return _sse()
client = _full_client(handler, cache=InMemoryCache())
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 1})
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 2})
second_same = await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 1})
assert calls == [1, 2] # 两个不同 seed 各自回源
assert second_same.cache_hit is True # 同 seed 才命中
@@ -11,6 +11,7 @@ DSN 走 .env `PGW_TELEMETRY_PG_DSN`,缺则 skip。该实例上有 app/chs_prod
from __future__ import annotations from __future__ import annotations
import asyncio import asyncio
import json
import os import os
from uuid import uuid4 from uuid import uuid4
@@ -39,6 +40,9 @@ _EXPECTED_COLUMNS = [
"error", "error",
"cost", "cost",
"created_at", "created_at",
"cached_prompt_tokens",
"model_reported",
"sampling",
] ]
# run 级前缀: 同库并存的其他运行(迁移批跑/另一开发机)互不可见 # run 级前缀: 同库并存的其他运行(迁移批跑/另一开发机)互不可见
@@ -100,6 +104,9 @@ async def _record_minimal(
"cache_hit": False, "cache_hit": False,
"error": None, "error": None,
"cost": None, "cost": None,
"cached_prompt_tokens": None,
"model_reported": None,
"sampling": None,
} }
fields.update(overrides) fields.update(overrides)
await recorder.record_llm_call(**fields) await recorder.record_llm_call(**fields)
@@ -115,6 +122,111 @@ async def _fetch(dsn: str, sql: str, *args):
await conn.close() await conn.close()
_LEGACY_DDL = """
CREATE TABLE {schema}.llm_calls (
call_id TEXT PRIMARY KEY,
parent_call_id TEXT,
session_id TEXT,
model TEXT NOT NULL,
provider TEXT NOT NULL,
source_name TEXT NOT NULL,
messages TEXT NOT NULL,
response TEXT NOT NULL,
thinking TEXT NOT NULL DEFAULT '',
prompt_tokens INTEGER NOT NULL,
completion_tokens INTEGER NOT NULL,
usage_source TEXT NOT NULL,
latency_ms INTEGER NOT NULL,
ttft_ms DOUBLE PRECISION,
max_inter_token_ms DOUBLE PRECISION,
cache_hit BOOLEAN NOT NULL DEFAULT FALSE,
error TEXT,
cost DOUBLE PRECISION,
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
)
"""
@pytest.fixture
async def legacy_schema(dsn):
"""在**自建的临时 schema** 里造一张 18 列旧表,验证补列(issue #3)。
绝不碰共享的 public.llm_calls: 用 search_path 把 recorder 指向临时 schema,
teardown 只 DROP 自己建的 schema。
"""
import asyncpg
name = f"pgwtest_{uuid4().hex[:8]}"
conn = await asyncpg.connect(dsn, timeout=10)
try:
await conn.execute(f"CREATE SCHEMA {name}")
await conn.execute(_LEGACY_DDL.format(schema=name))
finally:
await conn.close()
sep = "&" if "?" in dsn else "?"
yield f"{dsn}{sep}options=-csearch_path%3D{name}", name
conn = await asyncpg.connect(dsn, timeout=10)
try:
await conn.execute(f"DROP SCHEMA {name} CASCADE")
finally:
await conn.close()
class TestObservabilityColumns:
"""issue #3: 两列写入可回读,且已存在的 18 列旧表会被自动补列。"""
async def test_values_round_trip(self, dsn):
recorder = PostgresRecorder(dsn)
try:
await _record_minimal(recorder, call_id=_cid("hit"), cached_prompt_tokens=64)
await _record_minimal(recorder, call_id=_cid("zero"), cached_prompt_tokens=0)
await _record_minimal(recorder, call_id=_cid("model"), model_reported="MiniMax-01")
await _record_minimal(
recorder, call_id=_cid("samp"), sampling='{"seed": 42, "temperature": 0}'
)
rows = await _fetch(
dsn,
"SELECT call_id, cached_prompt_tokens, model_reported, sampling FROM llm_calls "
"WHERE call_id LIKE $1",
f"{_RUN_PREFIX}-%",
)
by_id = {r["call_id"]: r for r in rows}
assert by_id[_cid("hit")]["cached_prompt_tokens"] == 64
assert by_id[_cid("zero")]["cached_prompt_tokens"] == 0 # 真实零命中 ≠ NULL
assert by_id[_cid("model")]["cached_prompt_tokens"] is None
assert by_id[_cid("model")]["model_reported"] == "MiniMax-01"
# issue #4: PG 侧也须验非空 sampling 能读回原值(不只是列存在)
assert json.loads(by_id[_cid("samp")]["sampling"]) == {"seed": 42, "temperature": 0}
assert by_id[_cid("hit")]["sampling"] is None
finally:
await recorder.aclose()
async def test_legacy_table_is_upgraded_in_place(self, legacy_schema):
"""18 列旧表不补列的话,每行写入都会被逐行 warning 丢弃(遥测静默全失)。"""
schema_dsn, schema = legacy_schema
recorder = PostgresRecorder(schema_dsn)
try:
await _record_minimal(
recorder, call_id=_cid("legacy"), cached_prompt_tokens=7, model_reported="m-real"
)
cols = await _fetch(
schema_dsn,
"SELECT column_name FROM information_schema.columns "
"WHERE table_schema = $1 AND table_name = 'llm_calls' ORDER BY ordinal_position",
schema,
)
# ALTER 只能追加到末尾: 与新建库的列序一致才不会分叉
assert [r["column_name"] for r in cols] == _EXPECTED_COLUMNS
rows = await _fetch(
schema_dsn,
"SELECT cached_prompt_tokens, model_reported FROM llm_calls WHERE call_id = $1",
_cid("legacy"),
)
assert (rows[0]["cached_prompt_tokens"], rows[0]["model_reported"]) == (7, "m-real")
finally:
await recorder.aclose()
class TestSchema: class TestSchema:
async def test_schema_has_frozen_columns_in_order(self, dsn): async def test_schema_has_frozen_columns_in_order(self, dsn):
recorder = PostgresRecorder(dsn) recorder = PostgresRecorder(dsn)
+98
View File
@@ -53,6 +53,35 @@ class TestKeyFormula:
def test_any_dimension_change_changes_key(self, a, b): def test_any_dimension_change_changes_key(self, a, b):
assert build_cache_key(*a) != build_cache_key(*b) assert build_cache_key(*a) != build_cache_key(*b)
def test_empty_sampling_keeps_legacy_key(self):
"""空采样参数时键形逐字不变,存量缓存不被全量作废(issue #4 决策 C)。
golden 值取自加 sampling 维度之前的实现,不得随实现漂移。
"""
assert build_cache_key("qwen-max", [{"role": "user", "content": "hi"}], "proj", None) == (
"pgw:cache:c54544e8672f4c91373b4a72716a88497445b440b89445aa5379b356b228f58b"
)
assert build_cache_key("qwen-max", [{"role": "user", "content": "hi"}], "proj", "s1") == (
"pgw:cache:eed9cd9cc06acc0dedf4f337b74e06ed3482afdc30fa2acedd194f6cc1df33bf"
)
def test_differing_seed_changes_key(self):
"""issue #4 的直接回归: 5 个 seed 若共用一个 key,标准差会恒为 0。"""
k1 = build_cache_key("m", _MSGS, "proj", None, sampling={"seed": 1})
k2 = build_cache_key("m", _MSGS, "proj", None, sampling={"seed": 2})
assert k1 != k2
def test_sampling_key_order_irrelevant(self):
k1 = build_cache_key("m", _MSGS, "proj", None, sampling={"seed": 1, "temperature": 0})
k2 = build_cache_key("m", _MSGS, "proj", None, sampling={"temperature": 0, "seed": 1})
assert k1 == k2
def test_empty_sampling_equals_omitted(self):
"""空 dict 与不传须同键,否则升级后存量缓存全部 miss。"""
assert build_cache_key("m", _MSGS, "proj", None, sampling={}) == build_cache_key(
"m", _MSGS, "proj", None
)
def test_multimodal_part_digested_not_inlined(self): def test_multimodal_part_digested_not_inlined(self):
big_b64 = "data:image/png;base64," + "A" * 1_000_000 big_b64 = "data:image/png;base64," + "A" * 1_000_000
messages = [ messages = [
@@ -120,6 +149,32 @@ class TestCacheFlow:
assert second.call_id != first.call_id # 命中生成独立 cache_call_id assert second.call_id != first.call_id # 命中生成独立 cache_call_id
assert terminal.calls == 1 # 未再触达内层 assert terminal.calls == 1 # 未再触达内层
async def test_differing_sampling_does_not_hit(self):
"""issue #4 的中间件层回归: 逐 rollout 变 seed 必须回源,不得复用响应。"""
backend = InMemoryCache()
mw = _mw(backend)
terminal = _Terminal(_resp())
await mw(ChatRequest(messages=_MSGS, sampling={"seed": 1}), terminal)
await mw(ChatRequest(messages=_MSGS, sampling={"seed": 2}), terminal)
assert terminal.calls == 2 # 两次都回源
# 同 seed 才命中
third = await mw(ChatRequest(messages=_MSGS, sampling={"seed": 1}), terminal)
assert third.cache_hit is True and terminal.calls == 2
async def test_structured_injection_does_not_pollute_key(self):
"""CacheMW 读 sampling 而非 overlay: 结构化注入不该改变缓存身份。"""
backend = InMemoryCache()
mw = _mw(backend)
terminal = _Terminal(_resp())
await mw(ChatRequest(messages=_MSGS, sampling={"seed": 1}), terminal)
polluted = ChatRequest(
messages=_MSGS,
sampling={"seed": 1},
overlay={"seed": 1, "response_format": {"type": "json_object"}},
)
assert (await mw(polluted, terminal)).cache_hit is True
assert terminal.calls == 1
async def test_per_call_namespace_overrides_default(self): async def test_per_call_namespace_overrides_default(self):
backend = InMemoryCache() backend = InMemoryCache()
mw = _mw(backend) mw = _mw(backend)
@@ -150,6 +205,49 @@ class TestCacheFlow:
assert terminal.calls == 2 assert terminal.calls == 2
class TestObservabilityFieldsOnHit:
"""issue #3 决策 B1: 命中行原样回放,与 model/prompt_tokens 同一口径。"""
async def test_fields_replayed_on_hit(self):
backend = InMemoryCache()
mw = _mw(backend)
terminal = _Terminal(
_resp(cached_prompt_tokens=64, model_reported="MiniMax-Text-01-250321")
)
await mw(ChatRequest(messages=_MSGS), terminal)
hit = await mw(ChatRequest(messages=_MSGS), terminal)
assert hit.cache_hit is True
assert hit.cached_prompt_tokens == 64
assert hit.model_reported == "MiniMax-Text-01-250321"
async def test_legacy_cache_entry_without_new_keys_rehydrates(self):
"""旧格式条目(无这两个键)必须照常重建为 None,不得抛异常回源。"""
backend = InMemoryCache()
mw = _mw(backend)
key = build_cache_key("m", _MSGS, "proj", None)
legacy = {
"content": "legacy",
"thinking": "",
"model": "m",
"provider": "p",
"prompt_tokens": 1,
"completion_tokens": 2,
"latency_ms": 30,
"ttft_ms": 5.0,
"max_inter_token_ms": 2.0,
"cache_hit": False,
"call_id": "orig",
"source_name": "s1",
"cost": None,
"usage_source": "measured",
}
await backend.set(key, json.dumps(legacy), ttl_s=100)
terminal = _Terminal(_resp())
hit = await mw(ChatRequest(messages=_MSGS), terminal)
assert hit.content == "legacy" and terminal.calls == 0 # 真的走了缓存
assert hit.cached_prompt_tokens is None and hit.model_reported is None
class _BrokenBackend: class _BrokenBackend:
async def get(self, key): async def get(self, key):
raise ConnectionError("redis down") raise ConnectionError("redis down")
+101
View File
@@ -19,6 +19,7 @@ from polygateway.backends.memory.cache import InMemoryCache
from polygateway.backends.memory.limiter import InMemoryLimiter from polygateway.backends.memory.limiter import InMemoryLimiter
from polygateway.sources import RoundRobinSelector from polygateway.sources import RoundRobinSelector
from polygateway.structured.json_repair import JsonRepairStrategy from polygateway.structured.json_repair import JsonRepairStrategy
from polygateway.structured.native_schema import NativeSchemaStrategy
from polygateway.transports.openai_compat import OpenAICompatTransport from polygateway.transports.openai_compat import OpenAICompatTransport
from polygateway.types import ( from polygateway.types import (
BackpressurePolicy, BackpressurePolicy,
@@ -117,6 +118,106 @@ class TestChatEndToEnd:
await client.chat([{"role": "user", "content": "hi"}], structured="json") await client.chat([{"role": "user", "content": "hi"}], structured="json")
class TestSamplingOverlay:
"""调用级采样参数入口(issue #4 Task 3)。"""
def _capturing_client(self, captured, **overrides):
def handler(request):
captured.append(json.loads(request.content))
return _sse()
return _client(handler=handler, **overrides)
async def test_overlay_reaches_request_body(self):
captured = []
async with self._capturing_client(captured) as client:
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 42})
assert captured[0]["seed"] == 42
async def test_call_level_beats_config_level(self):
"""优先级: 调用级 > 配置级(设计决策 A)。"""
captured = []
source = _source(extra_body={"temperature": 0, "top_p": 0.9})
async with self._capturing_client(captured, sources=[source]) as client:
await client.chat([{"role": "user", "content": "hi"}], overlay={"temperature": 1})
assert captured[0]["temperature"] == 1 # 调用级覆盖
assert captured[0]["top_p"] == 0.9 # 配置级未被顶掉的键保留
async def test_structured_injection_beats_call_level(self):
"""结构化注入优先级最高: 它关系到响应能否被解析(设计决策 A)。"""
captured = []
client = self._capturing_client(captured, structured_strategy=NativeSchemaStrategy())
async with client:
await client.chat(
[{"role": "user", "content": "hi"}],
structured="json",
overlay={"response_format": {"type": "text"}},
)
assert captured[0]["response_format"] != {"type": "text"}
async def test_protected_key_rejected_before_onion(self):
"""保护键在进洋葱之前就报错,transport 一次都不该被碰到。"""
captured = []
async with self._capturing_client(captured) as client:
with pytest.raises(ValueError, match="stream"):
await client.chat([{"role": "user", "content": "hi"}], overlay={"stream": False})
assert captured == []
async def test_unserializable_value_rejected_before_onion(self):
"""裸 TypeError 会在 CacheMW 的降级 try 之外炸且无遥测(设计决策 B)。"""
captured = []
async with self._capturing_client(captured) as client:
with pytest.raises(ValueError, match="JSON"):
await client.chat(
[{"role": "user", "content": "hi"}], overlay={"temperature": object()}
)
assert captured == []
async def test_caller_dict_mutation_does_not_leak(self):
"""调用方逐次改 seed 复用同一 dict 是预期模式(设计决策 E)。"""
captured = []
caller_overlay = {"seed": 1}
async with self._capturing_client(captured) as client:
await client.chat([{"role": "user", "content": "hi"}], overlay=caller_overlay)
caller_overlay["seed"] = 2
await client.chat([{"role": "user", "content": "hi"}], overlay=caller_overlay)
assert [c["seed"] for c in captured] == [1, 2]
class TestModelFingerprint:
"""配置级采样参数须进缓存身份,否则改 temperature 后仍读旧缓存(决策 C)。"""
def test_empty_extra_body_keeps_legacy_fingerprint(self):
"""全源无 extra_body 时字面量与旧实现逐字相同,不触发存量缓存冷启动。"""
from polygateway.client import build_model_fingerprint
sources = [_source(), _source(name="qwen_2", model="qwen-plus")]
assert build_model_fingerprint(sources) == "qwen-max,qwen-plus"
def test_extra_body_changes_fingerprint(self):
from polygateway.client import build_model_fingerprint
plain = build_model_fingerprint([_source()])
tuned = build_model_fingerprint([_source(extra_body={"temperature": 0})])
assert plain != tuned
assert tuned.startswith("qwen-max|") # 旧字面量仍是前缀,便于人眼辨认
def test_source_rename_does_not_change_fingerprint(self):
"""指纹按 (model, extra_body) 而非源名: 改名不该误触全量冷启动。"""
from polygateway.client import build_model_fingerprint
a = build_model_fingerprint([_source(name="qwen_1", extra_body={"temperature": 0})])
b = build_model_fingerprint([_source(name="renamed", extra_body={"temperature": 0})])
assert a == b
def test_differing_extra_body_across_sources_is_distinguished(self):
from polygateway.client import build_model_fingerprint
a = build_model_fingerprint([_source(extra_body={"temperature": 0})])
b = build_model_fingerprint([_source(extra_body={"temperature": 1})])
assert a != b
class TestFactories: class TestFactories:
def test_from_env_assembles(self): def test_from_env_assembles(self):
client = GatewayClient.from_env("LLM", env=_ENV) client = GatewayClient.from_env("LLM", env=_ENV)
+29
View File
@@ -103,6 +103,35 @@ class TestSourceAggregation:
GatewaySettings.from_env("LLM", env=env) GatewaySettings.from_env("LLM", env=env)
class TestExtraBodyParsing:
"""配置级采样参数的 env 解析(issue #4 Task 2)。"""
def test_json_object_parsed(self):
env = _env(**{"LLM__QWEN__1__EXTRA_BODY": '{"temperature": 0, "seed": 42}'})
s = GatewaySettings.from_env("LLM", env=env)
assert s.sources[0].extra_body == {"temperature": 0, "seed": 42}
def test_absent_defaults_to_empty(self):
assert GatewaySettings.from_env("LLM", env=_env()).sources[0].extra_body == {}
def test_invalid_json_fails_loudly(self):
env = _env(**{"LLM__QWEN__1__EXTRA_BODY": "{invalid"})
with pytest.raises(ValueError, match="EXTRA_BODY"):
GatewaySettings.from_env("LLM", env=env)
def test_non_object_json_fails(self):
"""数组/标量都不是请求体片段,静默接受会让参数悄悄不生效。"""
env = _env(**{"LLM__QWEN__1__EXTRA_BODY": "[1, 2]"})
with pytest.raises(ValueError, match="JSON 对象"):
GatewaySettings.from_env("LLM", env=env)
def test_protected_key_rejected_through_assembly(self):
"""校验确实挂在装配路径上(而非只在 types.py 里孤立存在)。"""
env = _env(**{"LLM__QWEN__1__EXTRA_BODY": '{"model": "sneaky"}'})
with pytest.raises(ValueError, match="model"):
GatewaySettings.from_env("LLM", env=env)
class TestResilienceKeys: class TestResilienceKeys:
def test_flat_legacy_keys(self): def test_flat_legacy_keys(self):
s = GatewaySettings.from_env("LLM", env=_env()) s = GatewaySettings.from_env("LLM", env=_env())
+42
View File
@@ -4,11 +4,13 @@
VT adapters/embedding.py(归一化);库裁决见设计 §7.3 表。 VT adapters/embedding.py(归一化);库裁决见设计 §7.3 表。
""" """
import contextlib
import dataclasses import dataclasses
import json import json
import httpx import httpx
import pytest import pytest
from loguru import logger
from polygateway.errors import ( from polygateway.errors import (
RequestRejectedError, RequestRejectedError,
@@ -350,6 +352,46 @@ class TestEmbedTelemetry:
assert len(rec.rows[1]["messages"]) < 1000 # 长文本截断后入库 assert len(rec.rows[1]["messages"]) < 1000 # 长文本截断后入库
@contextlib.contextmanager
def _captured_warnings():
"""捕获库发出的 WARNING;loguru 不经标准 logging,pytest 的 caplog 抓不到。"""
messages: list[str] = []
sink_id = logger.add(messages.append, level="WARNING")
try:
yield messages
finally:
logger.remove(sink_id)
class TestExtraBodyStripped:
"""issue #4 决策 G: embedding 路径不消费 extra_body,剥离并 warning。"""
async def test_stripped_with_warning_but_assembly_succeeds(self):
"""报错会让下游整个装配起不来,而这条路径本无采样语义(人类拍板)。"""
with _captured_warnings() as warnings:
client, _ = _embed_client([_src(extra_body={"temperature": 0})], ["ok"])
assert client._sources[0].extra_body == {}
assert any("extra_body" in m for m in warnings)
assert any("dimensions" in m for m in warnings) # 文案须指路,不能只说不支持
await client.embed(["hi"]) # 装配后可正常工作
async def test_telemetry_never_records_a_parameter_that_was_not_sent(self):
"""剥离的真正理由: embed payload 硬编码 {model, input},不剥离则审计表
会显示这次调用带了 temperature=0——那是数据造假,比参数失效更坏。
"""
rec = _MemoryRecorder()
client, _ = _embed_client([_src(extra_body={"temperature": 0})], ["ok"], telemetry=rec)
await client.embed(["hi"])
assert rec.rows[0]["sampling"] is None
async def test_no_warning_without_extra_body(self):
with _captured_warnings() as warnings:
client, _ = _embed_client([_src()], ["ok"])
assert client._sources[0].extra_body == {}
assert not [m for m in warnings if "extra_body" in m]
class TestEmbeddingSettings: class TestEmbeddingSettings:
_ENV = { _ENV = {
"EMBED__QWEN__1__BASE_URL": "https://gw.example/v1", "EMBED__QWEN__1__BASE_URL": "https://gw.example/v1",
+22
View File
@@ -7,6 +7,7 @@ retry_exhausted/circuit_open/stalled 三组断言即设计 §6 ③ 的 G1 契约
import asyncio import asyncio
import pytest import pytest
from loguru import logger
from polygateway.backends.memory.breaker import InMemoryGate from polygateway.backends.memory.breaker import InMemoryGate
from polygateway.backends.memory.limiter import InMemoryLimiter from polygateway.backends.memory.limiter import InMemoryLimiter
@@ -377,6 +378,27 @@ class TestCheckHealth:
await task await task
class TestExtraBodyStripped:
"""issue #4 决策 G: OCR 路径只发 multipart 表单,剥离 extra_body 并 warning。"""
def test_stripped_with_warning_but_assembly_succeeds(self):
messages: list[str] = []
sink_id = logger.add(messages.append, level="WARNING")
try:
client, _, _ = _client([_src(extra_body={"temperature": 0})], ["text"])
finally:
logger.remove(sink_id)
assert client._sources[0].extra_body == {}
assert any("extra_body" in m for m in messages)
async def test_telemetry_never_records_a_parameter_that_was_not_sent(self):
"""不剥离则审计表会显示这次 OCR 带了 temperature=0——数据造假。"""
recorder = _MemoryRecorder()
client, _, _ = _client([_src(extra_body={"temperature": 0})], ["text"], telemetry=recorder)
await client.recognize_text(b"IMG")
assert recorder.rows[0]["sampling"] is None
class TestTelemetry: class TestTelemetry:
async def test_success_and_failure_recorded_without_image_bytes(self): async def test_success_and_failure_recorded_without_image_bytes(self):
recorder = _MemoryRecorder() recorder = _MemoryRecorder()
+155 -1
View File
@@ -36,7 +36,7 @@ def _source(**overrides):
return SourceConfig(**base) return SourceConfig(**base)
def _chunk(content=None, reasoning=None, usage=None): def _chunk(content=None, reasoning=None, usage=None, model=None):
delta = {} delta = {}
if content is not None: if content is not None:
delta["content"] = content delta["content"] = content
@@ -45,6 +45,8 @@ def _chunk(content=None, reasoning=None, usage=None):
body = {"choices": [{"delta": delta}]} if (delta or usage is None) else {"choices": []} body = {"choices": [{"delta": delta}]} if (delta or usage is None) else {"choices": []}
if usage is not None: if usage is not None:
body["usage"] = usage body["usage"] = usage
if model is not None:
body["model"] = model
return f"data: {json.dumps(body)}\n\n" return f"data: {json.dumps(body)}\n\n"
@@ -261,6 +263,137 @@ class TestEmptyCompletion:
await _complete(_transport_for(handler), _source()) await _complete(_transport_for(handler), _source())
class TestObservabilityFields:
"""issue #3: 供应商 prompt cache 命中数与 API 实际返回的模型版本串。
网关报文一律不可信: 形态异常只归 None,绝不因一个可观测字段打断调用。
"""
def _cached_usage(self, cached):
return {**_USAGE, "prompt_tokens_details": {"cached_tokens": cached}}
async def test_stream_reads_cached_tokens(self):
def handler(request):
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._cached_usage(128)))
result = await _complete(_transport_for(handler), _source())
assert result.cached_prompt_tokens == 128
async def test_non_stream_reads_cached_tokens(self):
def handler(request):
return httpx.Response(
200,
json={
"choices": [{"message": {"content": "42"}}],
"usage": self._cached_usage(128),
},
)
result = await _complete(_transport_for(handler), _source(), stream=False)
assert result.cached_prompt_tokens == 128
async def test_zero_cached_tokens_is_a_real_zero(self):
"""0(真实零命中)与 None(该源未上报)必须可区分——issue #3 的核心诉求。"""
def handler(request):
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._cached_usage(0)))
result = await _complete(_transport_for(handler), _source())
assert result.cached_prompt_tokens == 0
async def test_usage_without_details_is_none(self):
def handler(request):
return _sse_stream(_chunk(content="ok"), _chunk(usage=_USAGE))
result = await _complete(_transport_for(handler), _source())
assert result.cached_prompt_tokens is None
async def test_missing_usage_frame_is_none(self):
def handler(request):
return _sse_stream(_chunk(content="ok"))
result = await _complete(_transport_for(handler), _source())
assert result.cached_prompt_tokens is None
@pytest.mark.parametrize("bad", ["abc", -1, True, 1.5, None, [], {"x": 1}])
async def test_malformed_cached_tokens_degrade_to_none(self, bad):
"""`True` 必须排除: Python 里 isinstance(True, int) 为真。"""
def handler(request):
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._cached_usage(bad)))
result = await _complete(_transport_for(handler), _source())
assert result.cached_prompt_tokens is None
async def test_details_not_a_dict_is_none(self):
def handler(request):
usage = {**_USAGE, "prompt_tokens_details": "oops"}
return _sse_stream(_chunk(content="ok"), _chunk(usage=usage))
result = await _complete(_transport_for(handler), _source())
assert result.cached_prompt_tokens is None
async def test_non_stream_reads_reported_model(self):
def handler(request):
return httpx.Response(
200,
json={
"choices": [{"message": {"content": "42"}}],
"usage": _USAGE,
"model": "MiniMax-Text-01-250321",
},
)
result = await _complete(_transport_for(handler), _source(), stream=False)
assert result.model_reported == "MiniMax-Text-01-250321"
async def test_stream_keeps_the_first_reported_model(self):
"""末帧异常值不得覆盖首帧: 首次写入即固定。"""
def handler(request):
return _sse_stream(
_chunk(content="a", model="MiniMax-Text-01-250321"),
_chunk(content="b", model="something-else"),
_chunk(usage=_USAGE),
)
result = await _complete(_transport_for(handler), _source())
assert result.model_reported == "MiniMax-Text-01-250321"
async def test_empty_first_model_does_not_block_a_later_real_one(self):
"""首帧报空串不得锁死 sink: 守卫按"有效值"判断,否则真实版本会丢。"""
def handler(request):
return _sse_stream(
_chunk(content="a", model=""),
_chunk(content="b", model="MiniMax-Text-01-250321"),
_chunk(usage=_USAGE),
)
result = await _complete(_transport_for(handler), _source())
assert result.model_reported == "MiniMax-Text-01-250321"
@pytest.mark.parametrize("bad", [None, "", " ", 123, {}])
async def test_missing_or_malformed_model_is_none(self, bad):
def handler(request):
body = {"choices": [{"message": {"content": "42"}}], "usage": _USAGE}
if bad is not None:
body["model"] = bad
return httpx.Response(200, json=body)
result = await _complete(_transport_for(handler), _source(), stream=False)
assert result.model_reported is None
async def test_raw_payload_is_unchanged(self):
"""新字段是独立格子,不改动 raw 的既有内容。"""
def handler(request):
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._cached_usage(5)))
result = await _complete(_transport_for(handler), _source())
assert set(result.raw) == {"usage"}
class TestNonStreamFastPath: class TestNonStreamFastPath:
async def test_non_stream_parses_message(self): async def test_non_stream_parses_message(self):
def handler(request): def handler(request):
@@ -312,6 +445,27 @@ class TestRequestShaping:
) )
assert seen["response_format"] == {"type": "json_object"} assert seen["response_format"] == {"type": "json_object"}
async def test_extra_body_merged_and_outranked_by_overlay(self):
"""顺序即优先级: thinking profile → extra_body → overlay(issue #4)。"""
seen = {}
def handler(request):
seen.update(json.loads(request.content))
return _sse_stream(_chunk(content="x"), _chunk(usage=_USAGE))
await _complete(
_transport_for(handler),
_source(extra_body={"temperature": 0, "top_p": 0.9}),
overlay={"temperature": 1},
)
assert seen["temperature"] == 1 # 调用级覆盖配置级
assert seen["top_p"] == 0.9 # 未被顶掉的配置级键保留
async def test_extra_body_cannot_break_governed_keys(self):
"""治理键由 payload 骨架拥有;extra_body 的保护键在构造期已被拦下。"""
with pytest.raises(ValueError, match="stream"):
_source(extra_body={"stream": False})
class TestErrorTranslation: class TestErrorTranslation:
@pytest.mark.parametrize( @pytest.mark.parametrize(
+3
View File
@@ -114,6 +114,9 @@ class _DummyRecorder:
cache_hit, cache_hit,
error, error,
cost, cost,
cached_prompt_tokens,
model_reported,
sampling,
) -> None: ... ) -> None: ...
+68
View File
@@ -56,6 +56,74 @@ class TestPricingTable:
ModelPrice(input_per_1m=-1.0, output_per_1m=0.0) ModelPrice(input_per_1m=-1.0, output_per_1m=0.0)
class TestCachedInputTier:
"""issue #3: 供应商 prompt cache 命中部分按更低单价计费,不配则不猜折扣。"""
_CACHED = PricingTable(
{"m": ModelPrice(input_per_1m=10.0, output_per_1m=20.0, cached_input_per_1m=2.0)}
)
_PLAIN = PricingTable({"m": ModelPrice(input_per_1m=10.0, output_per_1m=20.0)})
def test_hit_is_billed_at_the_cached_rate(self):
# 1M prompt 中 600k 命中: 400k×10 + 600k×2 = 4.0 + 1.2
assert self._CACHED.cost("m", 1_000_000, 0, 600_000) == pytest.approx(5.2)
def test_without_the_tier_the_result_is_unchanged(self):
"""未配缓存档 = 退化为现状全额计价,绝不按经验折扣率猜(P5)。"""
full = self._PLAIN.cost("m", 1_000_000, 0)
assert self._PLAIN.cost("m", 1_000_000, 0, 600_000) == full == pytest.approx(10.0)
@pytest.mark.parametrize("cached", [None, 0])
def test_no_hit_is_billed_in_full(self, cached):
assert self._CACHED.cost("m", 1_000_000, 0, cached) == pytest.approx(10.0)
def test_negative_cached_is_billed_in_full(self):
"""负数命中数不得抬高成本: cost() 是公共方法,外部输入须校验后使用(P5)。"""
assert self._CACHED.cost("m", 1_000_000, 0, -500_000) == pytest.approx(10.0)
def test_cached_over_prompt_is_clamped_and_never_negative(self):
"""网关口径异常时按输入总数夹取: 全部按缓存价,不得算出负成本。"""
clamped = self._CACHED.cost("m", 1_000_000, 0, 5_000_000)
assert clamped == pytest.approx(2.0) and clamped >= 0
def test_legacy_three_arg_call_still_works(self):
"""embedding.py 的三参调用形态必须零改动可用。"""
assert self._CACHED.cost("m", 1_000_000, 0) == pytest.approx(10.0)
def test_from_file_accepts_and_validates_the_tier(self, tmp_path):
path = tmp_path / "p.json"
path.write_text(
json.dumps(
{"m": {"input_per_1m": 10.0, "output_per_1m": 20.0, "cached_input_per_1m": 2.0}}
),
encoding="utf-8",
)
assert PricingTable.from_file(path).cost("m", 1_000_000, 0, 1_000_000) == pytest.approx(2.0)
@pytest.mark.parametrize("bad", [-1.0, "x"])
def test_from_file_rejects_a_bad_tier(self, tmp_path, bad):
path = tmp_path / "bad.json"
path.write_text(
json.dumps(
{"m": {"input_per_1m": 1.0, "output_per_1m": 2.0, "cached_input_per_1m": bad}}
),
encoding="utf-8",
)
with pytest.raises(ValueError, match="cached_input_per_1m"):
PricingTable.from_file(path)
def test_legacy_price_file_without_the_tier_still_loads(self, tmp_path):
path = tmp_path / "old.json"
path.write_text(
json.dumps({"m": {"input_per_1m": 1.0, "output_per_1m": 2.0}}), encoding="utf-8"
)
assert PricingTable.from_file(path).cost("m", 1_000_000, 0, 500_000) == pytest.approx(1.0)
def test_negative_tier_rejected_on_construction(self):
with pytest.raises(ValueError):
ModelPrice(input_per_1m=1.0, output_per_1m=1.0, cached_input_per_1m=-0.1)
class _MemoryRecorder: class _MemoryRecorder:
def __init__(self): def __init__(self):
self.rows = [] self.rows = []
+29
View File
@@ -194,6 +194,35 @@ class TestSuccessPath:
assert (await limiter.source_stats("a")).tpm_used == 16 assert (await limiter.source_stats("a")).tpm_used == 16
class TestObservabilityPassthrough:
"""issue #3: transport 采到的两个可观测字段必须原样上浮到 LLMResponse。"""
async def test_fields_reach_the_response(self):
result = TransportResult(
content="ok",
thinking="",
prompt_tokens=10,
completion_tokens=5,
usage_source="measured",
ttft_ms=12.0,
max_inter_token_ms=3.0,
raw={},
cached_prompt_tokens=64,
model_reported="MiniMax-Text-01-250321",
)
mw, *_ = _harness([_src("a")], [result])
resp = await mw(_REQ)
assert resp.cached_prompt_tokens == 64
assert resp.model_reported == "MiniMax-Text-01-250321"
# model 仍是配置别名: 真实版本是旁证,不顶替溯源主字段
assert resp.model == "m"
async def test_absent_fields_stay_none(self):
mw, *_ = _harness([_src("a")], [_ok()])
resp = await mw(_REQ)
assert resp.cached_prompt_tokens is None and resp.model_reported is None
class TestRetryAndFailover: class TestRetryAndFailover:
async def test_transient_switches_source_then_succeeds(self): async def test_transient_switches_source_then_succeeds(self):
mw, _, _, transport, sleep, _ = _harness( mw, _, _, transport, sleep, _ = _harness(
+52
View File
@@ -177,3 +177,55 @@ class TestNativeOverlayFirstAttempt:
resp = await _mw()(ChatRequest(messages=_MSGS, structured=Verdict), terminal) resp = await _mw()(ChatRequest(messages=_MSGS, structured=Verdict), terminal)
with pytest.raises(dataclasses.FrozenInstanceError): with pytest.raises(dataclasses.FrozenInstanceError):
resp.structured_data = None resp.structured_data = None
class TestSamplingSnapshotInvariant:
"""地基不变式: `sampling` 跨洋葱层恒定,`overlay` 会被结构化注入(issue #4)。
缓存 key(决策 C)与三个遥测入口(决策 D)都建立在这条之上,而它此前只靠
"dataclasses.replace 恰好保留未提及字段"的约定成立,无任何机械执法。
这个测试是那份执法——它红了就意味着两个决策同时失效。
"""
async def test_sampling_survives_feedback_ladder_while_overlay_diverges(self):
caller_sampling = {"temperature": 0, "seed": 42}
# 先坏后好,强制走一次带反馈重问(重问会 replace messages)
terminal = ScriptedTerminal(["not json at all", '{"answer": 1, "reason": "r"}'])
mw = _mw(strategy=NativeSchemaStrategy(), max_retries=1)
await mw(
# overlay 与 sampling 传**同一个对象**,复现 client.py 的别名关系
# ——否则中间件就地改写 overlay 时不会波及 sampling,这条执法就是空的
ChatRequest(
messages=_MSGS,
structured=Verdict,
overlay=caller_sampling,
sampling=caller_sampling,
),
terminal,
)
assert len(terminal.requests) == 2 # 确实重问过
for seen in terminal.requests:
# ① 跨层恒定: 每次尝试看到的 sampling 与调用方传入的逐字相同
assert seen.sampling == caller_sampling
# ② 确实分叉: 同一时刻 overlay 已被注入 response_format
assert seen.overlay["response_format"]["type"] == "json_schema"
assert "response_format" not in seen.sampling
async def test_middleware_does_not_mutate_caller_mapping(self):
"""决策 E 的第二条约束: 中间件只能 replace 派生,不得就地改这两个 dict。
同样传同一对象: 生产中 overlay 与 sampling 是别名,任何对 overlay 的
就地改写都会同步毒化缓存 key 与遥测列。
"""
caller_sampling = {"seed": 7}
terminal = ScriptedTerminal(['{"answer": 1, "reason": "r"}'])
await _mw(strategy=NativeSchemaStrategy())(
ChatRequest(
messages=_MSGS,
structured=Verdict,
overlay=caller_sampling,
sampling=caller_sampling,
),
terminal,
)
assert caller_sampling == {"seed": 7} # 调用方的对象未被污染
+416 -10
View File
@@ -1,6 +1,7 @@
"""遥测子系统测试: SQLiteRecorder(18 列)+ TelemetryEmitter 单一 helper + TelemetryMW。""" """遥测子系统测试: SQLiteRecorder(21 列)+ TelemetryEmitter 单一 helper + TelemetryMW。"""
import asyncio import asyncio
import json
import sqlite3 import sqlite3
import subprocess import subprocess
from pathlib import Path from pathlib import Path
@@ -35,6 +36,9 @@ _EXPECTED_COLUMNS = [
"error", "error",
"cost", "cost",
"created_at", "created_at",
"cached_prompt_tokens",
"model_reported",
"sampling",
] ]
@@ -58,15 +62,17 @@ def _resp(**overrides):
return LLMResponse(**base) return LLMResponse(**base)
def _source(): def _source(**overrides):
return SourceConfig( base = {
name="s1", "name": "s1",
provider="p", "provider": "p",
base_url="https://gw.example/v1", "base_url": "https://gw.example/v1",
api_key="sk", "api_key": "sk",
model="m", "model": "m",
timeout_s=10.0, "timeout_s": 10.0,
) }
base.update(overrides)
return SourceConfig(**base)
# 输出单价 8 元/百万: 改前 `unavailable` 行按兜底的 0/4000 换算恰好是 0.032 # 输出单价 8 元/百万: 改前 `unavailable` 行按兜底的 0/4000 换算恰好是 0.032
@@ -93,6 +99,9 @@ async def _record_minimal(recorder, call_id="c1", **overrides):
"cache_hit": False, "cache_hit": False,
"error": None, "error": None,
"cost": None, "cost": None,
"cached_prompt_tokens": None,
"model_reported": None,
"sampling": None,
} }
fields.update(overrides) fields.update(overrides)
await recorder.record_llm_call(**fields) await recorder.record_llm_call(**fields)
@@ -134,6 +143,181 @@ class TestSQLiteRecorder:
await _record_minimal(recorder) # 不抛 await _record_minimal(recorder) # 不抛
recorder.close() recorder.close()
async def test_observability_columns_round_trip(self, tmp_path):
recorder = SQLiteRecorder(tmp_path / "t.db")
await _record_minimal(recorder, call_id="c-hit", cached_prompt_tokens=64)
await _record_minimal(recorder, call_id="c-zero", cached_prompt_tokens=0)
await _record_minimal(recorder, call_id="c-none", model_reported="MiniMax-Text-01")
recorder.close()
rows = dict(
sqlite3.connect(tmp_path / "t.db")
.execute("SELECT call_id, cached_prompt_tokens FROM llm_calls")
.fetchall()
)
assert rows["c-hit"] == 64
assert rows["c-zero"] == 0 # 真实零命中,读回仍是 0 而非 NULL
assert rows["c-none"] is None
async def test_sampling_column_round_trips(self, tmp_path):
"""issue #4: 采样参数落库,否则事后无法证明某批数据跑在什么温度下。"""
recorder = SQLiteRecorder(tmp_path / "t.db")
await _record_minimal(recorder, call_id="c-s", sampling='{"seed": 42, "temperature": 0}')
await _record_minimal(recorder, call_id="c-plain")
recorder.close()
rows = dict(
sqlite3.connect(tmp_path / "t.db")
.execute("SELECT call_id, sampling FROM llm_calls")
.fetchall()
)
assert json.loads(rows["c-s"]) == {"seed": 42, "temperature": 0}
assert rows["c-plain"] is None # 无采样参数为 NULL,便于 SQL 过滤
class TestSQLiteColumnBackfill:
"""issue #3: 已存在的 18 列旧表必须自动补列,否则每行写入都被丢弃。"""
_LEGACY_DDL = """
CREATE TABLE llm_calls (
call_id TEXT PRIMARY KEY,
parent_call_id TEXT,
session_id TEXT,
model TEXT NOT NULL,
provider TEXT NOT NULL,
source_name TEXT NOT NULL,
messages TEXT NOT NULL,
response TEXT NOT NULL,
thinking TEXT NOT NULL DEFAULT '',
prompt_tokens INTEGER NOT NULL,
completion_tokens INTEGER NOT NULL,
usage_source TEXT NOT NULL,
latency_ms INTEGER NOT NULL,
ttft_ms REAL,
max_inter_token_ms REAL,
cache_hit INTEGER NOT NULL DEFAULT 0,
error TEXT,
cost REAL,
created_at TEXT NOT NULL DEFAULT (datetime('now'))
);
"""
async def test_legacy_table_is_upgraded_in_place(self, tmp_path):
db = tmp_path / "legacy.db"
legacy = sqlite3.connect(db)
legacy.execute(self._LEGACY_DDL)
legacy.commit()
legacy.close()
recorder = SQLiteRecorder(db)
await _record_minimal(recorder, cached_prompt_tokens=7, model_reported="m-real")
recorder.close()
conn = sqlite3.connect(db)
cols = [r[1] for r in conn.execute("PRAGMA table_info(llm_calls)")]
assert cols == _EXPECTED_COLUMNS # ALTER 追加到末尾,与新建库列序一致
assert conn.execute(
"SELECT cached_prompt_tokens, model_reported FROM llm_calls"
).fetchone() == (7, "m-real")
async def test_backfill_failure_keeps_the_recorder_usable(self, tmp_path):
"""补列失败只能逐行降级,绝不能把 recorder 整体变成 no-op(设计 D1 纪律)。
把 llm_calls 建成同名 view: `CREATE TABLE IF NOT EXISTS` 遇 view 静默
no-op(不抛),随后的 ALTER 才抛 "Cannot add a column to a view"——正是
补列失败这条分支。`_conn` 必须保持非 None,否则整个 recorder 永久失能。
"""
db = tmp_path / "view.db"
conn = sqlite3.connect(db)
conn.execute("CREATE TABLE real_rows (call_id TEXT)")
conn.execute("CREATE VIEW llm_calls AS SELECT call_id FROM real_rows")
conn.commit()
conn.close()
recorder = SQLiteRecorder(db) # 不得抛
assert recorder._conn is not None # 补列失败 ≠ recorder 失能(D1 纪律)
await _record_minimal(recorder) # 不得抛
recorder.close()
class _FakePgConn:
"""记录执行过的语句;可让 ALTER 抛错以模拟权限不足。"""
def __init__(self, existing: list[str], *, fail_alter: bool = False):
self.existing = existing
self.fail_alter = fail_alter
self.statements: list[str] = []
async def execute(self, sql, *args):
self.statements.append(sql)
if sql.startswith("ALTER TABLE") and self.fail_alter:
raise RuntimeError("must be owner of table llm_calls")
async def fetch(self, sql, *args):
self.statements.append(sql)
return [{"attname": name} for name in self.existing]
class _FakePgPool:
def __init__(self, conn):
self._conn = conn
def acquire(self):
conn = self._conn
class _Ctx:
async def __aenter__(self):
return conn
async def __aexit__(self, *exc):
return False
return _Ctx()
class TestPostgresBackfillDiscipline:
"""PG 补列必须与 SQLite 侧对称: 失败只逐行降级,且稳态不抢排他锁(issue #3)。"""
_LEGACY = ["call_id", "cost", "created_at"]
_CURRENT = [
"call_id",
"cost",
"created_at",
"cached_prompt_tokens",
"model_reported",
"sampling",
]
def _recorder(self, conn):
from polygateway.telemetry.postgres import PostgresRecorder
return PostgresRecorder("postgresql://u:p@h:5432/polygateway", pool=_FakePgPool(conn))
async def test_alter_failure_does_not_disable_the_recorder(self):
"""ALTER 失败(如账号只有 INSERT 权限)不得置 _failed —— 那会让遥测全灭。"""
conn = _FakePgConn(self._LEGACY, fail_alter=True)
recorder = self._recorder(conn)
await _record_minimal(recorder) # 不得抛
assert recorder._failed is False
assert any(s.startswith("INSERT INTO llm_calls") for s in conn.statements)
async def test_no_alter_when_columns_already_exist(self):
"""ADD COLUMN IF NOT EXISTS 即使列已存在也会先抢 ACCESS EXCLUSIVE 锁,
而遥测是内联 await——稳态下必须一条 ALTER 都不发,否则每个进程的首次
写入都会去锁共享审计表。
"""
conn = _FakePgConn(self._CURRENT)
await _record_minimal(self._recorder(conn))
assert not [s for s in conn.statements if s.startswith("ALTER TABLE")]
async def test_missing_columns_are_added_once(self):
conn = _FakePgConn(self._LEGACY)
await _record_minimal(self._recorder(conn))
from polygateway.telemetry.postgres import _BACKFILL
altered = [s for s in conn.statements if s.startswith("ALTER TABLE")]
assert len(altered) == len(_BACKFILL) # 旧表缺全部补列,故一列一条 ALTER
assert all("IF NOT EXISTS" not in s for s in altered) # 探测已确认缺列,无需再判
class _MemoryRecorder: class _MemoryRecorder:
def __init__(self): def __init__(self):
@@ -143,6 +327,228 @@ class _MemoryRecorder:
self.rows.append(fields) self.rows.append(fields)
class TestEmitterRecorderContract:
"""emitter 的实参键集合必须与两个后端的 _COLUMNS 完全一致(issue #3)。
两个后端的 `row = tuple(fields[col] for col in _COLUMNS)` 都在 try **之外**,
emitter 漏传一个键就抛 KeyError,被 `_record` 的 except Exception 吞成 warning
→ 遥测静默全丢。而 8 个 `**fields` 形态的 fake 一个都拦不住,故显式断言。
"""
async def test_emitter_supplies_exactly_the_backend_columns(self):
from polygateway.telemetry.postgres import _COLUMNS as PG_COLUMNS
from polygateway.telemetry.sqlite import _COLUMNS as SQLITE_COLUMNS
rec = _MemoryRecorder()
await TelemetryEmitter(rec).emit_attempt(
request=_REQ,
source=_source(),
call_id="cid-1",
latency_ms=42,
response=_resp(),
error=None,
)
assert set(rec.rows[0]) == set(SQLITE_COLUMNS) == set(PG_COLUMNS)
@pytest.mark.parametrize("emit", ["attempt", "cache_hit", "terminal_failure"])
async def test_every_entry_point_supplies_the_same_keys(self, emit):
from polygateway.telemetry.sqlite import _COLUMNS as SQLITE_COLUMNS
rec = _MemoryRecorder()
emitter = TelemetryEmitter(rec)
if emit == "attempt":
await emitter.emit_attempt(
request=_REQ,
source=_source(),
call_id="c",
latency_ms=1,
response=None,
error="boom",
)
elif emit == "cache_hit":
await emitter.emit_cache_hit(request=_REQ, response=_resp())
else:
await emitter.emit_terminal_failure(
request=_REQ, call_id="c", latency_ms=1, error="dead"
)
assert set(rec.rows[0]) == set(SQLITE_COLUMNS)
class TestEmitterObservabilityFields:
"""issue #3: 三个入口各自的取值口径(设计 §5 表)。"""
async def test_attempt_carries_the_response_values(self):
rec = _MemoryRecorder()
await TelemetryEmitter(rec).emit_attempt(
request=_REQ,
source=_source(),
call_id="cid-1",
latency_ms=42,
response=_resp(cached_prompt_tokens=64, model_reported="m-real"),
error=None,
)
assert rec.rows[0]["cached_prompt_tokens"] == 64
assert rec.rows[0]["model_reported"] == "m-real"
async def test_failed_attempt_has_no_provider_facts(self):
rec = _MemoryRecorder()
await TelemetryEmitter(rec).emit_attempt(
request=_REQ,
source=_source(),
call_id="cid-2",
latency_ms=7,
response=None,
error="boom",
)
assert rec.rows[0]["cached_prompt_tokens"] is None
assert rec.rows[0]["model_reported"] is None
async def test_cache_hit_replays_the_recorded_values(self):
"""决策 B1: 命中行原样回放,故命中率统计必须带 WHERE cache_hit = false。"""
rec = _MemoryRecorder()
await TelemetryEmitter(rec).emit_cache_hit(
request=_REQ, response=_resp(cached_prompt_tokens=64, model_reported="m-real")
)
row = rec.rows[0]
assert row["cache_hit"] is True
assert row["cached_prompt_tokens"] == 64 and row["model_reported"] == "m-real"
async def test_terminal_failure_records_none(self):
rec = _MemoryRecorder()
await TelemetryEmitter(rec).emit_terminal_failure(
request=_REQ, call_id="c", latency_ms=1, error="dead"
)
assert rec.rows[0]["cached_prompt_tokens"] is None
assert rec.rows[0]["model_reported"] is None
class TestEmitterSamplingColumn:
"""issue #4: sampling 列在三个入口的口径(设计决策 D 表格)。
列语义 = 「调用方采样意图 ⊎ 生效源 extra_body」,**不含**结构化注入的
response_format(列名是采样参数,schema 不是;且数 KB schema 逐行落库会让
审计表无谓膨胀)。三入口若各读各的层,同一列在不同行含义就不同。
"""
_SAMPLED = ChatRequest(
messages=[{"role": "user", "content": "hi"}],
sampling={"seed": 42},
overlay={"seed": 42, "response_format": {"type": "json_object"}},
)
async def test_attempt_merges_source_extra_body(self):
rec = _MemoryRecorder()
await TelemetryEmitter(rec).emit_attempt(
request=self._SAMPLED,
source=_source(extra_body={"temperature": 0}),
call_id="c",
latency_ms=1,
response=_resp(),
error=None,
)
assert json.loads(rec.rows[0]["sampling"]) == {"seed": 42, "temperature": 0}
async def test_response_format_never_leaks_into_the_column(self):
"""三行都不得出现 response_format——它不是采样参数。"""
rec = _MemoryRecorder()
emitter = TelemetryEmitter(rec)
await emitter.emit_attempt(
request=self._SAMPLED,
source=_source(),
call_id="c",
latency_ms=1,
response=_resp(),
error=None,
)
await emitter.emit_cache_hit(request=self._SAMPLED, response=_resp())
await emitter.emit_terminal_failure(
request=self._SAMPLED, call_id="c", latency_ms=1, error="dead"
)
assert len(rec.rows) == 3
for row in rec.rows:
assert "response_format" not in row["sampling"]
@pytest.mark.parametrize("emit", ["cache_hit", "terminal_failure"])
async def test_sourceless_entries_record_call_level_only(self, emit):
"""两个最外层入口没有"生效源"可言,与 model/source_name 置空同一先例。"""
rec = _MemoryRecorder()
emitter = TelemetryEmitter(rec)
if emit == "cache_hit":
await emitter.emit_cache_hit(request=self._SAMPLED, response=_resp())
else:
await emitter.emit_terminal_failure(
request=self._SAMPLED, call_id="c", latency_ms=1, error="dead"
)
assert json.loads(rec.rows[0]["sampling"]) == {"seed": 42}
async def test_absent_sampling_is_null(self):
"""无采样参数时为 NULL,而非空字符串或 "{}"——便于 SQL 过滤。"""
rec = _MemoryRecorder()
await TelemetryEmitter(rec).emit_attempt(
request=_REQ,
source=_source(),
call_id="c",
latency_ms=1,
response=_resp(),
error=None,
)
assert rec.rows[0]["sampling"] is None
class TestCostWithCachedTier:
"""issue #3: 命中部分按缓存单价计费,避免 cost 系统性高估。"""
_TABLE = PricingTable(
{"m": ModelPrice(input_per_1m=10.0, output_per_1m=20.0, cached_input_per_1m=2.0)}
)
async def test_cached_hit_lowers_the_recorded_cost(self):
rec = _MemoryRecorder()
emitter = TelemetryEmitter(rec, pricing=self._TABLE)
full = _resp(prompt_tokens=1_000_000, completion_tokens=0)
await emitter.emit_attempt(
request=_REQ,
source=_source(),
call_id="c1",
latency_ms=1,
response=full,
error=None,
)
await emitter.emit_attempt(
request=_REQ,
source=_source(),
call_id="c2",
latency_ms=1,
response=_resp(
prompt_tokens=1_000_000, completion_tokens=0, cached_prompt_tokens=600_000
),
error=None,
)
assert rec.rows[0]["cost"] == pytest.approx(10.0)
assert rec.rows[1]["cost"] == pytest.approx(5.2) # 400k×10 + 600k×2
async def test_cache_hit_row_still_costs_zero(self):
"""缓存命中未产生新调用 → cost 恒 0.0,该短路必须排在任何换算之前。"""
rec = _MemoryRecorder()
await TelemetryEmitter(rec, pricing=self._TABLE).emit_cache_hit(
request=_REQ,
response=_resp(prompt_tokens=1_000_000, cached_prompt_tokens=600_000),
)
assert rec.rows[0]["cost"] == 0.0
async def test_unavailable_usage_still_costs_none(self):
rec = _MemoryRecorder()
await TelemetryEmitter(rec, pricing=self._TABLE).emit_attempt(
request=_REQ,
source=_source(),
call_id="c",
latency_ms=1,
response=_resp(usage_source="unavailable", cached_prompt_tokens=5),
error=None,
)
assert rec.rows[0]["cost"] is None
class TestEmitter: class TestEmitter:
async def test_attempt_success_row(self): async def test_attempt_success_row(self):
rec = _MemoryRecorder() rec = _MemoryRecorder()
+114
View File
@@ -49,6 +49,29 @@ class TestLLMResponse:
assert resp.usage_source == "measured" assert resp.usage_source == "measured"
assert resp.structured_data is None assert resp.structured_data is None
def test_observability_fields_default_to_none(self):
"""issue #3: None = 该源未上报,与"上报了但是 0"区分(0 是真实零命中)。"""
resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid")
assert resp.cached_prompt_tokens is None
assert resp.model_reported is None
filled = LLMResponse(
"c",
"t",
"m",
"p",
1,
2,
3,
None,
None,
False,
"cid",
cached_prompt_tokens=0,
model_reported="MiniMax-Text-01-250321",
)
assert filled.cached_prompt_tokens == 0 # 真实零命中,不得与 None 混同
assert filled.model_reported == "MiniMax-Text-01-250321"
def test_frozen(self): def test_frozen(self):
resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid") resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid")
with pytest.raises(dataclasses.FrozenInstanceError): with pytest.raises(dataclasses.FrozenInstanceError):
@@ -222,6 +245,8 @@ class TestAuxTypes:
raw={"id": "x"}, raw={"id": "x"},
) )
assert s.raw["id"] == "x" assert s.raw["id"] == "x"
# issue #3: 新字段带默认值,不填也能构造(OCR 等其他 transport 零改动)
assert s.cached_prompt_tokens is None and s.model_reported is None
class TestOcrTypes: class TestOcrTypes:
@@ -280,3 +305,92 @@ class TestOcrTypes:
with pytest.raises(TypeError): with pytest.raises(TypeError):
OcrTextResult(text="x") # 溯源件不可省略 OcrTextResult(text="x") # 溯源件不可省略
class TestSamplingValidation:
"""采样参数覆盖层的构造期校验(issue #4 设计决策 B)。"""
@pytest.mark.parametrize("key", ["model", "messages", "stream", "stream_options"])
def test_protected_keys_rejected(self, key):
"""保护键会击穿治理: 成本算错/口径失真/绕过看门狗与 usage 帧。"""
from polygateway.types import validate_request_overlay
with pytest.raises(ValueError) as exc:
validate_request_overlay({key: "x"}, origin="chat(overlay=...)")
assert key in str(exc.value)
assert "chat(overlay=...)" in str(exc.value) # 信息须能定位来源
def test_non_str_key_reports_key_problem(self):
"""非 str 键须报"键必须是 str",不能被 sort_keys 的比较错误误报成不可序列化。"""
from polygateway.types import validate_request_overlay
with pytest.raises(ValueError, match="str"):
validate_request_overlay({1: "a", "b": 2}, origin="test")
def test_unserializable_value_becomes_value_error(self):
"""裸 TypeError 会逃出 CacheMW 的降级 try 且一行遥测都没有(设计决策 B)。"""
from polygateway.types import validate_request_overlay
with pytest.raises(ValueError, match="JSON"):
validate_request_overlay({"temperature": object()}, origin="test")
def test_returns_independent_copy(self):
"""调用方逐次改 seed 复用同一 dict 是预期模式,不拷贝会有竞态(决策 E)。"""
from polygateway.types import validate_request_overlay
caller_dict = {"temperature": 0, "seed": 42}
validated = validate_request_overlay(caller_dict, origin="test")
caller_dict["seed"] = 43
assert validated == {"temperature": 0, "seed": 42}
def test_merge_prefers_call_level(self):
"""优先级: 调用级 > 配置级(设计决策 A)。"""
from polygateway.types import merge_sampling
merged = merge_sampling({"temperature": 0, "top_p": 1}, {"temperature": 1})
assert merged == {"temperature": 1, "top_p": 1}
def test_canonical_json_is_key_order_stable(self):
"""缓存 key 与遥测列共用同一序列化口径,键序不得影响结果。"""
from polygateway.types import canonical_sampling_json
assert canonical_sampling_json({"b": 1, "a": 2}) == canonical_sampling_json(
{"a": 2, "b": 1}
)
assert canonical_sampling_json({}) is None
class TestSourceConfigExtraBody:
"""配置级采样参数(issue #4 设计决策 A/E)。"""
def test_defaults_to_empty_and_is_read_only(self):
source = _make_source()
assert source.extra_body == {}
with pytest.raises(TypeError):
source.extra_body["temperature"] = 0 # MappingProxyType 只读
def test_protected_key_rejected_at_construction(self):
"""装配期报错,不放到运行时才炸(CLAUDE.md §4.5)。"""
with pytest.raises(ValueError, match="model"):
_make_source(extra_body={"model": "sneaky"})
def test_accepts_sampling_params(self):
source = _make_source(extra_body={"temperature": 0})
assert source.extra_body["temperature"] == 0
def test_replace_rebuilds_proxy(self):
"""决策 G 的剥离依赖 replace 能重跑 __post_init__ 且不递归。"""
source = _make_source(extra_body={"temperature": 0})
stripped = dataclasses.replace(source, extra_body={})
assert stripped.extra_body == {}
with pytest.raises(TypeError):
stripped.extra_body["x"] = 1
def test_no_longer_hashable_is_intentional(self):
"""加 mapping 字段的固有代价(裸 dict 亦然),库内无调用点会踩。
锁定为有意行为: 将来踩到的人不应把它当 bug""回去——要可变副本用
dict(source.extra_body),要改字段用 dataclasses.replace(设计 Task 1)。
"""
with pytest.raises(TypeError):
hash(_make_source())