Compare commits
68 Commits
| Author | SHA1 | Date | |
|---|---|---|---|
| ce630a37ef | |||
| dfda59fec2 | |||
| 15b9b02e96 | |||
| cce7562d07 | |||
| 2958dc8231 | |||
| a5ebf72f17 | |||
| 4516761dbe | |||
| b6e4cc3f3b | |||
| c31cc1adad | |||
| 6bb64ca938 | |||
| 152fa264ed | |||
| 6023d11bfb | |||
| b12bf6ce79 | |||
| b24e224beb | |||
| 09e77f11f8 | |||
| 0cc89fb03c | |||
| 20fd899d93 | |||
| 486809b08b | |||
| 58cb55b869 | |||
| 86fb4d5536 | |||
| 32d7869043 | |||
| 966d548245 | |||
| c2fcd5b1f8 | |||
| 0ed9dc107c | |||
| c4eda119ac | |||
| cd1a9520ff | |||
| 0aa7202c87 | |||
| 4841d901af | |||
| 30d7ffd94a | |||
| 037e7a011e | |||
| 0e2f734b0b | |||
| 7ccb25e8f1 | |||
| 42d16919fc | |||
| 005a90ca19 | |||
| 8a824e2000 | |||
| 8495cea5dc | |||
| abca723d3d | |||
| 63b85508c7 | |||
| 4e06d5e801 | |||
| 4e5a91d802 | |||
| 9e2d8ee43c | |||
| d1520cc0a5 | |||
| ab496bb298 | |||
| cd8bebba00 | |||
| 195454d2e3 | |||
| 42e429eb58 | |||
| 76e7d9594c | |||
| d8e8fd8124 | |||
| e5dbcf5d33 | |||
| 61231f7f6e | |||
| 4534444ad8 | |||
| 9a8f5cea5a | |||
| 637ac51754 | |||
| ac7c86fdee | |||
| fd7d9d330b | |||
| afd6101c08 | |||
| 726f26d8bd | |||
| c9fdff9d55 | |||
| a65b504a3d | |||
| 8c9e1179bc | |||
| b693d442f5 | |||
| 64d0fac879 | |||
| 8b8f396486 | |||
| b8f738f8cb | |||
| 91671a77df | |||
| f92065bc0b | |||
| f17044dead | |||
| f9995ef61b |
+8
-2
@@ -8,16 +8,20 @@ LLM__QWEN__1__BASE_URL=
|
||||
LLM__QWEN__1__API_KEY=
|
||||
LLM__QWEN__1__MODEL=
|
||||
LLM__QWEN__1__TIMEOUT_S=120
|
||||
# 可选(0 = 该闸不启用;TPM > 0 时 EST_TOKENS 必填 > 0):
|
||||
# 可选(0 = 该闸不启用):
|
||||
# LLM__QWEN__1__MAX_CONCURRENCY=8
|
||||
# LLM__QWEN__1__RPM=60
|
||||
# LLM__QWEN__1__TPM=100000
|
||||
# LLM__QWEN__1__EST_TOKENS=2000
|
||||
# LLM__QWEN__1__EST_TOKENS=2000 # 可选调优覆盖: TPM 入场预扣量;未填则库按 tpm//60 派生
|
||||
# LLM__QWEN__1__TTFT_TIMEOUT_S=30 # 须与 INTER_TOKEN 成对;0 < inter < ttft < timeout
|
||||
# LLM__QWEN__1__INTER_TOKEN_TIMEOUT_S=15
|
||||
# LLM__QWEN__1__ENABLE_THINKING=true # 三态: 缺省=不注入 / true=注入开启 / false=注入关闭
|
||||
# LLM__QWEN__1__MISSING_DONE=retry # SSE 缺 [DONE]: retry(默认) | salvage
|
||||
# LLM__QWEN__1__TRUST_ENV=true # false = 绕过本地代理(LAN 直连)
|
||||
# LLM__QWEN__1__EXTRA_BODY={"temperature":0} # 本源恒定的采样参数(JSON 对象串)
|
||||
# 并入请求体,优先级低于 chat(overlay=...);受控实验固定解码用它,免得漏传
|
||||
# 禁用键 model/messages/stream/stream_options(会击穿治理),配了直接报错
|
||||
# OCR/EMBED scope 不消费该键: 配了会被忽略并 warning(见 issue #4 决策 G)
|
||||
|
||||
# ══ scope 级全局闸(跨源合计;0/缺省 = 不启用)══
|
||||
# LLM__GLOBAL__MAX_CONCURRENCY=8
|
||||
@@ -54,6 +58,8 @@ PGW_TELEMETRY_BACKEND=none # sqlite | postgres | none(必填)
|
||||
# PGW_TELEMETRY_SQLITE_PATH=logs/telemetry.db # sqlite 时必填
|
||||
# PGW_TELEMETRY_PG_DSN=postgresql://user:pass@host:5432/polygateway # postgres 时必填;严禁指向在用业务库(实验室约定: 专用库 polygateway)
|
||||
# PGW_PRICING_PATH=config/prices.json # 可选: {"<model>": {"input_per_1m": x, "output_per_1m": y}};缺省 cost 恒 None
|
||||
# # 可选第三档 "cached_input_per_1m": z —— 供应商 prompt cache 命中部分的单价;
|
||||
# # 不填即命中部分也按 input 全额计(库不猜折扣率),cost 会偏高
|
||||
# PGW_CACHE_NAMESPACE=<项目名或租户前缀> # 缓存启用时必填(防跨项目毒化)
|
||||
# PGW_CACHE_TTL_S=604800 # 缓存启用时必填,须 > 0
|
||||
# PGW_STRUCTURED_MAX_RETRIES=2 # 缺省 2(M2.5);0 = 解析失败不重问(CHS 策略)
|
||||
|
||||
@@ -1,5 +1,86 @@
|
||||
# Changelog
|
||||
|
||||
## 1.0.5(2026-07-31)
|
||||
|
||||
采样参数透传(issue #4)。`chat()` 此前没有任何途径设置 `temperature` / `seed` / `max_tokens`——全库检索 `temperature` 零命中,`ChatRequest.overlay` 虽会被并进请求体却只由结构化中间件填充,调用方够不着。对受控实验而言这是阻塞性的:解码温度未知且可能随供应商默认值变化,每格配置跑 5 个 seed 报出的标准差无从解释。
|
||||
|
||||
### 新增(纯增,不破坏任何现有调用方)
|
||||
|
||||
- **`chat()` 新增 keyword-only 参数 `overlay: Mapping[str, Any] | None = None`**,承载逐次变化的采样参数(每个 rollout 不同的 `seed`)。带默认值的 keyword-only 参数不改变既有调用点。
|
||||
- **`SourceConfig` 新增 `extra_body` 字段**,对应环境键 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`(JSON **对象**串),承载全局恒定的参数(`temperature=0`)——免得每个调用点都要记得传,而漏传一次不会报错、只会让数字悄悄不可比。
|
||||
- **优先级为 结构化注入 > 调用级 `overlay` > 源级 `extra_body`。** 由现有层序天然给出,未引入新机制。
|
||||
- **遥测表 `llm_calls` 新增 `sampling` 列**,`TelemetryRecorder` 端口由 20 字段扩为 21;补列走 1.0.4 已建立的"先探测缺列再 ALTER、失败只逐行降级"套路。列语义是「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,**不含**结构化输出注入的 `response_format`(列名是采样参数,而数 KB 的 schema 逐行落库只会让审计表膨胀)。
|
||||
|
||||
### 下游请读
|
||||
|
||||
- **采样参数进缓存 key,所以逐次变化的 `seed` 天然全部 miss。** 这是正确语义而非缺陷:不进 key 的话,同 messages 跑 5 个 seed 会全部命中第一次的响应,标准差恒为 0 且不报错。代价是缓存对这条路径不再省钱。**不传采样参数时 key 逐字不变**,存量缓存不受影响。
|
||||
- **`model_fingerprint` 是集合级指纹,不是本次选中源的指纹。** 同 scope 下各源 `extra_body` 不同时,缓存仍可能返回另一源、另一组解码参数下产生的响应(这是既有取舍的延续,`model` 一直如此)。要求逐源可复现的实验应让每个源独享 scope 或 namespace。
|
||||
- **`{model, messages, stream, stream_options}` 是保护键,配了直接报 `ValueError`。** 它们由治理层拥有:`model` 被覆盖会让成本按错单价算,`stream`/`stream_options` 会绕过流式看门狗、丢掉 usage 帧。不可 JSON 序列化的值(如 numpy 标量)同样在进洋葱之前报错——否则会在缓存层的降级保护之外抛裸 `TypeError`,连一行遥测都留不下。
|
||||
- **`SourceConfig` 不再 hashable**,`dataclasses.asdict()` / `copy.deepcopy()` 也不再适用(加任何 mapping 字段的固有代价,裸 dict 亦然)。要可变副本用 `dict(source.extra_body)`,要改字段用 `dataclasses.replace(source, ...)`。
|
||||
- **OCR / embedding 路径不消费 `extra_body`**:配了会被**剥离并 warning**,装配照常成功。这两条路径的 transport 根本不发这个值(embed payload 硬编码 `{model, input}`、MonkeyOCR 只发 multipart 表单),剥离是为了让遥测不至于记录一个从未发出的参数。需要 `dimensions` 等 embedding 参数请提 issue。
|
||||
- **`enable_thinking` 对 `openai` / `minimax` 两个 provider 不产生任何效果**(它们的 thinking profile 两档皆空)。此前没有任何地方说明这一点,调用方可能以为自己关掉了推理。需要下发自定义参数请用 `extra_body`。
|
||||
|
||||
## 1.0.4(2026-07-31)
|
||||
|
||||
响应可观测字段扩展(issue #3)。下游 dissect 要把每次调用落成一行审计记录,其中两列拿不到值:供应商侧 prompt cache 命中了多少 token、这次调用实际跑的是哪个模型版本。前者关系到能否把「缓存命中率差异带来的成本」与「实验条件本身带来的成本」分开,后者关系到实验快照的可复现性。本次把两者暴露到公共类型与遥测表,并让成本换算认识缓存单价。
|
||||
|
||||
### 新增(纯增字段,不破坏任何现有调用方)
|
||||
|
||||
- **`LLMResponse` 新增 `cached_prompt_tokens: int | None` 与 `model_reported: str | None`。** 前者是供应商 prompt cache 命中的输入 token 数(OpenAI 兼容格式的 `usage.prompt_tokens_details.cached_tokens`),后者是 API 响应体里的 `model` 字段(与 `.env` 配的别名可能分叉——供应商把别名指向新权重时,只有它认得出真正跑的那个版本)。两者均带默认值 `None`,逐字段传参的 fake 构造零改动。
|
||||
- **`None` 与 `0` 是两回事,不可混同。** `None` = 该源不上报这个数(下游据此声明「本源不可做缓存成本校正」);`0` = 该源上报了一次真实零命中。网关报文一律不可信:形态异常(负数、字符串、`bool`、`prompt_tokens_details` 非 dict)一律归 `None` 且绝不抛异常——可观测字段缺失不得打断调用。
|
||||
- **遥测表 `llm_calls` 新增 `cached_prompt_tokens` 与 `model_reported` 两列**,`TelemetryRecorder` 端口由 18 字段扩为 20。两个后端在初始化期对**已存在的旧表幂等补列**——`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入都被逐行 warning 丢弃、遥测静默全失。两侧都是**先探测缺列、只在真缺列时才 ALTER**(SQLite 查 `PRAGMA table_info`,Postgres 查 `pg_attribute`):`ADD COLUMN IF NOT EXISTS` 即使列已存在也会先取 ACCESS EXCLUSIVE 锁,而遥测是内联 await,让每个进程的首次写入都去锁共享审计表会拖垮业务调用;稳态下一条 ALTER 都不会发。**补列失败只降级为逐行丢弃,绝不会让 recorder 整体失能**(应用账号只有 INSERT 权限时,`ALTER TABLE` 的 ownership 检查早于存在性判断,列齐全也会失败)。
|
||||
- **`PricingTable` 支持可选的缓存读取单价 `cached_input_per_1m`。** 配了该档且本次有命中时按 `(prompt - cached) × input + cached × cached_input` 分段计价,消除 cost 的系统性高估;**未配则不猜折扣率**,退化为现状全额输入价(P5 严禁默认值掩盖)。旧价格表文件与 embedding 侧的三参 `cost()` 调用零改动。命中数超过输入总数时按总数夹取并 warning,不产生负成本。
|
||||
|
||||
### 下游请读
|
||||
|
||||
- **`cache_hit` 与新字段是两个不同的东西。** `cache_hit` 指的始终是 **PolyGateway 自身的响应缓存**(未产生网关调用),而 `cached_prompt_tokens` 指的是**供应商服务器**复用了提示词前缀、那部分按更低单价计费——真实调用里天天发生,`cache_hit` 永远看不见它。字段名保持不变(改名会破坏迁移兼容),语义已在 docstring 中消歧。
|
||||
- **统计供应商缓存命中率必须写 `WHERE cache_hit = false`。** 缓存命中行的这两个字段是**原样回放**的历史值(与 `model`、`prompt_tokens` 同一口径:`CacheMW` 只覆写与本次调用相关的时序字段),计入会重复计数。这与 1.0.3 里 `cost` 缺口口径的坑是同一类。
|
||||
- 缓存命中行的 `cost` 仍恒为 `0.0`(未产生新调用),该短路排在任何单价换算之前,不受缓存单价档影响。
|
||||
- 旧格式的缓存条目(缺这两个键)照常可重建为 `None`,不会回源;历史遥测行的新列为 NULL。
|
||||
|
||||
## 1.0.3(2026-07-30)
|
||||
|
||||
`est_tokens` 解耦(issue #2):一个常量此前被派了两份对"保守"定义相反的差事——TPM 入场预扣(押多了只是慢,安全)与 usage 缺失时的用量兜底(按上界记账只会账单虚高)。本次把两者拆开。
|
||||
|
||||
### 行为收紧/变更(下游请读)
|
||||
|
||||
- **`usage_source` 新增第三个值 `unavailable`。** 值域由 `measured`/`estimated` 两态变三态:`unavailable` 表示用量信息不可得(usage 帧缺失、失败尝试、终态失败),`estimated` 收窄为"有实测数字但可信度降级"(只剩打捞路径这一个生产者:收到 usage 帧但流被截断)。历史库里既有的 `estimated` 行语义不变、读兼容;按 `usage_source` 分支的下游代码需要认识新值。OCR 成功行**不受影响**,仍是 `measured`(0 token 是事实而非未知)。
|
||||
- **用量不可得的行,`cost` 由数值变 NULL。** 此前 usage 帧缺失时库拿 `est_tokens`(按定义是最坏情形上界)当实测值,又整块塞进 `completion_tokens` 换算——输出单价通常是输入的数倍,实测双重高估约 26 倍;`est_tokens=0` 时则算出 `0.0`,让"免费"与"未知"在数据上不可区分。现在这类行如实记 `0/0` + `unavailable` + `cost=NULL`。`SUM(cost)` 天然跳过 NULL,账目缺口用 `WHERE usage_source = 'unavailable' AND cache_hit = false` 量化(**`cache_hit` 限定不可省**:缓存命中行未产生新调用,cost 仍是事实上的 `0.0`,本无缺口)。成本汇总若此前依赖"cost 非空"的隐含假设,请复核。
|
||||
- **`est_tokens` 由必填降为可选调优覆盖。** 装配校验 `tpm > 0 ⇒ est_tokens > 0` 已删除——它把供应商配额(运维能从配额页抄到)与库的实现细节(预扣量,无人能正确取值)绑死。未填时库按 `max(1, tpm // 60)` 派生("一次调用约占一秒钟的配额份额",尺度无关:任何配额规模都收敛到约 60 个在途)。字段与 `{SCOPE}__{PROVIDER}__{N}__EST_TOKENS` 环境键**保留不删不改名**,显式填值仍然优先。此前为绕开该校验而把 `tpm` 限死为 0 的调用方,现可填真实 TPM。
|
||||
|
||||
## 1.0.2(2026-07-30)
|
||||
|
||||
1.0.1 的续作:那一版把三条跨字段守卫收进构造期后,独立验证发现 `from_env` 上还留着同一类的 15 条校验与 4 条规范化,一并收拢。
|
||||
|
||||
### 修复
|
||||
|
||||
- **后端选择与条件必填项在任何构造路径上都校验。** 以下此前只有 `from_env` 拦得住,`from_settings()` 与直接构造一律放行:`limiter_backend`/`breaker_backend`/`cache_backend`/`telemetry_backend`/`selector`/`quota_full` 六个字段的合法域;取 `redis` 的后端必须有 `redis_url`;启用缓存必须有 `cache_namespace` 与正 `cache_ttl_s`;`telemetry_backend` 取 `sqlite`/`postgres` 时对应的路径/DSN 必填;`structured_max_retries` 非负;`scope` 非空。
|
||||
- **`client.py` 五处断言的前提现在真的成立。** `assert settings.redis_url is not None # 内部不变量: config 已校验` 之类的注释此前在 `from_settings` 路上是假的:断言开启时抛不含任何字段信息的 `AssertionError`,`python -O` 下断言被移除、错误退化为 redis 库抛出的连接串解析异常。注释已改为点明由哪个校验方法保证。
|
||||
- **构造路补齐了 `from_env` 一直在做的规范化**,两条装配路对同一输入产出同一个值:
|
||||
- `scope` 小写并去空白。它直接进 Redis key(`pgw:limit:{scope}:…`、`pgw:gate:{scope}:…`),此前一个进程走 `from_env("LLM")` 拿到 `llm`、另一个直接构造传 `"LLM"`,**同一逻辑 scope 的限流与熔断状态会分裂到两套命名空间**,各记各的配额与熔断状态,分布式治理静默失效且不报错。
|
||||
- `redis_url`、`pricing_path` 的空串归 `None`。留着空串会骗过 `is None` 判断,把错误推迟成 redis 客户端的连接串解析异常或 `Is a directory: '.'`。
|
||||
- Postgres DSN 剥掉 SQLAlchemy 驱动后缀(`postgresql+asyncpg://…` 的 `+asyncpg` asyncpg 不认)。这一条剥的时候会发一条 warning——库动了调用方给的值,不该静默;日志只出现 scheme 段,DSN 带密码,整串不进日志。经 `from_env` 装配的不受影响也不会有这条 warning(`_load_pg_dsn` 早就剥干净了)。
|
||||
- **`EmbeddingSettings` 的 `batch_size` / `expected_dim` 域校验也移入构造期**,此前只有 `EmbeddingSettings.from_env` 校验,直接构造出 `batch_size=-3` 要到 `EmbeddingClient` 构造时才 fail-loud。
|
||||
|
||||
### 行为收紧(下游请读)
|
||||
|
||||
同 1.0.1:经 `from_env()` 装配的调用方**不受影响**。手工构造 `GatewaySettings` 或对它 `dataclasses.replace` 的调用方,若配置组合非法,现在会在构造期抛 `ValueError` 并点出字段名,而不是留到运行时表现为静默不建后端、裸 `AssertionError` 或第三方库的天书报错。
|
||||
|
||||
**一处静默改值需要留意**:此前手工构造传 `scope="LLM"`(非全小写)的调用方,升级后 scope 会被规范化为 `llm`,**Redis key 随之从 `pgw:limit:LLM:…` 切到 `pgw:limit:llm:…`**。这正是本次要修的问题——旧行为下这批 key 与 `from_env` 装配的进程根本不在同一命名空间;但切换发生的那一刻,旧键上的在途租约会被遗弃,靠 TTL 自愈。滚动升级期间建议留意限流配额短暂偏松。
|
||||
|
||||
## 1.0.1(2026-07-30)
|
||||
|
||||
### 修复
|
||||
|
||||
- **装配守卫在任何构造路径上都生效,不再只在 `from_env` 上。** 三条跨字段不变量(源 `timeout_s` ≤ `lease_ttl_s`、`stall_window_s` ≥ 最大源 TTFT、`probe_ttl_s` ≥ 最慢源 `timeout_s` + 5)原先只在 `GatewaySettings.from_env` 里校验,而装配有两条官方路——走 `from_settings()` 或直接构造能装出违反不变量的配置且不报错,故障留到运行时才表现为:租约先于请求过期使并发悄悄超出配额、正常慢首包被误判卡死掐断、半开探针在途即被接管。守卫已收进 `GatewaySettings.__post_init__`,与 `types.py` 各子配置一致,三个 client(Gateway/Ocr/Embedding)的全部工厂一并覆盖。
|
||||
- 新增 `sources` 非空校验。此前零源配置只在 `from_env` 路径被拦,直接构造可装出必然选源失败的 client。
|
||||
|
||||
### 行为收紧(下游请读)
|
||||
|
||||
直接构造 `GatewaySettings` 或对它做 `dataclasses.replace` 时,若上述组合非法,**现在会在构造期抛 `ValueError`**,而不是留到运行时。经 `from_env()` 装配的调用方**不受影响**——那条路本就跑这些守卫。手工拼配置(如从 YAML 读出后构造)的调用方若此前撞上过上述任一故障,升级后会在启动时立即得到点名字段的报错。
|
||||
|
||||
守卫报错文案的**补救建议**改为点字段名(`lease_ttl_s`、`backpressure.stall_window_s`、`breaker.probe_ttl_s`)。原文案已点出字段名,但建议部分给的是环境变量键(如"调大 `PGW_LEASE_TTL_S`"),而不走 env 的调用方从没设过那些键。键名映射见 `.env.example` 与 wiki `参考-配置键`。
|
||||
|
||||
## 1.0.0(2026-07-22)
|
||||
|
||||
首个正式版。统一 LLM/VLM/OCR/Embedding 调度与中转库,治理单位为一次模型调用;经 GovDoc-SaaS 与 CHSAnalyzer 两个真实项目全量迁移验收(ARCHITECTURE §11)。
|
||||
|
||||
@@ -112,6 +112,7 @@ project_root/
|
||||
| 三项目迁移文档(ARCHITECTURE §11 的展开,库设计的常驻约束) | `research-wiki/migrations/`(govdoc-saas / video-tree-trm5 / chsanalyzer) |
|
||||
| 功能设计文档(每次实现新功能时新增) | `research-wiki/designs/` |
|
||||
| 实现计划 | `research-wiki/plans/` |
|
||||
| **用户文档站**(Gitea Wiki,Diátaxis 四区)结构/更新时机/写作纪律 | `research-wiki/docs-convention.md`;**发版或公共行为变更必须按其 §2 清单同步 wiki 与 CHANGELOG,版本 bump 提交不得裸发** |
|
||||
| 治理网关参考实现 | `reference/Video-Tree-TRM5/adapters/`(llm/breaker/streaming/redis_cache/telemetry) |
|
||||
| 分布式限流/熔断参考实现 | `reference/CHSAnalyzer/app/coordination/`(limiter+Lua/provider_gate)与 `app/providers/governance.py` |
|
||||
| 错误分类参考 | `reference/CHSAnalyzer/app/domain/errors.py` |
|
||||
|
||||
@@ -0,0 +1,203 @@
|
||||
# PolyGateway
|
||||
|
||||
实验室统一的大语言模型调度与中转库:LLM / VLM / OCR / Embedding 四类调用共用同一套生产级治理栈——多源多账号、限流、错误分类重试、熔断、响应缓存、流式看门狗、遥测与成本。治理单位是**一次模型调用**;任务编排、业务解析、图像预处理都留在业务侧。
|
||||
|
||||
> 由三个真实项目(GovDoc-SaaS / CHSAnalyzer / Video-Tree-TRM5)各自手写的治理栈提炼而来,并以"能否全量迁移回这三个项目"作为验收标准。v1.0.0 已通过 GovDoc 与 CHSAnalyzer 两项目的全量迁移验收(约 −6800 行项目侧治理代码由本库继任)。
|
||||
|
||||
## 为什么需要它
|
||||
|
||||
每个接入大模型的项目都会重写同一批东西:重试循环、429 处理、熔断器、SSE 解析、遥测埋点——写三遍就有三份 bug。本库把这些收敛为一份经过压测验证的实现:
|
||||
|
||||
| 能力 | 说明 |
|
||||
|---|---|
|
||||
| 多源多账号 | `{SCOPE}__{PROVIDER}__{N}__*` 配置任意多源;健康感知选源(EWMA×在途 P2C)自动避开坏源 |
|
||||
| 限流 | 并发/RPM/TPM × 全局/单源六道闸;TPM 预扣入场、按实际用量结算退款;Redis 后端跨进程原子(Lua) |
|
||||
| 错误分类重试 | 一切失败落入四分类(见下),由分类决定重试/换源/熔断;429 属 pushback 不消耗重试预算;退避含 jitter 且尊重 Retry-After |
|
||||
| 熔断 | 双通道(连续失败 + 失败率窗口,健康证据抑制误熔);半开单探针带租约(持有者死亡自动回收);epoch fencing 拒绝迟到写回;开路时长指数递增 |
|
||||
| 自适应并发 | AIMD:429 削减、成功缓升,防止打爆上游 |
|
||||
| 响应缓存 | Redis/内存;key 含 model + messages 摘要 + namespace/租户 + salt,多模态 content 先摘要再 hash(防毒化);可 per-call 绕过(科研重采样) |
|
||||
| 流式看门狗 | TTFT / inter-token / 总超时三层活性;thinking token 刷活性不计结果;截断流(缺 `[DONE]`)判瞬时不入缓存 |
|
||||
| 遥测与成本 | 每次调用(含缓存命中与失败)必录 18 字段;SQLite / Postgres 后端;按价格表折算成本;多模态内容摘要落库不存原图 |
|
||||
| 结构化输出 | json_repair 修复 / 原生 schema 双策略 + 校验失败有界带反馈重问 |
|
||||
| OCR | MonkeyOCR 双端点(文本转录 + 版面解析),bbox 数值防御下沉,逐源健康预检 `check_health()` |
|
||||
| Embedding | 分批、维度校验、与 chat 同一治理栈 |
|
||||
|
||||
**降级方向是铁律**:缓存/遥测后端掉线 → 静默降级(warning);限流/熔断后端掉线 → 报错而非放行(防击穿上游)。`asyncio.CancelledError` 全链路穿透,in-flight 资源在 finally 释放。
|
||||
|
||||
## 安装
|
||||
|
||||
发布在实验室 Gitea PyPI(公开包,匿名可装):
|
||||
|
||||
```bash
|
||||
pip install --extra-index-url https://gitea.iomgaa.online/api/packages/iomgaa/pypi/simple/ \
|
||||
"polygateway[redis,postgres,structured]==1.0.*"
|
||||
```
|
||||
|
||||
核心仅依赖 `httpx` + `pydantic`;按需选 extras:
|
||||
|
||||
| extra | 内容 | 何时需要 |
|
||||
|---|---|---|
|
||||
| `redis` | redis-py | Redis 限流/熔断/缓存后端 |
|
||||
| `postgres` | asyncpg | Postgres 遥测后端 |
|
||||
| `structured` | json-repair | 结构化输出的修复策略 |
|
||||
| `sdk` | openai | 可选的 SDK transport(默认手写 httpx,不需要) |
|
||||
|
||||
要求 Python ≥ 3.11。
|
||||
|
||||
## 快速开始
|
||||
|
||||
### 1. 配置 `.env`
|
||||
|
||||
```bash
|
||||
LLM__MINIMAX__1__BASE_URL=https://your-gateway/v1
|
||||
LLM__MINIMAX__1__API_KEY=sk-xxx
|
||||
LLM__MINIMAX__1__MODEL=MiniMax-M3
|
||||
LLM__MINIMAX__1__TIMEOUT_S=120
|
||||
LLM_MAX_RETRIES=3
|
||||
LLM_RETRY_BASE_DELAY=2.0
|
||||
LLM_RETRY_MAX_DELAY=30.0
|
||||
LLM_CIRCUIT_BREAKER_THRESHOLD=5
|
||||
LLM_CIRCUIT_BREAKER_COOLDOWN=60
|
||||
PGW_LIMITER_BACKEND=memory
|
||||
PGW_BREAKER_BACKEND=memory
|
||||
PGW_CACHE_BACKEND=none
|
||||
PGW_TELEMETRY_BACKEND=none
|
||||
```
|
||||
|
||||
缺任何关键键都会在装配时报错——本库禁止默认值兜底掩盖配置缺失。
|
||||
|
||||
### 2. 发起治理调用
|
||||
|
||||
```python
|
||||
from polygateway import GatewayClient
|
||||
|
||||
async def main() -> None:
|
||||
client = GatewayClient.from_env("LLM") # 读 .env 装配整套治理栈
|
||||
try:
|
||||
resp = await client.chat([{"role": "user", "content": "你好"}])
|
||||
print(resp.content, resp.source_name, resp.latency_ms)
|
||||
finally:
|
||||
await client.aclose() # 归还连接与治理后端资源
|
||||
```
|
||||
|
||||
`chat()` 原生接受 OpenAI 多模态 content 数组(`image_url` data URL),VLM 调用无需专门客户端;`session_id` / `parent_call_id` / `cache_salt` 关键字参数用于链路追踪与缓存控制;`overlay` 传采样参数(`temperature` / `seed` / `max_tokens` 等,恒定值宜配在源的 `EXTRA_BODY` 上)——它会进缓存 key,故逐次变化的 `seed` 天然不命中缓存。
|
||||
|
||||
### 3. OCR 与 Embedding
|
||||
|
||||
```python
|
||||
from polygateway import EmbeddingClient
|
||||
from polygateway.ocr import OcrClient
|
||||
|
||||
ocr = OcrClient.from_env("OCR") # OCR__MONKEY__1__* 多源
|
||||
text = await ocr.recognize_text(image_bytes) # 文本转录
|
||||
layout = await ocr.parse_layout(image_bytes) # 版面解析(带 bbox 的元素列表)
|
||||
health = await ocr.check_health() # 逐源预检 {"monkey_1": True, ...}
|
||||
|
||||
embed = EmbeddingClient.from_env("EMBED") # EMBED__*__* + EMBED__BATCH_SIZE
|
||||
vectors = (await embed.embed(["文本 a", "文本 b"])).vectors
|
||||
```
|
||||
|
||||
### 4. 业务侧异常处理
|
||||
|
||||
```python
|
||||
from polygateway import GatewayUnavailableError, RequestRejectedError
|
||||
|
||||
try:
|
||||
resp = await client.chat(messages)
|
||||
except GatewayUnavailableError as exc:
|
||||
# 整个 scope 暂时无源可用: 延期重投,不消耗业务失败预算
|
||||
schedule_retry(after_s=exc.retry_after_s) # exc.reason / exc.per_source_reasons 供诊断
|
||||
except RequestRejectedError:
|
||||
... # 请求本身有问题(400/格式拒绝): 不重试,直接失败
|
||||
```
|
||||
|
||||
## 错误模型(四分类)
|
||||
|
||||
一切失败在 transport 层翻译为四类之一,治理行为由分类决定,业务侧不需要判断状态码:
|
||||
|
||||
| 分类 | 含义 | 库内行为 |
|
||||
|---|---|---|
|
||||
| `TransientError` | 超时/5xx/网络抖动/截断流 | 换源重试 + 退避 |
|
||||
| `SourceDeadError` | 401/403/欠费(429+insufficient_quota) | 立即熔断该源 + 换源 |
|
||||
| `RequestRejectedError` | 400/内容拒绝/本地格式拒绝 | 不重试不换源,快速失败 |
|
||||
| `ResultInvalidError` | 调用成功但结果不合格(坏 JSON/维度不符/坏 bbox) | 不熔断("坏结果 ≠ 坏服务"),按策略有界重问或上抛 |
|
||||
|
||||
预算耗尽/全源熔断时抛 `GatewayUnavailableError` 族(`CircuitOpenError` / `AllSourcesExhausted`),携带 `scope` / `reason` / `retry_after_s` / `per_source_reasons`,供任务队列做延期重投。
|
||||
|
||||
## 配置参考
|
||||
|
||||
配置只有两条装配路径:`from_env()`(读 `.env`/环境变量)或构造函数全量注入(测试/高级);库内部任何组件不自读环境变量。键名全集见 [.env.example](.env.example),约定速览:
|
||||
|
||||
| 键形态 | 作用 |
|
||||
|---|---|
|
||||
| `{SCOPE}__{PROVIDER}__{N}__{FIELD}` | 第 N 个源;FIELD ∈ BASE_URL/API_KEY/MODEL/TIMEOUT_S/MAX_CONCURRENCY/RPM/TPM/EST_TOKENS/TTFT_TIMEOUT_S/INTER_TOKEN_TIMEOUT_S/ENABLE_THINKING/TRUST_ENV |
|
||||
| `{SCOPE}__GLOBAL__*` | scope 级全局限额(跨源并发/RPM/TPM) |
|
||||
| `{SCOPE}__RETRY__*` / `BREAKER__*` / `BACKPRESSURE__*` / `SELECTOR` | per-scope 韧性参数;缺省回落平铺键(`LLM_MAX_RETRIES` 等,兼容旧项目习惯) |
|
||||
| `PGW_LIMITER_BACKEND` / `PGW_BREAKER_BACKEND` | `memory`(单进程)或 `redis`(跨进程共享,需 `REDIS_URL`) |
|
||||
| `PGW_CACHE_BACKEND` | `none` / `redis`(需 `PGW_CACHE_NAMESPACE` + `PGW_CACHE_TTL_S`) |
|
||||
| `PGW_TELEMETRY_BACKEND` | `none` / `sqlite`(需 `PGW_TELEMETRY_SQLITE_PATH`)/ `postgres`(需 `PGW_TELEMETRY_PG_DSN`) |
|
||||
|
||||
`SCOPE` 是逻辑角色(LLM/VLM/OCR/EMBED/JUDGE/SEARCH…任意大写名),同一进程可按角色装配多个 client,各自独立配置与治理状态。
|
||||
|
||||
## 架构
|
||||
|
||||
端口适配器 + 中间件洋葱:决策逻辑一份,状态存储可插拔。
|
||||
|
||||
```mermaid
|
||||
graph LR
|
||||
A[业务代码] --> B[GatewayClient]
|
||||
B --> C[缓存 MW] --> D[遥测 MW] --> E[重试/选源/限流/熔断 MW]
|
||||
E --> F[Transport httpx]
|
||||
F --> G[(上游网关)]
|
||||
E -.端口.-> H[(内存 / Redis 后端)]
|
||||
D -.端口.-> I[(SQLite / Postgres)]
|
||||
```
|
||||
|
||||
| 模块 | 职责 |
|
||||
|---|---|
|
||||
| `types.py` / `errors.py` / `ports.py` | 内核:冻结类型、四分类异常、全部 Protocol(最内层,不依赖任何实现) |
|
||||
| `middleware/` | 治理算法(重试/限流/熔断/缓存/遥测),只面向端口 |
|
||||
| `transports/` | 协议细节:OpenAI 兼容 SSE、MonkeyOCR 双端点;错误翻译在此层 |
|
||||
| `backends/` | 限流/熔断/缓存的内存与 Redis 实现(同一契约测试套件双后端共用) |
|
||||
| `telemetry/` | SQLite / Postgres 遥测后端 |
|
||||
| `structured/` | 结构化输出策略 |
|
||||
|
||||
依赖纪律由 import-linter 机械化执法(`make lint`)。完整架构决策(D1-D14 含论证过程)见 [research-wiki/ARCHITECTURE.md](research-wiki/ARCHITECTURE.md)。
|
||||
|
||||
## 可靠性证据
|
||||
|
||||
行为不是宣称出来的,是压测出来的(数字见 `research-wiki/findings/`):
|
||||
|
||||
| 场景 | 结果 |
|
||||
|---|---|
|
||||
| 故障混编 soak(坏 key/黑洞/慢源/限流源混合,8000 调用) | 成功率 98.96%,坏源吸流被压制,真实源零误熔 |
|
||||
| OCR 故障池 soak(1500 调用,redis 双后端跨进程) | 成功率 99.73%,13 项不变量全过(租约归零/探针不悬挂/零取消泄漏等) |
|
||||
| 两项目全量迁移回归 | 原测试全绿 + 真实链路冒烟 + 50 样本批跑 100% 解析 |
|
||||
|
||||
时间语义测试(租约过期、窗口滚动、半开探针)全部真实等待不缩放;Redis/Postgres 测试打真实实验室后端,不 mock Lua。
|
||||
|
||||
## 开发
|
||||
|
||||
```bash
|
||||
conda create -n PolyGateway python=3.11 && conda activate PolyGateway
|
||||
make install # editable 安装(dev + 全部 extras)
|
||||
make test # pytest + 覆盖率(目标 ≥80%)
|
||||
make lint # ruff + import-linter
|
||||
make ci # 只读全量验证
|
||||
```
|
||||
|
||||
测试组织:`tests/{unit,integration,e2e}` + 双后端契约测试;并发/取消/降级方向是一等测试对象。压测 harness 在 `tools/soak/`。贡献流程与项目纪律见 [CLAUDE.md](CLAUDE.md)。
|
||||
|
||||
## 文档导航
|
||||
|
||||
| 想了解 | 看 |
|
||||
|---|---|
|
||||
| 全部架构决策及理由(单一事实源) | `research-wiki/ARCHITECTURE.md` |
|
||||
| 里程碑与状态 | `research-wiki/ROADMAP.md` |
|
||||
| 项目迁移指南(删除清单/组件映射/行为审计) | `research-wiki/migrations/` |
|
||||
| 每个功能的设计与验收记录 | `research-wiki/designs/`、`research-wiki/findings/` |
|
||||
| 版本变更 | [CHANGELOG.md](CHANGELOG.md) |
|
||||
|
||||
## 兼容性承诺
|
||||
|
||||
`LLMResponse` 等被下游消费的公共类型,字段**只增不删不改名**且新增字段必带默认值;`{SCOPE}__{PROVIDER}__{N}__{FIELD}` 与平铺韧性键名(`LLM_TIMEOUT` 等)沿用三项目既有习惯,不做破坏性改名。实验室内部库,随实验室项目需求演进。
|
||||
@@ -0,0 +1,6 @@
|
||||
{
|
||||
"MiniMax-M3": {
|
||||
"input_per_1m": 2.1,
|
||||
"output_per_1m": 8.4
|
||||
}
|
||||
}
|
||||
+1
-1
@@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"
|
||||
|
||||
[project]
|
||||
name = "polygateway"
|
||||
version = "1.0.0"
|
||||
version = "1.0.5"
|
||||
description = "PolyGateway:实验室统一的大语言模型(LLM/VLM/OCR)调度与中转库——多源、限流、重试、熔断、缓存、遥测"
|
||||
requires-python = ">=3.11"
|
||||
dependencies = [
|
||||
|
||||
@@ -302,7 +302,7 @@ flowchart TB
|
||||
### 4.4 一次调用的生命周期(walkthrough)
|
||||
|
||||
1. **缓存命中**: TelemetryMW 记录(cache_hit=True, latency_ms=0)→ CacheMW 返回,不触达任何更内层。
|
||||
2. **正常路径**: RetryMW 开始第一次尝试 → selector 选源(跳过冷却中的源)→ 该源熔断门(闭路)→ 限流 acquire permit(全局+该源,并发/RPM/TPM 三闸,token 按 `est_tokens` 预扣)→ transport 发请求、流式解析(看门狗包裹)、收 usage 帧 → permit 按实际 usage settle(多退少补)→ 回程写缓存 → 遥测记成功(含 ttft/max_inter_token/成本)。
|
||||
2. **正常路径**: RetryMW 开始第一次尝试 → selector 选源(跳过冷却中的源)→ 该源熔断门(闭路)→ 限流 acquire permit(全局+该源,并发/RPM/TPM 三闸,token 按**有效预扣量** `SourceConfig.effective_est_tokens()` 预扣,取值规则见 §7.7)→ transport 发请求、流式解析(看门狗包裹)、收 usage 帧 → permit 按实际 usage settle(多退少补)→ 回程写缓存 → 遥测记成功(含 ttft/max_inter_token/成本)。
|
||||
3. **瞬时错误**(超时/5xx/429/SSE 异常): transport 翻译为 `TransientError` → RetryMW 指数退避+jitter(取 Retry-After 提示与退避的较大值)后换源重试;每次尝试独立 call_id、独立过限流闸、失败即报熔断计数与遥测。
|
||||
4. **源死亡**(401/403/欠费): `SourceDeadError` → 该源熔断 force_open + 本地冷却备忘 → 立即换下一源,不退避等待。
|
||||
5. **请求被拒**(400/坏输入): `RequestRejectedError` → 不重试不换源,直接上抛;遥测记录。
|
||||
@@ -322,13 +322,36 @@ flowchart TB
|
||||
| `content` | str | 正式输出文本 |
|
||||
| `thinking` | str | 思考流内容(reasoning_content / think 标签,按 provider 注册表提取) |
|
||||
| `model` / `provider` | str | 溯源 |
|
||||
| `prompt_tokens` / `completion_tokens` | int | usage 帧读取;缺失时按估算标注 |
|
||||
| `prompt_tokens` / `completion_tokens` | int | usage 帧读取;帧缺失时记 `0/0` 并由 `usage_source` 标注不可得(不编造估算值,见下) |
|
||||
| `latency_ms` | int | 总延迟 |
|
||||
| `ttft_ms` / `max_inter_token_ms` | float? | 流式活性测量 |
|
||||
| `cache_hit` | bool | 是否缓存命中 |
|
||||
| `call_id` | str | UUID,每次**尝试**独立 |
|
||||
|
||||
新增字段(库扩展,全部带默认值): `source_name`(多源溯源)、`cost`(pricing 换算,可为 None)、`usage_source`(measured/estimated)、`structured_data`(D14 阶梯通过后的解析产物;不参与缓存序列化,命中时由 CacheMW 复用 strategy 零网络重建)。
|
||||
新增字段(库扩展,全部带默认值): `source_name`(多源溯源)、`cost`(pricing 换算,可为 None)、`usage_source`(三态,见下)、`structured_data`(D14 阶梯通过后的解析产物;不参与缓存序列化,命中时由 CacheMW 复用 strategy 零网络重建)、`cached_prompt_tokens` 与 `model_reported`(2026-07-31,issue #3,见下)。
|
||||
|
||||
**可观测字段(2026-07-31,issue #3;下游 dissect 的调用审计需求)**:
|
||||
|
||||
| 字段 | 含义 | 生产者 |
|
||||
|---|---|---|
|
||||
| `cached_prompt_tokens` | **供应商侧** prompt cache 命中的输入 token 数(OpenAI 兼容格式的 `usage.prompt_tokens_details.cached_tokens`)。`None` = 该源未上报;`0` = 上报了一次真实零命中——两者对下游处置不同(前者不可做缓存成本校正),故不可混同 | `openai_compat` 两条路径解析后经 `TransportResult` 上浮 |
|
||||
| `model_reported` | API 响应体里的 `model` 字段;`None` = 未上报。与 `model`(`.env` 配置别名)可能分叉——供应商把别名指向新权重时,实验复现必须认这个串 | 流式取首个含 `model` 的 chunk(首次写入即固定),非流式取 body 顶层 |
|
||||
|
||||
`cache_hit` 指的始终是 **PolyGateway 自身响应缓存**,与供应商 prompt cache 无关;两者语义不同但名字相近,docstring 已消歧(改名会破坏迁移兼容,故只注释)。
|
||||
|
||||
**缓存命中行的口径(决策 B1)**: 与 `model`/`prompt_tokens` 同一规则——`CacheMW._rehydrate` 只覆写与本次调用相关的时序字段,这两个新字段**原样回放**历史值。故**统计供应商缓存命中率必须写 `WHERE cache_hit = false`**,否则回放行会被重复计数(与 §5.1 `cost` 缺口口径同款教训)。
|
||||
|
||||
**`usage_source` 三态值域(2026-07-30,est_tokens 解耦设计;此前为 measured/estimated 两态)**:
|
||||
|
||||
| 值 | 含义 | 生产者 | cost |
|
||||
|---|---|---|---|
|
||||
| `measured` | usage 帧完整可信 | 正常路径;OCR 成功行(0 token 是**事实**而非未知) | 按 token 换算 |
|
||||
| `estimated` | 有实测数字但可信度降级 | 打捞路径(收到 usage 帧但流被截断,§7.1) | 按 token 换算 |
|
||||
| `unavailable` | 用量信息不可得 | usage 帧缺失、失败尝试、终态失败 | **NULL** |
|
||||
|
||||
值域在 `types.py` 以模块级 frozenset 常量 `USAGE_SOURCES` 落地,**仅约束库内生产侧**(所有写入点从该常量取值),不在 `LLMResponse`/`Usage`/`TransportResult` 上加 `__post_init__` 值域校验——它们是运行时构造点,裸 `ValueError` 不属 §6 四分类、`RetryMW` 不捕会逃出 `chat()`;且 `LLMResponse` 是三项目已消费的公共类型,新增运行时校验属下游可见行为变更。历史库里既有的 `estimated` 行在新值域中依然合法可读。
|
||||
|
||||
**cost 口径的不变式**: **产生了真实网关调用、但用量不可得的行 → `cost` 为 NULL**(不再算出一个假的 `0.0` 把"免费"与"未知"混为一谈)。**缓存命中行不在此列**——`cache_hit=True` 时 cost 仍为 `0.0`,因为未产生新调用,`0.0` 是事实而非未知;`TelemetryEmitter` 里 `unavailable → None` 的短路**插在 `cache_hit` 分支之后**正是为此。故账目缺口的度量口径必须写成 `WHERE usage_source = 'unavailable' AND cache_hit = false`,漏掉后半个条件会把本无缺口的缓存命中行灌进来,度量偏高。
|
||||
|
||||
**API 稳定性约定(2026-07-20,迁移文档反向约束)**: ① 公共类型新增字段必须带默认值——三项目测试中逐字段传参的 fake 构造才能零改动;② 错误四分类从 `polygateway` 顶层命名空间导出——业务侧步级重试要引用它们(GovDoc/Video-Tree 现有 `(TimeoutError, OSError)` 异常元组迁移后会**静默失效**,必须显式替换为库异常);③ `GatewayClient` 提供显式 `aclose()` 与 async context manager 生命周期 API;④ 被取消的调用尽力而为记遥测(error="cancelled",finally 中记录,绝不因遥测延迟取消传播,写失败静默)。
|
||||
|
||||
@@ -338,6 +361,8 @@ flowchart TB
|
||||
|
||||
**`chat()` 公共签名定稿(2026-07-20,GovDoc 迁移缺口 G1/G2)**: `chat(messages, *, session_id=None, parent_call_id=None, cache_salt=None, cache_namespace=None, structured=None, stream=True)`。要点: ① `session_id`/`parent_call_id` 与三项目现有 `LLMProvider.chat` Protocol 逐字兼容——这是"调用点零改动"承诺的前提;② **per-call `cache_namespace`**: GovDoc 是单 client 服务多租户、tenant 每请求变化,装配级 namespace 只是默认值,per-call 传入时覆盖并进入缓存 key(§7.5);③ `cache_salt` per-call 可传(Video-Tree 跨 epoch 重采样);④ `structured` 三档语义(D14),类型定稿 `type[BaseModel] | Literal["json"] | None`(M1 设计): 不传 = 原始文本,`"json"` = 仅修复,pydantic 模型 = 完整阶梯(修复+形态校验+有界带反馈重问)。
|
||||
|
||||
**`overlay` 追加(2026-07-31,issue #4)**: 签名末尾增 `overlay: Mapping[str, Any] | None = None`,承载采样参数(`temperature`/`seed`/`max_tokens` 等)。带默认值的 keyword-only 参数不改变既有调用点,"签名冻结"承诺不破。要点: ① 优先级 **结构化注入 > 调用级 overlay > 源级 `extra_body`**,由 `StructuredMW` 的 `{**request.overlay, **strategy_overlay}` 与 transport `_build_payload` 的 update 顺序天然给出,无新机制;② 保护键 `{model, messages, stream, stream_options}` 与不可 JSON 序列化的值在**进洋葱之前**报 `ValueError`(前者被覆盖会击穿成本换算/缓存口径/流式看门狗/usage 帧,后者会在 `CacheMW` 的降级 try 之外抛裸 `TypeError` 且一行遥测都没有);③ 同时填 `ChatRequest.sampling` 快照字段——`overlay` 在洋葱不同深度取值不同(内层含 `response_format`),缓存 key 与遥测需要一个跨层恒定的读取点,否则同一列在不同行口径分叉。
|
||||
|
||||
---
|
||||
|
||||
## 6. 错误模型
|
||||
@@ -381,7 +406,7 @@ flowchart TB
|
||||
|
||||
**职责**: 一次原始调用的全部协议细节——请求体组装(含 provider 注册表注入的 thinking 参数)、发送、流式 SSE 解析(增量 content/reasoning_content、usage 帧、[DONE] 检测)、HTTP/线路错误按 §6.2 翻译。**不含**重试/限流/缓存(那是中间件的事)。
|
||||
|
||||
- `OpenAICompatTransport`(默认): httpx.AsyncClient(每源一个,预配 Authorization 与分段超时),SSE 解析移植三项目的模块级纯函数;强制 `stream_options.include_usage`。**SSE 缺 [DONE] 语义(2026-07-20 M1 设计)**: per-source `missing_done: "retry" | "salvage"`,默认 retry(防截断响应进缓存被固化);零内容提前断流(early_eof)恒 retry 不可配;打捞路径强制 `usage_source="estimated"`。CHS 迁移配 salvage 保留其现状行为。看门狗活性口径: 任何增量(content 或 reasoning_content)都算 token——ttft = 首个任意 token,思考流刷新 inter_token 计时(CHS 迁移约束 R1)。**非流式快路径**: 短请求可配 `stream=False`(三项目都写死 stream=True 强迫短请求走 SSE+看门狗,库放开)。
|
||||
- `OpenAICompatTransport`(默认): httpx.AsyncClient(每源一个,预配 Authorization 与分段超时),SSE 解析移植三项目的模块级纯函数;强制 `stream_options.include_usage`。**SSE 缺 [DONE] 语义(2026-07-20 M1 设计)**: per-source `missing_done: "retry" | "salvage"`,默认 retry(防截断响应进缓存被固化);零内容提前断流(early_eof)恒 retry 不可配;打捞路径**仅在收到 usage 帧时**把 `measured` 降级为 `estimated`(**勘误 2026-07-30**: 原文"强制 estimated" 已改为有条件——没收到 usage 帧时用量本就是 `unavailable`,强制标 `estimated` 会让 `0/0` 被当作实测数字换算出一个假的 `0.0` 成本,§5.1)。CHS 迁移配 salvage 保留其现状行为。看门狗活性口径: 任何增量(content 或 reasoning_content)都算 token——ttft = 首个任意 token,思考流刷新 inter_token 计时(CHS 迁移约束 R1)。**非流式快路径**: 短请求可配 `stream=False`(三项目都写死 stream=True 强迫短请求走 SSE+看门狗,库放开)。
|
||||
- `OpenAISDKTransport`(可选 extra): 薄封装,`max_retries=0` 关掉 SDK 自带重试(治理归中间件),`extra_body`/`model_extra` 通道非标字段。
|
||||
- `MonkeyOcrTransport`: 见 §7.10。
|
||||
|
||||
@@ -412,11 +437,13 @@ flowchart TB
|
||||
|
||||
### 7.5 响应缓存
|
||||
|
||||
**key 公式**: `sha256(canonical_json({model, messages_digest, namespace, salt}))`,前缀 `pgw:cache:`。
|
||||
**key 公式**: `sha256(canonical_json({model, messages_digest, namespace, salt, sampling}))`,前缀 `pgw:cache:`。
|
||||
|
||||
- `messages_digest`: 文本部分原文参与;多模态 content part(base64 图像等)先各自 sha256 摘要再参与——修正 Video-Tree 把整段 base64 进 hash 的开销问题,且 key 稳定性不变。
|
||||
- `namespace`: 必填(项目名/租户 id),修正 GovDoc 缓存 key 缺租户隔离与多项目共用 Redis 时的互相毒化风险。
|
||||
- `salt`: 可选,跨 epoch 强制重采样(Video-Tree 需求)。
|
||||
- `sampling`(2026-07-31,issue #4): 调用级采样参数,**仅非空时参与**(注意与 `salt` 的"仅非 None"不同——空串是有意义的 salt,而空采样参数与不传无差别),故空 overlay 时旧键逐字不变、存量缓存不冷启动。读 `request.sampling` 而非 `request.overlay`,不依赖"CacheMW 恰在 StructuredMW 外侧"的层序巧合。**不进 key 的后果**: 同 messages 跑 5 个 seed 会全部命中第一次的响应,标准差恒为 0 且不报错——受控实验静默作废。源级 `extra_body` 同理并入 `model_fingerprint`(全源皆空时字面量不变,否则追加 `|sha256(...)`,摘要对象是各源 `(model, extra_body)` 的 canonical JSON 排序去重——按模型而非源名,改源名不误触冷启动)。
|
||||
- **两条已知副作用**: ① 逐 rollout 变化的 `seed` 进 key 后该路径天然全部 miss(正确语义,但缓存对它不再省钱);② `model_fingerprint` 是**集合级**指纹而非本次选中源的指纹,同 scope 各源 `extra_body` 不同时仍可能返回另一源的响应(既有取舍的延续,与 `model` 同),要求逐源可复现应让每源独享 scope 或 namespace。
|
||||
- value = `LLMResponse` 的 JSON;TTL 必填且 > 0(禁止永不过期,继承 Video-Tree 校验);Redis 不可用 → get 返回 None、set 吞异常记 warning(静默降级)。**只缓存成功响应**;`ResultInvalidError` 的原始响应不缓存(避免固化坏结果)。
|
||||
|
||||
### 7.6 流式活性看门狗
|
||||
@@ -425,17 +452,27 @@ flowchart TB
|
||||
|
||||
### 7.7 多源与选源
|
||||
|
||||
`SourceConfig`: name/provider/base_url/api_key/model/超时组/限额组(单源并发/RPM/TPM)/`est_tokens`(TPM 预扣常量,亦作 usage 缺失时的保守兜底,移植 CHS `config.py:55`;2026-07-20 缺口 G2 补)/enable_thinking。聚合自环境变量 `{SCOPE}__{PROVIDER}__{N}__{FIELD}`(§9)。`SourceSelector` 端口: `health_aware`(M2.5 新缺省: 成功率 EWMA / (1+在途) 的 P2C,0.05 探索地板,进程本地健康态,可选 `OutcomeAwareSelector` 扩展喂数)/ `round_robin` / `least_inflight`。**逻辑角色**: Video-Tree 式 SEARCH/JUDGE/VL/EVOLVE 多角色 = 命名的 client 配置组,`from_env()` 支持按角色前缀装配多个 client;禁止两个角色静默共享同一实例却在配置上看似独立(Video-Tree `evolve_llm = llm` 别名的教训——共享必须显式)。
|
||||
`SourceConfig`: name/provider/base_url/api_key/model/超时组/限额组(单源并发/RPM/TPM)/`est_tokens`(TPM 预扣量的**可选调优覆盖**,移植 CHS `config.py:55`;2026-07-20 缺口 G2 补,2026-07-30 由必填降为可选)/enable_thinking/`extra_body`(2026-07-31 issue #4: 本源恒定的采样参数,构造期校验保护键后转 `MappingProxyType`;**该字段令 SourceConfig 不再 hashable**——加任何 mapping 字段的固有代价,库内无以源作 dict key/set 元素的写法,要可变副本用 `dict(...)`、要改字段用 `dataclasses.replace`)。聚合自环境变量 `{SCOPE}__{PROVIDER}__{N}__{FIELD}`(§9)。
|
||||
|
||||
**TPM 有效预扣量(2026-07-30,est_tokens 解耦设计,G2 闭环)**: `try_acquire`(§7.3)传入的 est 来自 `SourceConfig.effective_est_tokens()` 这一份纯方法,五个调用点(`QuotaGate` 入场 + chat/embedding 各自的成功侧与失败侧结算)共用,保证预扣与结算恒取同一值(`delta == 0`,否则押金会被整笔退回、TPM 闸退化成进门即放行)。规则:显式 `est_tokens > 0` 则原样用;否则 `tpm > 0` 时派生 `max(1, tpm // 60)`;`tpm == 0`(该闸不启用)时为 0。
|
||||
|
||||
派生取 `tpm // 60` 的理由是**尺度无关**:任何配额规模都给出同一行为上限——"一次调用约占一秒钟的配额份额",故 `tpm=6000` 与 `tpm=600000` 都收敛到约 60 个在途。固定常量(如 1000)则与配额规模无关,在途上限随配额乱飘且取值无从解释。`est_tokens` **不再兼任 usage 缺失时的用量兜底**:那两份差事对"保守"的定义方向相反——限流语境下押多了只是慢(安全),计费语境下按上界记账只会系统性虚高(库把遥测拆成 prompt/completion 两列后又整块塞进 completion,而输出单价通常是输入的数倍,实测双重高估约 26 倍)。用量不可得现在如实记 `unavailable` + cost NULL(§5.1)。
|
||||
|
||||
**已知限制(既有行为,本次未修)**: 单源 `tpm == 0` 而 `{SCOPE}__GLOBAL__TPM > 0` 时,派生值为 0,全局 TPM 闸拿 0 预扣、入场保护形同虚设。修它需要把 `GlobalLimits` 注入 `QuotaGate`(改三处装配),属独立议题。`SourceSelector` 端口: `health_aware`(M2.5 新缺省: 成功率 EWMA / (1+在途) 的 P2C,0.05 探索地板,进程本地健康态,可选 `OutcomeAwareSelector` 扩展喂数)/ `round_robin` / `least_inflight`。**逻辑角色**: Video-Tree 式 SEARCH/JUDGE/VL/EVOLVE 多角色 = 命名的 client 配置组,`from_env()` 支持按角色前缀装配多个 client;禁止两个角色静默共享同一实例却在配置上看似独立(Video-Tree `evolve_llm = llm` 别名的教训——共享必须显式)。
|
||||
|
||||
**多 client 共享状态后端(2026-07-20,VT 迁移缺口 R5)**: 限流/熔断状态的 key 以 scope+source 为单位,与 client 实例解耦;多个逻辑角色的 client **显式注入同一个状态后端实例**时即共享全局并发/RPM/TPM 闸(Video-Tree `TREE_BUILD_API_CONCURRENCY` 跨 SEARCH+VL 共享 semaphore 的语义由此承接)。共享必须显式注入,禁止隐式全局。
|
||||
|
||||
### 7.8 遥测与成本
|
||||
|
||||
**必录字段**(继承三项目 15 字段规范): call_id、parent_call_id、session_id、model、provider、source_name、messages(JSON)、response、thinking、prompt_tokens、completion_tokens、usage_source、latency_ms、ttft_ms、max_inter_token_ms、cache_hit、error、**cost**。链路: `session_id`/`parent_call_id` 由调用方传入贯穿(agent step → LLM call)。`messages` 落库前对多模态 part 先摘要(与缓存 key 共用同一摘要函数,§7.5)——Video-Tree 现状 base64 整段进 SQLite 导致 db 膨胀(`llm.py:330`),库内修复(2026-07-20,VT 迁移缺口 R12)。
|
||||
**必录字段**(继承三项目 15 字段规范): call_id、parent_call_id、session_id、model、provider、source_name、messages(JSON)、response、thinking、prompt_tokens、completion_tokens、usage_source、latency_ms、ttft_ms、max_inter_token_ms、cache_hit、error、**cost**、**cached_prompt_tokens**、**model_reported**、**sampling**。
|
||||
|
||||
**`sampling` 列(2026-07-31,issue #4,端口 20 → 21)**: 列语义 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。三个 emit 入口口径必须各自定死,否则同一列在不同行含义不同: `emit_attempt`(RetryMW 调用,**唯一**有生效源者)并上 `source.extra_body`;`emit_cache_hit` / `emit_terminal_failure`(TelemetryMW 最外层调用)无 source 可言,只记调用级——与 `model`/`source_name` 在终态行置空是同一先例,且缓存命中行无损(`sampling` 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同)。三者统一读 `request.sampling` 而非 `request.overlay`(后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处未被污染,直接用必然三行分叉)。OCR/embedding 路径因决策 G 剥离 `extra_body`,该列恒 NULL。
|
||||
|
||||
(`cached_prompt_tokens`/`model_reported` 为 2026-07-31 issue #3 新增,端口由 18 字段扩为 20;两个后端在初始化期对已存在的旧表幂等补列——`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入都被逐行 warning 丢弃。补列一律**先探测缺列再 ALTER**(`ADD COLUMN IF NOT EXISTS` 即使列已存在也先取 ACCESS EXCLUSIVE 锁,而遥测内联 await,锁共享审计表会拖垮业务调用),且**失败只逐行降级、绝不置结构性失能标志**。新列在 DDL 里必须排在 `created_at` **之后**,与 `ALTER TABLE ADD COLUMN` 的追加位置一致,否则新建库与升级库的物理列序分叉)。链路: `session_id`/`parent_call_id` 由调用方传入贯穿(agent step → LLM call)。`messages` 落库前对多模态 part 先摘要(与缓存 key 共用同一摘要函数,§7.5)——Video-Tree 现状 base64 整段进 SQLite 导致 db 膨胀(`llm.py:330`),库内修复(2026-07-20,VT 迁移缺口 R12)。
|
||||
|
||||
- 后端: `SQLiteRecorder`(默认;WAL + busy_timeout、`INSERT OR IGNORE` 幂等、`asyncio.to_thread` 桥接、初始化/写入失败全降级不冒泡)与 `PostgresRecorder`。
|
||||
- **单一 helper 铁律**: 遥测调用点收敛为一个内部函数/上下文管理器;Video-Tree 与 GovDoc 各有 4-5 处逐字复制的 `record_llm_call(15 个参数)` 是本条的直接教训。
|
||||
- 成本: `pricing.py` 维护 model → (input 单价, output 单价) 表,遥测时换算 `cost` 字段;查不到价格记 None 并 warning,**不阻塞调用**。
|
||||
- 成本: `pricing.py` 维护 model → (input 单价, output 单价, **可选** cached_input 单价) 表,遥测时换算 `cost` 字段;查不到价格记 None 并 warning,**不阻塞调用**。缓存读取单价(2026-07-31,issue #3)只在配置了该档且本次有命中时启用,按 `(prompt - cached) × input + cached × cached_input` 分段计价;**未配该档绝不按经验折扣率猜**,退化为全额输入价(P5)。命中数超过输入总数时按总数夹取并 warning,不产生负成本。
|
||||
|
||||
### 7.9 结构化输出阶梯(D14)
|
||||
|
||||
@@ -491,6 +528,7 @@ src/polygateway/
|
||||
|
||||
- **载体**: `.env` + 环境变量(工程配置);缺失关键配置直接报错,严禁硬编码默认值兜底(三项目共同铁律)。**实现勘误(2026-07-20 M1,人类确认)**: 多源 `{SCOPE}__{PROVIDER}__{N}__{FIELD}` 是动态键族,pydantic-settings 的静态字段模型无法表达,故 `GatewaySettings` 为 frozen dataclass + python-dotenv(显式核心依赖)读取,fail-loud 校验语义与 pydantic-settings 一致。
|
||||
- **多源命名**: `{SCOPE}__{PROVIDER}__{N}__{FIELD}`(如 `LLM__QWEN__1__API_KEY`、`OCR__MONKEY__1__BASE_URL`),聚合为 `list[SourceConfig]`;SCOPE 支持逻辑角色前缀(§7.7)。
|
||||
- **`{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`(2026-07-31,issue #4)**: 值为 JSON **对象**串(数组/标量报错),解析为源级恒定采样参数。`_SOURCE_FIELDS` 是跨 scope 共用的一张表,故该键在 `OCR__`/`EMBED__` 下也语法合法,但那两条路径不消费它(embed payload 硬编码 `{model, input}`、MonkeyOCR 只发 multipart)——处置为**构造期剥离 + warning 放行**而非报错(2026-07-31 人类拍板: 这两条路径本无采样语义,配错后果远轻于 chat,不值得让下游装配起不来)。剥离本身是承重的: 不剥离则遥测 `sampling` 列会记录一个从未发出的参数(§7.8),那是数据造假而非参数失效。
|
||||
- **韧性参数键名**沿用三项目习惯(`LLM_TIMEOUT` / `LLM_MAX_RETRIES` / `LLM_RETRY_BASE_DELAY` / `LLM_RETRY_MAX_DELAY` / `LLM_CIRCUIT_BREAKER_THRESHOLD` / `LLM_CIRCUIT_BREAKER_COOLDOWN` / `LLM_TTFT_TIMEOUT` / `LLM_INTER_TOKEN_TIMEOUT`),降低三项目迁移改名成本。
|
||||
- **per-scope 韧性配置(2026-07-20,CHS 迁移缺口 G4)**: 韧性参数支持按 scope 覆盖——`{SCOPE}__RETRY__MAX_ATTEMPTS` / `{SCOPE}__BREAKER__FAIL_THRESHOLD` / `{SCOPE}__BREAKER__COOLDOWN_S` / `{SCOPE}__BACKPRESSURE__STALL_WINDOW_S` / `{SCOPE}__SELECTOR` / `{SCOPE}__GLOBAL__MAX_CONCURRENCY|RPM|TPM`(CHS 现状: VLM 与 OCR 两 scope 参数各异)。平铺键(`LLM_*`)是单 scope 场景的简写;两者并存时 scope 键优先。
|
||||
- **装配只有两条路**: `GatewayClient.from_env()`/`from_settings(settings)`(工厂,覆盖 90% 用户;补上三项目每次手写、GovDoc 缺失的"配置→client"一段)或构造函数全量依赖注入(测试/高级用户)。库内部任何组件**不得自读环境变量**(显式优于隐式)。
|
||||
|
||||
@@ -0,0 +1,151 @@
|
||||
# GatewaySettings 跨字段不变量守卫的生效范围
|
||||
|
||||
- **日期**: 2026-07-29;**状态**: **已批准并实施**(2026-07-29 人类门通过;§8 结论见文末)
|
||||
- **范围拍板**(用户 2026-07-29): 功能对齐社区 PR#1,但按本库规范重写;顺带销掉 PR#1 遗留的两个缺陷
|
||||
- **上游依据**: ARCHITECTURE §7.3 契约补强 G6(装配期守卫,"违反直接报错拒绝装配")、§9 配置聚合、CLAUDE.md §4.5(装配只有两条路)、`types.py` 同族 frozen dataclass 的既有校验笔迹
|
||||
|
||||
## 1. 缺陷取证(全部本地实测,worktree @ f76a89b 与 main 对照)
|
||||
|
||||
`GatewaySettings` 有三条**跨字段**不变量——单个字段合法、组合起来才非法,因此 `types.py` 各子配置的 `__post_init__` 管不到,只能在聚合层管:
|
||||
|
||||
| 不变量 | 现居位置 | 违反后的运行时后果 |
|
||||
|---|---|---|
|
||||
| 源 `timeout_s` ≤ `lease_ttl_s` | `_guard_lease`,仅 `from_env` 调用 | 租约先于请求过期,名额被放给他人 → 实际并发超配额,击穿网关 |
|
||||
| `backpressure.stall_window_s` ≥ 最大源 `ttft_timeout_s` | `_guard_stall`,仅 `from_env` 调用 | 正常慢首包被误判卡死掐断 |
|
||||
| `breaker.probe_ttl_s` ≥ 最慢源 `timeout_s` + 5 | `_load_breaker` 内联,仅 `from_env` 路径 | 半开探针在途即被接管(M2 设计 §3 原文) |
|
||||
|
||||
三条守卫都只挂在 `from_env` 上,而 CLAUDE.md §4.5 规定装配有**两条**官方路。走 `from_settings()` 能装出违反上述任一条的配置且不报错——类可以合法地存在于它自己 docstring 声称不可能的状态。
|
||||
|
||||
实测(在 PR#1 分支上,即已修前两条之后):
|
||||
|
||||
| 构造方式 | 结果 |
|
||||
|---|---|
|
||||
| `replace(base, breaker=replace(base.breaker, probe_ttl_s=1.0))`(最慢 timeout 120s) | **未拦截**,装配成功 |
|
||||
| `replace(base, sources=())` | `ValueError: max() arg is an empty sequence` —— 内置异常泄漏,既不点字段也不说原因 |
|
||||
|
||||
第一条说明 PR#1 的搬迁不完整:它的全部论证同等适用于 `probe_ttl_s`,却只搬了两条。第二条是 PR#1 **新引入**的失败模式——`max()` 此前只在 `_load_sources` 保证非空之后才执行,守卫上移到构造期后失去了这个前提。
|
||||
|
||||
另有一条隐性不变量此前从未表达:**`sources` 不得为空**。`from_env` 路径由 `_load_sources` 显式拦截,直接构造路径无人把关,零源的 client 装出来后选源必然失败。
|
||||
|
||||
## 2. 备选方案对比
|
||||
|
||||
| 方案 | 做法 | 权衡 |
|
||||
|---|---|---|
|
||||
| **A. `__post_init__` 集中校验(推荐)** | 三条跨字段守卫 + 空源检查全部收进 `GatewaySettings.__post_init__`,拆为 `_validate_sources/_validate_lease/_validate_stall/_validate_probe` 私有方法 | 与 `types.py` 同族五个 frozen dataclass 的既有笔迹完全一致;一处覆盖全部构造路径(六个工厂 + 直接构造 + `dataclasses.replace`);代价是收紧了构造承诺(见 §4) |
|
||||
| B. 各工厂入口显式调用 `settings.validate()` | 三个 client × 两个工厂,六处各加一行 | 不改构造承诺,零 breaking;但六处要永久保持同步,新增第四个 client 时必漏——正是"每个调用方各维护一份副本"的毛病挪进库里。且 `dataclasses.replace` 仍能绕过。**否决** |
|
||||
| C. 公共 `settings.validate()`,由调用方自愿调 | 提供校验入口,不强制 | 把类不变量降级成"建议";违反 P5 防御性(外部输入校验后使用)与 ARCHITECTURE §7.3"违反直接报错拒绝装配"。**否决** |
|
||||
|
||||
方案 A 与 `SourceConfig.__post_init__` 同构。选它的核心理由不是"少写五行",是**不变量的归属**:这三条约束是 `GatewaySettings` 这个类的定义的一部分,不是 `from_env` 这个函数的输入检查。放在函数里,类就失去了自我描述能力。
|
||||
|
||||
### 2.1 子决策:守卫的代码形态
|
||||
|
||||
`config.py` 现有 `_guard_lease(settings)` / `_guard_stall(settings)` 两个模块级函数,把自身实例传回给模块级函数是绕路。`types.py` 的既有做法是私有方法(`SourceConfig` 拆三个 `_validate_*`)。**改为私有方法**,与同族一致;模块级 `_guard_*` 一并删除(无其他调用点)。
|
||||
|
||||
### 2.2 子决策:`probe_ttl_s` 的派生逻辑留在哪
|
||||
|
||||
`_load_breaker` 对该字段做了两件事:未配置时**派生**(`max(2*slowest, cooldown_s, probe_floor)`,派生规则本身保证守卫恒成立)、显式配置时**校验**。派生需要读 env,必须留在 `_load_breaker`;校验上移到 `__post_init__` 后,`_load_breaker` 内联的那份校验删除(避免同一约束两处维护)。派生分支上移后仍恒过,无行为变化。
|
||||
|
||||
### 2.3 子决策:错误消息里是否列 env 键名
|
||||
|
||||
**不列。** 三条理由:(1) `types.py` 全部校验消息只点字段名,是既有笔迹;(2) 守卫现在服务两类调用方,env 键对手工拼 settings 的那类是不可执行的建议;(3) 键名的单一事实源是 `.env.example` 与 wiki `参考-配置键`,消息里复制一份即双处维护。消息格式沿用既有句式:`字段名(值)须 …;调大 X 或调小 Y`。
|
||||
|
||||
> 与 PR#1 的差异:PR#1 选择"点字段名 + 括号附 env 键",单行超 100 字符且把 `{SCOPE}__{PROVIDER}__{N}__TIMEOUT_S` 模板塞进运行时消息。本方案只留字段名。
|
||||
|
||||
## 3. 行为审计(逐条标注)
|
||||
|
||||
不是从 `reference/` 迁移,是既有模块的行为收紧,故审计对象为现有 `from_env` 路径的全部可观测行为:
|
||||
|
||||
| 现有行为 | 处置 |
|
||||
|---|---|
|
||||
| `from_env` 装配非法 lease/stall 组合 → `ValueError` | **保留**(改由 `__post_init__` 抛,时机提前到 `cls(...)` 那一行,对调用方不可见) |
|
||||
| `from_env` 配了过小 `PROBE_TTL_S` → `ValueError` | **保留**(同上,消息中不再含 env 键名 —— 有意变更,§2.3) |
|
||||
| `from_env` 未配 `PROBE_TTL_S` → 派生值 | **保留**,派生规则一字不改 |
|
||||
| `from_env` 未配任何源 → `ValueError: scope X 未配置任何源` | **保留**,`_load_sources` 的检查不动(它能给出键名模板,信息量高于构造期检查) |
|
||||
| 直接构造/`replace` 出非法组合 → 静默成功 | **有意替换**为构造期 `ValueError`(本设计的目的) |
|
||||
| 直接构造空 `sources` → 静默成功 | **有意替换**为构造期 `ValueError`,消息点明"至少一个源" |
|
||||
| 三条守卫的异常类型 `ValueError` | **保留**。装配期错误不入 `errors.py` 四分类(四分类描述的是一次调用的失败),与 `_load_sources`/`types.py` 既有装配错误一致 |
|
||||
| `GatewaySettings` 字段名与类型 | **不动**。迁移兼容约束(CLAUDE.md §4.3 例外条款)只增不删不改名,本次零字段变更 |
|
||||
|
||||
**有意放弃**:不提供 `strict=False` 之类的逃生开关。装出必然故障的配置没有正当用例。
|
||||
|
||||
## 4. 对下游的承诺变化(人类门要审的就是这条)
|
||||
|
||||
| 调用方式 | 影响 |
|
||||
|---|---|
|
||||
| `GatewayClient.from_env()` / `OcrClient.from_env()` / `EmbeddingClient.from_env()` | **零影响**,该路径本就跑这些守卫 |
|
||||
| `*.from_settings(settings)`,settings 来自 `from_env` | **零影响** |
|
||||
| 手工构造 `GatewaySettings(...)` 或 `dataclasses.replace(...)`,组合合法 | **零影响** |
|
||||
| 手工构造/`replace`,组合非法 | **行为变更**:构造期抛 `ValueError`,不再留到运行时表现为超配额/误判卡死/探针被接管 |
|
||||
|
||||
已知受影响的下游:CHSAnalyzer 重建中的 YAML → 直接构造 → `from_settings()` 路径(PR#1 提交者正是在此撞上的)。该路径若配置合法则不受影响,若非法则从"静默故障"变为"启动即报错"——方向是收益。
|
||||
|
||||
版本:**1.0.1**(patch,用户 2026-07-29 拍板)。设计初稿曾建议 minor(构造期新抛 `ValueError` 是可观测的收紧),用户判定受影响面仅限"手工拼出非法配置"这一本就故障的路径,按修复发 patch。CHANGELOG 必须把行为收紧单列小节,不能只混在"修复"里——patch 号不会给下游预警,changelog 是唯一的告知渠道。
|
||||
|
||||
发版时按 `docs-convention` §2 末行过发布清单;wiki `参考-配置键` 页现有表述("须 ≤ `PGW_LEASE_TTL_S`""须 ≥ 最大源 TTFT")与新行为一致,**无需改动内容**。
|
||||
|
||||
## 5. 非功能维度
|
||||
|
||||
| 维度 | 回答 |
|
||||
|---|---|
|
||||
| 并发与取消 | **不适用但需写明**:`__post_init__` 是同步纯计算(只读自身字段做比较),无 I/O、无 await、无锁,不存在取消穿透点。不引入任何全局状态,纯 asyncio 中立铁律不受影响 |
|
||||
| 降级方向 | 装配期校验属**准入侧**,按库铁律"报错而非放行"。无后端依赖,无降级分支 |
|
||||
| 幂等与重复 | `__post_init__` 不修改任何字段(frozen 也不允许),重复构造同一配置得同一结果;校验本身无副作用 |
|
||||
| 持久化与原子性 | 不适用,配置对象不落盘 |
|
||||
| 性能 | 每次构造增加三次 `max()` 遍历 sources(典型 1-4 个源)。`GatewaySettings` 只在装配期构造,不在请求路径上,可忽略 |
|
||||
|
||||
## 6. 测试策略
|
||||
|
||||
`tests/unit/test_config.py` 新增一个测试类,覆盖矩阵为 **4 条不变量 × 2 条构造路径**:
|
||||
|
||||
| 用例 | 断言 |
|
||||
|---|---|
|
||||
| 三条守卫各自:`dataclasses.replace` 构造出违反组合 | 抛 `ValueError`,消息含对应字段名 |
|
||||
| 三条守卫各自:边界值恰好相等(`timeout_s == lease_ttl_s` 等) | **构造成功**——守卫收紧的是错的那些,不是所有直接构造 |
|
||||
| `sources=()` | 抛 `ValueError`,消息点明"至少一个源",**且不是 `max() arg is an empty sequence`** |
|
||||
| `GatewayClient.from_settings(非法 settings)` | 抛 `ValueError`。**注意抛点**:方案 A 之下非法实例根本无法存在,异常发生在实参求值(构造 settings)那一刻,不在工厂内部——这正是构造期把关换来的性质,测试 docstring 须写明,以免后人误读为工厂自带校验 |
|
||||
| `OcrSettings` / `EmbeddingSettings` 直接构造包着非法 gateway | 抛 `ValueError`(证明三条 client 线一并覆盖) |
|
||||
| 既有 447 passed / 34 skipped | 全绿,零回归 |
|
||||
|
||||
TDD 顺序:先写测试跑出预期失败(预计 3 条守卫 + 空源 + from_settings 端到端 共失败 6 条以上),再实现,再全绿。测试不新增 mock,全部用既有 `_env()` helper 构造真实 settings 再派生。
|
||||
|
||||
## 7. 与 PR#1 的关系
|
||||
|
||||
功能对齐,不是推翻。PR#1 的问题诊断完全正确,本设计沿用其核心结论(守卫属于类不变量,应在构造期生效),差异集中在:
|
||||
|
||||
| 维度 | PR#1 | 本设计 |
|
||||
|---|---|---|
|
||||
| 覆盖的不变量 | 2 条 | 4 条(补 `probe_ttl_s`、空 sources) |
|
||||
| 代码形态 | 保留模块级 `_guard_*(settings)` | 改为 `_validate_*` 私有方法,同 `SourceConfig` |
|
||||
| docstring | 引用 `CLAUDE.md §4.5`(下游读者看不到该文件)、带论证口吻 | 只引 ARCHITECTURE §7.3 与自身概念,解释"为什么"不复述辩论 |
|
||||
| 错误消息 | 字段名 + 附 env 键模板 | 只点字段名(§2.3) |
|
||||
| 测试 | 4 条,全走 `replace`,其中 1 条同义反复 | 覆盖 4 不变量 × 2 路径 + 边界值 + 三条 client 线 |
|
||||
| 导入位置 | 两处函数内 `import dataclasses` | 文件顶部 |
|
||||
|
||||
合并后应关闭 PR#1 并在其中说明:诊断被采纳,实现按库内规范重写并扩展了覆盖范围。
|
||||
|
||||
## 8. 人类拍板结论(2026-07-29)
|
||||
|
||||
| 问题 | 结论 |
|
||||
|---|---|
|
||||
| 主决策 | **接受方案 A**,构造期强制,承诺收紧 |
|
||||
| 范围 | **全量**:`probe_ttl_s` 与空 sources 一并纳入 |
|
||||
| 消息文案 | **去掉 env 键名**,只点字段名(§2.3) |
|
||||
| 版本 | **1.0.1**(patch);初稿建议的 minor 被否,理由与代偿见 §4 |
|
||||
| PR#1 处置 | 重写合并后关闭并说明,诊断归功于提交者 |
|
||||
|
||||
## 9. 实施与验证留痕
|
||||
|
||||
实施于 `fix/settings-invariant-guards`(5 commits)。TDD 证据:新测试类先 **6 failed / 3 passed**(3 条为边界护栏,本就应过),实现后全绿。
|
||||
|
||||
独立 verifier(全新上下文)核验结论 **可以合并,无阻塞**,其中两项证据值得留档:
|
||||
|
||||
- **变异测试 12/12 全杀**:逐个破坏实现(删各 `_validate_*` 调用、`>`↔`>=`、`<`↔`<=`、删空源检查、`_PROBE_GRACE_S` 归零)均有测试失败,无一存活。边界侧用例(恰好相等必过)对每条守卫都真实有效,差一错误可捕获。
|
||||
- **无热路径回归**:库内**没有任何地方**构造或 `replace` `GatewaySettings`(`src/` 中 4 处 `dataclasses.replace` 全在 `middleware/structured.py`,作用于 `ChatRequest`/`LLMResponse`)。单次构造实测 1.45 µs,装配期一次性成本。`pickle`/`deepcopy` 不触发 `__post_init__`,只有 `replace` 触发——序列化往返既无额外开销也不构成二次守卫点。
|
||||
|
||||
### 9.1 verifier 发现的同族遗漏(范围外,另起任务)
|
||||
|
||||
`GatewaySettings` 仍有 **14 条校验只挂在 `from_env`**,直接构造/`replace` 全部放行,与本设计所修的是同一个 bug 类:`limiter/breaker/cache_backend=redis` 但 `redis_url=None`、`telemetry_backend=sqlite/postgres` 但 path/dsn 为 None、`selector`/`quota_full`/各 backend 的枚举合法性、`structured_max_retries` 负值、`scope` 空串等。
|
||||
|
||||
严重性高于本次所修的三条,因为 `client.py:262/282/302/312/316` 有 5 处 `assert ... # 内部不变量: config 已校验` **明文依赖这个前提**,而该前提在 `from_settings` 路上为假:断言开启时抛裸 `AssertionError`(不点字段不说原因),`python -O` 下断言消失、错误退化为 redis 库抛出的天书。后者同时违反 CLAUDE.md §4.3"禁止 assert 承担生产校验"。
|
||||
|
||||
**有意不纳入本次交付**(避免任务外扩张),另起任务处理。
|
||||
@@ -0,0 +1,152 @@
|
||||
# est_tokens 解耦设计(issue #2)
|
||||
|
||||
- **日期**: 2026-07-30
|
||||
- **触发**: Gitea issue #2《est_tokens 应由库按实测自估,而不是让调用方填一个没有正确取值的常量》
|
||||
- **档位**: 强制档(改公共 API 语义 + `usage_source` 公共值域 + 推翻一条已声明保留的迁移行为)→ 需人类审批门
|
||||
- **修订的权威文档**(经独立审查补全):
|
||||
- `ARCHITECTURE.md` §7.7 行 428(`SourceConfig.est_tokens` 描述)、§5.1 行 331(`usage_source` 值域)、**§4.4 行 305**("token 按 `est_tokens` 预扣")、**§7.1 行 384**("打捞路径强制 `usage_source="estimated"`",因 §3.2 #4 变为有条件)
|
||||
- `migrations/chsanalyzer.md` 行 151 与 G2(行 185)
|
||||
- **`.env.example` 行 11**("TPM > 0 时 EST_TOKENS 必填 > 0",约束已废除)
|
||||
|
||||
## 1. 问题:一个常量被派了两份互相矛盾的差事
|
||||
|
||||
`SourceConfig.est_tokens` 同时承担两个职责,而两者对"保守"的定义方向相反:
|
||||
|
||||
| 职责 | 语境 | "保守"意味着 | 填大的后果 |
|
||||
|---|---|---|---|
|
||||
| TPM 入场预扣 | 限流 | 多押金,宁可压吞吐也不击穿网关 | 安全(只是慢) |
|
||||
| usage 缺失时的用量兜底 | 计费 | **不存在保守方向** | 账单虚高 |
|
||||
|
||||
CHS 原版 `config.py:55` 把它定义为"须 ≥ 最坏情形 token"——按定义是**上界**。拿上界当实测值记账,必然系统性高估。库把遥测拆成 `prompt_tokens`/`completion_tokens` 两列后又把整个估值塞进 `completion`(`openai_compat.py:146`),而 `pricing.py:70-72` 按 `prompt×input价 + completion×output价` 换算,输出单价通常是输入的数倍——**双重高估**。
|
||||
|
||||
实测算例:`est_tokens=4000`,单价输入 1 元/百万、输出 8 元/百万,真实消耗 400+100:
|
||||
|
||||
| | 记账 token | cost |
|
||||
|---|---|---|
|
||||
| 真实 | 400 / 100 | 0.0012 元 |
|
||||
| 现状 | 0 / 4000 | 0.032 元(**26 倍**) |
|
||||
|
||||
第二个症状是装配约束:`types.py:125` 的 `tpm > 0 ⇒ est_tokens > 0` 把供应商配额(运维可从配额页抄到)与库的实现细节(预扣量,无人能正确取值)绑死。下游 CHSAnalyzer 删掉 `est_tokens` 配置项后,`tpm` 就再也不能填非 0,只能在自己的配置模型里把 `tpm` 限死为 0 绕开——库把内部细节泄漏进了配置面。
|
||||
|
||||
## 2. 备选方案对比
|
||||
|
||||
### 2.1 决策点一:usage 不可得时遥测记什么
|
||||
|
||||
| 方案 | 做法 | 权衡 |
|
||||
|---|---|---|
|
||||
| **A(选定)** | 记 `0/0`,`usage_source` 扩一个 `unavailable`,cost 记 NULL | 缺数据可被统计:`SUM(cost)` 跳过 NULL,`COUNT(*) WHERE usage_source='unavailable' AND cache_hit = false` 能量化账的缺口(**必须带 `cache_hit` 限定**:按 §3.2 #5 的裁决,缓存命中行可以既是 `unavailable` 又有 `cost=0.0`,它们本无账目缺口,不加限定就会灌水——与 §3.3 剔出 OCR 用的是同一把尺子)。代价:公共值域变更,需进 CHANGELOG,且该查询口径要一并写进 wiki(§8) |
|
||||
| B | 记 `0/0`,沿用 `estimated` | 改动最小(等于把 `est_tokens>0` 路径统一到 `est_tokens=0` 的现状行为)。**否决**:cost 算出 `0.0`,"免费"与"未知"在数据上不可区分,缺口不可量化 |
|
||||
| C | 保留 est 兜底,只修 `prompt`/`completion` 分配比例 | 保住 CHS"保守计量"意图。**否决**:比例是又一个没有正确取值的魔数,且未触及"拿上界当实测"这个根因,仍高估约 9 倍 |
|
||||
|
||||
### 2.2 决策点二:`est_tokens` 未填时的默认预扣量
|
||||
|
||||
先排除"不预扣":`try_acquire` 传 0 会让 TPM 窗口在请求飞出到 settle 回来的整段时间形同虚设,大批请求可同时入场,正是"防击穿网关"要防的场景,与 CLAUDE.md 降级方向铁律相悖。
|
||||
|
||||
| 方案 | 源甲 `tpm=6000` | 源乙 `tpm=600000` | 权衡 |
|
||||
|---|---|---|---|
|
||||
| **派生 `tpm//60`(选定)** | 押 100 → 60 个在途 | 押 10000 → 60 个在途 | 尺度无关:任何配额规模都给出同一行为上限,语义可写进 docstring("一次调用约占一秒钟的配额份额") |
|
||||
| 固定常量 1000 | 押金占配额 1/6 → 仅 6 个在途,小请求场景白慢数倍 | 押金占 1/600 → 600 个在途,大请求场景照样撞 429 | **否决**:常量与配额规模无关,在途上限随配额乱飘,无法解释取值 |
|
||||
|
||||
### 2.3 派生逻辑的落点
|
||||
|
||||
| 方案 | 权衡 |
|
||||
|---|---|
|
||||
| **`SourceConfig.effective_est_tokens()`(选定)** | 纯方法只读自身字段,落 `types.py` 内核不违反依赖铁律;零装配变更、零端口变更;5 个调用点(`QuotaGate` 入场 + retry/embedding 各自的成功侧与失败侧结算)共用一份 |
|
||||
| 注入 `GlobalLimits` 到 `QuotaGate`,派生取全局与单源 tpm 的较紧者 | 能覆盖"单源 `tpm=0` 而全局 `tpm>0`"的场景。**否决**:需改三处装配(`retry.py:186`/`embedding.py:116`/`ocr.py:119`),且它修的是一个**既有**缺口(见 §7),超出本任务范围 |
|
||||
| 派生下沉到两个 limiter 后端 | **否决**:`try_acquire(source_key, est_tokens)` 的入参会变成谎言(后端忽略它),且逻辑要写两遍,违反 D3"语义契约只有一份"与 P7"决策与存储分离" |
|
||||
|
||||
## 3. 选定方案
|
||||
|
||||
### 3.1 `usage_source` 三态值域
|
||||
|
||||
| 值 | 含义 | 生产者 | cost |
|
||||
|---|---|---|---|
|
||||
| `measured` | usage 帧完整可信 | 正常路径 | 按 token 换算 |
|
||||
| `estimated` | 有实测数字但可信度降级 | 打捞路径(收到 usage 帧但流被截断) | 按 token 换算 |
|
||||
| `unavailable` | 用量信息不可得 | usage 帧缺失、失败尝试、终态失败 | **NULL**(缓存命中行例外,见 §3.2 #5) |
|
||||
|
||||
`estimated` 保留且有真实生产者(打捞),同时保证历史库里既有的 `estimated` 行读兼容。
|
||||
|
||||
**不变式的准确表述**: 产生了真实网关调用、但用量不可得的行 → cost 为 NULL。缓存命中行不在此列(见 §3.2 #5)。
|
||||
|
||||
**值域的强制落点**: `types.py` 模块级 frozenset 常量,仅约束**库内生产侧**——所有写入 `usage_source` 的位置从该常量取值,测试断言库内产出恒在三态内。**不在 `LLMResponse`/`Usage`/`TransportResult` 等 frozen dataclass 上加 `__post_init__` 值域校验**,两条理由:① 它们是运行时构造点(如 `retry.py:418`),裸 `ValueError` 不属 `errors.py` 四分类,`RetryMW` 不捕它,会直接逃出 `chat()`,违反错误分类驱动铁律;② `LLMResponse` 是三项目已消费的公共类型,新增运行时校验是下游可见行为变更,超出本任务。故 §6 的值域测试断言"库内所有生产点的产出值落在三态内",而非"越界字符串被拒"。
|
||||
|
||||
### 3.2 逐处改动
|
||||
|
||||
| # | 位置 | 改动 |
|
||||
|---|---|---|
|
||||
| 1 | `types.py:125` | 删除 `tpm > 0 ⇒ est_tokens > 0`;`est_tokens` 保留字段、语义降为"可选调优覆盖" |
|
||||
| 2 | `types.py` `SourceConfig` | 新增 `effective_est_tokens()`:显式值 > 0 则原样返回;否则 `tpm > 0` 时返回 `max(1, tpm // 60)`,`tpm == 0` 时返回 0 |
|
||||
| 3 | `openai_compat.py:146,176` | 两处兜底改为 `(0, 0, "unavailable")` / `(0, "unavailable")`,不再读 `source.est_tokens` |
|
||||
| 4 | `openai_compat.py:336` | 打捞覆盖加条件:仅当 `usage_source == "measured"` 时降级为 `estimated`,否则保持 `unavailable`(否则 `0/0` 会被标 `estimated` 而算出假的 `0.0`) |
|
||||
| 5 | `middleware/telemetry.py:130-135` | cost 分支增加短路:`usage_source == "unavailable"` → `None`。**插在 `cache_hit` 分支之后**:缓存命中未产生新调用,`0.0` 是事实而非未知,既有"缓存命中 0.0"语义保持不动。故 `cache_hit=True` 且 `usage_source="unavailable"` 的行 cost 仍是 `0.0`,与 §3.1 不变式不冲突(那条只管产生了真实调用的行) |
|
||||
| 6 | `middleware/telemetry.py:58,100` | 失败尝试与终态失败的 `usage_source` 由 `estimated` 改 `unavailable`(用量确实不可得;这两行 cost 本已是 None,语义对齐不改金额) |
|
||||
| 7 | `middleware/ratelimit.py:26` | `source.est_tokens` → `source.effective_est_tokens()` |
|
||||
| 8 | `retry.py:370`、`embedding.py:294` | **失败侧**保守结算改用 `effective_est_tokens()`。必须同改:预扣派生值而结算退 `est_tokens=0` 会让 `delta` 为负、退掉全部押金,丢掉"失败可能已被计费"的保守意图 |
|
||||
| 9 | `retry.py:338`、`embedding.py:271` | **成功侧**结算:`usage_source == "unavailable"` 时按 `effective_est_tokens()` 结算,而非 `prompt+completion`(此时恒为 0)。**这条是保持既有行为、不是新增保守**:改前 `_resolve_usage` 恰好返回 `est_tokens`,使 `actual == 预扣量`、`delta == 0`、押金留存;#3 把它改成 `(0, 0)` 后若不同改,成功调用的押金会被整笔退回,对"从不返回 usage 帧的网关源"构成系统性 TPM 计量失效——闸门退化成进门即放行、出门即清账,正是降级方向铁律要防的击穿 |
|
||||
| 10 | `embedding.py:383,390` | 二值合并扩为三态:任一批 `unavailable` → 整体 `unavailable`;否则任一 `estimated` → `estimated`;否则 `measured`。同步更新 `types.py:273` 的行内注释 `# measured | estimated`,内核里不留与三态矛盾的注释 |
|
||||
| 11 | `embedding.py:397` `_total_cost` | 存在 `unavailable` 批时整体 cost 记 NULL(逐批求和会给出一个偏低却看似有效的金额) |
|
||||
|
||||
### 3.3 明确不改的
|
||||
|
||||
**非 dead 的瞬时失败路径**(`retry.py:369` 的 `if not dead` 分支)按预扣量做**限流**结算的行为保留——那是限流语境,保守方向正确(失败请求可能已被网关计费),且该值只流向 `_settle_and_release`,不进遥测。其余三条失败分支(`RequestRejectedError`/`ResultInvalidError`/`SourceDeadError`)的 `actual` 停在初值 0(`retry.py:329`),属既有行为,本次**不动**——#8 已把行号钉死,实现时不要顺手把这三条也改成保守结算。`SourceConfig.est_tokens` 字段与 `{SCOPE}__{PROVIDER}__{N}__EST_TOKENS` 环境键**保留不删不改名**(迁移兼容硬约束,ARCHITECTURE §5.1)。`RateLimiter` 端口签名不变。
|
||||
|
||||
**`ocr.py:411` 的 `usage_source="measured"` 保留不改**(初稿曾列为改动项,独立审查后剔出)。库既有立场是 OCR 的 0 token 属**事实**而非未知——`types.py:51` "token 用量;OCR 等无计费调用填 0"、`ocr.py:9` "settle 恒为 0(OCR 无 token 计费)"——故 `measured` 是准确陈述。改成 `unavailable` 还会反噬 §2.1 的核心度量:`COUNT(*) WHERE usage_source='unavailable'` 本用于量化账目缺口,灌进本无缺口的 OCR 行就失去意义。
|
||||
|
||||
## 4. 旧版行为审计(迁移保留项的推翻声明)
|
||||
|
||||
| 旧版行为 | 出处 | 本次处置 |
|
||||
|---|---|---|
|
||||
| usage 缺失按 `est_tokens` 估算并标 `estimated`,不静默用 0 | CHS `invokers.py:241-254`;`migrations/chsanalyzer.md:151` 标记为**保留** | **有意放弃**。理由:CHS 只记单个 `total_tokens`,不存在 prompt/completion 分配问题;库拆两列后无法忠实分配,且 `est_tokens` 按 CHS 自身定义是最坏情形上界。"保守"在限流语境安全、在计费语境只有错误一个方向 |
|
||||
| 缺失时不静默用 0(拒绝 VT 的"填 0 且不标注") | 同上;`m1-core-design.md:222` 行 10 | **保留**。本方案记 0 但带 `unavailable` 显式标记且 cost 为 NULL,反静默的原始意图完整保留——被放弃的只是"编一个数字"这个手段 |
|
||||
| `est_tokens` 作 TPM 入场预扣常量 | CHS `config.py:55` | **保留**,仅由必填降为可选覆盖 |
|
||||
| `tpm > 0 ⇒ est_tokens > 0` 装配校验 | `m1-core-plan.md:93` | **替换**为库内派生,校验删除 |
|
||||
| 打捞路径强制 `estimated` | `m1-core-design.md` §6 | **保留**,补一个前置条件(§3.2 #4) |
|
||||
| 遥测 `INSERT OR IGNORE` 幂等、写失败降级不冒泡、列只增 | `m1-core-design.md:218` | **保留**,本次无 DDL 变更 |
|
||||
|
||||
## 5. 非功能维度
|
||||
|
||||
**并发与取消**: `effective_est_tokens()` 是无状态纯方法(只读 frozen dataclass 字段),并发安全、无锁、可重复调用。本次改动不新增 `await` 点、不改变任何 `try/finally` 结构,取消穿透路径与 in-flight 释放语义原样不动。#8 与 #9 合起来保证**成功侧与非 dead 瞬时失败侧**的预扣与结算恒取同一派生值(`delta == 0`)——这是本设计里最容易漏的一致性约束(初稿只写了失败侧,独立审查发现成功侧缺口)。**取消 / RequestRejected / ResultInvalid / SourceDead 四侧不在此列**:它们的 `actual` 停在 `retry.py:329` 的初值 0、全额退回,属 §3.3 声明不动的既有行为。
|
||||
|
||||
**降级方向**: 不改变任何后端的降级方向。遥测侧仍是静默降级(`telemetry.py:161` 的 warning 不冒泡);限流侧仍是 `GovernanceBackendError` 上抛而非放行;TPM 计量不因 usage 帧缺失而静默失效(#9)。
|
||||
|
||||
否决 issue 建议的 p90 自估,主论据是 **`TelemetryRecorder` 目前是纯只写端口,自估需要新增读接口并强制所有后端(含 `none`)实现**,公共 API 扩张远大于它要省掉的一个可选字段,且尚无实测证据表明派生默认值不够用(§8)。初稿曾论证"那会把两条方向相反的降级铁律焊在一起",此论据经审查后**撤回**:p90 方案完全可以在遥测读失败时回退到纯派生值,限流侧仍能保持 fail-closed,故并非必然冲突。结论不变,理由收窄。
|
||||
|
||||
**幂等与重复**: `Permit.settle()`/`release()` 的幂等 flag 语义不变。#8 与 #9 使预扣与结算取自同一派生函数,同一请求重复结算仍是 no-op。
|
||||
|
||||
**持久化与原子性**: 无 DDL 变更(两 schema 的 `cost` 列已可空);无新增落盘点;Redis Lua 脚本不改(仍接收调用方算好的 est)。历史数据不迁移:旧行的 `estimated` 语义在新值域中依然合法可读。
|
||||
|
||||
## 6. 错误处理与测试策略
|
||||
|
||||
值域校验失败属配置/内部不变量违反 → `ValueError`(装配期 fail-loud),不进四分类运行时错误。本次不改变任何调用失败的分类归属。
|
||||
|
||||
| 测试 | 断言要点 | 文件 |
|
||||
|---|---|---|
|
||||
| 约束解绑 | `tpm=6000, est_tokens=0` 构造成功(改前抛 ValueError) | `tests/unit/test_types.py` |
|
||||
| 派生尺度无关 | `tpm=6000→100`、`tpm=600000→10000`、`tpm=0→0`、显式值优先、`tpm=30→max(1,·)` 不为 0 | 同上 |
|
||||
| cost 不再造假 | `est_tokens=4000` + usage 缺失 → `0/0/unavailable` 且 `record_llm_call` 收到 `cost=None`(改前 `0.032`) | `tests/unit/test_openai_compat.py`、`test_telemetry.py` |
|
||||
| 缓存命中不受牵连 | `cache_hit=True` 且 `unavailable` → cost 仍为 `0.0`(锁定 §3.2 #5 的分支次序) | `test_telemetry.py` |
|
||||
| 打捞前置条件 | 打捞 + usage 帧存在 → `estimated` 且 cost 非 None;打捞 + usage 缺失 → `unavailable` 且 cost 为 None(回归 §3.2 #4) | `test_openai_compat.py` |
|
||||
| **失败侧**结算不退多 | 未填 `est_tokens` 且 `tpm>0` 时失败请求,TPM 窗口残留量等于派生预扣量而非 0(回归 §3.2 #8) | `tests/contracts/test_limiter_contract.py` |
|
||||
| **成功侧**结算不退多 | usage 缺失的**成功**调用后,TPM 窗口残留量等于派生预扣量而非 0(回归 §3.2 #9,本设计最易漏的一条)。现有锚点 `test_retry.py:149` 的 `_src("a", tpm=1000, est_tokens=400)` 旁加一个 `est_tokens=0` + usage 缺失的用例 | `tests/unit/test_retry.py`、`test_limiter_contract.py` |
|
||||
| 三态合并 | 混合批 `measured+unavailable` → 整体 `unavailable` 且 cost 为 NULL | `tests/unit/test_embedding.py` |
|
||||
| OCR 不变 | OCR 成功行仍为 `measured` 且 settle 恒 0(防回归,锁定 §3.3 的剔出决定) | `tests/unit/test_ocr_client.py` |
|
||||
| 值域封闭 | 库内所有生产点的产出恒落在三态内;公共 dataclass 不因越界值抛异常(锁定 §3.1 的落点决定) | `test_types.py` |
|
||||
|
||||
限流侧断言随 `tests/contracts/test_limiter_contract.py` 同时覆盖内存与 Redis 两后端(Redis 走真实实例,遵守共享后端不并跑纪律)。
|
||||
|
||||
## 7. 已知限制(本次不修,显式声明)
|
||||
|
||||
单源 `tpm == 0` 而全局 `tpm > 0` 时,`effective_est_tokens()` 返回 0,全局 TPM 闸拿 0 预扣、入场保护形同虚设。**这是既有行为**(现状约束只管 `cfg.tpm > 0`,该场景下 `est_tokens=0` 本就合法),本方案不引入也不修复它。修它需要把 `GlobalLimits` 注入 `QuotaGate`(§2.3 备选二),属独立议题,建议另开 issue。
|
||||
|
||||
## 8. 下游影响与发布
|
||||
|
||||
`est_tokens` 从必填降为可选后,CHSAnalyzer 可删掉"`tpm` 必须为 0"的绕行校验并填真实 TPM。`usage_source` 出现第三个值、且不可得行的 cost 由数值变 NULL,是下游可见的行为变更:成本汇总若此前依赖"cost 非空"隐含假设需复核。按 `docs-convention.md` §2,发版须同步 CHANGELOG 与 wiki 的 usage/成本口径说明,并在 issue #2 回帖结论。
|
||||
|
||||
遥测驱动的自适应预估(issue 原建议)不在本次范围,待默认派生值在真实负载下出现实测问题后再评估。
|
||||
|
||||
## 9. 规模判定
|
||||
|
||||
改动面(独立审查后重算):**6 个源文件**(`types.py`、`transports/openai_compat.py`、`middleware/telemetry.py`、`middleware/ratelimit.py`、`middleware/retry.py`、`embedding.py`;`ocr.py` 已剔出)、**7 个测试文件**、**3 份权威文档**(ARCHITECTURE.md、`migrations/chsanalyzer.md`、`.env.example`),外加按 `docs-convention.md` §2 必须同步的 CHANGELOG 与用户文档站 wiki(版本 bump 不得裸发)。
|
||||
|
||||
属跨多文件功能 → 本设计经人类审批后须走 `writing-plans` 出实施计划,不得直接进实现。
|
||||
@@ -0,0 +1,164 @@
|
||||
# GatewaySettings 装配校验补齐(第二轮)
|
||||
|
||||
- **日期**: 2026-07-30;**状态**: **已批准并实施**(2026-07-30 人类门通过;§9 结论、§10 实施留痕)
|
||||
- **缘起**: [2026-07-29-settings-invariant-guards-design.md](2026-07-29-settings-invariant-guards-design.md) §9.1 —— 独立 verifier 在第一轮交付后发现,`from_env` 上还留着一批同族校验;本设计是那一轮的续作,**同一个 bug 类的剩余部分**
|
||||
- **上游依据**: 第一轮设计 §2 已批准的方案 A(不变量归属于类,不归属于某个工厂);CLAUDE.md §4.3(assert 仅用于内部不变量)、§4.5(装配只有两条路)
|
||||
|
||||
## 1. 待收拢的校验清单(逐条实测确认只在 `from_env` 生效)
|
||||
|
||||
### A. 枚举合法域(6 条)
|
||||
|
||||
| 字段 | 合法域 | 现居 |
|
||||
|---|---|---|
|
||||
| `limiter_backend` / `breaker_backend` | `{memory, redis}` | `_load_pgw`(经 `_load_choice`) |
|
||||
| `cache_backend` | `{redis, memory, none}` | `_load_pgw` 内联 |
|
||||
| `telemetry_backend` | `{sqlite, postgres, none}` | `_load_pgw` 内联 |
|
||||
| `selector` | `_SELECTORS` | `from_env` 调 `_load_choice` |
|
||||
| `quota_full` | `_QUOTA_FULL` | `from_env` 调 `_load_choice` |
|
||||
|
||||
直接构造传 `selector="random"` 或 `cache_backend="rediss"` 一律放行,后果是装配时落进 `_build_*` 的 else 分支或静默不建后端。
|
||||
|
||||
### B. 条件必填(7 条,跨字段)
|
||||
|
||||
| 条件 | 要求 | 违反后果 |
|
||||
|---|---|---|
|
||||
| `limiter_backend`/`breaker_backend`/`cache_backend` 取 `redis` | `redis_url` 非空 | **见 §2**,最严重 |
|
||||
| `cache_backend != "none"` | `cache_namespace` 非空 | 缓存 key 失去租户隔离——踩"无缓存毒化"铁律 |
|
||||
| `cache_backend != "none"` | `cache_ttl_s > 0` | `from_env` 明令禁止的"永不过期"从另一条路进来 |
|
||||
| `telemetry_backend == "sqlite"` | `telemetry_sqlite_path` 非空 | 断言炸或写空路径 |
|
||||
| `telemetry_backend == "postgres"` | `telemetry_pg_dsn` 非空 | 同上 |
|
||||
|
||||
### C. 标量域(2 条)
|
||||
|
||||
`structured_max_retries ≥ 0`;`scope` 非空(空 scope 会污染遥测与缓存命名空间)。
|
||||
|
||||
## 2. 为什么这批比第一轮更严重:`client.py` 的断言前提为假
|
||||
|
||||
`client.py` 有 5 处断言**明文声称这个前提已经成立**:
|
||||
|
||||
```python
|
||||
assert settings.redis_url is not None # 内部不变量: config 已校验
|
||||
```
|
||||
|
||||
位置:`client.py:262/282/302`(redis_url)、`:312`(pg_dsn)、`:316`(sqlite_path)。走 `from_settings` 时该注释是假的,verifier 实测:
|
||||
|
||||
| 运行方式 | 结果 |
|
||||
|---|---|
|
||||
| 断言开启 | `AssertionError()` —— 裸断言,不点字段、不说原因 |
|
||||
| `python -O` | 断言消失,退化为 redis 库的 `ValueError: Redis URL must specify one of the following schemes...` |
|
||||
|
||||
后者正是 CLAUDE.md §4.3 禁止的"assert 承担生产校验"。
|
||||
|
||||
**但注意结论的方向**:这 5 处 assert 本身不是要修的东西——它们要的前提是对的,错的是没人保证这个前提。§4 给出处置。
|
||||
|
||||
## 3. 方案
|
||||
|
||||
沿用第一轮已批准的方案 A,不重新论证:全部收进 `GatewaySettings.__post_init__`,新增三个私有方法与既有四个并列。
|
||||
|
||||
| 方法 | 覆盖 |
|
||||
|---|---|
|
||||
| `_validate_backends` | A 类 6 条枚举 + B 类 redis_url 三条件 |
|
||||
| `_validate_cache` | `cache_namespace` 非空、`cache_ttl_s > 0`(仅 `cache_backend != "none"` 时) |
|
||||
| `_validate_telemetry` | sqlite path / postgres dsn 条件必填 + §5 的 DSN 形态 |
|
||||
|
||||
标量两条(`structured_max_retries`、`scope`)并入 `_validate_sources` 改名后的 `_validate_identity`,与 `SourceConfig._validate_identity` 同名同职。
|
||||
|
||||
枚举合法域上提为模块级 frozenset 常量(`_LIMITER_BACKENDS` 等),`_load_pgw` 与 `__post_init__` 共用一份,消除现有的内联字面量重复。
|
||||
|
||||
**否决的替代**:在 `_build_limiter`/`_build_cache` 等工厂函数里逐个补显式检查。理由同第一轮 §2 方案 B——校验散落在消费点,每加一个后端就多一处要同步,且 `dataclasses.replace` 仍绕过。
|
||||
|
||||
## 4. 5 处 assert 的处置:**保留,不改**
|
||||
|
||||
修好构造期校验后,`settings.redis_url is not None` 就真的成了内部不变量——CLAUDE.md §4.3 原文"assert 仅用于内部不变量"说的正是这种用法,同时它给类型检查器收窄了 `str | None`。此时删掉 assert 反而丢失类型信息,改成 `raise` 则是在防御一个已被构造期排除的情况(死代码)。
|
||||
|
||||
**要改的是注释**:`# 内部不变量: config 已校验` 应点明由谁保证,例如 `# 内部不变量: GatewaySettings._validate_backends 已保证`。前一轮的教训就是这类注释会随时间变成谎言。
|
||||
|
||||
## 5. Postgres DSN:校验而非规范化(本轮唯一的新决策)
|
||||
|
||||
`_load_pg_dsn` 对 `from_env` 读到的 DSN 做了**规范化**:剥掉 SQLAlchemy 风格的 `+asyncpg` 驱动后缀(asyncpg 不认)。直接构造那条路不会剥,`postgresql+asyncpg://...` 会原样送进 asyncpg 然后在首次写遥测时才炸。
|
||||
|
||||
| 选项 | 权衡 |
|
||||
|---|---|
|
||||
| A. 构造期校验,含 `+driver` 即报错 | 显式,库不碰用户给的值;但两条装配路对同一输入接受度不同 |
|
||||
| B. 构造期静默剥后缀 | 两条路完全对齐;但 frozen 类在构造期悄悄改字段,调用方不知情 |
|
||||
| **C. 构造期剥后缀 + `logger.warning`(用户 2026-07-30 拍板)** | 两条路行为对齐,同时不静默——调用方在日志里看得见库动了他的值,想根治就自己改 DSN |
|
||||
|
||||
选 C。实现要点:`object.__setattr__` 改 frozen 字段(`SourceConfig` 无此先例,但 frozen 的约束是对**外部**不可变,构造期规范化是既有 dataclass 惯用法);warning 走 loguru(核心依赖,库内 `ocr.py:183`/`embedding.py:318` 同款用法)。
|
||||
|
||||
**warning 不会打扰 env 用户**:`_load_pg_dsn` 保留现有的剥离逻辑,`from_env` 传给构造函数时 DSN 已经干净,`__post_init__` 无事可做。只有手工构造传了带后缀的 DSN 才会触发。三项目 `.env` 里那些 SQLAlchemy 写法不会每次装配刷一条 warning。
|
||||
|
||||
代价是同一件事有两处剥离逻辑。用同一个模块级 helper `_strip_dsn_driver(dsn)` 供两处调用,避免实现分叉。
|
||||
|
||||
## 6. 行为审计
|
||||
|
||||
| 现有行为 | 处置 |
|
||||
|---|---|
|
||||
| `from_env` 对上述 15 条的校验与报错 | **全部保留**,时机提前到 `cls(...)`;`_load_*` 内联检查删除,避免同一约束两处维护 |
|
||||
| `_load_pg_dsn` 剥 `+driver` | **保留**,继续只在 env 路径生效(§5) |
|
||||
| `_load_choice` 的 `default` 语义(键缺失时取默认) | **保留**,那是 env 解析职责,不是不变量 |
|
||||
| `_load_breaker` 的有效阈值派生 `max(配置值, 源级并发×2)` | **有意保留在 env 层**(verifier 二次核验点名,记此备案免成"第五批")。它是**派生**不是校验/规范化:两路产出确实不同(env 装配 threshold=5/并发=100 得 200,直接构造得 5),但派生依赖的是"用户没显式表态时库替他选一个合理值"的 env 语义;代码构造那条路,调用方给什么就是什么表态。其跨字段下限风险由 `_validate_probe` 在构造期兜底 |
|
||||
| 直接构造出上述任一非法组合 → 静默成功 | **有意替换**为构造期 `ValueError` |
|
||||
| `client.py` 5 处 assert | **保留**,仅改注释(§4) |
|
||||
| 异常类型 | 一律 `ValueError`,与第一轮及既有装配错误一致 |
|
||||
|
||||
**有意放弃**:不校验 `pricing_path` 指向的文件是否存在(I/O 不属于配置校验,`PricingTable.from_file` 自会报错);不强制 `cache_backend == "none"` 时 namespace/ttl 必须为 None(多余字段无害)。
|
||||
|
||||
## 7. 非功能维度
|
||||
|
||||
与第一轮同构,不重复论证:`__post_init__` 纯同步计算无 I/O(不适用并发/取消/持久化);装配期属准入侧,报错不放行;`__post_init__` 不改字段故幂等。**性能**:新增约 10 次字符串比较,第一轮实测单次构造 1.45 µs 且库内无热路径构造 `GatewaySettings`,可忽略。
|
||||
|
||||
## 8. 测试策略
|
||||
|
||||
`tests/unit/test_config.py::TestCrossFieldInvariants` 扩充(不新建类,同族不变量归一处):
|
||||
|
||||
| 用例组 | 断言 |
|
||||
|---|---|
|
||||
| 6 条枚举各一条非法值 | 抛 `ValueError`,消息含字段名与合法域 |
|
||||
| redis_url 三条件(limiter/breaker/cache 各一) | 抛 `ValueError`,消息点明需要 `redis_url` |
|
||||
| cache namespace 缺失 / ttl ≤ 0 | 抛 `ValueError` |
|
||||
| telemetry sqlite path / pg dsn 缺失 | 抛 `ValueError` |
|
||||
| `structured_max_retries=-1`、`scope=""` | 抛 `ValueError` |
|
||||
| pg dsn 含 `+asyncpg`(直接构造) | 后缀被剥,字段值为干净 DSN,且发出一条 warning(用 `caplog`/loguru sink 断言) |
|
||||
| pg dsn 干净(直接构造)、或经 `from_env` 传入 | **不发** warning——env 路已在 `_load_pg_dsn` 剥过,不该刷噪音 |
|
||||
| 合法组合(每种 backend 组合各一) | 构造成功——收紧的是错的那些 |
|
||||
| **回归护栏**:`GatewayClient.from_settings` 走 redis 三后端的合法配置 | 装配成功,证明 assert 前提真的被保证了 |
|
||||
|
||||
TDD:先跑出红,预计 ≥14 条失败。要求同第一轮——每条实现改动都要有对应测试能杀死它。
|
||||
|
||||
版本:**1.0.2**(patch),CHANGELOG 同样单列"行为收紧"小节。
|
||||
|
||||
## 9. 人类拍板结论(2026-07-30)
|
||||
|
||||
| 问题 | 结论 |
|
||||
|---|---|
|
||||
| §5 DSN 处置 | **选 C**:构造期剥后缀 + `logger.warning`。不静默改用户的值,也不让两条装配路产出不一致 |
|
||||
| §4 assert 处置 | **保留,只改注释**,点明由哪个方法保证前提 |
|
||||
| 方案主体 | 沿用第一轮已批准的方案 A,无需重新论证 |
|
||||
| 版本 | 1.0.2(patch) |
|
||||
| **范围追加**(实施中经 verifier 发现后拍板) | G1-G4 四条同族遗漏一并纳入本轮;G1 的 scope 规范化取**静默**小写+strip(不告警——`from_env` 一直静默小写,scope 大小写不承载语义) |
|
||||
|
||||
## 10. 实施留痕
|
||||
|
||||
分支 `fix/settings-invariants-round-2`。TDD 两段:主体 15 条先 **16 failed**、G1-G4 追加 **9 failed**,实现后全绿(547 passed / 14 skipped,1.0.1 基线 516)。
|
||||
|
||||
### 10.1 独立 verifier 的关键发现
|
||||
|
||||
第一次核验判**有阻塞**,已修:
|
||||
|
||||
- **阻塞(本轮新引入)**:DSN 剥离的 warning 打印了完整连接串,**含明文密码**,而库内此前从无任何地方打印连接串——违反 P5。已改为只报 scheme 段变化,并补回归测试断言密码与 host/path 不进日志。
|
||||
- **变异测试 27/28 被杀**,唯一存活的是 `_load_pgw` 里 `PGW_CACHE_BACKEND` 域检查删掉后仍全绿(该 env 层 raise 零覆盖)。已补 `test_cache_backend_whitelist`,与既有 `test_telemetry_backend_whitelist` 对称。
|
||||
- **assert 处置经独立核验成立**:遍历所有可达构造路径均无法制造 assert 失败,`python -O` 下同样在构造期被拦(旧病症消失);唯一能触发的是 `object.__new__` 绕过 `__post_init__` 的人造路径,非公共 API。
|
||||
- **frozen 语义无副作用**:`object.__setattr__` 后 `hash`/相等性/集合去重正常,`replace` 幂等不重复告警,`pickle`/`deepcopy` 不触发 `__post_init__` 故不重复告警,对外仍抛 `FrozenInstanceError`。
|
||||
|
||||
### 10.2 G1-G4:第三批遗漏(已纳入本轮)
|
||||
|
||||
verifier 通读 `_load_*` 后发现,除设计 §1 的 15 条外还有四条**规范化**只在 env 路生效——与本轮所修的 DSN 是同一类:
|
||||
|
||||
| | 内容 | 危害 |
|
||||
|---|---|---|
|
||||
| G1 | `scope` 小写化 | **最严重**:scope 进 Redis key,大小写不一致使限流/熔断状态分裂到两套命名空间,分布式治理静默失效 |
|
||||
| G2 | `redis_url` 空串归 None | 空串骗过 `is None`,退化为 redis 客户端的连接串天书报错——正是本轮 CHANGELOG 声称已消除的那种 |
|
||||
| G3 | `pricing_path` 空串归 None | 退化为 `Is a directory: '.'` |
|
||||
| G4 | `EmbeddingSettings.batch_size`/`expected_dim` 域 | 该类无 `__post_init__`;晚一步到 client 构造才 fail-loud |
|
||||
|
||||
统一收进新增的 `GatewaySettings._normalize()`(在全部 `_validate_*` 之前跑)与 `EmbeddingSettings.__post_init__`。DSN 后缀因需看 backend 且需告警,规范化留在 `_validate_telemetry`。
|
||||
@@ -0,0 +1,164 @@
|
||||
# 响应可观测字段扩展设计(Issue #3)
|
||||
|
||||
- **日期**: 2026-07-31
|
||||
- **来源**: Gitea Issue #3(下游 dissect 审计需求)
|
||||
- **状态**: 已批准(2026-07-31,人类逐条确认 A2 / B1 / C1 / D1)
|
||||
- **触发档位**: 强制(变更 `types.py` 公共类型 + `ports.py` 端口签名 + 遥测持久化 schema)
|
||||
|
||||
## 1. 目标与非目标
|
||||
|
||||
| 项 | 内容 |
|
||||
|---|---|
|
||||
| 目标 1 | `LLMResponse` 暴露供应商侧 prompt cache 命中的输入 token 数 |
|
||||
| 目标 2 | `LLMResponse` 暴露 API 响应体实际返回的模型版本串 |
|
||||
| 目标 3 | 两字段同步落 `llm_calls` 遥测表(端口 18 → 20 字段) |
|
||||
| 目标 4 | `PricingTable` 支持可选的缓存读取单价,消除 cost 高估 |
|
||||
| 非目标 1 | 不改 `EmbeddingResponse` / OCR 响应——embedding 与 OCR 无 prompt cache 语义,且 issue 未提;`cost()` 新增参数带默认值,embedding 调用点(`embedding.py:419`)零改动 |
|
||||
| 非目标 2 | 不改 `cache_hit` 字段名/类型(破兼容),只在 docstring 消歧 |
|
||||
| 非目标 3 | 不为 `reasoning_tokens` 等其他 usage 细项开口(YAGNI,无下游需求) |
|
||||
|
||||
### 1.1 Issue 前提的一处修正
|
||||
|
||||
Issue 称「两者的数据都已经存在于 `TransportResult.raw` 里」。核查结果:
|
||||
|
||||
| 数据 | 实际所在 | 结论 |
|
||||
|---|---|---|
|
||||
| `usage.prompt_tokens_details.cached_tokens` | `raw={"usage": ...}`(流式 `openai_compat.py:362`、非流式 `:444`) | ✅ 已在 raw 内 |
|
||||
| 响应体顶层 `model` | **不在**。非流式 raw 只放 `body["usage"]`;流式 sink 只吸收 `usage` 与 `done` 两键(`_sse_delta`,`:44-47`),chunk 的 `model` 从未收集 | ❌ 需改 transport 采集 |
|
||||
|
||||
故本变更**不是纯字段暴露**,必须同时改 `transports/`。这决定了下面决策 A 的必要性。
|
||||
|
||||
## 2. 决策 A:字段的采集与传递路径
|
||||
|
||||
| 方案 | 做法 | 权衡 |
|
||||
|---|---|---|
|
||||
| A1 raw 约定键 | transport 往 `raw` 里塞 `{"model": ...}`;RetryMW 读 `raw.get("model")` 与 `raw["usage"]["prompt_tokens_details"]["cached_tokens"]` | 改动最小;但 `raw: dict[str, Any]` 变成隐式契约,键名靠约定;且 middleware 要懂 OpenAI 报文嵌套结构 |
|
||||
| A2 TransportResult 强类型字段(**推荐**) | `TransportResult` 追加 `cached_prompt_tokens: int \| None = None`、`model_reported: str \| None = None`;解析逻辑留在 `openai_compat.py`;RetryMW 直接搬运 | 报文格式知识不出 `transports/`,middleware 只做搬运,符合 P7(middleware 只依赖端口、不懂具体报文);两字段带默认值,`monkey_ocr` 的 OCR 结果类型不受影响 |
|
||||
| A3 middleware 解析 raw | RetryMW 内写 OpenAI 嵌套路径解析 | 把 provider 报文格式知识放进 middleware 层,新增非 OpenAI 兼容 transport 时会分叉;违反分层,否决 |
|
||||
|
||||
**选 A2**。`TransportResult` 是库内部流转类型(非三项目消费面),但仍按「新增必带默认值」处理,使 `openai_compat` 之外的构造点零改动;全库该类型仅 2 处构造(`openai_compat.py:354/436`)。
|
||||
|
||||
解析纪律(P5 一切外部输入校验后使用):`cached_tokens` 与 `model` 均来自网关响应,类型不可信。取值走防御 helper,不抛异常(可观测字段缺失绝不能打断主路径):
|
||||
|
||||
| 输入 | 结果 |
|
||||
|---|---|
|
||||
| `cached_tokens` 为非负 `int`(**含 `0`**) | 如实保留——`0` 是「该源上报了一次真实零命中」,与「未上报」的 `None` 语义不同,这正是本 issue 的核心诉求 |
|
||||
| `cached_tokens` 为负数 / 非 `int` / `bool` | `None`(`bool` 必须显式排除:`isinstance(True, int)` 在 Python 里为真) |
|
||||
| `usage` 或 `prompt_tokens_details` 非 dict | `None` |
|
||||
| `model` 为非空 `str` | 保留 |
|
||||
| `model` 为非 `str` / 空白串 | `None` |
|
||||
|
||||
## 3. 决策 B:缓存命中回放时两字段取什么值
|
||||
|
||||
| 方案 | LLMResponse 层 | 遥测层 | 权衡 |
|
||||
|---|---|---|---|
|
||||
| B1 原样回放(**推荐**) | 随缓存 JSON 回放原值 | 照记回放值 | 与既有口径一致——`CacheMW._rehydrate`(`cache.py:113-119`)只覆写与本次调用相关的时序字段(`latency_ms`/`ttft_ms`/`max_inter_token_ms`/`call_id`/`cache_hit`),`model`/`provider`/`prompt_tokens` 全部回放。新字段与它们同类(溯源 + 用量),按同一规则处理 |
|
||||
| B2 命中时置 None | 覆写为 None | NULL | 语义上「本次未打供应商,无供应商侧事实」也成立,但与同层的 `prompt_tokens` 回放行为不一致,下游要记两套规则 |
|
||||
| B3 混合 | `model_reported` 回放、`cached_prompt_tokens` 置 None | 同左 | 最难解释,否决 |
|
||||
|
||||
**选 B1**,并写入文档一条度量口径约束(与 `cost` 缺口口径同款教训,ARCHITECTURE §5.1):
|
||||
|
||||
> 统计供应商缓存命中率必须写 `WHERE cache_hit = false`——缓存命中行的 `cached_prompt_tokens` 是历史回放值,计入会重复计数。
|
||||
|
||||
`cost` 不受影响:遥测层 `cache_hit=True` 分支仍短路为 `0.0`,早于任何单价换算。
|
||||
|
||||
## 4. 决策 C:缓存读取单价(人类已选「增加可选档」)
|
||||
|
||||
| 方案 | 做法 | 权衡 |
|
||||
|---|---|---|
|
||||
| C1 ModelPrice 可选第三档(**推荐**) | `cached_input_per_1m: float \| None = None`;`cost()` 增可选参 `cached_prompt_tokens: int \| None = None` | 价格表旧文件零改动仍可加载;`embedding.py:419` 的三参调用零改动 |
|
||||
| C2 cost() 收 LLMResponse | 换算函数直接吃响应对象 | `pricing.py` 会反向依赖 `types.py` 且难以单测纯函数,否决 |
|
||||
|
||||
换算规则与退化路径:
|
||||
|
||||
| 条件 | 计价方式 |
|
||||
|---|---|
|
||||
| 配了 `cached_input_per_1m` 且本次 `cached_prompt_tokens` 为正 | `(prompt - cached) × input + cached × cached_input` |
|
||||
| 未配该档,或本次 `cached_prompt_tokens` 为 None/0 | 全额按 `input` 计(现状行为,不变) |
|
||||
| `cached > prompt`(网关口径异常) | 按 `cached = prompt` 夹取并记一次 warning;不抛异常、不产生负成本 |
|
||||
|
||||
**不猜折扣率**:未配置缓存档时绝不按「五分之一」之类经验值折算(P5 严禁默认值掩盖)。`from_file` 的 fail-loud 校验对新档同样适用:出现该键但非数或为负 → `ValueError`。
|
||||
|
||||
## 5. 决策 D:遥测表扩列的落地方式
|
||||
|
||||
人类确认「现在不存在必须保留的生产库」。但两个后端的 DDL 都是 `CREATE TABLE IF NOT EXISTS`,**已存在的开发库/下游库不会自动获得新列**,INSERT 会失败。两侧的失败形态都是**逐行 warning 丢弃**(SQLite `sqlite.py:93`;PG `postgres.py:127`——`_failed` 结构性标志只在 `_ensure_ready` 建池/建表失败时置位,与写入路径无关),即每一次调用的遥测行都丢,却不会有任何一次硬失败提示,与「遥测必录」相悖。
|
||||
|
||||
| 方案 | 做法 | 权衡 |
|
||||
|---|---|---|
|
||||
| D1 初始化期幂等补列(**推荐**) | DDL 加新列;初始化时按需 `ALTER TABLE ADD COLUMN`——PG 用原生 `IF NOT EXISTS`,SQLite 先查 `PRAGMA table_info` 再按需 ALTER | 旧库自动升列,新库无副作用;两处各约 5 行;补列失败沿用现有降级策略(warning,不冒泡) |
|
||||
| D2 只改 DDL,文档写「删表重建」 | 零代码 | 已建表的开发机/下游踩坑后只看到降级 warning,排查成本高;违反防御性 |
|
||||
| D3 引入迁移框架(alembic) | 正规版本化迁移 | 新增依赖,与「依赖极简」铁律冲突,规模严重不匹配,否决 |
|
||||
|
||||
**选 D1**。列类型:SQLite `cached_prompt_tokens INTEGER` / `model_reported TEXT`;PG `INTEGER` / `TEXT`。两列均可空(NULL = 该源未上报),不设 NOT NULL 与默认值——0 与 NULL 的区分正是本 issue 的核心诉求。
|
||||
|
||||
**D1 的实现纪律(必须钉进计划,否则补列会把降级放大成永久失能)**:
|
||||
|
||||
| 约束 | 原因 |
|
||||
|---|---|
|
||||
| SQLite 的 ALTER 必须用**独立 try**,且置于 `self._conn = conn` **之后** | `__init__` 现有 try 的最后一句才是 `self._conn = conn`(`sqlite.py:74-84`);ALTER 抛异常会让 `_conn` 停在 `None`,`record_llm_call` 首行即 return —— 整个 recorder 永久 no-op,比逐行丢弃严重得多 |
|
||||
| `duplicate column name` 视为成功吞掉 | `PRAGMA table_info` 探测 + ALTER 是 TOCTOU:多 worker 共用同一 db 文件时后到者必然撞上 |
|
||||
| 不得为补列加宽 `except` | `sqlite.py:93` 只捕 `(OSError, sqlite3.Error)`,取消是天然穿透的;PG 侧的 `except asyncio.CancelledError: raise` 必须留在最前 |
|
||||
| PG 用原生 `ALTER TABLE ... ADD COLUMN IF NOT EXISTS` | 无 TOCTOU;落在既有 `_init_lock` 保护的 `_ensure_ready` 内 |
|
||||
|
||||
端口 `TelemetryRecorder.record_llm_call` 由 18 字段扩为 20 字段(关键字参数),`ports.py:248` 的「18 字段冻结」注释与 ARCHITECTURE 相应表述同步更新。新增参数在 Protocol 上**不设默认值**——依据不是「漏改会报错」(本仓无 mypy,`make lint` 只有 ruff + import-linter,8 个测试 fake 全是 `**fields`,漏改根本不会自动红),而是**库外不存在第三方实现者**:三项目迁移文档明确删除各自的 TelemetryRecorder Protocol 与实现(`migrations/govdoc-saas.md:36`、`video-tree-trm5.md:36/51`),端口的唯一实现者就是库内两个后端,完整签名的成本为零。漏改的兜底靠 §8 的键集合断言测试,不靠类型检查。
|
||||
|
||||
## 6. 行为审计(既有行为逐条标注)
|
||||
|
||||
| 既有行为 | 处置 |
|
||||
|---|---|
|
||||
| `LLMResponse` 前 11 字段顺序即公共承诺 | **保留**,新字段追加到尾部(`structured_data` 之后) |
|
||||
| 缓存序列化 `_serialize` 用 `asdict` 后 pop 掉 `structured_data`、`_rehydrate` 按 `_RESPONSE_FIELDS` 过滤 | **保留**。新字段自动进出;旧缓存条目缺这两键时,`LLMResponse(**fields)` 靠默认值构造成功(向后兼容已验证) |
|
||||
| `cache_hit` 语义 = PolyGateway 自身响应缓存 | **保留**,仅补 docstring 消歧 |
|
||||
| `TelemetryEmitter` 单一 `_record` helper(遥测必录铁律:禁止复制参数列表) | **保留**,新字段只在 `_record` 增两个参数,三个 `emit_*` 入口各传一次 |
|
||||
| 失败尝试 / 终态失败行记 `usage_source="unavailable"` | **保留**,两个新字段在这些路径记 `None` |
|
||||
| `pricing.cost()` 是唯一换算点(注释语)| **修正**:实际有 `TelemetryEmitter` 与 `embedding.py:419` 两个调用点,顺带订正该 docstring(限于一行注释,不做结构重构) |
|
||||
| OCR / embedding 各自构造 `LLMResponse` | **保留**,两字段取默认 `None`(该路径无供应商 cache 概念) |
|
||||
|
||||
## 7. 非功能维度
|
||||
|
||||
| 维度 | 结论 |
|
||||
|---|---|
|
||||
| 并发与取消 | 纯数据字段,无新增 await 点、无共享状态。PG 补列在既有 `_init_lock` 保护的 `_ensure_ready` 内,并发首调用不会重复 ALTER;SQLite 的 `__init__` **不持** `_lock`(它只保护 `_write`/`close`),跨进程共库靠上面 D1 纪律里的「duplicate column 视为成功」兜底。取消穿透不变:PG 两处 `except asyncio.CancelledError: raise` 保持在最前,SQLite 侧只捕 `(OSError, sqlite3.Error)` 故天然穿透 |
|
||||
| 降级方向 | 遥测属「静默降级」侧:补列失败 → warning 并沿用既有逐行丢弃,绝不冒泡到调用方,也绝不让 recorder 整体失能(见 D1 纪律)。解析失败 → 字段记 `None`,不影响响应返回 |
|
||||
| 幂等与重复 | 补列幂等(PG `IF NOT EXISTS`;SQLite 先探测)。写入幂等性不变(`INSERT OR IGNORE` / `ON CONFLICT DO NOTHING` 按 `call_id`) |
|
||||
| 持久化与原子性 | 单行 INSERT 原子性不变;新增两列不参与主键与冲突判定。缓存 JSON 是整值覆写,无部分写入 |
|
||||
| 向后兼容 | 下游三项目 + dissect:纯增字段带默认值,逐字段传参的 fake 构造零改动;旧价格表文件、旧缓存条目、旧遥测表均可继续工作 |
|
||||
|
||||
## 8. 错误处理与测试策略
|
||||
|
||||
错误分类:本变更**不新增任何错误路径**。网关报文里这两项缺失或类型异常 → 记 `None`,不归入四分类(它们不是失败,是「该源没给」)。价格表配置错误仍走装配期 `ValueError`(fail-loud,不属运行时四分类)。
|
||||
|
||||
| 层 | 测试(先失败后通过) |
|
||||
|---|---|
|
||||
| types(unit) | 新字段默认值为 `None`;字段顺序不变(前 11 位置构造仍成立) |
|
||||
| transports(unit) | 用真实网关响应二次构造样本:① 流式含 `prompt_tokens_details.cached_tokens` → 解析出正整数;② 非流式同上;③ 无该键 → `None`;④ 值为 `"abc"`/负数 → `None` 不抛;⑤ 流式 chunk 的 `model` 被 sink 采集;⑥ 顶层无 `model` → `None` |
|
||||
| retry(unit) | `_build_response` 透传两字段;失败尝试路径不受影响 |
|
||||
| cache(unit) | ① 新字段随序列化往返;② **旧格式**缓存条目(缺这两键)仍能 rehydrate;③ 命中回放值符合 B1 |
|
||||
| pricing(unit) | ① 配缓存档 + 命中 → 成本低于全额;② 未配该档 → 与现状逐位相等;③ `cached > prompt` → 夹取且不为负;④ 三参旧调用签名仍可用(embedding 调用形态);⑤ 价格表含负缓存单价 → `ValueError` |
|
||||
| telemetry(integration) | ① 20 字段写入 SQLite/PG 成功并可读回;② **旧表**(18 列)在初始化后自动补列并写入成功;③ 补列失败时降级为 warning 且 recorder 仍能工作(SQLite `_conn` 不得因此为 None) |
|
||||
| 契约(**新增,不可省**) | 断言 `TelemetryEmitter` 传给 recorder 的实参键集合 == 两个后端的 `_COLUMNS`。理由:`row = tuple(fields[col] for col in _COLUMNS)` 位于两个后端的 try **之外**(`sqlite.py:90` / `postgres.py:121`),emitter 漏传新字段会抛 `KeyError`,被 `_record` 的 `except Exception` 吞成 warning → **静默丢遥测**。这是本变更最危险的失败形态,而现有 8 个 `**fields` 形态的 fake 一个都拦不住 |
|
||||
|
||||
> integration 层的 Redis/PG 测试遵守既有纪律:共享后端严禁并跑,`conda run -n PolyGateway --no-capture-output`。
|
||||
|
||||
## 9. 影响面清单
|
||||
|
||||
| 文件 | 改动 |
|
||||
|---|---|
|
||||
| `src/polygateway/types.py` | `LLMResponse` +2 字段;`TransportResult` +2 字段;`cache_hit` docstring 消歧 |
|
||||
| `src/polygateway/transports/openai_compat.py` | sink 采集 `model`;两处 `TransportResult` 构造填新字段;新增防御解析 helper |
|
||||
| `src/polygateway/middleware/retry.py` | `_build_response` 透传 2 字段 |
|
||||
| `src/polygateway/middleware/telemetry.py` | `_record` + 三个 `emit_*` 各透传 2 字段;cost 换算传入 `cached_prompt_tokens` |
|
||||
| `src/polygateway/pricing.py` | `ModelPrice` +可选档;`cost()` +可选参;`from_file` 校验;订正唯一换算点注释 |
|
||||
| `src/polygateway/ports.py` | `TelemetryRecorder` 18 → 20 字段 |
|
||||
| `src/polygateway/telemetry/{sqlite,postgres}.py` | DDL +2 列;`_COLUMNS` +2;初始化期幂等补列 |
|
||||
| `tests/` | 四处天然拦截点必须同步(漏改即红): 两个 `_record_minimal` 手写 18 键 dict(`unit/test_telemetry.py:76` 起、`integration/test_postgres_telemetry.py:81-105`)与两个 `_EXPECTED_COLUMNS` 列序断言(`unit/test_telemetry.py:18-40`、`integration/test_postgres_telemetry.py:22-41`);`unit/test_ports.py:96` 的全签名 fake 同步(它**不会**红,Protocol 的 isinstance 不校验签名);新增契约测试 |
|
||||
| `research-wiki/ARCHITECTURE.md` | §5.1 字段表 + 遥测表定义 + 「18 字段冻结」表述 |
|
||||
| 「18 字段冻结」的其余措辞点 | `ports.py:248`、`pricing.py:6`(币种说明里引用了该数字)、`telemetry/sqlite.py:87`、`tests/unit/test_telemetry.py:1` |
|
||||
| Wiki 站 + `CHANGELOG.md` | 按 `docs-convention.md` §2 清单同步(公共行为变更,版本 bump 不得裸发) |
|
||||
| `.env.example:56` | 该行内联注释是仓内**唯一**的价格表格式说明(无独立模板文件,`config/prices.json` 是未入库的本地文件),补 `cached_input_per_1m` 可选档 |
|
||||
|
||||
## 10. 审批记录
|
||||
|
||||
2026-07-31 人类逐条确认: **A2**(TransportResult 强类型字段)、**B1**(缓存命中原样回放 + 度量口径带 `cache_hit = false`)、**C1**(ModelPrice 可选缓存单价档)、**D1**(DDL 加列 + 初始化期幂等补列)。设计获批,进入 `writing-plans`。
|
||||
|
||||
版本号按 `1.1.0` 推进(纯增字段不破坏下游,但触及端口签名与表结构,minor 位比 patch 位更能提示下游);发版前若人类另有指示以指示为准。
|
||||
@@ -0,0 +1,234 @@
|
||||
# 采样参数透传设计(issue #4)
|
||||
|
||||
- **日期**: 2026-07-31
|
||||
- **状态**: 待人类审批
|
||||
- **触发**: issue #4 —— `chat()` 无法设置 `temperature`/`seed`/`max_tokens`,下游受控实验无法固定解码
|
||||
- **影响面**: `chat()` 公共签名、`SourceConfig` 公共类型、缓存 key 公式(ARCH §7.5)、遥测端口(20 → 21 字段)
|
||||
|
||||
---
|
||||
|
||||
## 1. 诉求与现状审计
|
||||
|
||||
下游 dissect 是一组受控实验:解码固定 `temperature=0`,每格配置跑 5 个 seed 报标准差。标准差必须只反映被研究的变量,不能混进解码随机性。
|
||||
|
||||
代码事实(本会话核实):
|
||||
|
||||
| 事实 | 位置 | 后果 |
|
||||
|---|---|---|
|
||||
| 全库 `temperature` 零命中 | `grep -rn temperature src/` | 解码跑在供应商默认值上,不可复现 |
|
||||
| `chat()` 签名无 overlay 入口 | `client.py:143-153` | 调用方够不着 `ChatRequest.overlay` |
|
||||
| `overlay` 唯一写入点是结构化中间件 | `middleware/structured.py:98` | 字段存在但只服务库内 |
|
||||
| `payload.update(overlay)` 是最后一步 | `transports/openai_compat.py:297` | overlay 可覆盖 `model`/`messages`/`stream`/`stream_options` |
|
||||
| 缓存 key 公式不含 overlay | `middleware/cache.py:52-64` | **见 §2 决策 C** |
|
||||
| `model_fingerprint` 只由源 `model` 名算 | `client.py:117` | 配置级采样参数变更不改 key |
|
||||
| minimax / openai profile 均 `thinking_off={}` | `providers.py:49,56` | `enable_thinking=False` 对两源均无效果 |
|
||||
|
||||
**issue 未提及但必须一并处理的**: 缓存与遥测的交互。不处理的话,failure mode 恰是 issue 自己最担心的那种——数字悄悄不可比,且不报错。
|
||||
|
||||
---
|
||||
|
||||
## 2. 设计决策
|
||||
|
||||
### 决策 A: 两层入口,合并优先级由现有层序天然给出
|
||||
|
||||
| 层 | 载体 | 用途 | 生效点 |
|
||||
|---|---|---|---|
|
||||
| 调用级 | `chat(..., overlay: Mapping[str, Any] \| None = None)` | 逐次变化(每 rollout 不同的 `seed`) | 填入 `ChatRequest` |
|
||||
| 配置级 | `SourceConfig.extra_body: Mapping[str, Any]` | 全局恒定(`temperature=0`) | transport `_build_payload` |
|
||||
|
||||
优先级 **结构化注入 > 调用级 > 配置级**,无需任何新机制:
|
||||
|
||||
```text
|
||||
_build_payload: payload{model,messages,stream} → thinking_profile
|
||||
→ source.extra_body ← 配置级(新增一行)
|
||||
→ overlay ← 调用级 ⊎ 结构化注入
|
||||
StructuredMW: {**request.overlay, **strategy_overlay} ← 结构化已在最右,天然最高
|
||||
```
|
||||
|
||||
配置级放在 transport 而非装配层合并,是因为 `extra_body` 是 per-source 的,选源在 RetryMW 之后才确定;放 transport 无需改动任何端口签名。
|
||||
|
||||
**`ChatRequest` 增第二个字段 `sampling: Mapping[str, Any] = field(default_factory=dict)`**(调用方原始采样意图的快照,库内中间件**永不修改**),与 `overlay`(请求体覆盖层,会被结构化注入)分开。`chat()` 同时填两者。理由是 `overlay` 在洋葱不同深度取值不同——`StructuredMW` 内侧含 `response_format`、外侧不含——缓存 key 与遥测若各自依赖"在哪一层读"就会口径分叉(见决策 C/D)。`sampling` 提供一个跨层恒定的读取点。
|
||||
|
||||
类型定死为 `Mapping` 而非 `dict[str, Any] | None`:空 dict 与 `None` 在此无语义差别(都是"没传采样参数"),多一种表示只会让 key 公式与 `merge()` 签名各选各的。因此决策 C 的 key 公式一律按**仅非空**参与(注意与同处的 `salt` 不同——`salt` 是"仅非 None",空串是有意义的 salt)。
|
||||
|
||||
### 决策 B: 保护键黑名单,构造期显式报错
|
||||
|
||||
`{model, messages, stream, stream_options}` 禁止出现在 overlay/extra_body 中。理由逐条:
|
||||
|
||||
| 键 | 被覆盖的后果 |
|
||||
|---|---|
|
||||
| `model` | 遥测记录的 model 与实际请求分叉 → 成本按错单价算 |
|
||||
| `messages` | 缓存 key 与遥测口径同时失真 |
|
||||
| `stream` | 绕过流式看门狗(TTFT/inter-token 三层超时全失效) |
|
||||
| `stream_options` | 丢 usage 帧 → 成本遥测归零、TPM 闸按预扣量结算失准 |
|
||||
|
||||
同一校验函数还必须验**值可 JSON 序列化**。理由:`CacheMW.__call__` 第 95 行的 `build_cache_key` 内部 `json.dumps`,**不在 `_safe_get`/`_safe_set` 的降级 try 内**;`TelemetryMW` 只捕 `GatewayUnavailableError`/`GovernanceBackendError`/`CancelledError`。调用方传 `{"temperature": np.float32(0)}`(温度扫描用 numpy 生成极自然)会抛裸 `TypeError`:不属四分类、一行遥测都没有、RetryMW 从未执行。构造期一次校验即可保住"overlay 错误全部发生在进洋葱之前"这条不变式。
|
||||
|
||||
校验函数落在 `types.py`(最内层,无依赖),两个入口各调一次:`chat()` 参数在进洋葱**之前**校验(与既有 `structured` 的 ImportError 同款先例),`SourceConfig.__post_init__` 在装配期校验(符合 §4.5「缺失/非法关键配置直接报错」)。抛裸 `ValueError`——这是调用方编程错误,不属 §6 四分类,不应被 RetryMW 当作可重试失败。
|
||||
|
||||
transport 不重复校验:三个 overlay 来源(chat 参数、SourceConfig 字段、库内策略)已全部在构造期收口,库内策略只注入 `response_format`(`json_repair.py:41` 恒空,`native_schema.py:21-29` 只产该键)。
|
||||
|
||||
### 决策 C: 调用级 overlay 进缓存 key —— 本设计的关键点
|
||||
|
||||
不做的话:同 messages 跑 5 个 seed,后 4 次命中第一次的缓存,返回同一 response,**标准差恒为 0**,实验静默作废。这正是「无缓存毒化」铁律的场景。
|
||||
|
||||
key 公式扩展(ARCH §7.5 需同步修订),读 `request.sampling` 而非 `request.overlay`——语义明确、不依赖"CacheMW 恰在 StructuredMW 外侧"这一层序巧合:
|
||||
|
||||
```text
|
||||
key_obj = {model, messages_digest, namespace, [salt], [sampling]}
|
||||
仅非 None 仅非空
|
||||
```
|
||||
|
||||
沿用 `salt` 的「仅非空时参与」写法,保证**空采样参数时旧键逐字不变**,不触发存量缓存全量冷启动。
|
||||
|
||||
配置级同理:`model_fingerprint` 从 `",".join(sorted(models))` 扩展为——所有源 `extra_body` 皆空时字面不变;否则追加 `"|" + sha256(...)`,摘要对象是「每个源的 `(model, extra_body)` 先各自 canonical-JSON 化成字符串,再排序去重」(dict 本身既不可排序也不可哈希,必须先序列化;`extra_body` 若存为 `MappingProxyType` 需 `dict(...)` 后再 `json.dumps`)。取 `(model, extra_body)` 而非 `(name, ...)`,语义是「本 scope 会用哪些(模型,解码参数)组合」,改源名不会误触冷启动。该计算在 `client.py:117` 且不在任何降级 try 内,写错即装配期崩——实施时须有直接单测。
|
||||
|
||||
**两条已知副作用(须写进 wiki)**:
|
||||
|
||||
1. 逐 rollout 变化的 `seed` 进 key 后,该路径**天然全部 miss**。这是正确语义而非缺陷,但下游要知道缓存对这条路径不再省钱。
|
||||
2. `model_fingerprint` 是**集合级**指纹,不是本次实际选中源的指纹。同 scope 下各源 `extra_body` 不同时,缓存仍可能返回另一源、另一组解码参数下产生的响应。这是既有取舍的延续(`cache.py:68-72` 对 `model` 已如此),不是本设计引入的新缺口,但"配置级采样参数进 key"容易被读成更强的保证,须写明边界。受控实验若要求逐源可复现,应让每个源独享 scope 或 namespace。
|
||||
|
||||
### 决策 D: 采样参数入遥测(端口 20 → 21 字段)
|
||||
|
||||
「实验可复现」的另一半是参数落库。不记的话,同 messages 不同输出在审计表里无法解释。与 issue #3 新增 `model_reported` 同类动机(供应商把别名指向新权重时,复现必须认真实串)。
|
||||
|
||||
**列语义定死**:`sampling: str | None` = 「调用方采样意图 ⊎ 生效源的 `extra_body`」的 canonical JSON,**不含库内结构化注入的 `response_format`**。两个理由:该列名叫采样参数,`response_format` 不是;schema 可达数 KB,逐行记会让审计表无谓膨胀。
|
||||
|
||||
`TelemetryEmitter` 有三个入口且都汇入同一个 `_record`(显式关键字参数,加列必须三处都传),必须逐个定死,否则同一列在不同行口径分叉——这正是 1.0.4 里 `cached_prompt_tokens` 不得不写"下游请读"警告的同类坑:
|
||||
|
||||
| 入口 | 调用者 | 有 `source`? | `sampling` 记什么 |
|
||||
|---|---|---|---|
|
||||
| `emit_attempt` | RetryMW(最内) | 有 | `merge(source.extra_body, request.sampling)` |
|
||||
| `emit_cache_hit` | TelemetryMW(最外) | **无** | 仅 `request.sampling` |
|
||||
| `emit_terminal_failure` | TelemetryMW | **无** | 仅 `request.sampling` |
|
||||
|
||||
后两行缺 `extra_body` 是**客观事实而非口径瑕疵**:它们没有"生效源"可言——与 `model`/`provider`/`source_name` 在终态行置空是同一先例。缓存命中行尤其无损:`sampling` 已进缓存 key,能命中就意味着历史那次的调用级采样参数与本次逐字相同;`extra_body` 亦已进 `model_fingerprint`,命中意味着源集合的配置指纹相同。
|
||||
|
||||
三个入口统一读 `request.sampling`(决策 A 的新字段)而非 `request.overlay`,是因为后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处则未被污染,直接用会让三行天然分叉。
|
||||
|
||||
**共用范围写清楚**:`types.py` 提供「2 参 dict 合并 + canonical 序列化」这一个原语,transport 与 emitter 共用它。**不追求统一到两者之上**——transport 是往更大的 payload 上依次 `update(thinking_profile) → update(extra_body) → update(overlay)`,emitter 算的是 `merge(extra_body, sampling)`,参与方与顺序本就不同,强行统一是错的。这不影响正确性:该列语义已定义为「调用方意图 ⊎ 生效源 `extra_body`」,而非 payload 的逐字回显。共用原语的目的只是让"合并语义与序列化口径"这一件事不出现两份实现。
|
||||
|
||||
两个后端按 issue #3 已建立的套路幂等补列:**先探测缺列再 ALTER**、失败只逐行降级不置结构性失能标志、新列排在 `created_at` 之后。
|
||||
|
||||
一次做完而非分两步:「能传参数但没记」的中间状态最危险——数据已产生且事后无法追溯,且分步要做两遍 DDL 迁移。
|
||||
|
||||
**OCR/Embedding 的 emit 调用点零改动**:`ocr.py:418` 与 `embedding.py:372` 也调 `emit_attempt` 且都传 `source`,只要 `sampling` 由 emitter 内部推导(而非作为新必填参数由调用者传入),这两处调用不动一行——反之立刻 TypeError,实施时必须走推导路线。两个文件本身仍有改动,即决策 G 的构造期剥离(它正是让这里的推导对 OCR/embedding 恒得 NULL 的前提)。
|
||||
|
||||
### 决策 E: 入参拷贝语义与两条只读约束
|
||||
|
||||
`chat()` 对传入 overlay 做**一次** `dict(overlay)` 浅拷贝,同一份快照对象同时填 `overlay` 与 `sampling` 两个字段(不做两份独立拷贝——它们在进入 `StructuredMW` 之前本就应当逐字相同,两份拷贝反而给"两者可以分叉"留了口子)。
|
||||
|
||||
issue 场景就是逐次改 `seed`——调用方复用同一 dict 对象改值是极可能的模式,不拷贝会出现「请求已发出、key 用了新 seed」的竞态。`ChatRequest` 虽 frozen 但 dict 是浅冻结,拦不住。`SourceConfig.extra_body` 在 `__post_init__` 转 `MappingProxyType` 同理(成本近零)。
|
||||
|
||||
拷贝之外的第二条约束:**任何中间件不得就地修改这两个 dict**,只能经 `dataclasses.replace` 派生新请求。现状已满足(`StructuredMW` 用 `{**a, **b}` 生成新 dict,`_build_payload` 只往 payload 上 `update`,全库无就地改写),本设计只是把它写成明文约束——决策 C 与 D 都建立在 `sampling` 跨层恒定之上,这条被破坏则两者同时失效(测试 #14 为此加机械执法)。
|
||||
|
||||
### 决策 F: 空 thinking profile 的诚实性缺口(issue 附带项)
|
||||
|
||||
`minimax` 与 `openai` 的 `thinking_on/thinking_off` 均为空字典。`providers.py:52` 那条「OpenAI 兼容基线,无已知注入差异」的注释在词法上属于紧随其后的 **minimax** 条目,`openai` 条目没有任何注释。所以现状是:已有的注释解释了"为何为空",但两个 provider 都没点明**后果**——`enable_thinking=False` 对它们不产生任何效果,调用方以为关掉了实际没关。
|
||||
|
||||
补的是这一句后果说明(覆盖两个 provider),不是重复已有的"为何为空"。不改行为:真需要关时经 `extra_body` 绕过。
|
||||
|
||||
### 决策 G: 非 chat 路径的 `extra_body` —— 剥离并 warning,不中断装配
|
||||
|
||||
`_SOURCE_FIELDS`(`config.py:33-47`)是**跨 scope 共用**的一张表,加了 `EXTRA_BODY` 之后 `OCR__MONKEY__1__EXTRA_BODY` / `EMBED__QWEN__1__EXTRA_BODY` 会被合法接受、进 `SourceConfig`、进遥测 `sampling` 列,但两条路径都不消费它:`monkey_ocr.py:225,247` 只发 multipart `files=`(**根本没有 JSON body**),`OpenAICompatTransport.embed`(`openai_compat.py:343`)payload 硬编码 `{"model", "input"}`。放任即**静默无效**,正是 §4.5 要禁的形态。
|
||||
|
||||
**处置(2026-07-31 人类拍板改此档)**:`EmbeddingClient` / `OcrClient` 构造期发现源带非空 `extra_body` → 记 warning 并 `dataclasses.replace(source, extra_body={})` **剥离后放行**,不抛异常。
|
||||
|
||||
剥离是这一档的**必要组成部分,不是顺手清理**。`ocr.py:390` 与 `embedding.py:350` 构造 `ChatRequest` 时不带 `sampling`,但传给 `emit_attempt` 的 `source` 是真实配置对象;若不剥离,决策 D 的 `merge(source.extra_body, request.sampling)` 会让遥测**记录一个从未发出的参数**——审计表显示该次 OCR 调用带了 `temperature=0`,实际请求体里没有。那不是"参数不生效",是遥测造假,污染的恰是事后复现的唯一依据。替代方案是在 emitter 里特判调用方身份,直接违背「遥测调用点收敛为单一 helper」铁律,否决。
|
||||
|
||||
剥离后该列在 OCR/embedding 行恒为 NULL,语义干净,emitter 零特判。
|
||||
|
||||
**被否决的原方案**: 装配期 `ValueError` 直接拒绝。理由是这两条路径本无采样语义,配错的后果远轻于 chat 路径,不值得让下游整个装配起不来。**残余风险须写进 wiki**: loguru warning 在生产中容易被淹没,运维可能仍以为参数生效——这是"不中断装配"换来的代价,故 warning 文案必须**指路**:`dimensions` 是 OpenAI embeddings 的正式参数,下游想调向量维度时会第一个撞上,文案应写明"embedding 路径暂不支持 `extra_body`,该配置已被忽略;需要 `dimensions` 等参数请提 issue"。
|
||||
|
||||
不顺手给 embed 加透传:embedding 没有采样一说,issue 也未提出诉求(YAGNI);真有需求时单独设计。
|
||||
|
||||
---
|
||||
|
||||
## 3. 关键岔路与否决记录
|
||||
|
||||
| 岔路 | 否决方 | 理由 |
|
||||
|---|---|---|
|
||||
| `chat()` 展开为 `temperature=`/`seed=`/`max_tokens=` 具名参数 | 否决 | 供应商私有参数无穷尽(`top_k`/`repetition_penalty`/`thinking_budget`),具名等于永久追加签名;且违背「深模块窄接口」(ARCH §132) |
|
||||
| 配置级放装配层全局字典而非 `SourceConfig` | 否决 | 采样参数与源强相关(不同供应商键名不同),全局字典会把无效键发给不认识它的源 |
|
||||
| overlay 不进缓存 key,靠调用方传 `cache_salt` 区分 | 否决 | 把毒化防护的责任推给调用方,漏传不报错——正是 issue 抱怨的失败形态 |
|
||||
| 采样参数不入遥测,由下游 run 快照自记 | 否决 | 见决策 D |
|
||||
| 缓存 key 与遥测都直接读 `request.overlay`,不加 `sampling` 字段 | 否决 | `overlay` 在洋葱不同深度取值不同(结构化注入),三个 emit 入口与 CacheMW 会各记各的,同一列口径分叉 |
|
||||
| `sampling` 列记「实际发出的完整合并结果」(含 `response_format`) | 否决 | 该列名为采样参数,schema 不是;且数 KB schema 逐行落库无谓膨胀 |
|
||||
| 给 embedding 路径也加 `extra_body` 透传 | 否决 | embedding 无采样一说,issue 未提诉求(决策 G) |
|
||||
| 非 chat 路径带 `extra_body` 时装配期 `ValueError` | 否决(人类拍板) | 这两条路径无采样语义,配错后果远轻于 chat,不值得让下游装配起不来;改为剥离 + warning |
|
||||
| 允许放行但**不剥离** `extra_body` | 否决 | 遥测会记录一个从未发出的参数(决策 D 的 merge 读 `source.extra_body`),是数据造假而非参数失效 |
|
||||
| 放行不剥离,改在 emitter 内特判 OCR/embedding 不记 | 否决 | emitter 是「遥测调用点收敛单一 helper」的产物,让它识别调用方身份是开倒车 |
|
||||
| transport 层再兜一次保护键校验 | 否决 | 三个入口已构造期收口,重复校验属 gold-plating |
|
||||
|
||||
---
|
||||
|
||||
## 4. 非功能维度
|
||||
|
||||
| 维度 | 回答 |
|
||||
|---|---|
|
||||
| **并发** | 无新增共享状态。`extra_body` 装配后只读(MappingProxyType);调用级 overlay 每调用独立拷贝,并发调用互不可见 |
|
||||
| **取消** | 无新增 await 点与等待循环,`CancelledError` 穿透路径完全不变 |
|
||||
| **降级方向** | 不涉及新后端。遥测新列写失败沿用既有逐行 warning 降级;缓存 key 变更不影响 Redis 掉线的静默降级方向。决策 G 的剥离 + warning 是**配置面**降级(装配期一次性、可复现、部署即暴露),与铁律里"限流/熔断后端不可用须报错"的**运行时**降级方向是两回事,不冲突 |
|
||||
| **幂等与重复** | 保护键校验是纯函数,重复调用安全;遥测补列先探测后 ALTER,重启幂等 |
|
||||
| **持久化与原子性** | 遥测单行写入,无部分写入风险。缓存 value 结构不变(`sampling` 只进遥测不进 `LLMResponse`,避免动已被三项目消费的公共类型) |
|
||||
| **重试交互** | overlay 在 RetryMW 循环外确定,换源重试时同一 overlay 应用到新源的 `extra_body` 之上——语义正确(调用级意图跨源保持) |
|
||||
| **限流交互** | overlay 里的 `max_tokens` 不影响入场预扣(取 `effective_est_tokens()`)。调用方把 `max_tokens` 抬到远超预扣量时 TPM 入场保护会短暂失真,结算侧(`retry.py:338-343`)按实测用量回填自愈。已知且可接受,不为此加机制 |
|
||||
|
||||
---
|
||||
|
||||
## 5. 错误处理与测试策略
|
||||
|
||||
**错误分类**: 保护键违规与 `EXTRA_BODY` JSON 解析失败均为裸 `ValueError`,发生在进入洋葱之前/装配期,不入四分类、不触发重试或熔断。运行时若供应商拒绝某个采样参数(如不支持 `seed`),网关返回 4xx,由既有 `RequestRejectedError` 路径处置——无需新增分类。
|
||||
|
||||
**测试清单**(每条须先失败后通过):
|
||||
|
||||
| # | 用例 | 层 |
|
||||
|---|---|---|
|
||||
| 1 | 同 messages 不同 `seed` → 两次 miss、两个不同 key(issue 场景直接回归) | unit |
|
||||
| 2 | 空 overlay 时 key 与旧实现逐字相同(防存量冷启动) | unit |
|
||||
| 3 | 全源 `extra_body` 为空时 fingerprint 与旧实现逐字相同 | unit |
|
||||
| 4 | 保护键:`chat(overlay={"stream": False})`、`SourceConfig(extra_body={"model": "x"})` 均 `ValueError` | unit |
|
||||
| 5 | 优先级:配置 `temperature=0` + 调用级 `temperature=1` → payload 为 1;结构化 `response_format` 覆盖调用级同名键 | unit |
|
||||
| 6 | 调用方在 `chat()` 返回前修改自己的 dict,不影响已发请求与已算 key(拷贝语义) | unit |
|
||||
| 7 | env 解析:`EXTRA_BODY` 合法 JSON 对象 → dict;非法 JSON / 非对象 → `ValueError` | unit |
|
||||
| 8 | 不可 JSON 序列化的值(如 `np.float32`)在 `chat()` 入口即 `ValueError`,不进洋葱 | unit |
|
||||
| 9 | 三个 emit 入口的 `sampling` 口径:attempt 含 `extra_body`、cache_hit 与 terminal 只含调用级、结构化注入的 `response_format` **三行都不出现** | unit |
|
||||
| 10 | `EmbeddingClient`/`OcrClient` 装配时源带 `extra_body` → 记 warning、装配成功、源上 `extra_body` 已被剥空,且该路径遥测 `sampling` 为 NULL(决策 G;后半段是防遥测造假的真正断言) | unit |
|
||||
| 11 | `_EXPECTED_COLUMNS` 断言更新后仍逐字匹配实际列序(见 §6,两处会直接红) | unit + integration |
|
||||
| 12 | 遥测 `sampling` 落库正确;两后端对既有旧表幂等补列 | integration |
|
||||
| 13 | 采样参数经全链路(chat → 选源 → transport payload)到达请求体 | integration |
|
||||
| 14 | **地基不变式**:走结构化重问阶梯(至少重问一次)后,RetryMW 每次尝试看到的 `request.sampling` 与 `chat()` 传入值逐字相同,且同一时刻 `request.overlay` 含 `response_format` | unit |
|
||||
|
||||
第 14 条是决策 C/D 共同的承重前提。它现在只靠"`dataclasses.replace` 恰好保留未提及字段"这一约定成立,无任何机械执法;缺这条测试则决策 E 的只读约束被破坏时不会有人发现。
|
||||
|
||||
第 2 条(空采样参数时旧键逐字不变)需自行先固化旧 key 值再比对——现有 `tests/unit/test_cache.py:39-54` 只有相等/不等与前缀断言,没有 golden hash 可依。
|
||||
|
||||
---
|
||||
|
||||
## 6. 配置与文档同步
|
||||
|
||||
env 键名沿用既有约定:`{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`,值为 JSON 对象串;`_SOURCE_FIELDS` 增一项、`_cast` 增 `json` 分支(解析失败与非 dict 均报错)。
|
||||
|
||||
同步清单(docs-convention §2):
|
||||
|
||||
| 目标 | 改什么 |
|
||||
|---|---|
|
||||
| ARCH §5.2 | `chat()` 签名定稿段追加 `overlay` 要点 |
|
||||
| ARCH §7.5 | key 公式补 `sampling` 项 + 两条已知副作用 |
|
||||
| ARCH §7.7 | 该节逐字段枚举 `SourceConfig` 构成(`ARCHITECTURE.md:452`),补 `extra_body` |
|
||||
| ARCH §7.8 | 必录字段 20 → 21 |
|
||||
| ARCH §9 | 配置面键族事实源(`:519-527`),登记 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY` |
|
||||
| `.env.example` | `client.py:247` docstring 声明它是键名清单的事实源,新键不写进去等于无处可查 |
|
||||
| `README.md:83` | 该行逐一列举 `chat()` 关键字参数,补 `overlay` |
|
||||
| wiki how-to | 增「固定解码参数」条目,写明 seed 进 key 导致缓存必 miss、以及 OCR/embedding 路径的 `extra_body` 会被忽略(仅 warning) |
|
||||
| CHANGELOG | 公共 API 新增 + 遥测端口扩列 |
|
||||
|
||||
## 7. 实施范围
|
||||
|
||||
`types.py`(保护键与 JSON 可序列化校验、合并纯函数、`ChatRequest.sampling`、`SourceConfig.extra_body`)、`client.py`(`chat()` 参数 + fingerprint)、`middleware/cache.py`(key 公式)、`transports/openai_compat.py`(`_build_payload` 一行)、`config.py`(env 解析)、`ports.py` + `middleware/telemetry.py` + `telemetry/{sqlite,postgres}.py`(第 21 字段与补列)、`ocr.py` + `embedding.py`(仅决策 G 的构造期剥离 + warning)、`providers.py`(注释)。
|
||||
|
||||
**测试侧必改**(否则直接红):`tests/unit/test_telemetry.py:18,113` 与 `tests/integration/test_postgres_telemetry.py:22,210,231` 的 `_EXPECTED_COLUMNS` 断言完整列表与列序。
|
||||
|
||||
无需改动:import-linter 契约(校验函数落最内层 `types.py`,分层关系不变)。
|
||||
|
||||
不做:给 embedding/OCR 加采样参数透传(决策 G)、任何任务外重构。
|
||||
@@ -0,0 +1,23 @@
|
||||
---
|
||||
type: design
|
||||
node_id: design:est-tokens-decoupling
|
||||
title: "est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)"
|
||||
date: 2026-07-30
|
||||
---
|
||||
|
||||
# est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)
|
||||
|
||||
全文见 [2026-07-30-est-tokens-decoupling-design.md](2026-07-30-est-tokens-decoupling-design.md)。缘起是 Gitea issue #2(下游 CHSAnalyzer 提出)。
|
||||
|
||||
- **缘起**: `SourceConfig.est_tokens` 被派了两份对"保守"定义相反的差事——TPM 入场预扣(多押金 = 安全)与 usage 帧缺失时的遥测用量兜底(计费没有安全方向)。CHS `config.py:55` 自己就把它定义为"须 ≥ 最坏情形 token"的**上界**,而库把遥测拆 `prompt`/`completion` 两列后又将整个估值塞进单价更贵的 `completion`(`openai_compat.py:146`),形成双重系统性高估:实测算例 26 倍。
|
||||
- **第二个症状**: `types.py:125` 的 `tpm > 0 ⇒ est_tokens > 0` 把供应商配额(可从配额页抄)与库的实现细节(无人能正确取值)绑死,下游删掉猜测项后 `tpm` 只能填 0,被迫在自己配置模型里加校验绕开。
|
||||
- **方案(两个决策点,人类审批)**: ① usage 不可得时记 `0/0` + `usage_source` 新增 `unavailable` + cost 记 NULL;② `est_tokens` 未填时由库派生 `tpm // 60`,字段降为可选调优覆盖(不删不改名,迁移兼容)。
|
||||
- **值域三态各有生产者**: `measured`(正常)、`estimated`(打捞路径——收到 usage 帧但流被截断,数字真实而可信度降级)、`unavailable`(用量不可得)。故 `estimated` 不是空值域,历史行亦读兼容。
|
||||
- **被否决 · usage 兜底记 0 但沿用 `estimated`**: cost 会算出 `0.0`,"免费"与"未知"在数据上不可区分,账目缺口无法量化。
|
||||
- **被否决 · 保留 est 兜底只修 prompt/completion 分配比例**: 比例是又一个没有正确取值的魔数,且未触及"拿上界当实测"的根因,仍高估约 9 倍。
|
||||
- **被否决 · 固定默认常量(如 1000)**: 与配额规模无关,在途上限随规模乱飘(`tpm=6000` 只剩 6 个在途、`tpm=600000` 放行 600 个)。派生值尺度无关且语义可文档化("一次调用约占一秒钟的配额份额")。
|
||||
- **被否决 · issue 原建议的遥测 p90 自估**: `TelemetryRecorder` 是纯只写端口,自估需新增读接口并强制所有后端(含 `none`)实现,公共 API 扩张远大于它要省掉的一个可选字段,且无实测证据表明派生默认值不够用。**注**: 初稿曾以"把两条方向相反的降级铁律焊在一起"为主论据,经独立审查撤回——p90 可在遥测读失败时回退纯派生值,限流侧仍能 fail-closed。
|
||||
- **被否决 · 派生逻辑取全局与单源 tpm 的较紧者**: 需改三处 `QuotaGate` 装配,且它修的是一个**既有**缺口(单源 `tpm=0` 而全局 `tpm>0` 时预扣为 0),属任务外,建议另开 issue。
|
||||
- **有意放弃的迁移保留项**: `migrations/chsanalyzer.md:151` 曾把"usage 缺失按 est 估算"列为保留(理由"保守计量")。本设计推翻:CHS 只记单个 `total_tokens` 不存在分配问题,而"保守"在计费语境只有错误一个方向。反静默的原始意图仍保留——被放弃的只是"编一个数字"这个手段。
|
||||
- **独立审查(两轮)抓出的两处实质缺陷**: ① 只改失败侧结算不够,`retry.py:338`/`embedding.py:271` 的**成功侧**取自同一返回值,改前 `actual` 恰等于预扣量使 `delta==0`,不同改则成功调用押金被整笔退回,对"从不回 usage 帧的网关源"构成系统性 TPM 失效;② OCR 行原判为"假陈述"是错的——`types.py:51` 与 `ocr.py:9` 明示 OCR 的 0 token 属**事实**,且改标 `unavailable` 会灌水本方案赖以成立的缺口度量,已剔出。
|
||||
- **待办**: 经 `writing-plans` 出实施计划;发版须同步 CHANGELOG 与 wiki(缺口查询口径须带 `AND cache_hit = false`),并回帖 issue #2。
|
||||
@@ -0,0 +1,40 @@
|
||||
---
|
||||
type: design
|
||||
node_id: design:response-observability-fields
|
||||
title: "响应可观测字段扩展(Issue #3)"
|
||||
date: 2026-07-31
|
||||
---
|
||||
|
||||
# 响应可观测字段扩展(Issue #3)
|
||||
|
||||
全文见 `2026-07-31-response-observability-fields-design.md`。来源: Gitea Issue #3(下游 dissect 的调用审计需求)。
|
||||
|
||||
## 选定方案
|
||||
|
||||
| 决策 | 选定 | 关键理由 |
|
||||
|---|---|---|
|
||||
| A 采集路径 | `TransportResult` 追加 `cached_prompt_tokens` / `model_reported` 强类型字段,解析留在 `openai_compat.py` | OpenAI 报文格式知识不出 `transports/`,middleware 只做搬运(P7) |
|
||||
| B 缓存命中语义 | 原样回放;度量口径必须带 `cache_hit = false` | 与 `_rehydrate` 既有口径一致——它只覆写时序字段,`model`/`prompt_tokens` 全回放 |
|
||||
| C 缓存单价 | `ModelPrice` 加可选 `cached_input_per_1m`,`cost()` 加可选参 | 旧价格表与 `embedding.py:419` 三参调用零改动;未配置该档时**不猜折扣率**,退化为全额计价 |
|
||||
| D 遥测扩列 | 端口 18 → 20 字段;DDL 加列 + 初始化期幂等补列 | `CREATE TABLE IF NOT EXISTS` 不会给旧库补列,INSERT 会**逐行 warning 丢弃**——遥测全失却无硬失败提示 |
|
||||
|
||||
## 被否决的备选
|
||||
|
||||
| 备选 | 否决原因 |
|
||||
|---|---|
|
||||
| A1 往 `raw` 里塞约定键 | `dict[str, Any]` 沦为隐式契约,且 middleware 要懂 OpenAI 嵌套结构 |
|
||||
| A3 middleware 内解析 raw | 报文格式知识进 middleware,新增非 OpenAI 兼容 transport 时会分叉,违反分层 |
|
||||
| B2 命中时置 None / B3 混合 | 与同层 `prompt_tokens` 的回放行为不一致,下游要记两套规则 |
|
||||
| C2 `cost()` 直接收 `LLMResponse` | `pricing.py` 会反向依赖 `types.py`,且纯函数难单测 |
|
||||
| D2 只改 DDL、文档写「删表重建」 | 已建表的开发机/下游只会看到降级 warning,排查成本高 |
|
||||
| D3 引入 alembic 迁移框架 | 新增依赖违反「依赖极简」铁律,规模严重不匹配 |
|
||||
|
||||
## 独立审查修正(2026-07-31)
|
||||
|
||||
Codex CLI 安装损坏(vendor 二进制缺失),改由全新上下文的 Claude subagent 审。三条问题全部核实属实并已折回设计:
|
||||
|
||||
1. PG 缺列时**不是**结构性短路,而是逐行 warning(`_failed` 仅在 `_ensure_ready` 置位)。
|
||||
2. SQLite 补列若塞进 `__init__` 现有 try,异常会让 `_conn` 停在 `None` → recorder 永久 no-op。已定纪律: 独立 try、置于 `self._conn = conn` 之后、duplicate column 视为成功。
|
||||
3. 「端口无默认值 → 漏改即报错」不成立(无 mypy,8 个 fake 全是 `**fields`)。改为新增「emitter 实参键集合 == `_COLUMNS`」契约测试兜底——否则 `KeyError` 会被 `_record` 的 `except Exception` 吞成 warning,静默丢遥测。
|
||||
|
||||
相关: [[m1-core-design]]、[[est-tokens-decoupling]]
|
||||
@@ -0,0 +1,17 @@
|
||||
---
|
||||
type: design
|
||||
node_id: design:sampling-params
|
||||
title: "采样参数透传设计(issue #4)"
|
||||
date: 2026-07-31
|
||||
---
|
||||
|
||||
# 采样参数透传设计(issue #4)
|
||||
|
||||
正文: `2026-07-31-sampling-params-design.md`。状态: 待人类审批。
|
||||
|
||||
- **选定方案**: 两层入口——调用级 `chat(..., overlay=)` 供逐 rollout 变化的 `seed`,配置级 `SourceConfig.extra_body`(env 键 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`,JSON 串)供恒定的 `temperature=0`。优先级 **结构化注入 > 调用级 > 配置级** 由现有层序天然给出,不加新机制。
|
||||
- **issue 未提但必须一并处理的四件事**: ① 采样参数进缓存 key(否则 5 个 seed 全命中同一缓存、标准差恒为 0,实验静默作废——「无缓存毒化」铁律);② 保护键黑名单 `{model, messages, stream, stream_options}` 与值可 JSON 序列化,均在构造期报错(覆盖它们会击穿流式看门狗、成本遥测与 TPM 结算;不可序列化的值会在 `CacheMW` 降级 try 之外抛裸 `TypeError`,一行遥测都没有);③ 采样参数入遥测(端口 20 → 21 字段,列名 `sampling`);④ 入参拷贝语义。
|
||||
- **关键结构决策**: `ChatRequest` 增 `sampling` 快照字段作为**跨洋葱层恒定的读取点**。`request.overlay` 在 `StructuredMW` 内侧含 `response_format`、外侧不含,缓存 key 与三个遥测 emit 入口若各读各的层就会口径分叉。`sampling` 列语义定死为「调用方意图 ⊎ 生效源 `extra_body`」,**不含**结构化注入。
|
||||
- **被否决备选及理由**: `chat()` 展开为 `temperature=`/`seed=` 具名参数(供应商私有参数无穷尽,等于永久追加签名,违「深模块窄接口」);配置级放装配层全局字典(采样参数与源强相关,会把无效键发给不认识它的源);overlay 不进 key 靠调用方传 `cache_salt`(把毒化防护责任推给调用方,漏传不报错——正是 issue 抱怨的失败形态);采样参数不入遥测由下游 run 快照自记(中间态数据不可追溯,且分两步要做两遍 DDL 迁移);缓存与遥测直接读 `request.overlay` 不加 `sampling` 字段(口径必分叉);`sampling` 记含 `response_format` 的完整合并结果(列名为采样参数,且数 KB schema 逐行落库无谓膨胀);给 embedding 加 `extra_body` 透传(embedding 无采样一说,装配期报错比静默无效更能指路);transport 层重复校验保护键(三入口已构造期收口,属 gold-plating)。
|
||||
- **附带修正**: `providers.py` 的 `minimax`/`openai` 空 thinking profile 补后果说明(`enable_thinking=False` 对两者不产生效果,调用方以为关掉了实际没关);`_SOURCE_FIELDS` 跨 scope 共用导致 `EXTRA_BODY` 在 OCR/EMBED scope 静默无效,改为构造期**剥离 + warning**(2026-07-31 人类拍板由原「装配期 `ValueError`」改此档: 这两条路径无采样语义,不值得让下游装配起不来)。**剥离不可省**——不剥离则遥测会记录一个从未发出的参数(决策 D 的 merge 读 `source.extra_body`,而 `monkey_ocr` 只发 multipart、`embed` payload 硬编码),那是数据造假而非参数失效;在 emitter 内特判调用方身份则违「遥测调用点收敛单一 helper」铁律。
|
||||
- **审查留痕**: Codex CLI 不可用(vendor 二进制缺失),改派全新上下文 subagent 两轮只读审查。首轮报 5 项必修(三个 emit 入口口径分叉、OCR/embedding 耦合、JSON 序列化缺口、注释归属写反、同步清单漏 4 处),逐条核实后全部采纳;次轮结论通过,其 5 条建议(承重不变式测试、`sampling` 类型定死、拷贝语义跟进、共用范围收窄、报错文案指路)亦已就地收进。
|
||||
@@ -0,0 +1,18 @@
|
||||
---
|
||||
type: design
|
||||
node_id: design:settings-invariant-guards
|
||||
title: "GatewaySettings 跨字段不变量守卫的生效范围"
|
||||
date: 2026-07-29
|
||||
---
|
||||
|
||||
# GatewaySettings 跨字段不变量守卫的生效范围
|
||||
|
||||
全文见 [2026-07-29-settings-invariant-guards-design.md](2026-07-29-settings-invariant-guards-design.md)。
|
||||
|
||||
- **缘起**: 社区 PR#1 指出装配守卫只挂在 `from_env`,走 CLAUDE.md §4.5 的另一条官方路 `from_settings()` 能装出违反类不变量的配置且不报错。诊断采纳,实现按库内规范重写并扩大覆盖。
|
||||
- **选定方案**: A——四条跨字段不变量(lease/stall/probe_ttl/sources 非空)全部收进 `GatewaySettings.__post_init__`,拆 `_validate_*` 私有方法,与 `types.py` 同族五个 frozen dataclass 的既有笔迹一致;模块级 `_guard_lease/_guard_stall` 删除。
|
||||
- **关键理由**: 这三条约束是**类的定义**的一部分,不是 `from_env` 的输入检查;放在函数里类就失去自我描述能力。构造期一处覆盖六个工厂 + 直接构造 + `dataclasses.replace`。
|
||||
- **被否决备选**: B 六个工厂各调 `validate()`(六处永久同步,新增 client 必漏,`replace` 仍绕过);C 公共 `validate()` 自愿调用(把不变量降级为建议,违反 P5 与 ARCH §7.3"拒绝装配")。
|
||||
- **补 PR#1 的两个缺口**(实测):`probe_ttl_s ≥ 最慢 timeout + 5` 仍只在 `from_env`(直接构造未拦截);守卫上移后 `sources=()` 泄漏内置异常 `max() arg is an empty sequence`。
|
||||
- **承诺变化**: 经 `from_env` 装配的调用方零影响;手工构造/`replace` 出非法组合者由静默故障改为构造期 `ValueError`。发版走 1.0.1(patch,用户拍板;CHANGELOG 单列"行为收紧"小节代替版本号预警),wiki `参考-配置键` 表述与新行为一致无需改。
|
||||
- **子决策**: 错误消息只点字段名不列 env 键(`types.py` 既有笔迹 + 键名单一事实源在 `.env.example`/wiki)。
|
||||
@@ -0,0 +1,17 @@
|
||||
---
|
||||
type: design
|
||||
node_id: design:settings-invariants-round-2
|
||||
title: "GatewaySettings 装配校验补齐(第二轮)"
|
||||
date: 2026-07-30
|
||||
---
|
||||
|
||||
# GatewaySettings 装配校验补齐(第二轮)
|
||||
|
||||
全文见 [2026-07-30-settings-invariants-round-2-design.md](2026-07-30-settings-invariants-round-2-design.md)。第一轮见 [settings-invariant-guards](settings-invariant-guards.md)。
|
||||
|
||||
- **缘起**: 第一轮交付后独立 verifier 发现 `from_env` 上还留着 15 条同族校验(枚举合法域 6、条件必填 7、标量域 2),`from_settings` 与直接构造全部放行。
|
||||
- **严重性高于第一轮**: `client.py:262/282/302/312/316` 有 5 处 `assert ... # 内部不变量: config 已校验` 明文依赖这个前提;实测断言开启抛裸 `AssertionError`,`python -O` 下退化为 redis 库天书。
|
||||
- **方案**: 沿用第一轮已批准的方案 A,不重新论证;新增 `_validate_backends/_validate_cache/_validate_telemetry`,枚举合法域上提为模块级常量供 `_load_pgw` 与构造期共用。
|
||||
- **assert 处置**: **保留不改**——前提一旦由构造期保证,它就是 CLAUDE.md §4.3 认可的内部不变量用法且给类型检查器收窄 `str | None`;只改那句会变成谎言的注释,点明由哪个方法保证。
|
||||
- **本轮唯一新决策**: `_load_pg_dsn` 剥 `+asyncpg` 驱动后缀是**规范化**不是校验,直接构造那条路不会剥。选校验拒绝(显式)而非构造期 `object.__setattr__` 剥后缀(在用户背后改 frozen 字段)。两条路接受度不同是有意的:env 路要吃三项目历史遗留的 SQLAlchemy DSN 写法,代码构造路没有历史包袱。
|
||||
- **版本**: 1.0.2(patch),CHANGELOG 单列"行为收紧"小节。
|
||||
@@ -0,0 +1,44 @@
|
||||
# 文档组织与维护约定(Gitea Wiki)
|
||||
|
||||
> **定位**: 用户文档站 = Gitea Wiki(`https://gitea.iomgaa.online/iomgaa/PolyGateway/wiki`);本文规定它的结构、更新时机与写作纪律。研发知识(设计/决策/验收)仍归 `research-wiki/`,两者职责不重叠。
|
||||
|
||||
## 1. 结构:Diátaxis 四区(2026-07-23 建站,17 页)
|
||||
|
||||
| 区 | 页面 | 职责(读者此刻要干什么) | 禁止 |
|
||||
|---|---|---|---|
|
||||
| 教程 | `教程-十分钟接入` | 新手被领着走通一遍 | 塞选项枚举与原理论述 |
|
||||
| 指南(How-to) | `指南-{多源与选源,限流与熔断,响应缓存,遥测与成本,结构化输出,OCR,Embedding,迁移既有项目}` | 一页一任务:配置片段+行为+坑 | 重复参考区的全量表 |
|
||||
| 参考 | `参考-{公共API,配置键,异常}` | 查表:签名/字段/键,**以源码实测为准** | 叙述与劝导 |
|
||||
| 解释 | `解释-{架构,错误四分类,治理行为,降级与取消}` | 讲为什么;机制挂回压测病灶 | 写成使用说明 |
|
||||
|
||||
导航:`Home.md`(按意图分流表)+ `_Sidebar.md`(全页目录);页间互链用 Gitea `[[双括号]]` 语法。
|
||||
|
||||
## 2. 更新时机(与代码变更绑定,发版检查清单)
|
||||
|
||||
| 变更类型 | 必须同步的页 |
|
||||
|---|---|
|
||||
| 新公共 API / 新能力 | 对应指南页(新增或扩写)+ `参考-公共API` + 侧边栏 + CHANGELOG |
|
||||
| 新增/改名配置键 | `参考-配置键` + 相关指南页的配置片段 + 主仓库 `.env.example` |
|
||||
| 治理行为变更(重试/熔断/选源语义) | `解释-治理行为` + 受影响指南页;若改公共承诺另走 brainstorming 流程 |
|
||||
| 新异常/分类语义调整 | `参考-异常` + `解释-错误四分类` |
|
||||
| **发版(任何版本号)** | `Home.md` 版本号与安装命令 + 主仓库 `CHANGELOG.md` + `README.md` 版本相关处;过一遍上面各行 |
|
||||
|
||||
**门**: 版本 bump 的提交不允许单独存在——同一次交付里必须包含对应的 wiki/CHANGELOG 同步(发布检查清单第一项)。
|
||||
|
||||
## 3. 写作纪律
|
||||
|
||||
- 中文;表格优先;单个代码块 ≤ 15 行;每个配置片段可直接复制运行。
|
||||
- **事实以源码为准**:参考区改动前先对照 `__init__.py` 导出面、`client.py`/`ocr.py`/`embedding.py` 签名与 `.env.example`;不确定就实测,不凭记忆写。
|
||||
- 深度内容(决策论证、迁移全文、验收数字)**只放指针**指向主仓库 `research-wiki/`,不复制——避免双处维护同一事实。
|
||||
- API 参考坚持**手写精选**(公共面小 + 只增不删承诺,手写比自动生成可读且低维护);若公共面显著膨胀再评估 mkdocstrings。
|
||||
|
||||
## 4. 更新操作
|
||||
|
||||
Wiki 是独立 git 仓库,两种改法:
|
||||
|
||||
```bash
|
||||
git clone https://gitea.iomgaa.online/iomgaa/PolyGateway.wiki.git # 批量改: clone→编辑→push
|
||||
# 或在 Gitea 网页 Wiki 页面上直接编辑(单页小改)
|
||||
```
|
||||
|
||||
文件名即页名(中文文件名);`Home.md` 是落地页,`_Sidebar.md` 是导航,新增页必须同步进侧边栏与 Home 分流表。凭据在本机 osxkeychain(git)与 `~/.pypirc`(twine)。
|
||||
@@ -90,6 +90,46 @@
|
||||
"id": "finding:m4-acceptance",
|
||||
"label": "M4 迁移验收(GovDoc+CHS)",
|
||||
"type": "finding"
|
||||
},
|
||||
{
|
||||
"id": "design:settings-invariant-guards",
|
||||
"label": "GatewaySettings 跨字段不变量守卫的生效范围",
|
||||
"type": "design"
|
||||
},
|
||||
{
|
||||
"id": "design:settings-invariants-round-2",
|
||||
"label": "GatewaySettings 装配校验补齐(第二轮)",
|
||||
"type": "design"
|
||||
},
|
||||
{
|
||||
"id": "design:est-tokens-decoupling",
|
||||
"label": "est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)",
|
||||
"type": "design"
|
||||
},
|
||||
{
|
||||
"id": "plan:est-tokens-decoupling",
|
||||
"label": "est_tokens 解耦实施计划",
|
||||
"type": "plan"
|
||||
},
|
||||
{
|
||||
"id": "design:response-observability-fields",
|
||||
"label": "响应可观测字段扩展(Issue #3)",
|
||||
"type": "design"
|
||||
},
|
||||
{
|
||||
"id": "plan:response-observability-fields",
|
||||
"label": "响应可观测字段扩展实现计划",
|
||||
"type": "plan"
|
||||
},
|
||||
{
|
||||
"id": "design:sampling-params",
|
||||
"label": "采样参数透传设计(issue #4)",
|
||||
"type": "design"
|
||||
},
|
||||
{
|
||||
"id": "plan:sampling-params-plan",
|
||||
"label": "采样参数透传实现计划(issue #4)",
|
||||
"type": "plan"
|
||||
}
|
||||
],
|
||||
"links": [
|
||||
@@ -155,6 +195,34 @@
|
||||
"relation": "implements",
|
||||
"evidence": "T0-T14 逐节实现设计 §4-§10/§15",
|
||||
"added": "2026-07-22T09:33:05.357964+00:00"
|
||||
},
|
||||
{
|
||||
"source": "design:est-tokens-decoupling",
|
||||
"target": "design:m1-core-design",
|
||||
"relation": "refines",
|
||||
"evidence": "精化 M1 冻结的 est_tokens 双职责语义: 保留 TPM 预扣、推翻 usage 缺失按 est 兜底(m1-core-design.md:59,222),改记 0/0 + unavailable + cost NULL",
|
||||
"added": "2026-07-30T09:33:55.383401+00:00"
|
||||
},
|
||||
{
|
||||
"source": "plan:est-tokens-decoupling",
|
||||
"target": "design:est-tokens-decoupling",
|
||||
"relation": "implements",
|
||||
"evidence": "5 任务实现设计 §3.2 的 11 条改动项;任务排序经中间态破窗分析(先加能力→切调用点→三态生效→解绑约束)",
|
||||
"added": "2026-07-30T09:39:26.442986+00:00"
|
||||
},
|
||||
{
|
||||
"source": "plan:response-observability-fields",
|
||||
"target": "design:response-observability-fields",
|
||||
"relation": "implements",
|
||||
"evidence": "计划 T1-T7 逐条实现设计的 A2/B1/C1/D1 四个决策",
|
||||
"added": "2026-07-31T11:10:03.872049+00:00"
|
||||
},
|
||||
{
|
||||
"source": "plan:sampling-params-plan",
|
||||
"target": "design:sampling-params",
|
||||
"relation": "implements",
|
||||
"evidence": "11 个任务逐条覆盖设计的决策 A-G 与 §5 的 14 条测试清单",
|
||||
"added": "2026-07-31T16:59:35.657367+00:00"
|
||||
}
|
||||
]
|
||||
}
|
||||
+20
-4
@@ -1,18 +1,28 @@
|
||||
# Research Wiki 索引
|
||||
|
||||
> 自动生成,更新时间:2026-07-22 14:36 UTC
|
||||
> 自动生成,更新时间:2026-08-01 01:58 UTC
|
||||
|
||||
## design (10)
|
||||
## design (20)
|
||||
- [2026-07-20-m1-core-design](designs/2026-07-20-m1-core-design.md) `design:2026-07-20-m1-core-design`
|
||||
- [2026-07-20-m2-distributed-design](designs/2026-07-20-m2-distributed-design.md) `design:2026-07-20-m2-distributed-design`
|
||||
- [2026-07-21-m25-resilience-design](designs/2026-07-21-m25-resilience-design.md) `design:2026-07-21-m25-resilience-design`
|
||||
- [2026-07-21-m3-ocr-design](designs/2026-07-21-m3-ocr-design.md) `design:2026-07-21-m3-ocr-design`
|
||||
- [2026-07-22-m4-migration-design](designs/2026-07-22-m4-migration-design.md) `design:2026-07-22-m4-migration-design`
|
||||
- [2026-07-29-settings-invariant-guards-design](designs/2026-07-29-settings-invariant-guards-design.md) `design:2026-07-29-settings-invariant-guards-design`
|
||||
- [2026-07-30-est-tokens-decoupling-design](designs/2026-07-30-est-tokens-decoupling-design.md) `design:2026-07-30-est-tokens-decoupling-design`
|
||||
- [2026-07-30-settings-invariants-round-2-design](designs/2026-07-30-settings-invariants-round-2-design.md) `design:2026-07-30-settings-invariants-round-2-design`
|
||||
- [2026-07-31-response-observability-fields-design](designs/2026-07-31-response-observability-fields-design.md) `design:2026-07-31-response-observability-fields-design`
|
||||
- [2026-07-31-sampling-params-design](designs/2026-07-31-sampling-params-design.md) `design:2026-07-31-sampling-params-design`
|
||||
- [est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2)](designs/est-tokens-decoupling.md) `design:est-tokens-decoupling`
|
||||
- [GatewaySettings 装配校验补齐(第二轮)](designs/settings-invariants-round-2.md) `design:settings-invariants-round-2`
|
||||
- [GatewaySettings 跨字段不变量守卫的生效范围](designs/settings-invariant-guards.md) `design:settings-invariant-guards`
|
||||
- [M1 核心里程碑设计:公共签名冻结与治理栈落地](designs/m1-core-design.md) `design:m1-core-design`
|
||||
- [M2 分布式:Redis 治理后端+背压+Postgres 遥测+pricing+Embedding+压测 harness](designs/m2-distributed.md) `design:m2-distributed`
|
||||
- [M2.5 治理韧性: 半死源隔离与健康感知调度](designs/m25-resilience.md) `design:m25-resilience`
|
||||
- [M3 OCR 端口族设计](designs/m3-ocr.md) `design:m3-ocr`
|
||||
- [M4 迁移验证设计(GovDoc→CHS,发 v1.0)](designs/m4-migration.md) `design:m4-migration`
|
||||
- [响应可观测字段扩展(Issue #3)](designs/response-observability-fields.md) `design:response-observability-fields`
|
||||
- [采样参数透传设计(issue #4)](designs/sampling-params.md) `design:sampling-params`
|
||||
|
||||
## finding (11)
|
||||
- [2026-07-20-m2-soak-workload](findings/2026-07-20-m2-soak-workload.md) `finding:2026-07-20-m2-soak-workload`
|
||||
@@ -27,20 +37,26 @@
|
||||
- [P6 混合浸泡首跑基线与记分板三重伪击穿修复](findings/p6-soak-baseline.md) `finding:p6-soak-baseline`
|
||||
- [P7 OCR soak 验收: 99.73% 与 13 不变量全 PASS](findings/p7-ocr-soak.md) `finding:p7-ocr-soak`
|
||||
|
||||
## plan (10)
|
||||
## plan (16)
|
||||
- [2026-07-20-m1-core-plan](plans/2026-07-20-m1-core-plan.md) `plan:2026-07-20-m1-core-plan`
|
||||
- [2026-07-20-m2-distributed-plan](plans/2026-07-20-m2-distributed-plan.md) `plan:2026-07-20-m2-distributed-plan`
|
||||
- [2026-07-21-m25-resilience-plan](plans/2026-07-21-m25-resilience-plan.md) `plan:2026-07-21-m25-resilience-plan`
|
||||
- [2026-07-21-m3-ocr-plan](plans/2026-07-21-m3-ocr-plan.md) `plan:2026-07-21-m3-ocr-plan`
|
||||
- [2026-07-22-m4-migration-plan](plans/2026-07-22-m4-migration-plan.md) `plan:2026-07-22-m4-migration-plan`
|
||||
- [2026-07-30-est-tokens-decoupling-plan](plans/2026-07-30-est-tokens-decoupling-plan.md) `plan:2026-07-30-est-tokens-decoupling-plan`
|
||||
- [2026-07-31-response-observability-fields](plans/2026-07-31-response-observability-fields.md) `plan:2026-07-31-response-observability-fields`
|
||||
- [2026-07-31-sampling-params](plans/2026-07-31-sampling-params.md) `plan:2026-07-31-sampling-params`
|
||||
- [est_tokens 解耦实施计划](plans/est-tokens-decoupling.md) `plan:est-tokens-decoupling`
|
||||
- [M1 核心里程碑实现计划](plans/m1-core-plan.md) `plan:m1-core-plan`
|
||||
- [M2 分布式实现计划](plans/m2-distributed.md) `plan:m2-distributed`
|
||||
- [M2.5 治理韧性实现计划](plans/m25-resilience.md) `plan:m25-resilience`
|
||||
- [M3 OCR 实现计划](plans/m3-ocr.md) `plan:m3-ocr`
|
||||
- [M4 迁移实现计划(T0-T14)](plans/m4-migration.md) `plan:m4-migration`
|
||||
- [响应可观测字段扩展实现计划](plans/response-observability-fields.md) `plan:response-observability-fields`
|
||||
- [采样参数透传实现计划(issue #4)](plans/sampling-params-plan.md) `plan:sampling-params-plan`
|
||||
|
||||
## schema (1)
|
||||
- [表结构: llm_calls(遥测 18 字段)](schemas/llm-calls.md) `schema:llm-calls`
|
||||
- [表结构: llm_calls(遥测 21 字段)](schemas/llm-calls.md) `schema:llm-calls`
|
||||
|
||||
## metric (2)
|
||||
- [OCR 治理调用成功率与错误分类分布](metrics/ocr-call-success.md) `metric:ocr-call-success`
|
||||
|
||||
@@ -46,3 +46,30 @@
|
||||
- [2026-07-22 09:33 UTC] 重建索引: 32 篇页面
|
||||
- [2026-07-22 14:36 UTC] 新增 finding: M4 迁移验收(GovDoc+CHS) (finding:m4-acceptance)
|
||||
- [2026-07-22 14:36 UTC] 重建索引: 34 篇页面
|
||||
- [2026-07-30 03:44 UTC] 新增 design: GatewaySettings 跨字段不变量守卫的生效范围 (design:settings-invariant-guards)
|
||||
- [2026-07-30 03:44 UTC] 重建索引: 36 篇页面
|
||||
- [2026-07-30 04:44 UTC] 新增 design: GatewaySettings 装配校验补齐(第二轮) (design:settings-invariants-round-2)
|
||||
- [2026-07-30 04:44 UTC] 重建索引: 38 篇页面
|
||||
- [2026-07-30 09:32 UTC] 新增 design: est_tokens 解耦: 拆分限流预扣与遥测用量兜底(issue #2) (design:est-tokens-decoupling)
|
||||
- [2026-07-30 09:33 UTC] 重建索引: 40 篇页面
|
||||
- [2026-07-30 09:33 UTC] 新增边: design:est-tokens-decoupling --refines--> design:m1-core-design
|
||||
- [2026-07-30 09:33 UTC] 重建索引: 40 篇页面
|
||||
- [2026-07-30 09:39 UTC] 新增 plan: est_tokens 解耦实施计划 (plan:est-tokens-decoupling)
|
||||
- [2026-07-30 09:39 UTC] 新增边: plan:est-tokens-decoupling --implements--> design:est-tokens-decoupling
|
||||
- [2026-07-30 09:39 UTC] 重建索引: 42 篇页面
|
||||
- [2026-07-31 08:35 UTC] 新增 design: 响应可观测字段扩展(Issue #3) (design:response-observability-fields)
|
||||
- [2026-07-31 08:37 UTC] 重建索引: 44 篇页面
|
||||
- [2026-07-31 11:10 UTC] 新增 plan: 响应可观测字段扩展实现计划 (plan:response-observability-fields)
|
||||
- [2026-07-31 11:10 UTC] 新增边: plan:response-observability-fields --implements--> design:response-observability-fields
|
||||
- [2026-07-31 11:10 UTC] 重建索引: 46 篇页面
|
||||
- [2026-07-31 11:11 UTC] 重建索引: 46 篇页面
|
||||
- [2026-07-31 12:25 UTC] 重建索引: 46 篇页面
|
||||
- [2026-07-31 15:57 UTC] 新增 design: 采样参数透传设计(issue #4) (design:sampling-params)
|
||||
- [2026-07-31 15:57 UTC] 重建索引: 48 篇页面
|
||||
- [2026-07-31 16:00 UTC] 重建索引: 48 篇页面
|
||||
- [2026-07-31 16:31 UTC] 重建索引: 48 篇页面
|
||||
- [2026-07-31 16:59 UTC] 新增 plan: 采样参数透传实现计划(issue #4) (plan:sampling-params-plan)
|
||||
- [2026-07-31 16:59 UTC] 新增边: plan:sampling-params-plan --implements--> design:sampling-params
|
||||
- [2026-07-31 16:59 UTC] 重建索引: 50 篇页面
|
||||
- [2026-07-31 17:01 UTC] 重建索引: 50 篇页面
|
||||
- [2026-08-01 01:58 UTC] 重建索引: 50 篇页面
|
||||
|
||||
@@ -92,7 +92,7 @@
|
||||
|
||||
| 项目键(.env.example 实测) | 库对应 | 差异 |
|
||||
|---|---|---|
|
||||
| `{SCOPE}__{PROVIDER}__{N}__{FIELD}` | 同名继承 | 无;`EST_TOKENS` 见 ⚠️ G2 |
|
||||
| `{SCOPE}__{PROVIDER}__{N}__{FIELD}` | 同名继承 | 无;`EST_TOKENS` 键保留不改名,但 2026-07-30 起由必填降为可选(G2 已闭,详见该行) |
|
||||
| `{SCOPE}__GLOBAL__MAX_CONCURRENCY/RPM/TPM`(config.py:249-271) | 全局闸限额 | ARCH §9 未定义 GLOBAL 段命名,M2 设计须定(建议原样继承) |
|
||||
| `{SCOPE}__SELECTOR`(round_robin/least_inflight) | `SourceSelector` 策略选择 | 命名待 M1/M2 定稿,建议继承 |
|
||||
| `{SCOPE}__RETRY__MAX_ATTEMPTS/BACKOFF_BASE_S/BACKOFF_MAX_S` | RetryPolicy | ⚠️ G4:ARCH §9 只列平铺 `LLM_MAX_RETRIES` 等键,无 per-scope 形态 |
|
||||
@@ -148,7 +148,7 @@ stack = ExtractionProviderStack(
|
||||
| Retry-After 仅支持秒数形态(invokers.py:127-141) | HTTP-date 返回 None | **有意放弃** date 形态(ARCH §6.2 同款) |
|
||||
| 429 body 细分 insufficient_quota → SourceDead(invokers.py:144-166) | 欠费≠限流 | **保留**(ARCH §6.1 已承诺) |
|
||||
| 零 content 提前结束 → Transient "early_eof";有 content 缺 [DONE] → 打捞并埋点 "missing_done"(invokers.py:306-313) | 线路级异常定性(D2 的核心价值) | **保留** |
|
||||
| usage 缺失按 est_tokens 估算并标 `estimated`,不静默用 0(invokers.py:241-254) | 保守计量 | **保留**(`usage_source` 已进 ARCH §5.1;依赖 G2) |
|
||||
| usage 缺失按 est_tokens 估算并标 `estimated`,不静默用 0(invokers.py:241-254) | 保守计量 | **有意放弃**(2026-07-30 推翻原"保留"判定,est_tokens 解耦设计 §4)。理由:CHS 只记单个 `total_tokens`,不存在 prompt/completion 分配问题;库拆成两列后无法忠实分配,而 `est_tokens` 按 CHS 自身定义(config.py:55)是**最坏情形上界**——"保守"在限流语境安全(押多了只是慢),在计费语境只有虚高一个方向。库改为如实记 `0/0` + `usage_source="unavailable"` + cost NULL(ARCH §5.1)。**"不静默用 0"的原始意图完整保留**:被放弃的只是"编一个数字"这个手段,缺失依然有显式标注且可被 `WHERE usage_source='unavailable' AND cache_hit = false` 量化 |
|
||||
| reasoning_content 刷新活性但不计入结果;ttft=首个任意 token(invokers.py:55-79, 336-364) | 防 thinking 模型被看门狗误杀 | **替换+增强**:库把 thinking 收进 `LLMResponse.thinking`(不再丢弃);活性语义必须保留(反向约束 M1) |
|
||||
| enable_thinking=True 不注入参数、False 注入关闭参数(invokers.py:230-238) | 与 D11 注册表"声明注入方式"方向相反 | **替换**(provider 注册表须支持"注入关闭参数"形态) |
|
||||
| 图片 magic bytes 探测,非 PNG/JPEG 抛 RequestRejected(invokers.py:116-123) | 本地快速拒绝 | **保留**(移入库 transport) |
|
||||
@@ -182,7 +182,7 @@ stack = ExtractionProviderStack(
|
||||
| R5 | RequestRejected 二分(真实响应记成功/本地拒绝释放探针);换源重试跨源计数口径 | M2 | 须进 M2 设计 |
|
||||
| R6 | OCR ZIP 协议 + bbox 数值防御下沉;OCR Usage=0;glm 白名单预留 | M3 | §7.10 已覆盖 |
|
||||
| **G1** | ✅ 已闭(M3 核实): 库 `GatewayUnavailableError` 一族自 M1 起携 `scope/reason/retry_after_s/per_source_reasons`(errors.py:74-105),chat/embedding/OCR 三循环抛出点均已填充且有契约测试钉住;项目侧仅剩约 10 行翻译 shim(库异常 → ProviderUnavailableError)或 tracking.py 直接 except 库异常 | M2 | 已闭 |
|
||||
| **G2** | ⚠️ `est_tokens`(TPM 预扣常量 + usage 缺失兜底,config.py:55)不在 ARCH §7.7 SourceConfig 字段清单;§7.3 `try_acquire(source, est_tokens)` 的 est 来源未定义 | M2 | **架构缺口**,修订 §7.7 |
|
||||
| **G2** | ✅ 已闭(2026-07-30 核实): `est_tokens` 已进 ARCH §7.7 SourceConfig 字段清单,且 §7.3 `try_acquire` 的 est 来源已定义为 `SourceConfig.effective_est_tokens()`(显式值优先,否则按 `tpm // 60` 派生)。双职责一并拆开:该字段只剩 TPM 预扣的可选调优覆盖,usage 缺失不再由它兜底(见 §7 行 151 的推翻判定) | M2 | 已闭 |
|
||||
| **G3** | ⚠️ §4.3 层序图文矛盾:图示 熔断→限流→重试(重试最内),但理由要求"每次重试重新过限流闸"且熔断/限流是 per-source 的、选源在重试循环内(governance.py:120-167 实践为每次尝试执行 选源→冷却备忘→permit→熔断门)。洋葱不澄清"逐次准入"机制则多源语义无法成立 | M2 | **架构缺口**,澄清 §4.3/§4.4 |
|
||||
| **G4** | ⚠️ per-scope 韧性配置命名(`{SCOPE}__RETRY__*`/`BREAKER__*`/`BACKPRESSURE__*`/`SELECTOR`/`GLOBAL__*`)未进 ARCH §9,现文只有平铺 `LLM_*` 键;CHSAnalyzer 的 VLM/OCR 两 scope 参数各异,平铺键无法表达 | M2 | **架构缺口**,修订 §9 |
|
||||
| **G5** | ⚠️ 半开探针租约 TTL(探针持有者死亡后 TTL 过期自动可再探,scripts.py:96-105)与 `release_probe` 操作未见于 ARCH §7.4(只写单探针/epoch fencing);缺失则探针死锁 | M2 | **架构缺口**,修订 §7.4 |
|
||||
|
||||
@@ -130,7 +130,7 @@ loop = AgentLoop(client, max_steps=...,
|
||||
| B10 | 缓存命中也记遥测(cache_hit=True, latency_ms=0)(client.py:309-331);每次 attempt 独立 call_id(client.py:337);thinking 帧 content 优先于 reasoning_content(client.py:79-91) | **保留**(库同款语义) |
|
||||
| B11 | SSE 流提前断开且未见 `[DONE]` 时正常返回:`usage_sink["done"]` 写入后无人检查(client.py:115-117),截断响应被当成功**并写入缓存** | **修复**: 库把"断流无 [DONE]"定性 TransientError(§6.1),且坏结果不进缓存 |
|
||||
| B12 | provider 差异靠字符串猜: `"deepseek" in provider`/`"qwen" in provider` 注入 thinking 参数(client.py:139-144)、`<think>` 剥离(client.py:348) | **替换**: provider 注册表(D11) |
|
||||
| B13 | usage 帧缺失时 prompt/completion_tokens 落 0(client.py:354-355),无标注 | **升级**: 库 `usage_source=measured/estimated` |
|
||||
| B13 | usage 帧缺失时 prompt/completion_tokens 落 0(client.py:354-355),无标注 | **升级**: 库 `usage_source` 三态 `measured/estimated/unavailable`(2026-07-30 est_tokens 解耦后由两态扩为三态)。GovDoc 原行为"落 0 且无标注"中的落 0 反而与库一致,被升级的是**标注**:缺失行记 `unavailable` 且 cost 为 NULL,缺口可被 `WHERE usage_source='unavailable' AND cache_hit = false` 量化 |
|
||||
| B14 | 遥测 schema 无 source_name/cost/usage_source(telemetry_sqlite.py:29-48) | **升级**: 库超集 schema;GovDoc 骨架期无生产遥测数据,直接换新库文件,不做数据迁移 |
|
||||
| B15 | 双层重试: 治理层 max_retries + AgentLoop 步级 step_retries(loop.py:308-356,默认延迟 (20,40)s) | **有意保留**(业务侧任务级重试,ARCHITECTURE §7.2 允许留在库外),但须按 §4 改 retryable_exceptions,否则静默失效 |
|
||||
| B16 | `CancelledError` 穿透重试循环(client.py:396 `except Exception` 天然放行),取消的调用**不记遥测** | **保留**穿透;取消是否记遥测库未定义,见 §9-R6 |
|
||||
|
||||
@@ -0,0 +1,227 @@
|
||||
# est_tokens 解耦实施计划
|
||||
|
||||
- **目标**: 把 `SourceConfig.est_tokens` 的两个职责(TPM 入场预扣 / usage 缺失时的遥测用量兜底)拆开,并解绑 `tpm > 0 ⇒ est_tokens > 0` 装配约束。
|
||||
- **方案概述**: usage 不可得时遥测记 `0/0` 并标新值 `unavailable`、cost 落 NULL(不再拿限流押金编计费数字);`est_tokens` 未填时由库按 `tpm // 60` 派生预扣量,字段降为可选调优覆盖(保留不删不改名)。全部依据已批准设计 `research-wiki/designs/2026-07-30-est-tokens-decoupling-design.md`,其 §3.2 的 11 条改动项已钉死行号。
|
||||
- **涉及技术**: Python 3.11 frozen dataclass、pytest(含 `tests/contracts` 双后端契约测试)、真实 Redis(integration/contracts)、pydantic-settings 不涉及改动。
|
||||
- **溯源**: Gitea issue #2;wiki 实体 `design:est-tokens-decoupling`。
|
||||
|
||||
## 1. 任务排序的硬约束(先读这节再动手)
|
||||
|
||||
三处改动互相牵制,**顺序错了会引入押金泄漏或计费造假**,且中间态不报错、只静默偏差:
|
||||
|
||||
| 若单独先做 | 后果 |
|
||||
|---|---|
|
||||
| 先改 usage 兜底为 `(0, 0)`,结算点还没切派生值 | 成功调用 `actual = 0` 而入场押了 `est_tokens`,`delta` 为负 → **押金整笔退回**,TPM 闸退化成进门即放行 |
|
||||
| 先切入场(`ratelimit.py:26`)+ 解绑约束,而结算点还没切 | `est_tokens=0` 的源入场押派生值、结算退 0 → 同样泄漏。**注意机制**:若**只**解绑约束而 `ratelimit.py` 一行未动,后果不是泄漏而是入场**完全不预扣**(仍传 `est_tokens=0`)——那是设计 §2.2 已否决的"不预扣"退化形态。两者都要避免,故 T4 必须在 T2 之后 |
|
||||
| 先改 `openai_compat.py:176` 而 `_merge` 还是二值逻辑 | embedding 的 `unavailable` 批被 `any(== "estimated")` 判 False 从而**误标 `measured`**,cost 照算 |
|
||||
|
||||
因此排序为:**先加能力(零行为变更)→ 再把所有调用点切到新能力(此时等价,因显式值优先)→ 再让三态生效 → 最后解绑约束**。Task 1-2 完成后行为逐字不变,Task 3 才是行为变更主体,Task 4 才让派生值真正启用。**不要合并或调换 Task 2 / 3 / 4 的顺序。**
|
||||
|
||||
## 2. 文件结构
|
||||
|
||||
| 文件 | 职责 | 涉及任务 |
|
||||
|---|---|---|
|
||||
| `src/polygateway/types.py` | 新增 `SourceConfig.effective_est_tokens()` 派生方法与 `usage_source` 三态值域常量;删除 `_validate_gates` 里的条件必填;修 `EmbeddingTransportResult` 行内注释 | T1、T3、T4 |
|
||||
| `src/polygateway/middleware/ratelimit.py` | `QuotaGate.try_acquire` 入场预扣改用派生值 | T2 |
|
||||
| `src/polygateway/middleware/retry.py` | 成功侧(`:338`)与失败侧(`:370`)结算改用派生值 | T2 |
|
||||
| `src/polygateway/embedding.py` | 同上(`:271`/`:294`);`_merge` 三态合并;`_total_cost` 存在不可得批时整体 NULL | T2、T3 |
|
||||
| `src/polygateway/transports/openai_compat.py` | 两处 usage 兜底改 `unavailable`;打捞覆盖加前置条件 | T3 |
|
||||
| `src/polygateway/middleware/telemetry.py` | cost 短路(插在 `cache_hit` 之后);失败尝试与终态失败改标 `unavailable` | T3 |
|
||||
| `src/polygateway/ocr.py` | **不改**(设计 §3.3 已剔出),仅加防回归测试 | T3 |
|
||||
| `research-wiki/ARCHITECTURE.md` | §7.7 行 428、§5.1 行 331、§4.4 行 305、§7.1 行 384 | T5 |
|
||||
| `research-wiki/migrations/chsanalyzer.md` | 行 151 保留项改判为有意放弃;G2(行 185)标注已解决 | T5 |
|
||||
| `.env.example` | 行 11 删除"TPM > 0 时 EST_TOKENS 必填 > 0" | T5 |
|
||||
| `CHANGELOG.md` | 行为变更小节(值域新增 + cost 口径 + 约束解绑) | T5 |
|
||||
|
||||
测试文件:`tests/unit/test_types.py`、`test_retry.py`、`test_openai_compat.py`、`test_telemetry.py`、`test_embedding.py`、`test_ocr_client.py`、`tests/contracts/test_limiter_contract.py`。
|
||||
|
||||
## 3. 保真校验
|
||||
|
||||
本计划**不新增**任何 `reference/` 移植代码,但触及 ARCHITECTURE §1.4 关键资产索引中的"遥测口径"与"限流结算",且**有意推翻**一条已声明保留的迁移行为(CHS `invokers.py:241-254` 的"usage 缺失按 est 估算",见设计 §4)。故设以下检查点,每个任务完成时逐条确认:
|
||||
|
||||
1. `Permit.settle()` 的"多退少补 + 幂等 flag"语义不得改变;`release()` 的 finally 必然执行不得改变。
|
||||
2. **不得触碰** `backends/redis/limiter.py` 的 Lua 脚本与 `backends/memory/limiter.py` 的窗口/租约算法——本计划只改传入 `try_acquire` 的**数值来源**,不改闸门算法。
|
||||
3. 不得改变 `errors.py` 四分类归属,不得新增运行时异常类型;值域违反不走异常路径(设计 §3.1 已裁决)。
|
||||
4. `ocr.py` 的 `settle` 恒 0 与 `usage_source="measured"` 保持不变。
|
||||
5. 遥测 18 字段冻结不变、无 DDL 变更(两 schema 的 `cost` 列已可空)。
|
||||
|
||||
## 4. 任务清单
|
||||
|
||||
### T1 — 加派生能力与值域常量(零行为变更)
|
||||
|
||||
- [x] **改** `src/polygateway/types.py`
|
||||
|
||||
新增模块级值域常量与 `SourceConfig` 方法。派生按**源自身 tpm**,全局 tpm 不参与(设计 §7 已声明为既有限制、本次不修):
|
||||
|
||||
```python
|
||||
USAGE_SOURCES = frozenset({"measured", "estimated", "unavailable"})
|
||||
"""usage_source 值域;仅约束库内生产侧取值,不在 frozen dataclass 上做运行时校验。"""
|
||||
|
||||
_EST_TOKENS_QUOTA_DIVISOR = 60
|
||||
"""未显式配置时的预扣量除数: 假定一次调用约占一秒钟的 TPM 配额份额。"""
|
||||
```
|
||||
|
||||
```python
|
||||
def effective_est_tokens(self) -> int:
|
||||
"""TPM 入场预扣量: 显式配置优先,否则按 tpm 派生(设计 §2.2)。"""
|
||||
if self.est_tokens > 0:
|
||||
return self.est_tokens
|
||||
if self.tpm > 0:
|
||||
return max(1, self.tpm // _EST_TOKENS_QUOTA_DIVISOR)
|
||||
return 0
|
||||
```
|
||||
|
||||
**验收标准**: 方法存在且为纯函数(不读全局、不 await);此任务**不修改任何调用点**,全库行为逐字不变。
|
||||
|
||||
**测试要求**(先失败后通过:方法不存在时 `AttributeError`)——`tests/unit/test_types.py`:
|
||||
- `tpm=6000, est_tokens=0` → `100`;`tpm=600000, est_tokens=0` → `10000`(尺度无关:两者在途上限同为 60)
|
||||
- `tpm=30, est_tokens=0` → `1`(下界不塌到 0)
|
||||
- `tpm=0, est_tokens=0` → `0`(TPM 闸未启用,不预扣)
|
||||
- `tpm=6000, est_tokens=4000` → `4000`(显式值优先于派生)
|
||||
- `USAGE_SOURCES` 恰为三元集合
|
||||
|
||||
**值域封闭的两条实质断言**(设计 §6 要求;缺了它们 `USAGE_SOURCES` 会沦为零消费者的死常量,且 §3.1 的落点裁决无回归保护):
|
||||
- **生产侧封闭**: 参数化覆盖库内全部 `usage_source` 生产点(`_resolve_usage`、`_resolve_embedding_usage`、`_merge`、`TelemetryEmitter.emit_*`),断言产出恒 ∈ `USAGE_SOURCES`。此断言在 T1 阶段即可写(此时产出仅 `measured`/`estimated`),T3 完成后自动覆盖 `unavailable`
|
||||
- **不做运行时校验**: `LLMResponse(usage_source="garbage")` 构造**不抛异常**——锁定设计 §3.1 的裁决(公共 frozen dataclass 不加 `__post_init__` 值域校验,否则裸 `ValueError` 不属四分类、会逃出 `chat()`)。没有这条,后人很容易顺手补上校验而击穿 `chat()`
|
||||
|
||||
**验证**: `conda run --no-capture-output -n PolyGateway pytest tests/unit/test_types.py -v` → 全 PASS
|
||||
|
||||
### T2 — 五个调用点切到派生值(零行为变更)
|
||||
|
||||
此时 `est_tokens > 0` 仍是必填(约束未解绑),故 `effective_est_tokens()` 恒返回显式值,**行为与改前逐字相同**。这一步只是把数值来源换掉,为 T3 铺路。
|
||||
|
||||
- [x] **改** `src/polygateway/middleware/ratelimit.py:26` — `source.est_tokens` → `source.effective_est_tokens()`
|
||||
- [x] **改** `src/polygateway/middleware/retry.py:370`(失败侧,`if not dead` 分支内)→ `source.effective_est_tokens()`
|
||||
- [x] **改** `src/polygateway/embedding.py:294`(失败侧)→ 同上
|
||||
- [x] **改** `src/polygateway/middleware/retry.py:338`(成功侧)— 加不可得分支:
|
||||
|
||||
```python
|
||||
if result.usage_source == "unavailable":
|
||||
actual = source.effective_est_tokens()
|
||||
else:
|
||||
actual = result.prompt_tokens + result.completion_tokens
|
||||
```
|
||||
|
||||
- [x] **改** `src/polygateway/embedding.py:271`(成功侧)— 同构,`actual = result.prompt_tokens` 落在 else 分支。
|
||||
|
||||
`TransportResult.usage_source`(`types.py:75`)与 `EmbeddingTransportResult.usage_source`(`types.py:273`)均为必填字段,在这两处的 `result` 局部变量上直接可读。
|
||||
|
||||
**验收标准**: 五处均不再直接读 `source.est_tokens`;新分支在本任务中**永不触发**(尚无 `unavailable` 生产者),现有测试全绿即证明零行为变更。**不要**在本任务修改 `retry.py:329` 的 `actual = 0` 初值,也不要动 `RequestRejectedError`/`ResultInvalidError`/`SourceDeadError` 三条失败分支(设计 §3.3:它们的 `actual` 停在 0 属既有行为)。
|
||||
|
||||
**测试要求**(回归保护,先失败后通过不适用于零行为变更任务,故以"现有测试不得回归 + 新增等价性断言"为门):
|
||||
- 新增 `tests/unit/test_retry.py`:`tpm=1000, est_tokens=400` 的源在 usage 正常返回时,settle 收到的 `actual` 等于实测 token 之和(锁定 else 分支);现有 `test_settle_uses_actual_usage` 必须继续通过
|
||||
- `tests/contracts/test_limiter_contract.py` 全绿(双后端)
|
||||
|
||||
**验证**:
|
||||
```
|
||||
conda run --no-capture-output -n PolyGateway pytest tests/unit tests/contracts -v
|
||||
```
|
||||
→ 全 PASS。**Redis 契约测试用真实 Redis db3,不得与其他 Redis 测试并跑**(时序隔离)。
|
||||
|
||||
### T3 — 值域三态生效(行为变更主体)
|
||||
|
||||
- [x] **改** `src/polygateway/transports/openai_compat.py:146` — 兜底不再读 `est_tokens`:
|
||||
|
||||
```python
|
||||
return 0, 0, "unavailable"
|
||||
```
|
||||
|
||||
- [x] **改** `src/polygateway/transports/openai_compat.py:176` — embedding 兜底同理 `return 0, "unavailable"`
|
||||
- [x] **改** `src/polygateway/transports/openai_compat.py:336` — 打捞覆盖加前置条件(否则 `0/0` 会被标 `estimated` 而算出假的 `0.0`):
|
||||
|
||||
```python
|
||||
if salvaged and usage_source == "measured":
|
||||
usage_source = "estimated" # 收到 usage 帧但流被截断: 数字真实、可信度降级
|
||||
```
|
||||
|
||||
- [x] **改** `src/polygateway/middleware/telemetry.py:130-135` — cost 短路,**插在 `cache_hit` 分支之后**(缓存命中未产生新调用,`0.0` 是事实):
|
||||
|
||||
```python
|
||||
if cache_hit:
|
||||
cost: float | None = 0.0
|
||||
elif usage_source == "unavailable":
|
||||
cost = None # 用量不可得: 宁可算不出成本,也不算错成本
|
||||
elif error is None and model and self._pricing is not None:
|
||||
cost = self._pricing.cost(model, prompt_tokens, completion_tokens)
|
||||
else:
|
||||
cost = None
|
||||
```
|
||||
|
||||
- [x] **改** `src/polygateway/middleware/telemetry.py:58` 与 `:100` — 失败尝试与终态失败的 `usage_source` 由 `"estimated"` 改 `"unavailable"`(cost 本已是 None,不改金额)
|
||||
- [x] **改** `src/polygateway/embedding.py:383,390` — 二值合并扩三态(优先级:任一不可得 → 整体不可得):
|
||||
|
||||
```python
|
||||
sources = {o.result.usage_source for o in outcomes}
|
||||
if "unavailable" in sources:
|
||||
merged_source = "unavailable"
|
||||
elif "estimated" in sources:
|
||||
merged_source = "estimated"
|
||||
else:
|
||||
merged_source = "measured"
|
||||
```
|
||||
|
||||
- [x] **改** `src/polygateway/embedding.py:397` `_total_cost` — 存在 `unavailable` 批时整体返回 `None`(逐批求和会给出偏低却看似有效的金额)
|
||||
- [x] **改** `src/polygateway/types.py:273` — 行内注释 `# measured | estimated` → 三态(内核里不留矛盾注释)
|
||||
- [x] **改** `src/polygateway/transports/openai_compat.py:142` 与 `:172` — 两个函数的中文 docstring 仍写着"缺失/非法按 `est_tokens` 保守兜底并标 `estimated`",改完不改就留下两句主动陈述旧行为的文档(与 `types.py:273` 同一把尺子)
|
||||
|
||||
**验收标准**: 全库不再有任何位置把 `est_tokens` 写进遥测用量;`ocr.py` 一字未动。
|
||||
|
||||
**测试要求**(先失败后通过):
|
||||
- `test_openai_compat.py`:`est_tokens=4000` + usage 帧缺失 → `(0, 0, "unavailable")`(**改前返回 `(0, 4000, "estimated")`,故先失败**;现有用例 `test_usage_missing_falls_back_to_est` 需改写)
|
||||
- **改写** `tests/unit/test_embedding.py:105 test_missing_usage_falls_back_estimated` — 现断言 `prompt_tokens == 7 and usage_source == "estimated"`(夹具 `est_tokens=7`),改 `openai_compat.py:176` 后必然变红,须改为 `(0, "unavailable")`。这是一条**位于 `test_embedding.py` 里的 transport 级用例**,容易在只盯 `test_openai_compat.py` 时漏掉
|
||||
- `test_openai_compat.py`:打捞 + usage 帧存在 → `estimated` **且 cost 非 None**;打捞 + usage 缺失 → `unavailable` **且 cost 为 None**。cost 配套断言不可省——#4 的真正目的就是防 `0/0` 被算成假的 `0.0`,只断言 `usage_source` 钉不住它
|
||||
- `test_telemetry.py`:成功行 `usage_source="unavailable"` → `record_llm_call` 收到 `cost=None`(改前按 4000×输出单价算出 `0.032`)
|
||||
- `test_telemetry.py`:`cache_hit=True` 且 `unavailable` → cost 仍为 `0.0`(锁定分支次序)
|
||||
- `test_telemetry.py`:失败尝试与终态失败行 `usage_source == "unavailable"`
|
||||
- `test_embedding.py`:混合批 `measured + unavailable` → 整体 `unavailable` 且 `cost is None`(改前误标 `measured`)
|
||||
- `test_ocr_client.py`:OCR 成功行仍为 `measured` 且 settle 恒 0(**防回归**,锁定设计 §3.3 的剔出决定)
|
||||
|
||||
**验证**: `conda run --no-capture-output -n PolyGateway pytest tests/unit -v` → 全 PASS;`conda run --no-capture-output -n PolyGateway pytest tests/contracts -v` → 全 PASS(T2 的结算分支此时首次被激活,契约测试须复跑)
|
||||
|
||||
### T4 — 解绑装配约束(派生值真正启用)
|
||||
|
||||
- [x] **改** `src/polygateway/types.py:125-126` — 删除:
|
||||
|
||||
```python
|
||||
if self.tpm > 0 and self.est_tokens <= 0:
|
||||
raise ValueError("启用 TPM 闸时 est_tokens 必须 > 0(入场预扣依据)")
|
||||
```
|
||||
|
||||
`_validate_gates` 的其余部分(`timeout_s > 0`、四个限额非负)**保留不动**。`est_tokens` 字段本身与 `{SCOPE}__{PROVIDER}__{N}__EST_TOKENS` 环境键保留不删不改名(迁移兼容硬约束);`config.py:40` 的键映射无需改动。
|
||||
|
||||
**验收标准**: `tpm=6000, est_tokens=0` 可构造;该源入场预扣 100,**成功侧与非 dead 瞬时失败侧**按 100 结算(delta=0)。**取消 / RequestRejected / ResultInvalid / SourceDead 四侧维持既有的 `actual=0` 全额退回**——`retry.py:355-359` 的取消分支不给 `actual` 赋值、停在 `:329` 初值,这是设计 §3.3 声明不动的既有行为,**不要**为了凑"三侧一致"去改它。
|
||||
|
||||
**测试要求**(先失败后通过:改前构造即抛 `ValueError`):
|
||||
- **改写** `tests/unit/test_types.py:94 test_tpm_requires_est_tokens` — 它现在断言 `_make_source(tpm=10000, est_tokens=0)` 抛 `ValueError`,删约束后必然变红。保留后半条正向断言(`est_tokens=800` 仍原样返回),把前半条改为"构造成功且 `effective_est_tokens()` 返回派生值"
|
||||
- `test_retry.py`:**成功侧**——未填 `est_tokens`、`tpm>0`、usage 帧缺失的成功调用后,TPM 窗口残留量等于派生预扣量而非 0(**这是设计中最易漏的一条**,回归 §3.2 #9;在 `test_retry.py:149` 的 `_src("a", tpm=1000, est_tokens=400)` 旁加 `est_tokens=0` 用例)
|
||||
- `test_retry.py`:**失败侧**——同配置的非 dead 瞬时失败调用后,窗口残留量同为派生预扣量(回归 §3.2 #8)
|
||||
- `tests/contracts/test_limiter_contract.py`:**只加后端级断言**——传入派生值时双后端的结算口径一致。**不要**在契约文件里写端到端用例:该文件直接驱动 limiter(形如 `limiter.try_acquire("s1", 0)`),不经 `QuotaGate`/`RetryMW`,照字面写会产出 `try_acquire(src.effective_est_tokens())` + `settle(同值)` 的退化用例——只测了后端算术,没测调用点是否真的切了派生值。上面两条端到端断言的载体是 `retry.py`,放 `test_retry.py`(内存后端)
|
||||
|
||||
**验证**: `make ci`(即 check + test,含 import-linter 契约)→ 全 PASS。**不要**在外层再套 `conda run`:`Makefile` 的 `check`/`test` 目标内部已各自 `conda run -n $(ENV)`,嵌套后外层的 `--no-capture-output` 也管不到内层缓冲
|
||||
|
||||
### T5 — 权威文档与发布物同步
|
||||
|
||||
- [x] **改** `research-wiki/ARCHITECTURE.md` 四处:§7.7 行 428(`est_tokens` 描述:可选调优覆盖 + 派生规则,删去"亦作 usage 缺失时的保守兜底")、§5.1 行 331(`usage_source` 三态 + cost NULL 口径)、§4.4 行 305("token 按 `est_tokens` 预扣" → 按有效预扣量)、§7.1 行 384(打捞路径强制 `estimated` → 仅在收到 usage 帧时降级)
|
||||
- [x] **改** `research-wiki/migrations/chsanalyzer.md`:行 151 由"保留"改判"**有意放弃**"并写入设计 §4 的理由(CHS 只记单个 `total_tokens` 不存在分配问题;保守在计费语境无安全方向);G2(行 185)标注已由本设计解决
|
||||
- [x] **改** `.env.example` 行 11:删除"TPM > 0 时 EST_TOKENS 必填 > 0",改注为"可选;未填则库按 tpm 派生"
|
||||
- [x] **改** `CHANGELOG.md`:新增"行为收紧/变更"小节三条——`usage_source` 新增 `unavailable`、用量不可得行 cost 由数值变 NULL、`est_tokens` 降为可选
|
||||
- [x] **改** wiki 用户文档站(按 `docs-convention.md` §2):usage/成本口径说明须写明缺口查询为 `WHERE usage_source='unavailable' AND cache_hit = false`(**必须带 `cache_hit` 限定**:缓存命中行按裁决 cost 为 `0.0` 且标 `unavailable`,本无账目缺口,不加限定则度量偏高)
|
||||
- [x] **回帖** Gitea issue #2:结论与下游可删绕行校验的时点
|
||||
|
||||
**验收标准**: 全库 grep `EST_TOKENS 必填`、`est_tokens` 兜底相关表述无残留;ARCHITECTURE.md 无自相矛盾表述。
|
||||
|
||||
**测试要求**: 纯文档,无行为测试。以 `grep` 输出为验收证据。
|
||||
|
||||
**验证**: `make ci` → PASS;`grep -rn "EST_TOKENS 必填" . --exclude-dir=.git` → 无输出
|
||||
|
||||
## 5. 完成判定
|
||||
|
||||
- [x] T1-T5 全部 checkbox 勾选,每个任务一次语义化提交(`commit` skill)
|
||||
- [x] `make ci` 全绿(含 ruff、import-linter 洋葱契约、pytest 覆盖率)
|
||||
- [x] 设计 §6 测试表的 10 行断言全部有对应测试且可出示"先失败后通过"证据(T2 的零行为变更任务以"现有测试不回归 + 等价性断言"替代)
|
||||
- [x] 派新上下文 verifier subagent 独立验证(`verification-before-completion`,里程碑级/跨多文件硬门)
|
||||
- [x] 版本 bump 与 CHANGELOG 同步发布(不得裸发)
|
||||
|
||||
## 6. 明确不做
|
||||
|
||||
派生值取全局与单源 tpm 较紧者(需改三处 `QuotaGate` 装配,修的是既有缺口,设计 §7 已声明另开 issue);遥测驱动的 p90 自适应预估(设计 §5 已否决,待实测证据);`ocr.py` 的 usage 标记(设计 §3.3 已剔出);`retry.py` 另外三条失败分支的 `actual` 初值。
|
||||
@@ -0,0 +1,224 @@
|
||||
# 实现计划: 响应可观测字段扩展(Issue #3)
|
||||
|
||||
- **目标**: 让 `LLMResponse` 与遥测表如实暴露「供应商 prompt cache 命中的输入 token 数」与「API 实际返回的模型版本串」。
|
||||
- **方案概述**: 报文解析留在 `transports/`(新增两个强类型字段随 `TransportResult` 上浮),`RetryMW` 只搬运;遥测端口由 18 字段扩到 20 并给两个后端加幂等补列;`PricingTable` 增加可选缓存单价档消除 cost 高估。缓存命中行按既有口径原样回放。
|
||||
- **依据设计**: `research-wiki/designs/2026-07-31-response-observability-fields-design.md`(2026-07-31 已获人类批准,决策 A2/B1/C1/D1)。
|
||||
- **涉及技术**: Python 3.11 frozen dataclass、httpx SSE 解析、sqlite3、asyncpg、pytest。
|
||||
- **保真校验**: 本计划**不涉及** `reference/` 参考实现迁移,保真校验不适用。但遥测后端属 ARCHITECTURE §1.4 资产,T5 明确约束「不得改变既有降级语义」。
|
||||
|
||||
## 文件结构
|
||||
|
||||
| 文件 | 职责 | 本次改动 |
|
||||
|---|---|---|
|
||||
| `src/polygateway/types.py` | 冻结公共类型 | `LLMResponse` / `TransportResult` 各 +2 字段;`cache_hit` docstring 消歧 |
|
||||
| `src/polygateway/transports/openai_compat.py` | OpenAI 兼容报文解析 | 防御解析 helper;SSE sink 采集 `model`;两处 `TransportResult` 构造填新字段 |
|
||||
| `src/polygateway/middleware/retry.py` | 尝试循环 | `_build_response` 搬运两字段 |
|
||||
| `src/polygateway/middleware/cache.py` | 响应缓存 | **零代码改动**(自动透传),仅补测试固化行为 |
|
||||
| `src/polygateway/pricing.py` | 单价换算 | `ModelPrice` +可选档;`cost()` +可选参;`from_file` 校验 |
|
||||
| `src/polygateway/ports.py` | 端口契约 | `TelemetryRecorder` 18 → 20 字段 |
|
||||
| `src/polygateway/telemetry/{sqlite,postgres}.py` | 遥测后端 | DDL +2 列;`_COLUMNS` +2;初始化期幂等补列 |
|
||||
| `src/polygateway/middleware/telemetry.py` | 遥测唯一调用点 | `_record` 与三个 `emit_*` 搬运两字段;cost 换算传入缓存 token |
|
||||
|
||||
字段定义(全库唯一权威,后续任务一律引用此处):
|
||||
|
||||
```python
|
||||
# LLMResponse 与 TransportResult 尾部,同名同类型同默认值
|
||||
cached_prompt_tokens: int | None = None # 供应商 prompt cache 命中的输入 token;None = 该源未上报
|
||||
model_reported: str | None = None # API 响应体的 model 字段;None = 未上报
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## T1. 类型层加字段
|
||||
|
||||
- [ ] **改**: `src/polygateway/types.py`
|
||||
|
||||
**行为**: 在 `LLMResponse` 尾部(`structured_data` 之后)与 `TransportResult` 尾部(`raw` 之后)各追加上面两个字段。`cache_hit` 的语义在 `LLMResponse` docstring 中写明是「**PolyGateway 自身响应缓存**命中,与供应商 prompt cache 无关,后者见 `cached_prompt_tokens`」。
|
||||
|
||||
**验收**: 前 11 个字段的顺序与名字一字不动;新字段有默认值,`LLMResponse(...)` 按前 11 位置参数构造仍成立;`TransportResult` 现有两处构造(`openai_compat.py:354/436`)不传新字段也能构造。
|
||||
|
||||
**测试**(`tests/unit/test_types.py`): ① 不传新字段时两个类型的新字段均为 `None`;② 按位置构造 `LLMResponse` 的前 11 字段仍可用(迁移兼容承诺)。
|
||||
|
||||
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_types.py -v` → PASS。
|
||||
|
||||
**提交**: `feat: add cached prompt tokens and reported model to response types`
|
||||
|
||||
## T2. transport 采集与防御解析
|
||||
|
||||
- [ ] **改**: `src/polygateway/transports/openai_compat.py`
|
||||
|
||||
**行为**分三处:
|
||||
|
||||
1. 新增两个模块级防御 helper(网关返回一律不可信,解析失败**返回 None,不抛异常**):
|
||||
|
||||
```python
|
||||
def _coerce_cached_tokens(usage: Any) -> int | None:
|
||||
"""从 usage.prompt_tokens_details.cached_tokens 取非负整数;任何形态异常 → None。"""
|
||||
|
||||
def _coerce_model_reported(value: Any) -> str | None:
|
||||
"""响应体 model 字段: 非空 str 才收,其余(含空串/非 str)→ None。"""
|
||||
```
|
||||
|
||||
`_coerce_cached_tokens` 需容忍:`usage` 为 None、`prompt_tokens_details` 缺失或非 dict、`cached_tokens` 为 `bool`/`str`/负数/浮点。`bool` 必须排除(Python 中 `isinstance(True, int)` 为真)。**`0` 必须如实保留而非归 None**——真实零命中与未上报是两回事,这是 issue 的核心诉求。
|
||||
|
||||
2. 流式路径:`_sse_delta`(`:44-47`)当前只把 `usage` 旁路进 sink。补一条——chunk 里出现 `model` 时写 `usage_sink["model"]`(**首次写入即固定**,后续 chunk 不覆盖,避免末帧异常值污染)。`_stream_once` 的 `TransportResult` 构造(`:354`)填 `cached_prompt_tokens=_coerce_cached_tokens(sink.get("usage"))`、`model_reported=_coerce_model_reported(sink.get("model"))`。
|
||||
|
||||
3. 非流式路径:`TransportResult` 构造(`:436`)填 `_coerce_cached_tokens(body.get("usage"))` 与 `_coerce_model_reported(body.get("model"))`。
|
||||
|
||||
**验收**: `raw` 的内容保持原样不动(新字段是独立格子,不是杂物袋的扩充);OCR 与 embedding 的解析路径一行不改。
|
||||
|
||||
**测试**(`tests/unit/test_openai_compat.py`,用现有 fake 响应二次构造):
|
||||
|
||||
| 用例 | 期望 |
|
||||
|---|---|
|
||||
| 非流式 usage 含 `prompt_tokens_details.cached_tokens: 128` | `cached_prompt_tokens == 128` |
|
||||
| 流式 usage 帧同上 | 同上 |
|
||||
| 无 `prompt_tokens_details` / usage 帧缺失 | `None` |
|
||||
| `cached_tokens` 为 `"abc"` / `-1` / `True` / `1.5` / `[]` / dict | `None`,且**不抛异常** |
|
||||
| `cached_tokens` 为 `0` | `0`(真实零命中,**不得**归 None) |
|
||||
| `prompt_tokens_details` 非 dict | `None` |
|
||||
| 非流式 body 含 `model: "MiniMax-Text-01-250321"` | `model_reported` 为该串 |
|
||||
| 流式首个含 model 的 chunk 后又出现不同 model | 取**首个** |
|
||||
| body 无 `model` / `model` 为 `""` | `None` |
|
||||
|
||||
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_openai_compat.py -v` → PASS(新增用例先失败后通过)。
|
||||
|
||||
**提交**: `feat: collect provider cache tokens and reported model in transport`
|
||||
|
||||
## T3. RetryMW 搬运与缓存回放固化
|
||||
|
||||
- [ ] **改**: `src/polygateway/middleware/retry.py`
|
||||
|
||||
**行为**: `_build_response`(`:419-438`)追加 `cached_prompt_tokens=result.cached_prompt_tokens`、`model_reported=result.model_reported`。`model=source.model` **保持不变**——别名仍是主字段,真实版本是旁证(设计非目标 2)。
|
||||
|
||||
`middleware/cache.py` **不改一行**:`_RESPONSE_FIELDS` 由 `dataclasses.fields(LLMResponse)` 动态生成(`:26`)、`_serialize` 用 `asdict`(`:139`),新字段自动进出;决策 B1 要求命中时原样回放,而 `_rehydrate` 的覆写清单(`:113-119`)本就不含新字段,零改动即是正确行为。本任务用测试把它钉死。
|
||||
|
||||
**测试**:
|
||||
|
||||
- `tests/unit/test_retry.py`: transport 返回带两字段的 `TransportResult` → `chat()` 返回的 `LLMResponse` 上两字段一致;transport 未上报时为 `None`。
|
||||
- `tests/unit/test_cache.py`: ① 带两字段的响应写入缓存再命中,回放值与原值相等且 `cache_hit=True`;② **旧格式兼容**——手工构造缺这两个键的缓存 JSON 塞进后端,命中后能正常 rehydrate 且两字段为 `None`(不得抛异常回源)。
|
||||
|
||||
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_retry.py tests/unit/test_cache.py -v` → PASS。
|
||||
|
||||
**提交**: `feat: carry the new observability fields through retry and cache`
|
||||
|
||||
## T4. 缓存读取单价
|
||||
|
||||
- [ ] **改**: `src/polygateway/pricing.py`
|
||||
|
||||
**行为**:
|
||||
|
||||
```python
|
||||
class ModelPrice: # 追加第三档,可选
|
||||
cached_input_per_1m: float | None = None
|
||||
|
||||
def cost(self, model: str, prompt_tokens: int, completion_tokens: int,
|
||||
cached_prompt_tokens: int | None = None) -> float | None:
|
||||
```
|
||||
|
||||
换算规则(设计 §4):配了缓存档**且** `cached_prompt_tokens` 为正 → `(prompt - cached) × input + cached × cached_input`;否则全额按 `input`(现状,逐位不变)。`cached > prompt` 时按 `prompt` 夹取并 `logger.warning` 一次(沿用 `_warned` 的去重思路,按 model 去重,防日志风暴),**绝不产生负成本**。
|
||||
|
||||
`from_file` 的 fail-loud 扩展:条目出现 `cached_input_per_1m` 键时必须可转 float 且非负,否则 `ValueError`;不出现该键 = 合法(旧价格表零改动)。`__post_init__` 同步校验非负。
|
||||
|
||||
顺带订正 `PricingTable` docstring(`pricing.py:36`)那句「cost() 是全库唯一换算点(经 TelemetryEmitter)」——实际有 `TelemetryEmitter`(`middleware/telemetry.py:137`)与 `embedding.py:419` 两个调用点(设计 §6 行为审计已声明)。**只改这一行注释,不做任何结构重构**。
|
||||
|
||||
**验收**: `embedding.py:419` 的三参调用形态**一行不改**仍可用;未配缓存档时,任意输入下 `cost()` 结果与改前逐位相等。
|
||||
|
||||
**测试**(`tests/unit/test_pricing.py`): ① 配缓存档 + 命中 → 成本严格低于全额且等于手算值;② 未配缓存档 + 命中 → 与不传该参数结果相等;③ `cached > prompt` → 结果等于全部按缓存价、非负、有 warning;④ `cached_prompt_tokens=None/0` → 全额;⑤ 三参旧调用签名可用;⑥ 价格表含 `cached_input_per_1m: -1` 或 `"x"` → `from_file` 抛 `ValueError`;⑦ 无该键的旧价格表照常加载。
|
||||
|
||||
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_pricing.py -v` → PASS。
|
||||
|
||||
**提交**: `feat: support a cached input price tier in the pricing table`
|
||||
|
||||
## T5. 端口扩字段与后端补列
|
||||
|
||||
- [ ] **改**: `src/polygateway/ports.py`、`src/polygateway/telemetry/sqlite.py`、`src/polygateway/telemetry/postgres.py`
|
||||
|
||||
**行为**:
|
||||
|
||||
1. `TelemetryRecorder.record_llm_call`(`ports.py:250-271`)在 `cost` 之后追加 `cached_prompt_tokens: int | None` 与 `model_reported: str | None`,**不设默认值**(设计 §5:库外无第三方实现者)。同步该 Protocol 的「18 字段冻结」docstring。
|
||||
|
||||
2. 两个后端的 `_DDL` 加列(SQLite `INTEGER`/`TEXT`;PG `INTEGER`/`TEXT`,均可空、无默认值)。**新列在 DDL 里必须放在 `created_at` 之后(即表的最末尾),不得插在 `cost` 之后**——旧表走 `ALTER TABLE ADD COLUMN` 只能追加到末尾,若新建库把新列插在 `created_at` 前面,两条路径的物理列序就会分叉,而 `tests/integration/test_postgres_telemetry.py:117-128` 的 `test_schema_has_frozen_columns_in_order` 按 `ordinal_position` 逐位断言,且该表是与真实批跑共享的表、**严禁 DROP/TRUNCATE**(文件头隔离纪律),分叉后没有合规修法。
|
||||
|
||||
`_COLUMNS` 在 `"cost"` 之后追加同名两项即可——`_INSERT` 是显式列名拼装(`sqlite.py:62-65`/`postgres.py:67-71`),`_COLUMNS` 只需与自身的 `row = tuple(...)` 自洽,**与 DDL 物理列序无关**。
|
||||
|
||||
3. **幂等补列**,按设计 D1 纪律执行:
|
||||
|
||||
- **SQLite**(`sqlite.py:74-84`):补列代码必须放在 `self._conn = conn` **之后**、用**独立 try**,且**首行必须守卫 `if self._conn is None: return`**——初始化 try 吞掉失败时 `self._conn` 仍是 `None`(局部 `conn` 甚至未绑定),无守卫的补列块会抛 `AttributeError`/`NameError`,这两者不被 `sqlite3.Error` 捕获,会直接逃出 `__init__`,打破「初始化失败静默降级」的对外契约(既有测试 `tests/unit/test_telemetry.py:132-135` `test_unwritable_path_degrades_silently` 会红)。守卫之后:`PRAGMA table_info(llm_calls)` 取现有列名集合,缺哪列补哪列;捕获 `sqlite3.Error` 时消息含 `duplicate column` 视为成功(多进程共库的 TOCTOU),其余记 warning。**绝不允许**因补列失败把 `self._conn` 置回 `None`——那会让整个 recorder 永久 no-op。
|
||||
- **Postgres**(`postgres.py:_ensure_ready` 内、`_DDL` 执行之后):两条 `ALTER TABLE llm_calls ADD COLUMN IF NOT EXISTS ...`,共享既有 `_init_lock` 与 `except asyncio.CancelledError: raise` 结构。
|
||||
|
||||
**验收(降级语义不得改变)**: SQLite 侧 `except` 不得加宽(取消天然穿透);PG 侧 `CancelledError` 分支保持在最前;写入失败仍是逐行 warning 丢弃,不冒泡。
|
||||
|
||||
**必须同步改的测试(共 5 处)**:
|
||||
|
||||
| 位置 | 内容 | 漏改会怎样 |
|
||||
|---|---|---|
|
||||
| `tests/unit/test_telemetry.py:76` 起 `_record_minimal` | 手写 18 键 dict | **红**(`KeyError`) |
|
||||
| `tests/integration/test_postgres_telemetry.py:81-105` `_record_minimal` | 同上 | **红** |
|
||||
| `tests/unit/test_telemetry.py:18-40` `_EXPECTED_COLUMNS` | 19 项列序断言(含 `created_at`) | **红**;新列追加到 `created_at` **之后** |
|
||||
| `tests/integration/test_postgres_telemetry.py:22-41` `_EXPECTED_COLUMNS` | 同上 | **红**;同上 |
|
||||
| `tests/unit/test_ports.py:96` `_DummyRecorder` | 唯一写全签名的 fake | **不会红**(它只被 `:131` 的 `isinstance` 使用,`runtime_checkable` Protocol 只校验方法名不校验签名),但仍应同步以免误导后来者 |
|
||||
|
||||
前四处是本次仅有的天然拦截点;端口加参数**不会**带来编译期保护(本仓无 mypy,其余 8 个 fake 全是 `**fields`)。
|
||||
|
||||
**测试**:
|
||||
|
||||
- `tests/unit/test_telemetry.py`(SQLite):① 20 字段写入后可读回两个新列的值(含 `None`);② **旧表升级**——先用 18 列 DDL 手工建表,再实例化 `SQLiteRecorder`,写入成功且新列有值;③ **补列失败路径**(设计 §8 第 ③ 条,最危险的分支,不可用成功路径顶替)——构造一个 ALTER 必然失败的场景(把 `llm_calls` 建成同名 view,或注入在 ALTER 上抛 `sqlite3.OperationalError` 的连接),断言构造**不抛异常**、`recorder._conn` 仍非 `None`、后续 `record_llm_call` 不抛(降级为逐行 warning);④ 初始化路径不可写时仍静默降级(`test_unwritable_path_degrades_silently` 保持绿)。
|
||||
- `tests/integration/test_postgres_telemetry.py`:① 20 字段写入 PG 并 `SELECT` 回读;② 18 列旧表经初始化后自动补列并写入成功。
|
||||
|
||||
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_telemetry.py tests/unit/test_ports.py -v` → PASS;PG 部分 `conda run -n PolyGateway --no-capture-output pytest tests/integration/test_postgres_telemetry.py -v` → PASS。**PG/Redis 属共享后端,严禁与其他会话或钩子测试并跑**,起跑前确认无并发占用。
|
||||
|
||||
**提交**: **与 T6 合并为一次提交**,不得单独落地。理由:T5 落地后 emitter 仍只传 18 键,后端的 `row = tuple(fields[col] for col in _COLUMNS)` 会抛 `KeyError`,被 `_record` 的 `except Exception`(`middleware/telemetry.py:164-165`)吞成 warning → **该 commit 处于全量遥测静默丢失的状态**,且现有测试无一能捕获。提交信息见 T6。
|
||||
|
||||
## T6. Emitter 搬运与契约测试(与 T5 同一次提交)
|
||||
|
||||
- [ ] **改**: `src/polygateway/middleware/telemetry.py`
|
||||
|
||||
**行为**: `_record`(`:108-125`)新增两个参数并透传给 `record_llm_call`;三个入口各自提供取值——
|
||||
|
||||
| 入口 | `cached_prompt_tokens` | `model_reported` |
|
||||
|---|---|---|
|
||||
| `emit_attempt` | `response.cached_prompt_tokens if response else None` | 同左 |
|
||||
| `emit_cache_hit` | `response.cached_prompt_tokens`(B1 原样回放) | 同左 |
|
||||
| `emit_terminal_failure` | `None` | `None` |
|
||||
|
||||
cost 换算(`:137`)改为把 `cached_prompt_tokens` 传进 `self._pricing.cost(...)`。`cache_hit → 0.0` 与 `usage_source == "unavailable" → None` 两条短路的**先后顺序一字不动**(ARCHITECTURE §5.1 cost 口径不变式)。
|
||||
|
||||
**测试**(`tests/unit/test_telemetry.py`):
|
||||
|
||||
- **契约测试(不可省)**: 用记录 kwargs 的 fake recorder 跑一次 `emit_attempt`,断言 `set(kwargs) == set(sqlite._COLUMNS) == set(postgres._COLUMNS)`。理由:`row = tuple(fields[col] for col in _COLUMNS)` 位于两个后端 try 之外(`sqlite.py:90`/`postgres.py:121`),emitter 漏传字段会抛 `KeyError` 并被 `_record` 的 `except Exception` 吞成 warning → 静默丢遥测;现有 8 个 `**fields` 形态的 fake 一个都拦不住。
|
||||
- 三个入口各记一行,断言新字段取值符合上表。
|
||||
- cost 回归:配了缓存档且响应带 `cached_prompt_tokens` → 落库 cost 低于全额;缓存命中行 cost 仍为 `0.0`;`unavailable` 行仍为 `None`。
|
||||
|
||||
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_telemetry.py -v` → PASS;随后 `make ci` 全绿(含 ruff 与 import-linter)。
|
||||
|
||||
**提交**(含 T5 全部改动): `feat: record the observability fields end to end through telemetry`
|
||||
|
||||
## T7. 文档同步与发版
|
||||
|
||||
- [ ] **改**: `research-wiki/ARCHITECTURE.md`、`CHANGELOG.md`、`.env.example`、`pyproject.toml`、`src/polygateway/__init__.py`、四处「18 字段冻结」措辞点、Gitea Wiki 站
|
||||
|
||||
**行为**:
|
||||
|
||||
1. `ARCHITECTURE.md`:§5.1 新增字段表补两行;**§7.8「必录字段」的行内清单**(`:452`)补两项——该文件不含字面「18 字段冻结」,§7.8 与 D8(`:202`)才是遥测字段的落点;§7.8 末条「`pricing.py` 维护 model →(input 单价, output 单价)表」同步第三档。补一条度量口径警示(与 cost 缺口同款):**统计供应商缓存命中率必须带 `WHERE cache_hit = false`**,否则缓存回放行会被重复计入。
|
||||
2. 代码里的「18 字段/18 列」措辞共 **6 处**,全部订正(`grep -rn "18 字段\|18 列" src/ tests/` 可复核):`ports.py:248`、`middleware/telemetry.py:31`、`pricing.py:6`、`telemetry/sqlite.py:87`、`telemetry/postgres.py:9`(「18 列 schema 与 SQLite 版同名同序」)、`tests/unit/test_telemetry.py:1`。
|
||||
3. `.env.example:56` 是仓内**唯一**的价格表格式说明(无独立模板文件),补 `cached_input_per_1m` 可选档与「不填即全额计价、库不猜折扣率」的说明。
|
||||
4. 版本 bump `1.0.3` → `1.1.0`,**两处必须同步**(`pyproject.toml:7` 与 `src/polygateway/__init__.py:34`;`tests/unit/test_package.py:11` 会断言二者相等)。
|
||||
5. `CHANGELOG.md` 顶部新增 `## 1.1.0` 段,沿用既有写法(先讲问题、再讲变更、点明下游要读什么):两个新字段的语义与 `None`/`0` 之别、`cache_hit` 与供应商 prompt cache 的区分、遥测表新增两列与自动补列、价格表可选缓存档、度量口径的 `cache_hit = false` 约束。
|
||||
6. Gitea Wiki 站(需单独 `git clone https://gitea.iomgaa.online/iomgaa/PolyGateway.wiki.git`)按 `docs-convention.md` §2 清单同步:`参考-公共API`(LLMResponse 字段表)、`参考-配置键`(价格表格式)、`指南-遥测与成本`(新列与成本校正口径)、`Home.md` 版本号与安装命令、`_Sidebar.md` 如有结构变化。
|
||||
|
||||
**验收**: 版本 bump 的提交**不允许单独存在**(docs-convention §2 门),必须与 wiki/CHANGELOG 同步在同一次交付内。
|
||||
|
||||
**验证**: `conda run -n PolyGateway --no-capture-output pytest tests/unit/test_package.py -v` → PASS;`make ci` 全绿。
|
||||
|
||||
**提交**: `chore: release 1.1.0 with the response observability fields`
|
||||
|
||||
---
|
||||
|
||||
## 合并前门(逐条对应 CLAUDE.md §3)
|
||||
|
||||
- [ ] 每个行为变更都有「先失败后通过」的测试证据(T1-T6 各自的新增用例)。
|
||||
- [ ] `make ci` 全绿(ruff + import-linter + pytest + 覆盖率)。
|
||||
- [ ] 派**全新上下文**的 verifier subagent 独立验证(跨多文件,`verification-before-completion` 强制档)。
|
||||
- [ ] 合并前整分支代码审查(`requesting-code-review`)。
|
||||
- [ ] Gitea Issue #3 的关闭说明:两个字段的最终名字与语义、缓存命中行的回放口径、遥测新列与补列行为。
|
||||
@@ -0,0 +1,396 @@
|
||||
# 实现计划: 采样参数透传(issue #4)
|
||||
|
||||
- **设计**: `research-wiki/designs/2026-07-31-sampling-params-design.md`(2026-07-31 人类批准)
|
||||
- **分支**: `feat/issue-4-sampling-params`
|
||||
- **目标**: 让下游能固定解码参数(`temperature`/`seed`/`max_tokens`),且不破坏缓存隔离与遥测诚实性。
|
||||
- **方案概述**: `chat()` 增 keyword-only `overlay` 参数(调用级),`SourceConfig` 增 `extra_body` 字段(配置级)。`ChatRequest` 增 `sampling` 快照字段作为跨洋葱层恒定读取点,供缓存 key 与遥测消费。遥测端口 20 → 21 字段。
|
||||
- **技术**: Python 3.11+,frozen dataclass,`MappingProxyType`,sqlite3 / asyncpg DDL 幂等补列。
|
||||
|
||||
**保真校验**: 本计划不涉及 `reference/` 参考实现迁移,保真校验不适用。
|
||||
|
||||
---
|
||||
|
||||
## 1. 文件结构
|
||||
|
||||
| 文件 | 职责变更 |
|
||||
|---|---|
|
||||
| `src/polygateway/types.py` | 新增 `validate_request_overlay()` 与 `merge_sampling()` 两个纯函数;`ChatRequest.sampling` 字段;`SourceConfig.extra_body` 字段与构造期校验 |
|
||||
| `src/polygateway/client.py` | `chat()` 增 `overlay` 参数;`model_fingerprint` 计算纳入 `extra_body` |
|
||||
| `src/polygateway/middleware/cache.py` | `build_cache_key()` 增 `sampling` 入参并纳入 key |
|
||||
| `src/polygateway/transports/openai_compat.py` | `_build_payload` 在 thinking profile 之后、overlay 之前应用 `source.extra_body` |
|
||||
| `src/polygateway/config.py` | `_SOURCE_FIELDS` 增 `EXTRA_BODY`;`_cast` 增 `json` 分支 |
|
||||
| `src/polygateway/ports.py` | `TelemetryRecorder.record_llm_call` 增第 21 参 `sampling` |
|
||||
| `src/polygateway/middleware/telemetry.py` | 三个 emit 入口按设计表格产出 `sampling`;`_record` 透传 |
|
||||
| `src/polygateway/telemetry/sqlite.py` | DDL / `_BACKFILL_COLUMNS` / `_COLUMNS` 增 `sampling` |
|
||||
| `src/polygateway/telemetry/postgres.py` | DDL / `_BACKFILL` / `_COLUMNS` 增 `sampling` |
|
||||
| `src/polygateway/ocr.py` / `embedding.py` | 构造期剥离 `extra_body` + warning(决策 G) |
|
||||
| `src/polygateway/providers.py` | minimax/openai 空 thinking profile 补后果注释(决策 F) |
|
||||
| `.env.example` / `README.md` / `CHANGELOG.md` / `research-wiki/ARCHITECTURE.md` | 文档同步(设计 §6) |
|
||||
|
||||
**各任务需新增的 import**(现状核实,不加即 NameError):
|
||||
|
||||
| 文件 | 需新增 |
|
||||
|---|---|
|
||||
| `types.py` | `from collections.abc import Mapping`、`from types import MappingProxyType`、`import json`。**该文件无 `from __future__ import annotations`**,注解在类体求值,`Mapping` 必须真导入 |
|
||||
| `client.py` | `import json`、`import hashlib` |
|
||||
| `config.py` | `import json` |
|
||||
| `ocr.py` / `embedding.py` | `import dataclasses`(现只有 `from dataclasses import dataclass`)、`from loguru import logger`(若未导入) |
|
||||
| `middleware/telemetry.py` | `merge_sampling`/`canonical_sampling_json` 需**运行时**导入(现对 `polygateway.types` 只在 `TYPE_CHECKING` 下导入) |
|
||||
|
||||
**关键接口**(跨任务消费,此处定死):
|
||||
|
||||
```python
|
||||
# types.py —— 两个纯函数 + 两个字段
|
||||
_PROTECTED_OVERLAY_KEYS = frozenset({"model", "messages", "stream", "stream_options"})
|
||||
|
||||
def validate_request_overlay(overlay: Mapping[str, Any], *, origin: str) -> dict[str, Any]:
|
||||
"""校验采样参数覆盖层并返回浅拷贝;origin 用于错误信息定位来源。
|
||||
|
||||
保护键会击穿治理(model→成本算错、messages→缓存与遥测口径失真、
|
||||
stream/stream_options→绕过看门狗与 usage 帧);值必须 JSON 可序列化,
|
||||
否则会在 CacheMW 的降级 try 之外抛裸 TypeError(设计 §决策 B)。
|
||||
"""
|
||||
|
||||
def merge_sampling(extra_body: Mapping[str, Any], sampling: Mapping[str, Any]) -> dict[str, Any]:
|
||||
"""合并配置级与调用级采样参数(调用级优先);两者皆空返回空 dict。"""
|
||||
|
||||
def canonical_sampling_json(merged: Mapping[str, Any]) -> str | None:
|
||||
"""遥测列与缓存 key 共用的序列化口径;空 mapping → None。"""
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ChatRequest:
|
||||
...
|
||||
overlay: dict[str, Any] = field(default_factory=dict)
|
||||
sampling: Mapping[str, Any] = field(default_factory=dict) # 新增
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class SourceConfig:
|
||||
...
|
||||
extra_body: Mapping[str, Any] = field(default_factory=dict) # 新增,__post_init__ 转 MappingProxyType
|
||||
```
|
||||
|
||||
```python
|
||||
# middleware/cache.py —— 签名扩展(sampling 为 keyword-only)
|
||||
# 默认值用 None 而非 {}: dict 字面量作默认参数会被 ruff B006 拦下
|
||||
def build_cache_key(
|
||||
model_fingerprint: str,
|
||||
messages: list[dict[str, Any]],
|
||||
namespace: str,
|
||||
salt: str | None,
|
||||
*,
|
||||
sampling: Mapping[str, Any] | None = None,
|
||||
) -> str: ...
|
||||
```
|
||||
|
||||
```python
|
||||
# client.py —— chat() 新签名
|
||||
async def chat(
|
||||
self, messages: list[dict[str, Any]], *,
|
||||
session_id: str | None = None, parent_call_id: str | None = None,
|
||||
cache_salt: str | None = None, cache_namespace: str | None = None,
|
||||
structured: type[BaseModel] | Literal["json"] | None = None,
|
||||
stream: bool = True,
|
||||
overlay: Mapping[str, Any] | None = None, # 新增
|
||||
) -> LLMResponse: ...
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. 任务清单
|
||||
|
||||
任务按依赖排序;每个任务一次提交、独立可验证。每个任务合并前必须出示**先失败后通过**的测试证据(先写测试跑红,再实现跑绿)。
|
||||
|
||||
统一验证命令前缀:`conda run -n PolyGateway --no-capture-output pytest`。
|
||||
|
||||
> **共享后端纪律**: 涉及 Redis/Postgres 的 integration 测试严禁与其他会话并跑(含 git 钩子触发的测试)。Task 7、Task 11 受此约束。
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 1: `types.py` 内核 —— 校验与合并纯函数 + 两个新字段
|
||||
|
||||
**文件**: 改 `src/polygateway/types.py`;测试 `tests/unit/test_types.py`
|
||||
|
||||
**实现行为**:
|
||||
|
||||
1. `validate_request_overlay(overlay, *, origin)`,**校验顺序即下列顺序**:
|
||||
- 键必须是 `str`,否则 `ValueError`(canonical JSON 要求)。**必须排在序列化试探之前**——`{1: "a", "b": 2}` 在 `sort_keys=True` 下抛的是 `TypeError: '<' not supported between 'str' and 'int'`,若先试序列化会被误报成"值不可 JSON 序列化",指错方向;
|
||||
- 命中 `_PROTECTED_OVERLAY_KEYS` 任一键 → `ValueError`,信息含 origin、违规键名、以及**为什么**(如 `stream` 会绕过流式看门狗);
|
||||
- 对整个 mapping 做 `json.dumps(..., sort_keys=True)` 试序列化,`TypeError` → 转 `ValueError` 并指出该值不可 JSON 序列化(信息提示改用 `float(x)` 等原生类型);
|
||||
- 返回 `dict(overlay)` 浅拷贝。
|
||||
2. `merge_sampling(extra_body, sampling)` → `{**extra_body, **sampling}`(调用级优先)。
|
||||
3. `canonical_sampling_json(merged)` → 空则 `None`,否则 `json.dumps(merged, sort_keys=True, ensure_ascii=False)`。
|
||||
4. `ChatRequest` 增 `sampling` 字段(见 §1 关键接口)。
|
||||
5. `SourceConfig` 增 `extra_body` 字段;`__post_init__` 新增 `_validate_extra_body()`:调 `validate_request_overlay(self.extra_body, origin=f"SourceConfig({self.name}).extra_body")`,再 `object.__setattr__(self, "extra_body", MappingProxyType(dict(...)))`(frozen dataclass 需用 `object.__setattr__`)。
|
||||
|
||||
**已知后果(必须显式接受,不是疏漏)**: `SourceConfig` 加 mapping 字段后**不再 hashable**(`hash()` → `TypeError`),且因 `MappingProxyType` 不可 pickle,`dataclasses.asdict()` / `copy.deepcopy()` 也会失败。
|
||||
|
||||
- 不可 hash 是**加任何 mapping 字段的固有代价**,与是否用 `MappingProxyType` 无关(裸 `dict` 同样不可 hash),无法规避;
|
||||
- 库内当前无调用点会踩:`asdict` 只用于 `LLMResponse`/`EmbeddingResponse`(`cache.py:140`),全库无 `set(sources)` 或以源作 dict key 的写法;
|
||||
- 保留 `MappingProxyType` 而非裸 dict,是因为决策 E 的只读约束值得这个代价;下游要可变副本用 `dict(source.extra_body)`,要改字段用 `dataclasses.replace(source, ...)`(已验证可行,会重跑 `__post_init__` 重新包 proxy,不递归)。
|
||||
|
||||
**验收标准**: 四个保护键各自触发 `ValueError` 且信息含原因;非 str 键报的是"键必须是 str"而非"不可序列化";`{"temperature": object()}` 类不可序列化值报 `ValueError` 而非 `TypeError`;合法 `{"temperature": 0, "seed": 42}` 通过并返回独立副本(改原 dict 不影响返回值);`SourceConfig.extra_body` 构造后为 `MappingProxyType` 且不可改。
|
||||
|
||||
**测试要求**: 新增 `tests/unit/test_types.py::TestSamplingValidation`,覆盖上述每条。不可序列化值用 `object()` 实例即可,不引入 numpy 依赖。**另加一条锁定测试**:`pytest.raises(TypeError): hash(source_config)`,把"不再 hashable"钉成有意行为——否则将来有人踩到时会以为是 bug 并"修"回去。
|
||||
|
||||
**验证**: `pytest tests/unit/test_types.py -v` → 全 PASS
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 2: `config.py` —— `EXTRA_BODY` env 解析
|
||||
|
||||
**文件**: 改 `src/polygateway/config.py`;测试 `tests/unit/test_config.py`
|
||||
|
||||
**实现行为**:
|
||||
- `_SOURCE_FIELDS` 增 `"EXTRA_BODY": ("extra_body", "json")`;
|
||||
- `_cast` 增 `json` 分支:`json.loads` 失败 → `ValueError`(沿用既有 `配置 {key} 解析失败: {exc}` 包装);解析结果**非 dict** → `ValueError`,信息说明必须是 JSON 对象(而非数组/标量)。
|
||||
|
||||
**验收标准**: `LLM__QWEN__1__EXTRA_BODY={"temperature":0}` → `SourceConfig.extra_body == {"temperature": 0}`;`{invalid` → `ValueError`;`[1,2]` → `ValueError`;`{"model":"x"}` → `ValueError`(经 Task 1 的 `SourceConfig.__post_init__` 保护键校验)。
|
||||
|
||||
**测试要求**: 新增 4 个 case 覆盖上述。**注意**: 这里同时验证了 Task 1 的校验确实挂在装配路径上。
|
||||
|
||||
**验证**: `pytest tests/unit/test_config.py -v` → 全 PASS
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 3: `chat()` 入口 + transport 应用 + fingerprint
|
||||
|
||||
**文件**: 改 `src/polygateway/client.py`、`src/polygateway/transports/openai_compat.py`;测试 `tests/unit/test_client.py`、`tests/unit/test_openai_compat.py`
|
||||
|
||||
**实现行为**:
|
||||
|
||||
1. `chat()` 增 `overlay` 参数(见 §1 签名)。进洋葱**之前**:
|
||||
```python
|
||||
validated = validate_request_overlay(overlay or {}, origin="chat(overlay=...)")
|
||||
```
|
||||
同一份 `validated` 对象同时填 `ChatRequest.overlay` 与 `.sampling`(设计决策 E:一次拷贝、两个字段指向同一快照,不做两份独立拷贝)。
|
||||
2. `_build_payload`:在 thinking profile 之后、`payload.update(overlay)` 之前插入 `payload.update(source.extra_body)`。**顺序即优先级,不可调换**。
|
||||
3. `model_fingerprint`(`client.py:117`)改为:
|
||||
```python
|
||||
fingerprint = ",".join(sorted({s.model for s in sources}))
|
||||
marks = sorted({json.dumps([s.model, dict(s.extra_body)], sort_keys=True, ensure_ascii=False)
|
||||
for s in sources if s.extra_body})
|
||||
if marks:
|
||||
fingerprint += "|" + hashlib.sha256("".join(marks).encode()).hexdigest()
|
||||
```
|
||||
全源 `extra_body` 皆空时字面量与旧实现**逐字相同**。`dict(...)` 是因为 `MappingProxyType` 不能直接进 `json.dumps`。
|
||||
|
||||
**验收标准**: 配置 `temperature=0` + 调用级 `temperature=1` → payload 中为 1;结构化注入的 `response_format` 覆盖调用级同名键;保护键在 `chat()` 入口即 `ValueError`(未进洋葱,可用 mock handler 断言未被调用);全源无 `extra_body` 时 fingerprint 与旧值逐字相同;有 `extra_body` 时不同;改源 `name` 不改变 fingerprint。
|
||||
|
||||
**测试要求**: 覆盖设计 §5 测试 #3、#4(chat 侧)、#5、#6(拷贝语义:调用方在 `chat()` 返回后修改自己的 dict,不影响已构造的 request)、#8(不可 JSON 序列化的值在 `chat()` 入口即 `ValueError`,断言洋葱 handler 未被调用)。
|
||||
|
||||
**验证**: `pytest tests/unit/test_client.py tests/unit/test_openai_compat.py -v` → 全 PASS
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 4: 缓存 key 纳入 `sampling`
|
||||
|
||||
**文件**: 改 `src/polygateway/middleware/cache.py`;测试 `tests/unit/test_cache.py`
|
||||
|
||||
**实现行为**:
|
||||
- `build_cache_key` 增 keyword-only `sampling` 参数(见 §1 签名),非空时以 `"sampling"` 键并入 `key_obj`(**仅非空参与**,与 `salt` 的"仅非 None"不同——见设计决策 A 末段);
|
||||
- `CacheMW.__call__` 传 `sampling=request.sampling`(**不是 `request.overlay`**——后者在此层虽尚未被结构化注入污染,但读 `sampling` 才是语义正确且不依赖层序巧合的写法)。
|
||||
|
||||
**验收标准**:
|
||||
- 同 messages、不同 `seed` → 两个不同 key,第二次 miss(**issue 场景的直接回归**);
|
||||
- 空 `sampling` 时 key 与旧实现**逐字相同**——测试须先把旧实现的 key 值固化为常量再比对(现有 `tests/unit/test_cache.py:39-54` 只有相等/不等断言,无 golden hash 可依);
|
||||
- 同 `sampling` 不同键序 → 同一 key(canonical 序列化)。
|
||||
|
||||
**测试要求**: 覆盖设计 §5 测试 #1、#2。golden hash 的取法:在改动前先运行一次现有 `build_cache_key` 打印结果,写死进测试。
|
||||
|
||||
**验证**: `pytest tests/unit/test_cache.py -v` → 全 PASS
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 5: 地基不变式回归(承重)
|
||||
|
||||
**文件**: 测试 `tests/unit/test_structured.py`(或就近的洋葱集成测试文件)
|
||||
|
||||
**实现行为**: 纯测试任务,不改产品代码。
|
||||
|
||||
落点:`tests/unit/test_structured.py` 里既有的 `ScriptedTerminal` 恰好站在 RetryMW 的位置(`client.py:91` 的 `terminal = RetryMW(...)`,StructuredMW 是最内中间件),扩写它即可,**无需搭全洋葱**。
|
||||
|
||||
断言:走结构化重问阶梯(强制至少重问一次,用先返回坏 JSON 再返回好 JSON 的 scripted terminal)后——
|
||||
1. terminal 每次收到的 `request.sampling` 与**构造 `ChatRequest` 时传入的 `sampling`** 逐字相同;
|
||||
2. 同一时刻 `request.overlay` **含** `response_format`(证明两者确实分叉,`sampling` 不是冗余字段)。
|
||||
|
||||
**为什么单列一个任务**: 决策 C 与 D 都建立在"`sampling` 跨层恒定"之上,而这条目前只靠"`dataclasses.replace` 恰好保留未提及字段"的约定成立,无任何机械执法。这条测试同时钉死决策 A 的"库内中间件永不修改"与决策 E 的只读约束。缺它则约束被破坏时无人发现。
|
||||
|
||||
**验收标准**: 该测试在故意把 `structured.py` 的 `replace` 改成重建 `ChatRequest`(丢掉 `sampling`)时**必须变红**——实施时须实际验证这一点,否则测试是空的。
|
||||
|
||||
**验证**: `pytest tests/unit/test_structured.py -v` → 全 PASS,且上述"故意破坏"实验红过一次
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 6: 遥测端口扩至 21 字段 + 三入口口径
|
||||
|
||||
**文件**: 改 `src/polygateway/ports.py`、`src/polygateway/middleware/telemetry.py`;测试 `tests/unit/test_telemetry.py`
|
||||
|
||||
**实现行为**:
|
||||
|
||||
1. `ports.TelemetryRecorder.record_llm_call` 增第 21 参 `sampling: str | None`(排在 `model_reported` 之后)。
|
||||
2. `TelemetryEmitter._record` 增同名参数并透传给 recorder。
|
||||
3. 三个入口按设计决策 D 的表格产出(**不含**结构化注入的 `response_format`):
|
||||
|
||||
| 入口 | `sampling` 取值 |
|
||||
|---|---|
|
||||
| `emit_attempt` | `canonical_sampling_json(merge_sampling(source.extra_body, request.sampling))` |
|
||||
| `emit_cache_hit` | `canonical_sampling_json(request.sampling)` |
|
||||
| `emit_terminal_failure` | `canonical_sampling_json(request.sampling)` |
|
||||
|
||||
后两者无 `source` 可言(由最外层 TelemetryMW 调用),与 `model`/`provider`/`source_name` 在终态行置空是同一先例。
|
||||
|
||||
**关键约束**: `sampling` 必须由 emitter **内部推导**,**不得**作为新必填参数由调用者传入——否则 `ocr.py:418` 与 `embedding.py:372` 立刻 TypeError。
|
||||
|
||||
**验收标准**: 三个入口各自的 `sampling` 值符合上表;`response_format` **三行都不出现**;`request.sampling` 与 `source.extra_body` 皆空时为 `None`。
|
||||
|
||||
**测试要求**: 覆盖设计 §5 测试 #9。用 fake recorder 捕获 kwargs 断言。
|
||||
|
||||
**验证**: `pytest tests/unit/test_telemetry.py -v` → 全 PASS
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 7: 两个遥测后端落列 + 幂等补列
|
||||
|
||||
**文件**: 改 `src/polygateway/telemetry/sqlite.py`、`src/polygateway/telemetry/postgres.py`;测试 `tests/unit/test_telemetry.py`、`tests/integration/test_postgres_telemetry.py`
|
||||
|
||||
**实现行为**(逐字沿用 issue #3 建立的套路):
|
||||
|
||||
- **sqlite.py**: DDL 在 `model_reported` **之后**加 `sampling TEXT`;`_BACKFILL_COLUMNS` 追加 `("sampling", "TEXT")`;`_COLUMNS` 末尾追加 `"sampling"`。
|
||||
- **postgres.py**: DDL 同位置加 `sampling TEXT`;`_BACKFILL` 追加 `("sampling", "ALTER TABLE llm_calls ADD COLUMN sampling TEXT")`;`_COLUMNS` 末尾追加。
|
||||
- 两处 `record_llm_call(**fields)` 按 `_COLUMNS` 取值,**无需改动**。
|
||||
|
||||
**硬约束**: 新列必须排在 `created_at` **之后**(两文件既有注释已说明理由:旧表只能 ALTER 追加到末尾,新建库若插在前面,两条路径物理列序分叉)。补列一律**先探测缺列再 ALTER**;失败只逐行降级,**绝不置结构性失能标志**(postgres 的 `_failed`)。
|
||||
|
||||
**连带必改**(不改则直接红):
|
||||
|
||||
| 位置 | 改什么 | 不改的后果 |
|
||||
|---|---|---|
|
||||
| `tests/unit/test_telemetry.py:78-102` 的 `_record_minimal()` | `fields` dict 加 `"sampling": None` | **两侧所有落库测试全红**:`sqlite.py:126` / `postgres.py:161` 的 `row = tuple(fields[col] for col in _COLUMNS)` **在 try 之外**,`_COLUMNS` 加列后抛裸 `KeyError: 'sampling'` 冒泡出 `record_llm_call` |
|
||||
| `tests/integration/test_postgres_telemetry.py:88-109` 的 `_record_minimal()` | 同上 | 同上 |
|
||||
| `tests/unit/test_telemetry.py:18` 的 `_EXPECTED_COLUMNS` | 追加 `"sampling"` | 列序断言红 |
|
||||
| `tests/integration/test_postgres_telemetry.py:22` 的 `_EXPECTED_COLUMNS` | 追加(另见 `:210,231` 引用点) | 列序断言红 |
|
||||
| `tests/unit/test_ports.py:95-119` 的 `_DummyRecorder.record_llm_call` | 显式 20 参签名同步为 21 | **不会红**(`runtime_checkable` 的 isinstance 只查方法存在不查签名),但会与端口脱节,顺带同步 |
|
||||
| `sqlite.py:123` docstring、`test_telemetry.py:1` 文案 | "20 字段" → "21 字段" | 无功能影响,文案与事实脱节 |
|
||||
|
||||
**验收标准**: 新建库列序正确;对**已存在的 20 列旧表**能幂等补列且补后列序与新建库一致;重复初始化不报错;补列失败(模拟只有 INSERT 权限)时仅 warning、后续写入不被禁用。
|
||||
|
||||
**测试要求**: 覆盖设计 §5 测试 #11、#12。Postgres 部分是 integration,**须独占 PG `polygateway` 库时序,严禁并跑**。
|
||||
|
||||
**验证**:
|
||||
- `pytest tests/unit/test_telemetry.py -v` → 全 PASS
|
||||
- `pytest tests/integration/test_postgres_telemetry.py -v` → 全 PASS(确认无其他会话在用 PG)
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 8: 决策 G —— OCR/embedding 构造期剥离 + warning
|
||||
|
||||
**文件**: 改 `src/polygateway/ocr.py`、`src/polygateway/embedding.py`;测试 `tests/unit/test_ocr_client.py`、`tests/unit/test_embedding.py`
|
||||
|
||||
**实现行为**: 两个 `__init__` 在既有校验块(`quota_full` 域校验附近)之后、`self._sources = list(sources)` 之前:
|
||||
|
||||
```python
|
||||
stripped = []
|
||||
for src in sources:
|
||||
if src.extra_body:
|
||||
logger.warning(
|
||||
"{} 路径暂不支持 extra_body,源 {} 的该配置已被忽略"
|
||||
"(需要 dimensions 等参数请提 issue): {}",
|
||||
<"embedding"|"OCR">, src.name, dict(src.extra_body),
|
||||
)
|
||||
src = dataclasses.replace(src, extra_body={})
|
||||
stripped.append(src)
|
||||
self._sources = stripped
|
||||
```
|
||||
|
||||
**剥离不是顺手清理,是承重的**: 不剥离则 Task 6 的 `merge_sampling(source.extra_body, ...)` 会让遥测**记录一个从未发出的参数**——`monkey_ocr.py:225,247` 只发 multipart `files=`(根本没有 JSON body),`openai_compat.py:343` 的 embed payload 硬编码 `{"model","input"}`。那是数据造假而非参数失效。替代方案(emitter 内特判调用方身份)违「遥测调用点收敛单一 helper」铁律,已否决。
|
||||
|
||||
**验收标准**: 带 `extra_body` 的源 → 装配**成功**(不抛异常)、记一条 warning、`client._sources` 上 `extra_body` 为空;该路径遥测 `sampling` 列为 `None`;不带 `extra_body` 时无 warning。
|
||||
|
||||
**测试要求**: 覆盖设计 §5 测试 #10。**后半段(遥测 `sampling` 为 None)是防遥测造假的真正断言,不可省**——只断言"装配成功 + 有 warning"是不够的。用 `caplog`/loguru 捕获断言 warning 存在。
|
||||
|
||||
**验证**: `pytest tests/unit/test_ocr_client.py tests/unit/test_embedding.py -v` → 全 PASS
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 9: 决策 F —— 空 thinking profile 的后果注释
|
||||
|
||||
**文件**: 改 `src/polygateway/providers.py`
|
||||
|
||||
**实现行为**: 给 `openai`(`:46-51`)与 `minimax`(`:53-58`)两个 profile 各补一句**后果**说明:`enable_thinking=False` 对本 provider 不产生任何效果,需要关闭推理请用 `SourceConfig.extra_body`。
|
||||
|
||||
**注意**: `:52` 那条既有注释(「OpenAI 兼容基线,无已知注入差异」)在词法上属于紧随其后的 **minimax** 条目,`openai` 条目**没有**任何注释。补的是"后果"而非重复"为何为空"——不要写出与既有注释重复或矛盾的内容。
|
||||
|
||||
**验收标准**: 两个 profile 都能让读者明白 `enable_thinking=False` 对它们无效。纯注释变更,无行为变化。
|
||||
|
||||
**测试要求**: 无(纯注释)。此任务不单独提交,与 Task 10 合并提交。
|
||||
|
||||
**验证**: `make lint` → PASS
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 10: 文档同步(设计 §6 清单)
|
||||
|
||||
**文件**: 改 `.env.example`、`README.md`、`CHANGELOG.md`、`research-wiki/ARCHITECTURE.md`
|
||||
|
||||
| 目标 | 具体改动 |
|
||||
|---|---|
|
||||
| `.env.example` | 在 `LLM__QWEN__1__TRUST_ENV` 注释行(`:20`)后加 `# LLM__QWEN__1__EXTRA_BODY={"temperature":0}` 及说明(JSON 对象串;保护键会报错;OCR/EMBED scope 会被忽略并 warning)。`client.py:251` docstring 声明本文件是键名清单事实源,漏写等于新键无处可查 |
|
||||
| `README.md:83` | 该行逐一列举 `chat()` 关键字参数,补 `overlay` 及一句用途 |
|
||||
| ARCH §5.2 | `chat()` 签名定稿段追加 `overlay` 要点(带默认值的 keyword-only,不破坏"调用点零改动"承诺) |
|
||||
| ARCH §7.5 | key 公式补 `sampling` 项 + 两条已知副作用(seed 进 key 导致该路径必 miss;`model_fingerprint` 是集合级指纹,同 scope 各源 `extra_body` 不同时仍可能跨源命中) |
|
||||
| ARCH §7.7(`:451`) | 该节逐字段枚举 `SourceConfig` 构成(`name/provider/.../enable_thinking`),补 `extra_body` |
|
||||
| ARCH §7.8(`:463`) | 必录字段 20 → 21,补 `sampling` 及其列语义(不含 `response_format`) |
|
||||
| ARCH §9(`:519-527`) | 配置面键族事实源,登记 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY` |
|
||||
| `CHANGELOG.md` | 公共 API 新增(`chat(overlay=)`、`SourceConfig.extra_body`)+ 遥测端口扩列 |
|
||||
|
||||
**Gitea Wiki 同步**(`docs-convention.md` §2,CLAUDE.md §6 标为硬门)。本次同时命中该表两行:
|
||||
|
||||
| 命中行 | 必同步页 |
|
||||
|---|---|
|
||||
| 新公共 API / 新能力 | 对应指南页(新增「固定解码参数」内容,落在 `指南-遥测与成本` 或新页)+ `参考-公共API`(`chat()` 签名、`SourceConfig.extra_body`)+ `_Sidebar.md` + CHANGELOG |
|
||||
| 新增配置键 | `参考-配置键`(登记 `{SCOPE}__{PROVIDER}__{N}__EXTRA_BODY`)+ 相关指南页的配置片段 + `.env.example` |
|
||||
|
||||
指南页必须写明三条坑:① `seed` 逐次变化时该路径缓存**必 miss**;② `model_fingerprint` 是集合级指纹,同 scope 各源 `extra_body` 不同时仍可能跨源命中(要逐源可复现需每源独享 scope 或 namespace);③ OCR/EMBED scope 的 `EXTRA_BODY` 会被忽略并 warning。
|
||||
|
||||
**验收标准**: 每条都能在文件中指到具体位置;ARCH 的改动与设计文档不矛盾;wiki 两行清单逐页落实。
|
||||
|
||||
**测试要求**: 无(纯文档)。与 Task 9 合并提交。
|
||||
|
||||
**验证**: `make lint` → PASS
|
||||
|
||||
---
|
||||
|
||||
### - [ ] Task 11: 全链路集成验证与合并前检查
|
||||
|
||||
**文件**: 测试 `tests/integration/`(就近文件或新增)
|
||||
|
||||
**实现行为**: 端到端断言采样参数经 `chat()` → 选源 → transport payload 到达请求体(设计 §5 测试 #13),用 fake HTTP 层捕获实际 payload。
|
||||
|
||||
**合并前门(逐条出示证据)**:
|
||||
1. `make ci` → 全绿(`make lint` + `make test` + 覆盖率)
|
||||
2. import-linter 契约无新违规(校验函数落最内层 `types.py`,分层关系不变)
|
||||
3. 设计 §5 的 14 条测试全部有对应实现,逐条对应到具体测试函数名
|
||||
4. 派**全新上下文** verifier subagent 独立验证(CLAUDE.md §3.2 里程碑级/跨多文件硬门)
|
||||
|
||||
**验证**:
|
||||
- `make ci` → 全 PASS(**不要**在外面套 `conda run`:`Makefile` 每条 target 内部已是 `conda run -n PolyGateway ...`,嵌套会让内层输出被缓冲)
|
||||
- verifier 报告无 blocking 问题
|
||||
|
||||
---
|
||||
|
||||
## 3. 提交节奏
|
||||
|
||||
| 提交 | 内容 |
|
||||
|---|---|
|
||||
| 1 | Task 1(types 内核) |
|
||||
| 2 | Task 2(env 解析) |
|
||||
| 3 | Task 3(chat 入口 + transport + fingerprint) |
|
||||
| 4 | Task 4(缓存 key) |
|
||||
| 5 | Task 5(地基不变式测试) |
|
||||
| 6 | Task 6(遥测三入口) |
|
||||
| 7 | Task 7(两后端落列) |
|
||||
| 8 | Task 8(决策 G) |
|
||||
| 9 | Task 9 + 10(注释与文档) |
|
||||
| 10 | Task 11(集成验证,如有修补) |
|
||||
|
||||
每次提交调 `commit` skill。Task 1-4 是 issue 诉求的最小闭环;Task 5-8 是设计中"issue 未提但必须处理"的部分,**不可跳过**。
|
||||
@@ -0,0 +1,17 @@
|
||||
---
|
||||
type: plan
|
||||
node_id: plan:est-tokens-decoupling
|
||||
title: "est_tokens 解耦实施计划"
|
||||
date: 2026-07-30
|
||||
---
|
||||
|
||||
# est_tokens 解耦实施计划
|
||||
|
||||
全文见 [2026-07-30-est-tokens-decoupling-plan.md](2026-07-30-est-tokens-decoupling-plan.md)。实现设计 [est-tokens-decoupling](../designs/est-tokens-decoupling.md)。
|
||||
|
||||
- **5 个任务**: T1 加派生能力与三态值域常量(零行为变更)→ T2 五个入场/结算点切到派生值(零行为变更,因显式值优先)→ T3 值域三态生效(行为变更主体)→ T4 解绑 `tpm > 0 ⇒ est_tokens > 0`(派生值真正启用)→ T5 权威文档、CHANGELOG、wiki 与 issue 回帖。
|
||||
- **排序是硬约束,不可调换**: 三处改动互相牵制且中间态**静默偏差、不报错**。先改 usage 兜底为 `(0,0)` 而结算点未切派生值 → 成功调用押金整笔退回(TPM 闸退化成进门即放行);先解绑约束而结算点未切 → 同样泄漏;先改 `openai_compat.py:176` 而 `_merge` 仍是二值 `any(=="estimated")` → `unavailable` 批被误标 `measured` 且 cost 照算。
|
||||
- **T2 的等价性是安全阀**: 约束未解绑时 `effective_est_tokens()` 恒返回显式值,故 T1/T2 后行为逐字不变,现有测试全绿即为证明;T3 才是唯一的行为变更点。
|
||||
- **保真校验(不新增移植,但触及关键资产)**: 不得改 Redis Lua 与内存后端的窗口/租约算法(只改传入 `try_acquire` 的数值来源)、不得改 `settle` 多退少补与幂等语义、不得改错误四分类归属、`ocr.py` 一字不动、遥测 18 字段冻结且无 DDL。
|
||||
- **最易漏的测试**: T4 的"成功侧结算不退多"——未填 `est_tokens` 且 usage 帧缺失的**成功**调用后,TPM 窗口残留须等于派生预扣量而非 0。这正是独立审查在设计阶段抓出的缺陷,实施阶段必须有回归钉死。
|
||||
- **发布口径**: 缺口度量必须写成 `WHERE usage_source='unavailable' AND cache_hit = false`——缓存命中行按裁决 cost 为 `0.0` 且标 `unavailable`,本无账目缺口,不加限定则度量偏高。
|
||||
@@ -0,0 +1,30 @@
|
||||
---
|
||||
type: plan
|
||||
node_id: plan:response-observability-fields
|
||||
title: 响应可观测字段扩展实现计划
|
||||
date: 2026-07-31
|
||||
---
|
||||
|
||||
# 响应可观测字段扩展实现计划
|
||||
|
||||
全文见 `2026-07-31-response-observability-fields.md`。实现 [[response-observability-fields]] 设计(A2/B1/C1/D1)。
|
||||
|
||||
## 任务序列
|
||||
|
||||
| 任务 | 内容 | 提交 |
|
||||
|---|---|---|
|
||||
| T1 | `types.py` 两个类型各 +2 字段;`cache_hit` docstring 消歧 | 独立 |
|
||||
| T2 | `openai_compat.py` 防御解析 + SSE sink 采集 `model` + 两处构造填值 | 独立 |
|
||||
| T3 | `retry.py` 搬运;`cache.py` 零改动但用测试固化 B1 回放语义 | 独立 |
|
||||
| T4 | `pricing.py` 可选缓存单价档 + 夹取防负 | 独立 |
|
||||
| T5+T6 | 端口 18→20、两后端 DDL 加列与幂等补列、emitter 搬运、契约测试 | **必须合一次提交** |
|
||||
| T7 | ARCHITECTURE §7.8 / CHANGELOG / `.env.example:56` / 6 处「18 字段」措辞 / 版本 1.1.0 / Gitea Wiki 站 | 独立 |
|
||||
|
||||
## 独立审查抓出的四个坑(已折回计划)
|
||||
|
||||
1. **DDL 新列必须放在 `created_at` 之后**(表末尾)。旧表走 `ALTER ADD COLUMN` 只能追加到末尾,若新建库把新列插在 `created_at` 前,两条路径列序分叉 —— 而 `test_schema_has_frozen_columns_in_order` 按 `ordinal_position` 逐位断言,且该 PG 表与真实批跑共享、严禁 DROP,分叉后无合规修法。
|
||||
2. **SQLite 补列块首行必须守卫 `if self._conn is None: return`**。否则初始化失败时补列块抛 `AttributeError`/`NameError`(不被 `sqlite3.Error` 捕获)逃出 `__init__`,打破「初始化失败静默降级」契约。
|
||||
3. **T5 与 T6 不得分开提交**。中间状态下 emitter 只传 18 键,后端抛 `KeyError` 被吞成 warning,该 commit 全量遥测静默丢失。
|
||||
4. **天然拦截点是四处而非三处**:两个 `_record_minimal` + 两个 `_EXPECTED_COLUMNS`;`test_ports.py:96` 的全签名 fake **不会**红(Protocol 的 isinstance 不校验签名),不能当作覆盖保证。
|
||||
|
||||
相关: [[response-observability-fields]]、[[est-tokens-decoupling]]
|
||||
@@ -0,0 +1,17 @@
|
||||
---
|
||||
type: plan
|
||||
node_id: plan:sampling-params-plan
|
||||
title: "采样参数透传实现计划(issue #4)"
|
||||
date: 2026-07-31
|
||||
---
|
||||
|
||||
# 采样参数透传实现计划(issue #4)
|
||||
|
||||
正文: `2026-07-31-sampling-params.md`。实现 `design:sampling-params`。
|
||||
|
||||
- **任务数**: 11 个,每个一次提交、独立可验证。Task 1-4 是 issue 诉求的最小闭环;Task 5-8 是设计中「issue 未提但必须处理」的部分(地基不变式、遥测三入口、两后端落列、决策 G 剥离),不可跳过。
|
||||
- **关键接口已在计划 §1 定死**: `validate_request_overlay()` / `merge_sampling()` / `canonical_sampling_json()` 三个纯函数落 `types.py`(最内层),`ChatRequest.sampling`、`SourceConfig.extra_body` 两个新字段,`build_cache_key()` 与 `chat()` 的新签名。
|
||||
- **审查暴露的执行陷阱(已写进计划)**: ① 两个 `_record_minimal()` 的硬编码 20 键 fields dict 必须同步,否则 `row = tuple(fields[col] for col in _COLUMNS)`(在 try 之外)抛裸 `KeyError` 让两侧落库测试全红;② `SourceConfig` 加 mapping 字段后不再 hashable、`asdict`/`deepcopy` 失效——已核实库内无调用点会踩,作为已知后果显式接受并加锁定测试;③ 各文件需新增的 import 逐一列出(`types.py` 无 `from __future__ import annotations`,注解在类体求值);④ `validate_request_overlay` 的校验顺序必须先查 str 键再试序列化,否则非 str 键会被误报成"值不可序列化"。
|
||||
- **测试证据门**: 每个任务合并前须出示先失败后通过的证据。Task 5 单列一条地基不变式回归——决策 C/D 都建立在「`sampling` 跨层恒定」之上,而这条目前只靠 `dataclasses.replace` 的约定,无机械执法;该测试须在故意破坏 `structured.py` 时验证过确实变红。
|
||||
- **共享后端纪律**: Task 7、Task 11 涉及 PG `polygateway` 库,严禁与其他会话并跑(含 git 钩子触发的测试)。
|
||||
- **审查留痕**: Codex CLI 不可用(vendor 二进制缺失),派全新上下文 subagent 只读审查。报 5 项必修(两处测试文件路径不存在、`_record_minimal` 漏项、Gitea Wiki 同步漏整块、`SourceConfig` 可哈希性后果未声明),逐条核实后全部采纳;5 条建议(import 清单、校验顺序、Task 5 落点表述、`make ci` 勿嵌套 `conda run`、`test_ports.py` 的 `_DummyRecorder` 同步)亦已收进。
|
||||
@@ -1,11 +1,11 @@
|
||||
---
|
||||
type: schema
|
||||
node_id: schema:llm-calls
|
||||
title: "表结构: llm_calls(遥测 18 字段)"
|
||||
title: "表结构: llm_calls(遥测 21 字段)"
|
||||
date: 2026-07-20
|
||||
---
|
||||
|
||||
# 表结构: llm_calls(遥测 18 字段)
|
||||
# 表结构: llm_calls(遥测 21 字段)
|
||||
|
||||
|
||||
## 列定义(冻结,M1 设计 §4.4 / ARCH §7.8)
|
||||
@@ -16,14 +16,63 @@ date: 2026-07-20
|
||||
| parent_call_id / session_id | TEXT | 调用链路(agent step → LLM call) |
|
||||
| model / provider / source_name | TEXT NOT NULL | 溯源;model 由旧 Protocol 的 model_name 更名(VT 迁移 §8) |
|
||||
| messages / response / thinking | TEXT NOT NULL | messages 落库前多模态 part 摘要(与缓存 key 共用 digest_messages) |
|
||||
| prompt_tokens / completion_tokens | INTEGER NOT NULL | usage 帧;缺失按 est 兜底 |
|
||||
| usage_source | TEXT NOT NULL | measured / estimated |
|
||||
| prompt_tokens / completion_tokens | INTEGER NOT NULL | usage 帧;帧缺失记 0/0(不编造估值,由 usage_source 标注) |
|
||||
| usage_source | TEXT NOT NULL | measured / estimated / unavailable(2026-07-30 起三态,见下) |
|
||||
| latency_ms | INTEGER NOT NULL | 尝试耗时;缓存命中 0 |
|
||||
| ttft_ms / max_inter_token_ms | REAL | 流式活性测量 |
|
||||
| cache_hit | INTEGER NOT NULL DEFAULT 0 | 命中标记 |
|
||||
| error | TEXT | 异常信息;取消记 "cancelled" |
|
||||
| cost | REAL | M1 恒 NULL,M2 pricing 换算 |
|
||||
| cost | REAL | M2 起 pricing 换算;`usage_source='unavailable'` 的真实调用行为 NULL(缓存命中行例外,仍为 0.0) |
|
||||
| created_at | TEXT NOT NULL DEFAULT (datetime('now')) | 落库时刻 |
|
||||
| cached_prompt_tokens | INTEGER | 供应商 prompt cache 命中的输入 token(2026-07-31,issue #3);NULL = 该源未上报,`0` = 上报了真实零命中,两者不可混同 |
|
||||
| model_reported | TEXT | API 响应体实际返回的 model;NULL = 未上报。与 `model`(配置别名)可能分叉 |
|
||||
| sampling | TEXT | 本次调用的采样参数 canonical JSON(2026-07-31,issue #4);NULL = 未传。见下方口径 |
|
||||
|
||||
## usage/成本口径(2026-07-30,est_tokens 解耦)
|
||||
|
||||
| usage_source | 含义 | 生产者 | cost |
|
||||
|---|---|---|---|
|
||||
| `measured` | usage 帧完整可信 | 正常路径;OCR 成功行(0 token 是事实) | 按 token 换算 |
|
||||
| `estimated` | 有实测数字但可信度降级 | 打捞路径(收到 usage 帧但流被截断) | 按 token 换算 |
|
||||
| `unavailable` | 用量信息不可得 | usage 帧缺失、失败尝试、终态失败 | NULL |
|
||||
|
||||
`SUM(cost)` 天然跳过 NULL,故账单汇总不再被虚构的估值污染;账目缺口的度量口径固定为 `WHERE usage_source = 'unavailable' AND cache_hit = false`。**`cache_hit` 限定不可省**:缓存命中行未产生新调用,cost 是事实上的 `0.0` 而非未知,本无账目缺口,漏掉该条件会让缺口度量偏高。
|
||||
|
||||
## 供应商 prompt cache 口径(2026-07-31,issue #3)
|
||||
|
||||
新增两列排在 `created_at` **之后**——旧表只能经 `ALTER TABLE ADD COLUMN` 追加到末尾,DDL 里若插在前面,新建库与升级库的物理列序会分叉(列序断言无合规修法)。两个后端在初始化期幂等补列:`CREATE TABLE IF NOT EXISTS` 不会给旧表加列,不补则每行写入被逐行 warning 丢弃、遥测静默全失。两侧都**先探测缺列再 ALTER**(`ADD COLUMN IF NOT EXISTS` 即使列已存在也先取 ACCESS EXCLUSIVE 锁,遥测是内联 await,锁共享审计表会拖垮业务调用),且**补列失败只降级为逐行丢弃,绝不让 recorder 整体失能**——两侧纪律必须对称。
|
||||
|
||||
`cache_hit` 指 **PolyGateway 自身响应缓存**,与供应商 prompt cache 是两回事。缓存命中行的这两列是**原样回放**的历史值(与 `model`/`prompt_tokens` 同一口径),故命中率度量口径固定为:
|
||||
|
||||
```sql
|
||||
SELECT SUM(cached_prompt_tokens)::float / NULLIF(SUM(prompt_tokens), 0)
|
||||
FROM llm_calls WHERE cache_hit = false AND cached_prompt_tokens IS NOT NULL;
|
||||
```
|
||||
|
||||
`WHERE cache_hit = false` 不可省,理由与上面 cost 缺口口径同源:回放行计入即重复计数。
|
||||
|
||||
## 采样参数口径(2026-07-31,issue #4)
|
||||
|
||||
`sampling` 列 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化输出注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。补列纪律与 issue #3 两列逐字相同(排在末尾、先探测再 ALTER、失败只逐行降级)。
|
||||
|
||||
三个 emit 入口的取值必须各自定死,否则同一列在不同行含义不同:
|
||||
|
||||
| 入口 | 调用者 | 有生效源? | 记什么 |
|
||||
|---|---|---|---|
|
||||
| `emit_attempt` | RetryMW(最内) | 有 | `merge(source.extra_body, request.sampling)` |
|
||||
| `emit_cache_hit` | TelemetryMW(最外) | 无 | 仅 `request.sampling` |
|
||||
| `emit_terminal_failure` | TelemetryMW | 无 | 仅 `request.sampling` |
|
||||
|
||||
后两行缺 `extra_body` 是客观事实而非口径瑕疵——它们没有"生效源"可言,与 `model`/`source_name` 在终态行置空是同一先例;缓存命中行亦无损:`sampling` 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同。三者统一读 `request.sampling` 而非 `request.overlay`(后者在 RetryMW 处已被结构化注入污染、在 TelemetryMW 处未被污染,直接用必然三行分叉)。
|
||||
|
||||
OCR / embedding 路径的该列**恒为 NULL**:两条路径的 transport 不发 `extra_body`(embed payload 硬编码 `{model, input}`、MonkeyOCR 只发 multipart),故其源在构造期就被剥离——不剥离则该列会记录一个从未发出的参数,那是数据造假而非参数失效。
|
||||
|
||||
复现某批实验的解码条件:
|
||||
|
||||
```sql
|
||||
SELECT DISTINCT sampling FROM llm_calls
|
||||
WHERE session_id = $1 AND cache_hit = false AND error IS NULL;
|
||||
```
|
||||
|
||||
## 埋点位置(单一 helper 铁律)
|
||||
|
||||
|
||||
@@ -31,7 +31,7 @@ from polygateway.types import (
|
||||
SourceConfig,
|
||||
)
|
||||
|
||||
__version__ = "1.0.0"
|
||||
__version__ = "1.0.5"
|
||||
|
||||
__all__ = [
|
||||
"DEFAULT_PROFILES",
|
||||
|
||||
+47
-10
@@ -9,6 +9,8 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import hashlib
|
||||
import json
|
||||
import random
|
||||
import time
|
||||
from typing import TYPE_CHECKING, Any, Literal, TypeVar
|
||||
@@ -32,7 +34,7 @@ from polygateway.sources import (
|
||||
SourceCooldownMemo,
|
||||
)
|
||||
from polygateway.transports.openai_compat import OpenAICompatTransport
|
||||
from polygateway.types import ChatRequest, LLMResponse
|
||||
from polygateway.types import ChatRequest, LLMResponse, validate_request_overlay
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from collections.abc import Awaitable, Iterable, Mapping
|
||||
@@ -59,6 +61,29 @@ if TYPE_CHECKING:
|
||||
_T = TypeVar("_T")
|
||||
|
||||
|
||||
def build_model_fingerprint(sources: Iterable[SourceConfig]) -> str:
|
||||
"""缓存 key 的模型身份: 多源 scope = 排序去重的 model 合集。
|
||||
|
||||
配置级采样参数(`extra_body`)必须参与,否则把 temperature 从 0 改成 1
|
||||
后重启仍会读到旧缓存(issue #4 设计决策 C)。全源 `extra_body` 皆空时
|
||||
字面量与历史实现逐字相同,不触发存量缓存冷启动。
|
||||
"""
|
||||
fingerprint = ",".join(sorted({s.model for s in sources}))
|
||||
# 按 (model, extra_body) 而非源名摘要: 语义是"本 scope 会用哪些
|
||||
# (模型, 解码参数)组合",改源名不该误触全量冷启动
|
||||
marks = sorted(
|
||||
{
|
||||
json.dumps([s.model, dict(s.extra_body)], sort_keys=True, ensure_ascii=False)
|
||||
for s in sources
|
||||
if s.extra_body
|
||||
}
|
||||
)
|
||||
if marks:
|
||||
digest = hashlib.sha256("".join(marks).encode("utf-8")).hexdigest()
|
||||
fingerprint = f"{fingerprint}|{digest}"
|
||||
return fingerprint
|
||||
|
||||
|
||||
class GatewayClient:
|
||||
"""统一治理入口;构造函数全量注入(测试/高级),工厂覆盖 90% 场景。"""
|
||||
|
||||
@@ -113,12 +138,11 @@ class GatewayClient:
|
||||
if cache is not None:
|
||||
if cache_namespace is None or cache_ttl_s is None:
|
||||
raise ValueError("启用缓存必须提供 cache_namespace 与 cache_ttl_s")
|
||||
# 多源 scope 的 key 身份 = 排序去重的 model 合集;源集合变化 → 一次性冷启动
|
||||
fingerprint = ",".join(sorted({s.model for s in sources}))
|
||||
# 多源 scope 的 key 身份;源集合或其 extra_body 变化 → 一次性冷启动
|
||||
middlewares.append(
|
||||
CacheMW(
|
||||
backend=cache,
|
||||
model_fingerprint=fingerprint,
|
||||
model_fingerprint=build_model_fingerprint(sources),
|
||||
default_namespace=cache_namespace,
|
||||
ttl_s=cache_ttl_s,
|
||||
strategy=structured_strategy,
|
||||
@@ -150,12 +174,23 @@ class GatewayClient:
|
||||
cache_namespace: str | None = None,
|
||||
structured: type[BaseModel] | Literal["json"] | None = None,
|
||||
stream: bool = True,
|
||||
overlay: Mapping[str, Any] | None = None,
|
||||
) -> LLMResponse:
|
||||
"""一次治理调用(签名冻结,ARCH §5.2;与三项目 LLMProvider 协议兼容)。"""
|
||||
"""一次治理调用(签名冻结,ARCH §5.2;与三项目 LLMProvider 协议兼容)。
|
||||
|
||||
`overlay` 是采样参数覆盖层(`temperature`/`seed`/`max_tokens` 等),优先级
|
||||
高于源级 `extra_body`、低于结构化输出的注入。带默认值的 keyword-only
|
||||
参数不影响既有调用点(issue #4)。
|
||||
"""
|
||||
if structured is not None and not self._structured_available:
|
||||
raise ImportError(
|
||||
"结构化输出未启用: 安装 pip install 'polygateway[structured]' 后重新装配"
|
||||
)
|
||||
# 进洋葱之前校验并拷贝: 保护键/不可序列化值在此收口(否则会在 CacheMW
|
||||
# 的降级 try 之外抛裸 TypeError);拷贝防调用方复用同一 dict 逐次改 seed
|
||||
# 造成的竞态。同一份快照填 overlay 与 sampling——前者会被结构化注入,
|
||||
# 后者跨层恒定,供缓存 key 与遥测读取(设计决策 A/B/E)
|
||||
sampling = validate_request_overlay(overlay or {}, origin="chat(overlay=...)")
|
||||
request = ChatRequest(
|
||||
messages=messages,
|
||||
session_id=session_id,
|
||||
@@ -164,6 +199,8 @@ class GatewayClient:
|
||||
cache_namespace=cache_namespace,
|
||||
structured=structured,
|
||||
stream=stream,
|
||||
overlay=sampling,
|
||||
sampling=sampling,
|
||||
)
|
||||
return await self._handler(request)
|
||||
|
||||
@@ -259,7 +296,7 @@ def _build_limiter(settings: GatewaySettings, sources: list[SourceConfig]) -> Ra
|
||||
if settings.limiter_backend == "redis":
|
||||
from polygateway.backends.redis.limiter import RedisLimiter
|
||||
|
||||
assert settings.redis_url is not None # 内部不变量: config 已校验
|
||||
assert settings.redis_url is not None # 内部不变量: _validate_backends 已保证
|
||||
return RedisLimiter.from_url(
|
||||
settings.redis_url,
|
||||
scope=settings.scope,
|
||||
@@ -279,7 +316,7 @@ def _build_breaker(settings: GatewaySettings) -> ProviderGate:
|
||||
if settings.breaker_backend == "redis":
|
||||
from polygateway.backends.redis.breaker import RedisGate
|
||||
|
||||
assert settings.redis_url is not None # 内部不变量: config 已校验
|
||||
assert settings.redis_url is not None # 内部不变量: _validate_backends 已保证
|
||||
return RedisGate.from_url(settings.redis_url, config=settings.breaker, scope=settings.scope)
|
||||
return InMemoryGate(config=settings.breaker)
|
||||
|
||||
@@ -299,7 +336,7 @@ def _build_cache(settings: GatewaySettings) -> CacheBackend | None:
|
||||
return InMemoryCache()
|
||||
from polygateway.backends.redis_cache import RedisCache
|
||||
|
||||
assert settings.redis_url is not None # 内部不变量: config 已校验
|
||||
assert settings.redis_url is not None # 内部不变量: _validate_backends 已保证
|
||||
return RedisCache.from_url(settings.redis_url)
|
||||
|
||||
|
||||
@@ -309,11 +346,11 @@ def _build_telemetry(settings: GatewaySettings) -> TelemetryRecorder | None:
|
||||
if settings.telemetry_backend == "postgres":
|
||||
from polygateway.telemetry.postgres import PostgresRecorder
|
||||
|
||||
assert settings.telemetry_pg_dsn is not None # 内部不变量: config 已校验
|
||||
assert settings.telemetry_pg_dsn is not None # 内部不变量: _validate_telemetry 已保证
|
||||
return PostgresRecorder(settings.telemetry_pg_dsn)
|
||||
from polygateway.telemetry.sqlite import SQLiteRecorder
|
||||
|
||||
assert settings.telemetry_sqlite_path is not None # 内部不变量: config 已校验
|
||||
assert settings.telemetry_sqlite_path is not None # 内部不变量: _validate_telemetry 已保证
|
||||
return SQLiteRecorder(settings.telemetry_sqlite_path)
|
||||
|
||||
|
||||
|
||||
+173
-44
@@ -11,11 +11,13 @@ fail-loud 校验语义与 pydantic-settings 一致。
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
import os
|
||||
from dataclasses import dataclass
|
||||
from typing import TYPE_CHECKING
|
||||
|
||||
from dotenv import dotenv_values
|
||||
from loguru import logger
|
||||
|
||||
from polygateway.types import (
|
||||
BackpressurePolicy,
|
||||
@@ -43,14 +45,22 @@ _SOURCE_FIELDS: dict[str, tuple[str, str]] = {
|
||||
"ENABLE_THINKING": ("enable_thinking", "bool"),
|
||||
"MISSING_DONE": ("missing_done", "str"),
|
||||
"TRUST_ENV": ("trust_env", "bool"),
|
||||
"EXTRA_BODY": ("extra_body", "json"),
|
||||
}
|
||||
_RESERVED_SEGMENTS = frozenset({"GLOBAL", "RETRY", "BREAKER", "BACKPRESSURE"})
|
||||
_SELECTORS = frozenset({"round_robin", "least_inflight", "health_aware"})
|
||||
_QUOTA_FULL = frozenset({"wait", "fail_fast"})
|
||||
# 后端合法域: env 解析与构造期校验共用一份定义,避免两处分叉
|
||||
_LIMITER_BACKENDS = frozenset({"memory", "redis"})
|
||||
_BREAKER_BACKENDS = frozenset({"memory", "redis"})
|
||||
_CACHE_BACKENDS = frozenset({"redis", "memory", "none"})
|
||||
_TELEMETRY_BACKENDS = frozenset({"sqlite", "postgres", "none"})
|
||||
_REDIS_DEPENDENT_BACKENDS = ("limiter_backend", "breaker_backend", "cache_backend")
|
||||
# 背压默认(M1 仅 poll 生效;CHS _BACKOFF_S=0.05 同源)
|
||||
_DEFAULT_STALL_WINDOW_S = 300.0
|
||||
_DEFAULT_POLL_INTERVAL_S = 0.05
|
||||
_DEFAULT_LEASE_TTL_S = 1500.0 # CHS _DEFAULT_LEASE_TTL_MS 同源
|
||||
_PROBE_GRACE_S = 5.0 # 半开探针租约相对最慢调用的清理宽限(CHS container.py:274-275)
|
||||
|
||||
|
||||
def _cast(raw: str, kind: str, key: str) -> object:
|
||||
@@ -66,6 +76,12 @@ def _cast(raw: str, kind: str, key: str) -> object:
|
||||
if lowered in ("0", "false", "no", "off"):
|
||||
return False
|
||||
raise ValueError(f"非法布尔值: {raw!r}")
|
||||
if kind == "json":
|
||||
# JSONDecodeError 是 ValueError 子类,复用下方的统一包装
|
||||
parsed = json.loads(raw)
|
||||
if not isinstance(parsed, dict):
|
||||
raise ValueError(f"必须是 JSON 对象(而非数组/标量): {raw!r}")
|
||||
return parsed
|
||||
return raw
|
||||
except ValueError as exc:
|
||||
raise ValueError(f"配置 {key} 解析失败: {exc}") from exc
|
||||
@@ -88,7 +104,16 @@ def _require(env: Mapping[str, str], *keys: str) -> tuple[str, str]:
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class GatewaySettings:
|
||||
"""一个 scope 的完整装配配置;构造经 from_env 聚合并通过全部守卫。"""
|
||||
"""一个 scope 的完整装配配置;**任何**构造路径都通过全部装配守卫(ARCH §7.3)。
|
||||
|
||||
守卫校验的是**跨字段**不变量: 单看一个字段都合法,组合起来才会在运行时
|
||||
咬人(租约先于请求过期、正常慢首包被误判卡死、半开探针在途被接管)。
|
||||
types.py 各子配置的 `__post_init__` 只看得见自己的字段,故由本类把关。
|
||||
|
||||
放在 `__post_init__` 而非某个工厂里: 这些约束是本类定义的一部分,不是
|
||||
某个入口的输入检查。挂在构造期,直接构造、`dataclasses.replace` 与全部
|
||||
装配工厂一并覆盖;挂在工厂里则每加一个工厂就多一处要同步。
|
||||
"""
|
||||
|
||||
scope: str
|
||||
sources: tuple[SourceConfig, ...]
|
||||
@@ -111,6 +136,123 @@ class GatewaySettings:
|
||||
structured_max_retries: int
|
||||
lease_ttl_s: float
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
self._normalize()
|
||||
self._validate_identity()
|
||||
self._validate_backends()
|
||||
self._validate_cache()
|
||||
self._validate_telemetry()
|
||||
self._validate_lease()
|
||||
self._validate_stall()
|
||||
self._validate_probe()
|
||||
|
||||
def _normalize(self) -> None:
|
||||
"""把 `from_env` 一直在做的规范化补到构造路上,两条路必须产出同一个值。
|
||||
|
||||
`scope` 最要紧: 它直接进 Redis key(`pgw:limit:{scope}:…`/`pgw:gate:{scope}:…`)。
|
||||
一个进程走 `from_env("LLM")` 拿到 "llm"、另一个直接构造传 "LLM",同一逻辑
|
||||
scope 的限流与熔断状态会分裂到两套命名空间,各记各的,治理静默失效且不报错。
|
||||
|
||||
空串归 None 同理: 留着空串会骗过 `is None` 判断,把错误推迟到 redis 客户端
|
||||
抛连接串解析异常。`telemetry_pg_dsn` 的驱动后缀因为要看 backend 且需告警,
|
||||
规范化留在 `_validate_telemetry`。
|
||||
"""
|
||||
normalized_scope = self.scope.strip().lower()
|
||||
if normalized_scope != self.scope:
|
||||
object.__setattr__(self, "scope", normalized_scope)
|
||||
for field in ("redis_url", "pricing_path"):
|
||||
if getattr(self, field) == "":
|
||||
object.__setattr__(self, field, None)
|
||||
|
||||
def _validate_identity(self) -> None:
|
||||
"""本类自身字段的基本域: 空 scope 会污染遥测与缓存命名空间;零源必然选源失败。"""
|
||||
if not self.scope.strip():
|
||||
raise ValueError("GatewaySettings.scope 不能为空")
|
||||
if not self.sources:
|
||||
raise ValueError("GatewaySettings.sources 不能为空: 至少一个源")
|
||||
if self.structured_max_retries < 0:
|
||||
raise ValueError(f"structured_max_retries 不能为负: {self.structured_max_retries}")
|
||||
|
||||
def _validate_backends(self) -> None:
|
||||
"""后端选择必须落在合法域内,取 redis 的还必须有连接串。
|
||||
|
||||
域外取值此前只有 `from_env` 拦得住,直接构造会一路走到 `client.py` 的
|
||||
`_build_*`,落进 else 分支静默不建后端,或撞上那里的断言。
|
||||
"""
|
||||
for field, allowed in (
|
||||
("limiter_backend", _LIMITER_BACKENDS),
|
||||
("breaker_backend", _BREAKER_BACKENDS),
|
||||
("cache_backend", _CACHE_BACKENDS),
|
||||
("telemetry_backend", _TELEMETRY_BACKENDS),
|
||||
("selector", _SELECTORS),
|
||||
("quota_full", _QUOTA_FULL),
|
||||
):
|
||||
value = getattr(self, field)
|
||||
if value not in allowed:
|
||||
raise ValueError(f"{field} 非法值 {value!r};允许: {sorted(allowed)}")
|
||||
on_redis = [f for f in _REDIS_DEPENDENT_BACKENDS if getattr(self, f) == "redis"]
|
||||
if on_redis and self.redis_url is None:
|
||||
raise ValueError(f"{'、'.join(on_redis)} 取 redis 时必须提供 redis_url")
|
||||
|
||||
def _validate_cache(self) -> None:
|
||||
"""启用缓存必须有命名空间与正 TTL(缺命名空间即失去租户隔离,会毒化缓存)。"""
|
||||
if self.cache_backend == "none":
|
||||
return
|
||||
if not self.cache_namespace:
|
||||
raise ValueError("启用缓存时 cache_namespace 不能为空: 缓存 key 靠它做租户隔离")
|
||||
if self.cache_ttl_s is None or self.cache_ttl_s <= 0:
|
||||
raise ValueError(f"cache_ttl_s 必须 > 0(禁止永不过期): {self.cache_ttl_s}")
|
||||
|
||||
def _validate_telemetry(self) -> None:
|
||||
"""遥测后端各自的落点必填;顺带剥掉 asyncpg 不认的 SQLAlchemy 驱动后缀。
|
||||
|
||||
剥而不是拒: 两条装配路对同一 DSN 应产出同一结果。但不静默——`from_env`
|
||||
那条路在 `_load_pg_dsn` 就剥干净了,能走到这里的只有手工构造的调用方,
|
||||
他有权知道库动了他给的值。
|
||||
"""
|
||||
if self.telemetry_backend == "sqlite" and not self.telemetry_sqlite_path:
|
||||
raise ValueError("telemetry_backend=sqlite 时必须提供 telemetry_sqlite_path")
|
||||
if self.telemetry_backend != "postgres":
|
||||
return
|
||||
if not self.telemetry_pg_dsn:
|
||||
raise ValueError("telemetry_backend=postgres 时必须提供 telemetry_pg_dsn")
|
||||
stripped = _strip_dsn_driver(self.telemetry_pg_dsn)
|
||||
if stripped != self.telemetry_pg_dsn:
|
||||
# 只报 scheme 段: DSN 带密码,整串不得进日志(P5 敏感信息只走 .env)
|
||||
logger.warning(
|
||||
"telemetry_pg_dsn 的 scheme 含 asyncpg 不认的驱动后缀,已由 {} 剥为 {}",
|
||||
self.telemetry_pg_dsn.partition("://")[0],
|
||||
stripped.partition("://")[0],
|
||||
)
|
||||
object.__setattr__(self, "telemetry_pg_dsn", stripped)
|
||||
|
||||
def _validate_lease(self) -> None:
|
||||
"""调用超时须 ≤ permit 租约 TTL,防租约先于请求过期使并发超出配额。"""
|
||||
slowest = max(s.timeout_s for s in self.sources)
|
||||
if slowest > self.lease_ttl_s:
|
||||
raise ValueError(
|
||||
f"源最大 timeout_s({slowest})超过 permit 租约 lease_ttl_s"
|
||||
f"({self.lease_ttl_s});调大 lease_ttl_s 或调小源的 timeout_s"
|
||||
)
|
||||
|
||||
def _validate_stall(self) -> None:
|
||||
"""stall 窗口须 ≥ 最慢源 TTFT 上限,防把正常慢首包误判为卡死。"""
|
||||
ttfts = [s.ttft_timeout_s for s in self.sources if s.ttft_timeout_s is not None]
|
||||
if ttfts and self.backpressure.stall_window_s < max(ttfts):
|
||||
raise ValueError(
|
||||
f"backpressure.stall_window_s({self.backpressure.stall_window_s})须 ≥ "
|
||||
f"最大源 ttft_timeout_s({max(ttfts)});调大 stall_window_s 或调小 ttft_timeout_s"
|
||||
)
|
||||
|
||||
def _validate_probe(self) -> None:
|
||||
"""半开探针租约须撑过一次最慢调用,否则探针在途即被接管(M2 设计 §3)。"""
|
||||
floor = max(s.timeout_s for s in self.sources) + _PROBE_GRACE_S
|
||||
if self.breaker.probe_ttl_s < floor:
|
||||
raise ValueError(
|
||||
f"breaker.probe_ttl_s({self.breaker.probe_ttl_s})须 ≥ 最慢源 "
|
||||
f"timeout_s + {_PROBE_GRACE_S}({floor});调大 probe_ttl_s 或调小源的 timeout_s"
|
||||
)
|
||||
|
||||
@classmethod
|
||||
def from_env(
|
||||
cls,
|
||||
@@ -119,7 +261,7 @@ class GatewaySettings:
|
||||
*,
|
||||
env_file: str = ".env",
|
||||
) -> GatewaySettings:
|
||||
"""聚合 env(缺省 .env + os.environ,后者优先)并执行装配守卫。"""
|
||||
"""聚合 env(缺省 .env + os.environ,后者优先);守卫由 `__post_init__` 执行。"""
|
||||
if env is None:
|
||||
env = {
|
||||
k: v for k, v in {**dotenv_values(env_file), **os.environ}.items() if v is not None
|
||||
@@ -129,7 +271,7 @@ class GatewaySettings:
|
||||
global_limits = _load_global_limits(scope_u, env)
|
||||
retry = _load_retry(scope_u, env)
|
||||
breaker = _load_breaker(scope_u, env, sources, global_limits)
|
||||
settings = cls(
|
||||
return cls(
|
||||
scope=scope_u.lower(),
|
||||
sources=tuple(sources),
|
||||
global_limits=global_limits,
|
||||
@@ -140,9 +282,6 @@ class GatewaySettings:
|
||||
quota_full=_load_choice(env, f"{scope_u}__QUOTA_FULL", _QUOTA_FULL, "wait"),
|
||||
**_load_pgw(env),
|
||||
)
|
||||
_guard_lease(settings)
|
||||
_guard_stall(settings)
|
||||
return settings
|
||||
|
||||
|
||||
def _load_sources(scope: str, env: Mapping[str, str]) -> list[SourceConfig]:
|
||||
@@ -217,16 +356,12 @@ def _load_breaker(
|
||||
if concurrency > 0:
|
||||
threshold = max(threshold, concurrency * 2)
|
||||
slowest = max(s.timeout_s for s in sources)
|
||||
probe_floor = slowest + 5.0 # CHS container.py:274-275: 最慢调用 + 清理宽限
|
||||
probe_floor = slowest + _PROBE_GRACE_S
|
||||
probe = _first(env, f"{scope}__BREAKER__PROBE_TTL_S")
|
||||
if probe is not None:
|
||||
# 配置值不在此校验: 探针租约下限是跨字段不变量,由 GatewaySettings._validate_probe
|
||||
# 统一把关(否则直接构造那条装配路会绕过)
|
||||
probe_ttl_s = float(_cast(probe[1], "float", probe[0]))
|
||||
# 装配守卫(M2 设计 §3): 探针租约必须撑过一次最慢调用,否则半开探针在途即被接管
|
||||
if probe_ttl_s < probe_floor:
|
||||
raise ValueError(
|
||||
f"probe_ttl_s({probe_ttl_s})须 ≥ 最大源 timeout_s + 5({probe_floor});"
|
||||
f"调大 {probe[0]} 或调小源超时"
|
||||
)
|
||||
else:
|
||||
# 派生规则: 探针租约须撑过一次最慢调用,且不短于冷却期(第三项保证守卫恒成立)
|
||||
probe_ttl_s = max(2 * slowest, cooldown_s, probe_floor)
|
||||
@@ -277,17 +412,15 @@ def _load_choice(env: Mapping[str, str], key: str, allowed: frozenset[str], defa
|
||||
|
||||
|
||||
def _load_pgw(env: Mapping[str, str]) -> dict[str, object]:
|
||||
limiter_backend = _load_choice(
|
||||
env, "PGW_LIMITER_BACKEND", frozenset({"memory", "redis"}), "memory"
|
||||
)
|
||||
breaker_backend = _load_choice(
|
||||
env, "PGW_BREAKER_BACKEND", frozenset({"memory", "redis"}), "memory"
|
||||
)
|
||||
# 合法域与构造期守卫共用常量;此处的检查保留是为了报错能点出 env 键名,
|
||||
# 构造期那道点的是字段名(两类调用方各看得懂自己那套)
|
||||
limiter_backend = _load_choice(env, "PGW_LIMITER_BACKEND", _LIMITER_BACKENDS, "memory")
|
||||
breaker_backend = _load_choice(env, "PGW_BREAKER_BACKEND", _BREAKER_BACKENDS, "memory")
|
||||
_, cache_backend = _require(env, "PGW_CACHE_BACKEND")
|
||||
_, telemetry_backend = _require(env, "PGW_TELEMETRY_BACKEND")
|
||||
if cache_backend not in ("redis", "memory", "none"):
|
||||
if cache_backend not in _CACHE_BACKENDS:
|
||||
raise ValueError(f"PGW_CACHE_BACKEND 非法值 {cache_backend!r}")
|
||||
if telemetry_backend not in ("sqlite", "postgres", "none"):
|
||||
if telemetry_backend not in _TELEMETRY_BACKENDS:
|
||||
raise ValueError(f"PGW_TELEMETRY_BACKEND 非法值 {telemetry_backend!r}")
|
||||
redis_url = env.get("REDIS_URL") or None
|
||||
if "redis" in (limiter_backend, breaker_backend) and redis_url is None:
|
||||
@@ -309,13 +442,22 @@ def _load_pgw(env: Mapping[str, str]) -> dict[str, object]:
|
||||
}
|
||||
|
||||
|
||||
def _load_pg_dsn(env: Mapping[str, str]) -> str:
|
||||
"""读取 Postgres DSN 并剥 SQLAlchemy 风格驱动后缀(asyncpg 不认 `+driver`)。"""
|
||||
_, dsn = _require(env, "PGW_TELEMETRY_PG_DSN")
|
||||
def _strip_dsn_driver(dsn: str) -> str:
|
||||
"""剥 SQLAlchemy 风格的 `+driver` 后缀(asyncpg 不认);已干净的原样返回。"""
|
||||
scheme, sep, rest = dsn.partition("://")
|
||||
return f"{scheme.partition('+')[0]}{sep}{rest}"
|
||||
|
||||
|
||||
def _load_pg_dsn(env: Mapping[str, str]) -> str:
|
||||
"""读取 Postgres DSN 并剥驱动后缀。
|
||||
|
||||
env 路在此剥干净,构造期那道就无事可做——三项目 `.env` 里的 SQLAlchemy
|
||||
写法不会每次装配都刷一条 warning。
|
||||
"""
|
||||
_, dsn = _require(env, "PGW_TELEMETRY_PG_DSN")
|
||||
return _strip_dsn_driver(dsn)
|
||||
|
||||
|
||||
def _load_cache_keys(
|
||||
env: Mapping[str, str], cache_backend: str, redis_url: str | None
|
||||
) -> dict[str, object]:
|
||||
@@ -339,26 +481,6 @@ def _load_structured_retries(env: Mapping[str, str]) -> int:
|
||||
return value
|
||||
|
||||
|
||||
def _guard_lease(settings: GatewaySettings) -> None:
|
||||
"""装配守卫: 调用超时须 ≤ permit 租约 TTL,防租约先于请求过期(ARCH §7.3)。"""
|
||||
slowest = max(s.timeout_s for s in settings.sources)
|
||||
if slowest > settings.lease_ttl_s:
|
||||
raise ValueError(
|
||||
f"源最大 timeout_s({slowest})超过 permit 租约 TTL({settings.lease_ttl_s});"
|
||||
f"调大 PGW_LEASE_TTL_S 或调小超时"
|
||||
)
|
||||
|
||||
|
||||
def _guard_stall(settings: GatewaySettings) -> None:
|
||||
"""装配守卫: stall 窗口须 ≥ 最慢源 TTFT 上限,防把正常慢首包误判为卡死(ARCH §7.3)。"""
|
||||
ttfts = [s.ttft_timeout_s for s in settings.sources if s.ttft_timeout_s is not None]
|
||||
if ttfts and settings.backpressure.stall_window_s < max(ttfts):
|
||||
raise ValueError(
|
||||
f"stall_window_s({settings.backpressure.stall_window_s})须 ≥ 最大源 "
|
||||
f"ttft_timeout_s({max(ttfts)});调大 BACKPRESSURE__STALL_WINDOW_S 或调小 TTFT"
|
||||
)
|
||||
|
||||
|
||||
def _load_lease_ttl(env: Mapping[str, str]) -> float:
|
||||
found = _first(env, "PGW_LEASE_TTL_S")
|
||||
return float(_cast(found[1], "float", found[0])) if found else _DEFAULT_LEASE_TTL_S
|
||||
@@ -378,6 +500,13 @@ class EmbeddingSettings:
|
||||
normalize: bool = False
|
||||
expected_dim: int | None = None
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
"""自身字段的域校验;内嵌的 gateway 由 `GatewaySettings.__post_init__` 自己把关。"""
|
||||
if self.batch_size < 1:
|
||||
raise ValueError(f"EmbeddingSettings.batch_size 必须 ≥ 1: {self.batch_size}")
|
||||
if self.expected_dim is not None and self.expected_dim < 1:
|
||||
raise ValueError(f"EmbeddingSettings.expected_dim 必须 ≥ 1: {self.expected_dim}")
|
||||
|
||||
@classmethod
|
||||
def from_env(
|
||||
cls,
|
||||
|
||||
@@ -41,7 +41,12 @@ from polygateway.middleware.ratelimit import QuotaGate
|
||||
from polygateway.middleware.retry import _failure_reason, backoff_delay
|
||||
from polygateway.middleware.telemetry import TelemetryEmitter
|
||||
from polygateway.sources import SourceCooldownMemo
|
||||
from polygateway.types import ChatRequest, EmbeddingResponse, LLMResponse
|
||||
from polygateway.types import (
|
||||
ChatRequest,
|
||||
EmbeddingResponse,
|
||||
LLMResponse,
|
||||
strip_unsupported_extra_body,
|
||||
)
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from collections.abc import Awaitable, Callable, Mapping
|
||||
@@ -111,7 +116,9 @@ class EmbeddingClient:
|
||||
if expected_dim is not None and expected_dim < 1:
|
||||
raise ValueError("expected_dim 必须 ≥ 1")
|
||||
self._scope = scope
|
||||
self._sources = list(sources)
|
||||
# embed payload 硬编码 {model, input},带 extra_body 的源必须先剥离,
|
||||
# 否则遥测会记录一个从未发出的采样参数(issue #4 决策 G)
|
||||
self._sources = strip_unsupported_extra_body(list(sources), path="embedding")
|
||||
self._selector = selector
|
||||
self._quota = QuotaGate(limiter)
|
||||
self._breaker = BreakerGate(breaker)
|
||||
@@ -268,6 +275,10 @@ class EmbeddingClient:
|
||||
source_name=source.name,
|
||||
operation="embedding",
|
||||
)
|
||||
if result.usage_source == "unavailable":
|
||||
# 与 RetryMW 同口径: 用量不可得时按入场预扣量结算(设计 §3.2 #9)
|
||||
actual = source.effective_est_tokens()
|
||||
else:
|
||||
actual = result.prompt_tokens
|
||||
await self._record_quietly(self._breaker.record_success(entry))
|
||||
await self._record_quietly(self._quota.mark_progress())
|
||||
@@ -291,7 +302,8 @@ class EmbeddingClient:
|
||||
reasons[source.name] = reason
|
||||
await self._record_quietly(self._breaker.record_failure(entry, reason, dead))
|
||||
if not dead:
|
||||
actual = source.est_tokens # 保守: 失败请求可能已被网关计费(CHS 同款)
|
||||
# 保守: 失败请求可能已被网关计费(CHS 同款);与入场预扣同源取值
|
||||
actual = source.effective_est_tokens()
|
||||
await self._emit(batch, source, call_id, started, session_id, parent_call_id, error=exc)
|
||||
return _FailedBatch(exc, immediate=dead)
|
||||
finally:
|
||||
@@ -380,14 +392,21 @@ class EmbeddingClient:
|
||||
vectors = [_l2_normalize(v) for v in vectors]
|
||||
first = outcomes[0]
|
||||
prompt_tokens = sum(o.result.prompt_tokens for o in outcomes)
|
||||
estimated = any(o.result.usage_source == "estimated" for o in outcomes)
|
||||
# 三态合并优先级(解耦设计 §3.2 #10): 任一批不可得 → 整体不可得
|
||||
sources = {o.result.usage_source for o in outcomes}
|
||||
if "unavailable" in sources:
|
||||
merged_source = "unavailable"
|
||||
elif "estimated" in sources:
|
||||
merged_source = "estimated"
|
||||
else:
|
||||
merged_source = "measured"
|
||||
return EmbeddingResponse(
|
||||
vectors=vectors,
|
||||
dim=first.result.dim,
|
||||
model=first.source.model,
|
||||
provider=first.source.provider,
|
||||
prompt_tokens=prompt_tokens,
|
||||
usage_source="estimated" if estimated else "measured",
|
||||
usage_source=merged_source,
|
||||
latency_ms=sum(o.latency_ms for o in outcomes),
|
||||
call_id=first.call_id,
|
||||
source_name=first.source.name,
|
||||
@@ -395,8 +414,15 @@ class EmbeddingClient:
|
||||
)
|
||||
|
||||
def _total_cost(self, outcomes: list[_BatchOutcome]) -> float | None:
|
||||
"""全批成本;任一批用量不可得则整体记 NULL(解耦设计 §3.2 #11)。
|
||||
|
||||
逐批求和会把不可得的批当 0 计入,给出一个偏低却看似有效的金额——
|
||||
与"宁可算不出成本,也不算错成本"的不变式相悖。
|
||||
"""
|
||||
if self._pricing is None:
|
||||
return None
|
||||
if any(o.result.usage_source == "unavailable" for o in outcomes):
|
||||
return None
|
||||
costs = [self._pricing.cost(o.source.model, o.result.prompt_tokens, 0) for o in outcomes]
|
||||
known = [c for c in costs if c is not None]
|
||||
return sum(known) if known else None
|
||||
|
||||
@@ -20,6 +20,8 @@ from loguru import logger
|
||||
from polygateway.types import ChatRequest, LLMResponse
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from collections.abc import Mapping
|
||||
|
||||
from polygateway.ports import CacheBackend, CallNext, StructuredOutputStrategy
|
||||
|
||||
_KEY_PREFIX = "pgw:cache:"
|
||||
@@ -50,9 +52,19 @@ def _digest_part(part: Any) -> Any:
|
||||
|
||||
|
||||
def build_cache_key(
|
||||
model_fingerprint: str, messages: list[dict[str, Any]], namespace: str, salt: str | None
|
||||
model_fingerprint: str,
|
||||
messages: list[dict[str, Any]],
|
||||
namespace: str,
|
||||
salt: str | None,
|
||||
*,
|
||||
sampling: Mapping[str, Any] | None = None,
|
||||
) -> str:
|
||||
"""缓存 key 公式;salt 仅非 None 时参与(VT 旧键语义: 不传 salt 键形不变)。"""
|
||||
"""缓存 key 公式;salt 仅非 None 时参与(VT 旧键语义: 不传 salt 键形不变)。
|
||||
|
||||
`sampling` 仅**非空**时参与(与 salt 的"仅非 None"不同——空串是有意义的
|
||||
salt,而空采样参数与不传无语义差别)。它必须进 key: 否则同 messages 跑 5 个
|
||||
seed 会全部命中第一次的响应,标准差恒为 0 且不报错(issue #4 决策 C)。
|
||||
"""
|
||||
key_obj: dict[str, Any] = {
|
||||
"model": model_fingerprint,
|
||||
"messages": digest_messages(messages),
|
||||
@@ -60,6 +72,8 @@ def build_cache_key(
|
||||
}
|
||||
if salt is not None:
|
||||
key_obj["salt"] = salt
|
||||
if sampling:
|
||||
key_obj["sampling"] = dict(sampling)
|
||||
payload = json.dumps(key_obj, sort_keys=True, ensure_ascii=False)
|
||||
return _KEY_PREFIX + hashlib.sha256(payload.encode("utf-8")).hexdigest()
|
||||
|
||||
@@ -92,7 +106,15 @@ class CacheMW:
|
||||
|
||||
async def __call__(self, request: ChatRequest, call_next: CallNext) -> LLMResponse:
|
||||
namespace = request.cache_namespace or self._namespace
|
||||
key = build_cache_key(self._fingerprint, request.messages, namespace, request.cache_salt)
|
||||
# 读 sampling 而非 overlay: 语义明确,且不依赖"CacheMW 恰在 StructuredMW
|
||||
# 外侧"这一层序巧合——结构化注入不该改变缓存身份(设计决策 C)
|
||||
key = build_cache_key(
|
||||
self._fingerprint,
|
||||
request.messages,
|
||||
namespace,
|
||||
request.cache_salt,
|
||||
sampling=request.sampling,
|
||||
)
|
||||
cached = await self._safe_get(key)
|
||||
if cached is not None:
|
||||
hit = self._rehydrate(cached, request)
|
||||
|
||||
@@ -23,7 +23,7 @@ class QuotaGate:
|
||||
|
||||
async def try_acquire(self, source: SourceConfig) -> Permit | None:
|
||||
try:
|
||||
return await self._limiter.try_acquire(source.name, source.est_tokens)
|
||||
return await self._limiter.try_acquire(source.name, source.effective_est_tokens())
|
||||
except GovernanceBackendError:
|
||||
raise
|
||||
except Exception as exc:
|
||||
|
||||
@@ -335,6 +335,11 @@ class RetryMW:
|
||||
overlay=request.overlay,
|
||||
call_id=call_id,
|
||||
)
|
||||
if result.usage_source == "unavailable":
|
||||
# 用量不可得时按入场预扣量结算(delta==0),否则押金会被整笔退回,
|
||||
# 对"从不返回 usage 帧"的源等于 TPM 闸失效(设计 §3.2 #9)
|
||||
actual = source.effective_est_tokens()
|
||||
else:
|
||||
actual = result.prompt_tokens + result.completion_tokens
|
||||
await self._record_quietly(self._breaker.record_success(entry))
|
||||
await self._record_quietly(self._quota.mark_progress())
|
||||
@@ -367,7 +372,8 @@ class RetryMW:
|
||||
self._pacer.on_backpressure(source.name)
|
||||
await self._record_quietly(self._breaker.record_failure(entry, reason, dead))
|
||||
if not dead:
|
||||
actual = source.est_tokens # 保守: 失败请求可能已被网关计费(CHS 同款)
|
||||
# 保守: 失败请求可能已被网关计费(CHS 同款);与入场预扣同源取值
|
||||
actual = source.effective_est_tokens()
|
||||
await self._emit(request, source, call_id, started, error=exc)
|
||||
return _Failed(exc, immediate=dead)
|
||||
finally:
|
||||
@@ -430,6 +436,8 @@ class RetryMW:
|
||||
source_name=source.name,
|
||||
cost=None,
|
||||
usage_source=result.usage_source,
|
||||
cached_prompt_tokens=result.cached_prompt_tokens,
|
||||
model_reported=result.model_reported,
|
||||
)
|
||||
|
||||
async def _settle_and_release(self, permit: Permit, actual: int) -> None:
|
||||
|
||||
@@ -18,6 +18,7 @@ from loguru import logger
|
||||
|
||||
from polygateway.errors import GatewayUnavailableError, GovernanceBackendError
|
||||
from polygateway.middleware.cache import digest_messages
|
||||
from polygateway.types import canonical_sampling_json, merge_sampling
|
||||
|
||||
if TYPE_CHECKING:
|
||||
from collections.abc import Callable
|
||||
@@ -28,7 +29,7 @@ if TYPE_CHECKING:
|
||||
|
||||
|
||||
class TelemetryEmitter:
|
||||
"""从请求与结果组装 18 字段并写入 recorder;一切写失败降级 warning。"""
|
||||
"""从请求与结果组装 21 字段并写入 recorder;一切写失败降级 warning。"""
|
||||
|
||||
def __init__(self, recorder: TelemetryRecorder, *, pricing: PricingTable | None = None) -> None:
|
||||
self._recorder = recorder
|
||||
@@ -44,7 +45,7 @@ class TelemetryEmitter:
|
||||
response: LLMResponse | None,
|
||||
error: str | None,
|
||||
) -> None:
|
||||
"""逐次尝试记录(RetryMW 调用);失败尝试 usage 按 estimated 记 0。"""
|
||||
"""逐次尝试记录(RetryMW 调用);失败尝试无用量可言,记 0 并标 unavailable。"""
|
||||
await self._record(
|
||||
request=request,
|
||||
call_id=call_id,
|
||||
@@ -55,12 +56,16 @@ class TelemetryEmitter:
|
||||
thinking=response.thinking if response else "",
|
||||
prompt_tokens=response.prompt_tokens if response else 0,
|
||||
completion_tokens=response.completion_tokens if response else 0,
|
||||
usage_source=response.usage_source if response else "estimated",
|
||||
usage_source=response.usage_source if response else "unavailable",
|
||||
latency_ms=latency_ms,
|
||||
ttft_ms=response.ttft_ms if response else None,
|
||||
max_inter_token_ms=response.max_inter_token_ms if response else None,
|
||||
cache_hit=False,
|
||||
error=error,
|
||||
cached_prompt_tokens=response.cached_prompt_tokens if response else None,
|
||||
model_reported=response.model_reported if response else None,
|
||||
# 唯一有"生效源"的入口,故是唯一能并上 extra_body 的(设计决策 D)
|
||||
sampling=canonical_sampling_json(merge_sampling(source.extra_body, request.sampling)),
|
||||
)
|
||||
|
||||
async def emit_cache_hit(self, *, request: ChatRequest, response: LLMResponse) -> None:
|
||||
@@ -81,6 +86,13 @@ class TelemetryEmitter:
|
||||
max_inter_token_ms=None,
|
||||
cache_hit=True,
|
||||
error=None,
|
||||
# 决策 B1: 与 model/prompt_tokens 同一口径,原样回放历史值。
|
||||
# 统计供应商缓存命中率必须带 WHERE cache_hit = false,否则重复计数。
|
||||
cached_prompt_tokens=response.cached_prompt_tokens,
|
||||
model_reported=response.model_reported,
|
||||
# 由最外层 TelemetryMW 调用,手上没有 source。缓存命中行无损:
|
||||
# sampling 已进缓存 key,能命中即意味调用级参数与历史那次逐字相同
|
||||
sampling=canonical_sampling_json(request.sampling),
|
||||
)
|
||||
|
||||
async def emit_terminal_failure(
|
||||
@@ -97,12 +109,16 @@ class TelemetryEmitter:
|
||||
thinking="",
|
||||
prompt_tokens=0,
|
||||
completion_tokens=0,
|
||||
usage_source="estimated",
|
||||
usage_source="unavailable",
|
||||
latency_ms=latency_ms,
|
||||
ttft_ms=None,
|
||||
max_inter_token_ms=None,
|
||||
cache_hit=False,
|
||||
error=error,
|
||||
cached_prompt_tokens=None,
|
||||
model_reported=None,
|
||||
# 无具体源,与 model/provider/source_name 置空同一先例(设计决策 D)
|
||||
sampling=canonical_sampling_json(request.sampling),
|
||||
)
|
||||
|
||||
async def _record(
|
||||
@@ -123,14 +139,23 @@ class TelemetryEmitter:
|
||||
max_inter_token_ms: float | None,
|
||||
cache_hit: bool,
|
||||
error: str | None,
|
||||
cached_prompt_tokens: int | None,
|
||||
model_reported: str | None,
|
||||
sampling: str | None,
|
||||
) -> None:
|
||||
try:
|
||||
# 成本换算(M2 §6): 成功行按单价换算;缓存命中 0.0(未产生新调用);
|
||||
# 失败/终态行 None;未注入价格表 = 恒 None(M1 现状)
|
||||
if cache_hit:
|
||||
cost: float | None = 0.0
|
||||
elif usage_source == "unavailable":
|
||||
# 用量不可得: 宁可算不出成本,也不算错成本(解耦设计 §3.1 不变式)。
|
||||
# 必须排在 cache_hit 之后——缓存命中未产生新调用,0.0 是事实而非未知
|
||||
cost = None
|
||||
elif error is None and model and self._pricing is not None:
|
||||
cost = self._pricing.cost(model, prompt_tokens, completion_tokens)
|
||||
cost = self._pricing.cost(
|
||||
model, prompt_tokens, completion_tokens, cached_prompt_tokens
|
||||
)
|
||||
else:
|
||||
cost = None
|
||||
# messages 落库前多模态摘要,与缓存 key 共用同一函数(VT R12)
|
||||
@@ -154,6 +179,9 @@ class TelemetryEmitter:
|
||||
cache_hit=cache_hit,
|
||||
error=error,
|
||||
cost=cost,
|
||||
cached_prompt_tokens=cached_prompt_tokens,
|
||||
model_reported=model_reported,
|
||||
sampling=sampling,
|
||||
)
|
||||
except asyncio.CancelledError:
|
||||
raise
|
||||
|
||||
@@ -44,6 +44,7 @@ from polygateway.types import (
|
||||
OcrLayoutResult,
|
||||
OcrTextResult,
|
||||
Usage,
|
||||
strip_unsupported_extra_body,
|
||||
)
|
||||
|
||||
if TYPE_CHECKING:
|
||||
@@ -113,7 +114,9 @@ class OcrClient:
|
||||
if quota_full not in ("wait", "fail_fast"):
|
||||
raise ValueError(f"quota_full 必须是 wait|fail_fast: {quota_full!r}")
|
||||
self._scope = scope
|
||||
self._sources = list(sources)
|
||||
# MonkeyOCR 只发 multipart 表单,带 extra_body 的源必须先剥离,否则
|
||||
# 遥测会记录一个从未发出的采样参数(issue #4 决策 G)
|
||||
self._sources = strip_unsupported_extra_body(list(sources), path="OCR")
|
||||
self._selector = selector
|
||||
self._feed_health = isinstance(selector, OutcomeAwareSelector)
|
||||
self._quota = QuotaGate(limiter)
|
||||
|
||||
@@ -245,7 +245,11 @@ class StructuredOutputStrategy(Protocol):
|
||||
|
||||
@runtime_checkable
|
||||
class TelemetryRecorder(Protocol):
|
||||
"""遥测后端;18 字段冻结(M1 设计 §4.4),唯一调用点是 TelemetryEmitter。"""
|
||||
"""遥测后端;20 字段冻结(M1 设计 §4.4 + issue #3),唯一调用点是 TelemetryEmitter。
|
||||
|
||||
新增参数不设默认值: 库外无第三方实现者(三项目迁移时删除了各自的同名
|
||||
Protocol),完整签名的成本为零,而少写一列会被 emitter 的降级吞成 warning。
|
||||
"""
|
||||
|
||||
async def record_llm_call(
|
||||
self,
|
||||
@@ -268,4 +272,7 @@ class TelemetryRecorder(Protocol):
|
||||
cache_hit: bool,
|
||||
error: str | None,
|
||||
cost: float | None,
|
||||
cached_prompt_tokens: int | None,
|
||||
model_reported: str | None,
|
||||
sampling: str | None,
|
||||
) -> None: ...
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
**零内置单价**: 实验室走中转网关,计费非官方牌价;库内硬编码单价表
|
||||
必然过时并掩盖真实成本(P5 严禁默认值掩盖错误)。价格一律由使用方
|
||||
提供——JSON 文件(`PGW_PRICING_PATH`)或 dict 注入;币种由使用方全表
|
||||
统一口径,库不设币种字段(18 字段冻结)。查不到的 model → cost=None
|
||||
统一口径,库不设币种字段(20 字段冻结)。查不到的 model → cost=None
|
||||
且每 model 仅首次 warning(防日志风暴),不阻塞调用。
|
||||
"""
|
||||
|
||||
@@ -22,22 +22,33 @@ if TYPE_CHECKING:
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class ModelPrice:
|
||||
"""每百万 token 的输入/输出单价(币种由使用方口径统一)。"""
|
||||
"""每百万 token 的输入/输出单价(币种由使用方口径统一)。
|
||||
|
||||
`cached_input_per_1m` 是可选的**缓存读取单价**(issue #3): 供应商 prompt
|
||||
cache 命中的那部分输入按更低单价计费。不填即不启用——库绝不按经验折扣率
|
||||
猜一个数(P5 严禁默认值掩盖),未填时全额按 `input_per_1m` 计。
|
||||
"""
|
||||
|
||||
input_per_1m: float
|
||||
output_per_1m: float
|
||||
cached_input_per_1m: float | None = None
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
if self.input_per_1m < 0 or self.output_per_1m < 0:
|
||||
raise ValueError("单价不能为负")
|
||||
if self.cached_input_per_1m is not None and self.cached_input_per_1m < 0:
|
||||
raise ValueError("缓存读取单价不能为负")
|
||||
|
||||
|
||||
class PricingTable:
|
||||
"""model → 单价 的只读表;cost() 是全库唯一换算点(经 TelemetryEmitter)。"""
|
||||
"""model → 单价 的只读表;cost() 有两个调用点: `TelemetryEmitter`(chat 主路径)
|
||||
与 `embedding.py` 的批量换算。"""
|
||||
|
||||
def __init__(self, prices: Mapping[str, ModelPrice]) -> None:
|
||||
self._prices = dict(prices)
|
||||
self._warned: set[str] = set()
|
||||
# 独立集合: 与"未知 model"的告警去重键分开,避免 model 名恰好撞上时互相抑制
|
||||
self._warned_clamp: set[str] = set()
|
||||
|
||||
@classmethod
|
||||
def from_file(cls, path: Path | str) -> PricingTable:
|
||||
@@ -53,21 +64,61 @@ class PricingTable:
|
||||
for model, entry in data.items():
|
||||
if not isinstance(entry, dict) or not {"input_per_1m", "output_per_1m"} <= set(entry):
|
||||
raise ValueError(f"价格表 {p} 条目 {model!r} 须含 input_per_1m 与 output_per_1m")
|
||||
cached_raw = entry.get("cached_input_per_1m")
|
||||
try:
|
||||
cached = None if cached_raw is None else float(cached_raw)
|
||||
except (TypeError, ValueError) as exc:
|
||||
raise ValueError(
|
||||
f"价格表 {p} 条目 {model!r} 的 cached_input_per_1m 必须是数字: {cached_raw!r}"
|
||||
) from exc
|
||||
if cached is not None and cached < 0:
|
||||
raise ValueError(f"价格表 {p} 条目 {model!r} 的 cached_input_per_1m 不能为负")
|
||||
prices[model] = ModelPrice(
|
||||
input_per_1m=float(entry["input_per_1m"]),
|
||||
output_per_1m=float(entry["output_per_1m"]),
|
||||
cached_input_per_1m=cached,
|
||||
)
|
||||
return cls(prices)
|
||||
|
||||
def cost(self, model: str, prompt_tokens: int, completion_tokens: int) -> float | None:
|
||||
"""换算一次调用成本;未知 model 记 None 并仅首次 warning。"""
|
||||
def cost(
|
||||
self,
|
||||
model: str,
|
||||
prompt_tokens: int,
|
||||
completion_tokens: int,
|
||||
cached_prompt_tokens: int | None = None,
|
||||
) -> float | None:
|
||||
"""换算一次调用成本;未知 model 记 None 并仅首次 warning。
|
||||
|
||||
`cached_prompt_tokens` 是供应商 prompt cache 命中的输入 token 数
|
||||
(issue #3);仅当该 model 配了 `cached_input_per_1m` 时才分段计价,
|
||||
否则全额按输入价——不猜折扣率。参数带默认值: embedding 侧的三参调用
|
||||
形态不受影响。
|
||||
"""
|
||||
price = self._prices.get(model)
|
||||
if price is None:
|
||||
if model not in self._warned:
|
||||
self._warned.add(model)
|
||||
logger.warning("pricing 表无 model {!r} 的单价,cost 记 None", model)
|
||||
return None
|
||||
return (
|
||||
prompt_tokens / 1_000_000 * price.input_per_1m
|
||||
+ completion_tokens / 1_000_000 * price.output_per_1m
|
||||
billed_input = prompt_tokens / 1_000_000 * price.input_per_1m
|
||||
# 负数按"无命中"处理: cost() 是公共方法,不能假定调用方已过 transport 的校验
|
||||
if price.cached_input_per_1m is not None and (cached_prompt_tokens or 0) > 0:
|
||||
cached = self._clamp_cached(model, prompt_tokens, cached_prompt_tokens)
|
||||
billed_input = (prompt_tokens - cached) / 1_000_000 * price.input_per_1m + (
|
||||
cached / 1_000_000 * price.cached_input_per_1m
|
||||
)
|
||||
return billed_input + completion_tokens / 1_000_000 * price.output_per_1m
|
||||
|
||||
def _clamp_cached(self, model: str, prompt_tokens: int, cached: int) -> int:
|
||||
"""命中数按输入总数夹取: 网关口径异常不得算出负成本(每 model 只警告一次)。"""
|
||||
if cached <= prompt_tokens:
|
||||
return cached
|
||||
if model not in self._warned_clamp:
|
||||
self._warned_clamp.add(model)
|
||||
logger.warning(
|
||||
"model {!r} 上报的缓存命中 {} 超过输入总数 {},按总数夹取计价",
|
||||
model,
|
||||
cached,
|
||||
prompt_tokens,
|
||||
)
|
||||
return prompt_tokens
|
||||
|
||||
@@ -19,6 +19,10 @@ class ProviderProfile:
|
||||
True/False 时并入请求体的参数片段(None 时二者都不注入,用模型默认);
|
||||
strip_think_tags 声明响应 content 需剥离 ``<think>`` 标签(qwen 系);
|
||||
supports_native_schema 供 D14 阶梯选择原生 response_format 策略。
|
||||
|
||||
注: 某个 provider 的两档若皆为空字典(如 openai/minimax),说明该 provider
|
||||
无已知的推理开关参数——此时 `enable_thinking` 对它**不产生任何效果**,
|
||||
而非静默生效。需要下发自定义参数时用 `SourceConfig.extra_body`。
|
||||
"""
|
||||
|
||||
name: str
|
||||
@@ -43,13 +47,16 @@ DEFAULT_PROFILES: Mapping[str, ProviderProfile] = MappingProxyType(
|
||||
thinking_off={"thinking": {"type": "disabled"}},
|
||||
strip_think_tags=False,
|
||||
),
|
||||
# 两档皆空 ⇒ `enable_thinking` 对本 provider **不产生任何效果**(调用方
|
||||
# 以为关掉了实际没关)。真需要控制推理时经 `SourceConfig.extra_body` 下发
|
||||
"openai": ProviderProfile(
|
||||
name="openai",
|
||||
thinking_on={},
|
||||
thinking_off={},
|
||||
strip_think_tags=False,
|
||||
),
|
||||
# OpenAI 兼容基线,无已知注入差异;reasoning_content 由 transport 通用处理
|
||||
# OpenAI 兼容基线,无已知注入差异;reasoning_content 由 transport 通用处理。
|
||||
# 同上: 两档皆空 ⇒ `enable_thinking` 对 MiniMax 源不产生任何效果
|
||||
"minimax": ProviderProfile(
|
||||
name="minimax",
|
||||
thinking_on={},
|
||||
|
||||
@@ -6,7 +6,7 @@
|
||||
① 结构性失败(建池/建表)→ warning 一次后永久降级(池置 None 短路);
|
||||
② 运行时单条写失败 → 逐条 warning 丢弃,不降级不重试(连接抖动由
|
||||
asyncpg 池自恢复;避免浸泡开头一次抖动导致后续全程失遥测)。
|
||||
构造不连库(lazy),18 列 schema 与 SQLite 版同名同序。
|
||||
构造不连库(lazy),20 列 schema 与 SQLite 版同名同序。
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
@@ -39,10 +39,26 @@ CREATE TABLE IF NOT EXISTS llm_calls (
|
||||
cache_hit BOOLEAN NOT NULL DEFAULT FALSE,
|
||||
error TEXT,
|
||||
cost DOUBLE PRECISION,
|
||||
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
|
||||
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
|
||||
cached_prompt_tokens INTEGER,
|
||||
model_reported TEXT,
|
||||
sampling TEXT
|
||||
);
|
||||
"""
|
||||
|
||||
# 新列排在 created_at 之后: 与旧表 ALTER 追加的位置一致(见 sqlite.py 同款注释)
|
||||
_BACKFILL = (
|
||||
("cached_prompt_tokens", "ALTER TABLE llm_calls ADD COLUMN cached_prompt_tokens INTEGER"),
|
||||
("model_reported", "ALTER TABLE llm_calls ADD COLUMN model_reported TEXT"),
|
||||
("sampling", "ALTER TABLE llm_calls ADD COLUMN sampling TEXT"),
|
||||
)
|
||||
|
||||
# 探测现有列;尊重 search_path(to_regclass 按当前 search_path 解析)
|
||||
_EXISTING_COLUMNS = (
|
||||
"SELECT attname FROM pg_attribute "
|
||||
"WHERE attrelid = to_regclass('llm_calls') AND attnum > 0 AND NOT attisdropped"
|
||||
)
|
||||
|
||||
_COLUMNS = (
|
||||
"call_id",
|
||||
"parent_call_id",
|
||||
@@ -62,6 +78,9 @@ _COLUMNS = (
|
||||
"cache_hit",
|
||||
"error",
|
||||
"cost",
|
||||
"cached_prompt_tokens",
|
||||
"model_reported",
|
||||
"sampling",
|
||||
)
|
||||
|
||||
_INSERT = (
|
||||
@@ -104,6 +123,7 @@ class PostgresRecorder:
|
||||
self._pool = await asyncpg.create_pool(self._dsn, timeout=10)
|
||||
async with self._pool.acquire() as conn:
|
||||
await conn.execute(_DDL)
|
||||
await self._backfill_columns(conn)
|
||||
self._schema_ready = True
|
||||
return self._pool
|
||||
except asyncio.CancelledError:
|
||||
@@ -113,6 +133,29 @@ class PostgresRecorder:
|
||||
logger.warning("Postgres 遥测初始化失败,后续记录降级为 no-op: {}", exc)
|
||||
return None
|
||||
|
||||
async def _backfill_columns(self, conn: object) -> None:
|
||||
"""给已存在的旧表补新列(issue #3);**先探测再 ALTER,失败绝不置 `_failed`**。
|
||||
|
||||
两条纪律各有实测理由:
|
||||
① 不置 `_failed`: 应用账号只有 INSERT 权限时,`ALTER TABLE` 的 ownership
|
||||
检查早于 `IF NOT EXISTS` 的存在性判断——列明明齐全也会失败。置位会让
|
||||
整个 recorder 永久 no-op,与「补列失败只降级为逐行丢弃」的承诺相悖
|
||||
(SQLite 侧同款守卫,两侧必须对称)。
|
||||
② 先探测: `ADD COLUMN IF NOT EXISTS` 即便列已存在,也会**先取 ACCESS
|
||||
EXCLUSIVE 锁**再判存在性(实测会被一个开着的读事务阻塞)。遥测是内联
|
||||
await,让每个进程的首次写入都去抢共享审计表的排他锁,等于用记录基础设施
|
||||
拖垮业务调用。探测走 ACCESS SHARE,稳态下一条 ALTER 都不会发。
|
||||
"""
|
||||
try:
|
||||
existing = {row["attname"] for row in await conn.fetch(_EXISTING_COLUMNS)} # type: ignore[attr-defined]
|
||||
for column, statement in _BACKFILL:
|
||||
if column not in existing:
|
||||
await conn.execute(statement) # type: ignore[attr-defined]
|
||||
except asyncio.CancelledError:
|
||||
raise
|
||||
except Exception as exc:
|
||||
logger.warning("Postgres 遥测补列失败(写入将逐行降级): {}", exc)
|
||||
|
||||
async def record_llm_call(self, **fields: object) -> None:
|
||||
"""写一行遥测;单条失败逐条 warning 丢弃(两级降级之二),绝不冒泡。"""
|
||||
pool = await self._ensure_ready()
|
||||
|
||||
@@ -34,10 +34,21 @@ CREATE TABLE IF NOT EXISTS llm_calls (
|
||||
cache_hit INTEGER NOT NULL DEFAULT 0,
|
||||
error TEXT,
|
||||
cost REAL,
|
||||
created_at TEXT NOT NULL DEFAULT (datetime('now'))
|
||||
created_at TEXT NOT NULL DEFAULT (datetime('now')),
|
||||
cached_prompt_tokens INTEGER,
|
||||
model_reported TEXT,
|
||||
sampling TEXT
|
||||
);
|
||||
"""
|
||||
|
||||
# 新列必须排在 created_at 之后: 旧表只能经 ALTER 追加到末尾,新建库若把它们
|
||||
# 插在前面,两条路径的物理列序会分叉(列序断言测试无合规修法)。
|
||||
_BACKFILL_COLUMNS = (
|
||||
("cached_prompt_tokens", "INTEGER"),
|
||||
("model_reported", "TEXT"),
|
||||
("sampling", "TEXT"),
|
||||
)
|
||||
|
||||
_COLUMNS = (
|
||||
"call_id",
|
||||
"parent_call_id",
|
||||
@@ -57,6 +68,9 @@ _COLUMNS = (
|
||||
"cache_hit",
|
||||
"error",
|
||||
"cost",
|
||||
"cached_prompt_tokens",
|
||||
"model_reported",
|
||||
"sampling",
|
||||
)
|
||||
|
||||
_INSERT = (
|
||||
@@ -82,9 +96,37 @@ class SQLiteRecorder:
|
||||
self._conn = conn
|
||||
except (OSError, sqlite3.Error) as exc:
|
||||
logger.warning("SQLite 遥测初始化失败,后续记录降级为 no-op: {}", exc)
|
||||
self._backfill_columns()
|
||||
|
||||
def _backfill_columns(self) -> None:
|
||||
"""给已存在的旧表补新列(issue #3);独立 try,失败只降级为逐行丢弃。
|
||||
|
||||
必须放在 `self._conn` 赋值**之后**并先判空: 初始化失败时连接为 None,
|
||||
无守卫的补列会抛 AttributeError 逃出 `__init__`,把"静默降级"变成崩溃。
|
||||
补列失败也绝不清空 `self._conn`——那会让整个 recorder 永久 no-op,
|
||||
比逐行丢弃严重得多。
|
||||
"""
|
||||
if self._conn is None:
|
||||
return
|
||||
try:
|
||||
existing = {row[1] for row in self._conn.execute("PRAGMA table_info(llm_calls)")}
|
||||
except sqlite3.Error as exc:
|
||||
logger.warning("SQLite 遥测列探测失败(写入将逐行降级): {}", exc)
|
||||
return
|
||||
for column, decl in _BACKFILL_COLUMNS:
|
||||
if column in existing:
|
||||
continue
|
||||
# 逐列独立 try: 一列撞上 duplicate 不得让后面的列漏补
|
||||
try:
|
||||
self._conn.execute(f"ALTER TABLE llm_calls ADD COLUMN {column} {decl}")
|
||||
self._conn.commit()
|
||||
except sqlite3.Error as exc:
|
||||
# duplicate column: 多进程共库时后到者必然撞上,属预期竞态,视为成功
|
||||
if "duplicate column" not in str(exc).lower():
|
||||
logger.warning("SQLite 遥测补列失败(写入将逐行降级): {}", exc)
|
||||
|
||||
async def record_llm_call(self, **fields: object) -> None:
|
||||
"""写一行遥测;字段集合即 18 字段冻结签名(ports.TelemetryRecorder)。"""
|
||||
"""写一行遥测;字段集合即 21 字段冻结签名(ports.TelemetryRecorder)。"""
|
||||
if self._conn is None:
|
||||
return
|
||||
row = tuple(fields[col] for col in _COLUMNS)
|
||||
|
||||
@@ -45,6 +45,12 @@ def _sse_delta(chunk: dict[str, Any], usage_sink: dict[str, Any]) -> tuple[bool,
|
||||
"""从 chunk 提取增量: (True, content) 或 (False, reasoning);usage 帧旁路进 sink。"""
|
||||
if chunk.get("usage"):
|
||||
usage_sink["usage"] = chunk["usage"]
|
||||
if "model" not in usage_sink:
|
||||
# 首个**有效**值即固定: 末帧的异常值不得覆盖它;但首帧报空串也不能锁死
|
||||
# sink——否则后续真实版本会丢(issue #3)
|
||||
reported = _coerce_model_reported(chunk.get("model"))
|
||||
if reported is not None:
|
||||
usage_sink["model"] = reported
|
||||
choices = chunk.get("choices") or []
|
||||
if not choices:
|
||||
return None
|
||||
@@ -138,12 +144,60 @@ def _strip_think(content: str) -> tuple[str, str]:
|
||||
return _THINK_PATTERN.sub("", content).strip(), match.group(1).strip()
|
||||
|
||||
|
||||
def _resolve_usage(usage: dict[str, Any], source: SourceConfig) -> tuple[int, int, str]:
|
||||
"""usage 帧读取;缺失/非法按 est_tokens 保守兜底并标 estimated(CHS invokers.py:241)。"""
|
||||
def _resolve_usage(usage: dict[str, Any]) -> tuple[int, int, str]:
|
||||
"""usage 帧读取;缺失/非法记 0/0 并标 unavailable(est_tokens 解耦设计 §3.2 #3)。
|
||||
|
||||
不再拿 `est_tokens` 兜底: 它按 CHS 定义是"最坏情形上界",拿上界当实测值
|
||||
只会系统性高估账单;宁可把用量记成显式的"不可得"(cost 随之为 NULL),
|
||||
让缺口可被统计,也不编一个看似有效的数字。用量口径自此不依赖源配置,
|
||||
故不再收 `SourceConfig`。
|
||||
"""
|
||||
prompt, completion = usage.get("prompt_tokens"), usage.get("completion_tokens")
|
||||
if isinstance(prompt, int) and isinstance(completion, int) and prompt + completion > 0:
|
||||
return prompt, completion, "measured"
|
||||
return 0, source.est_tokens, "estimated"
|
||||
return 0, 0, "unavailable"
|
||||
|
||||
|
||||
def _coerce_cached_tokens(usage: Any) -> int | None:
|
||||
"""取 usage.prompt_tokens_details.cached_tokens(issue #3);形态异常一律 None。
|
||||
|
||||
`0` 与 `None` 必须可区分: 前者是"该源上报了一次真实零命中",后者是"该源
|
||||
不报这个数",下游对两者的处置不同(后者不可做缓存成本校正)。故只把
|
||||
**负数与非整数**归 None,`0` 如实保留。`bool` 显式排除——isinstance(True, int)
|
||||
在 Python 里为真,放行会把 `True` 记成 1 个命中 token。
|
||||
"""
|
||||
if not isinstance(usage, dict):
|
||||
return None
|
||||
details = usage.get("prompt_tokens_details")
|
||||
if not isinstance(details, dict):
|
||||
return None
|
||||
cached = details.get("cached_tokens")
|
||||
if isinstance(cached, bool) or not isinstance(cached, int) or cached < 0:
|
||||
return None
|
||||
return cached
|
||||
|
||||
|
||||
def _coerce_model_reported(value: Any) -> str | None:
|
||||
"""取响应体的 model 字段(issue #3);非 str 或空白串一律 None,收口时去空白。
|
||||
|
||||
去空白不是洁癖: 下游拿这个串做实验快照的 key,`" m "` 与 `"m"` 会造成假分叉。
|
||||
"""
|
||||
if not isinstance(value, str) or not value.strip():
|
||||
return None
|
||||
return value.strip()
|
||||
|
||||
|
||||
def _resolve_stream_usage(sink: dict[str, Any], salvaged: bool) -> tuple[int, int, str]:
|
||||
"""流式用量口径: 打捞路径把 measured 降级为 estimated,unavailable 原样保留。
|
||||
|
||||
前置条件不可省(解耦设计 §3.2 #4): usage 帧本就缺失时 `0/0` 会被洗成
|
||||
`estimated`,进而按 token 换算出一个假的 `0.0` 成本。
|
||||
"""
|
||||
prompt, completion, usage_source = _resolve_usage(sink.get("usage") or {})
|
||||
if salvaged and usage_source == "measured":
|
||||
# 收到 usage 帧但流被截断: 数字真实、可信度降级(M1 设计 §6)
|
||||
usage_source = "estimated"
|
||||
return prompt, completion, usage_source
|
||||
|
||||
|
||||
def _extract_vectors(
|
||||
@@ -168,12 +222,12 @@ def _extract_vectors(
|
||||
return vectors
|
||||
|
||||
|
||||
def _resolve_embedding_usage(data: dict[str, Any], source: SourceConfig) -> tuple[int, str]:
|
||||
"""usage 读取;缺失/非法按 est_tokens 保守兜底并标 estimated(与 chat 同口径)。"""
|
||||
def _resolve_embedding_usage(data: dict[str, Any]) -> tuple[int, str]:
|
||||
"""usage 读取;缺失/非法记 0 并标 unavailable(与 chat 同口径,设计 §3.2 #3)。"""
|
||||
prompt = (data.get("usage") or {}).get("prompt_tokens")
|
||||
if isinstance(prompt, int) and prompt > 0:
|
||||
return prompt, "measured"
|
||||
return source.est_tokens, "estimated"
|
||||
return 0, "unavailable"
|
||||
|
||||
|
||||
def _parse_embedding_payload(
|
||||
@@ -186,7 +240,7 @@ def _parse_embedding_payload(
|
||||
except json.JSONDecodeError as exc:
|
||||
raise ResultInvalidError(f"{source.name} embedding 响应非 JSON: {exc}", **ctx) from exc
|
||||
vectors = _extract_vectors(data, source, expected_count, ctx)
|
||||
prompt_tokens, usage_source = _resolve_embedding_usage(data, source)
|
||||
prompt_tokens, usage_source = _resolve_embedding_usage(data)
|
||||
return EmbeddingTransportResult(
|
||||
vectors=vectors,
|
||||
dim=len(vectors[0]),
|
||||
@@ -240,6 +294,9 @@ class OpenAICompatTransport:
|
||||
payload.update(profile.thinking_on)
|
||||
elif source.enable_thinking is False:
|
||||
payload.update(profile.thinking_off)
|
||||
# 顺序即优先级(issue #4 设计决策 A): 配置级 extra_body 在前,调用级
|
||||
# overlay(含结构化注入)在后覆盖之。两行不可调换
|
||||
payload.update(source.extra_body)
|
||||
payload.update(overlay)
|
||||
return payload
|
||||
|
||||
@@ -331,9 +388,7 @@ class OpenAICompatTransport:
|
||||
salvaged = self._check_done(sink, content_parts, thinking_parts, source)
|
||||
content, thinking = self._finalize_text(content_parts, thinking_parts, profile)
|
||||
self._reject_empty_completion(content, source)
|
||||
prompt, completion, usage_source = _resolve_usage(sink.get("usage") or {}, source)
|
||||
if salvaged:
|
||||
usage_source = "estimated" # 打捞路径强制 estimated(设计 §6)
|
||||
prompt, completion, usage_source = _resolve_stream_usage(sink, salvaged)
|
||||
return TransportResult(
|
||||
content=content,
|
||||
thinking=thinking,
|
||||
@@ -343,6 +398,8 @@ class OpenAICompatTransport:
|
||||
ttft_ms=ttft_ms,
|
||||
max_inter_token_ms=(max_gap if ttft_ms is not None else None),
|
||||
raw={"usage": sink.get("usage")},
|
||||
cached_prompt_tokens=_coerce_cached_tokens(sink.get("usage")),
|
||||
model_reported=_coerce_model_reported(sink.get("model")),
|
||||
)
|
||||
|
||||
def _check_done(
|
||||
@@ -415,7 +472,7 @@ class OpenAICompatTransport:
|
||||
[message.get("content") or ""], [message.get("reasoning_content") or ""], profile
|
||||
)
|
||||
self._reject_empty_completion(content, source)
|
||||
prompt, completion, usage_source = _resolve_usage(body.get("usage") or {}, source)
|
||||
prompt, completion, usage_source = _resolve_usage(body.get("usage") or {})
|
||||
return TransportResult(
|
||||
content=content,
|
||||
thinking=thinking,
|
||||
@@ -425,6 +482,8 @@ class OpenAICompatTransport:
|
||||
ttft_ms=None,
|
||||
max_inter_token_ms=None,
|
||||
raw={"usage": body.get("usage")},
|
||||
cached_prompt_tokens=_coerce_cached_tokens(body.get("usage")),
|
||||
model_reported=_coerce_model_reported(body.get("model")),
|
||||
)
|
||||
|
||||
async def aclose(self) -> None:
|
||||
|
||||
+129
-3
@@ -4,11 +4,72 @@
|
||||
fake,字段顺序即公共承诺;新增字段只增不删且必带默认值。
|
||||
"""
|
||||
|
||||
import dataclasses
|
||||
import json
|
||||
from collections.abc import Mapping
|
||||
from dataclasses import dataclass, field
|
||||
from types import MappingProxyType
|
||||
from typing import Any
|
||||
|
||||
from loguru import logger
|
||||
|
||||
_MISSING_DONE_DOMAIN = frozenset({"retry", "salvage"})
|
||||
|
||||
_PROTECTED_OVERLAY_KEYS: Mapping[str, str] = MappingProxyType(
|
||||
{
|
||||
"model": "会让遥测记录的 model 与实际请求分叉,成本按错单价换算",
|
||||
"messages": "会同时破坏缓存 key 与遥测的 messages 口径",
|
||||
"stream": "会绕过流式活性看门狗(TTFT/inter-token 超时全部失效)",
|
||||
"stream_options": "会丢 usage 帧,导致成本遥测归零、TPM 闸按预扣量结算失准",
|
||||
}
|
||||
)
|
||||
"""禁止出现在采样参数覆盖层里的键: 它们由治理层拥有,被覆盖即击穿治理。"""
|
||||
|
||||
USAGE_SOURCES = frozenset({"measured", "estimated", "unavailable"})
|
||||
"""usage_source 值域;仅约束库内生产侧取值,不在 frozen dataclass 上做运行时校验。"""
|
||||
|
||||
_EST_TOKENS_QUOTA_DIVISOR = 60
|
||||
"""未显式配置时的预扣量除数: 假定一次调用约占一秒钟的 TPM 配额份额。"""
|
||||
|
||||
|
||||
def validate_request_overlay(overlay: Mapping[str, Any], *, origin: str) -> dict[str, Any]:
|
||||
"""校验采样参数覆盖层并返回浅拷贝;origin 用于把错误指回配置/调用点。
|
||||
|
||||
两类校验缺一不可(issue #4 设计决策 B):保护键会击穿治理;不可 JSON
|
||||
序列化的值会在 `CacheMW` 的降级 try **之外**抛裸 `TypeError`——那条路径
|
||||
不属错误四分类、`TelemetryMW` 也不捕,结果是一行遥测都没有就崩了。
|
||||
两者都在进洋葱之前收口,故抛裸 `ValueError`(调用方编程错误,不可重试)。
|
||||
"""
|
||||
# Phase 1: 键形态——必须先于序列化试探,否则非 str 键会因 sort_keys 的
|
||||
# 比较失败被误报成"值不可序列化",把人指向错误的方向
|
||||
for key in overlay:
|
||||
if not isinstance(key, str):
|
||||
raise ValueError(f"{origin} 的键必须是 str: {key!r}(canonical JSON 要求)")
|
||||
# Phase 2: 保护键
|
||||
for key, reason in _PROTECTED_OVERLAY_KEYS.items():
|
||||
if key in overlay:
|
||||
raise ValueError(f"{origin} 不得覆盖 {key!r}: {reason}")
|
||||
# Phase 3: 值可序列化(缓存 key 与遥测列都要 json.dumps)
|
||||
try:
|
||||
json.dumps(dict(overlay), sort_keys=True, ensure_ascii=False)
|
||||
except (TypeError, ValueError) as exc:
|
||||
raise ValueError(
|
||||
f"{origin} 的值必须可 JSON 序列化(如 numpy 标量请先转 float/int): {exc}"
|
||||
) from exc
|
||||
return dict(overlay)
|
||||
|
||||
|
||||
def merge_sampling(extra_body: Mapping[str, Any], sampling: Mapping[str, Any]) -> dict[str, Any]:
|
||||
"""合并配置级与调用级采样参数;调用级优先(issue #4 设计决策 A)。"""
|
||||
return {**extra_body, **sampling}
|
||||
|
||||
|
||||
def canonical_sampling_json(merged: Mapping[str, Any]) -> str | None:
|
||||
"""缓存 key 与遥测 sampling 列共用的序列化口径;空 mapping → None。"""
|
||||
if not merged:
|
||||
return None
|
||||
return json.dumps(dict(merged), sort_keys=True, ensure_ascii=False)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LLMResponse:
|
||||
@@ -24,12 +85,20 @@ class LLMResponse:
|
||||
ttft_ms: float | None
|
||||
max_inter_token_ms: float | None
|
||||
cache_hit: bool
|
||||
"""**PolyGateway 自身响应缓存**命中(未产生网关调用);与供应商侧 prompt
|
||||
cache 无关,后者见 `cached_prompt_tokens`。"""
|
||||
call_id: str
|
||||
# —— 库新增(只增不删,必带默认值;迁移兼容硬约束)——
|
||||
source_name: str = ""
|
||||
cost: float | None = None
|
||||
usage_source: str = "measured"
|
||||
structured_data: Any | None = None
|
||||
cached_prompt_tokens: int | None = None
|
||||
"""供应商 prompt cache 命中的输入 token 数(issue #3);None = 该源未上报,
|
||||
与"上报了但是 0"(真实零命中)区分——两者对下游的处置不同。"""
|
||||
model_reported: str | None = None
|
||||
"""API 响应体里的 model 字段;None = 未上报。与 `model`(配置别名)可能
|
||||
分叉——供应商把别名指向新权重时,实验复现必须认这个串。"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
@@ -44,6 +113,12 @@ class ChatRequest:
|
||||
structured: Any | None = None
|
||||
stream: bool = True
|
||||
overlay: dict[str, Any] = field(default_factory=dict)
|
||||
sampling: Mapping[str, Any] = field(default_factory=dict)
|
||||
"""调用方采样意图的快照,库内中间件**永不修改**(issue #4 设计决策 A)。
|
||||
|
||||
与 `overlay` 分开是因为后者会被结构化中间件注入 `response_format`,在洋葱
|
||||
不同深度取值不同;缓存 key 与三个遥测入口需要一个跨层恒定的读取点,否则
|
||||
同一列在不同行口径分叉。"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
@@ -76,6 +151,9 @@ class TransportResult:
|
||||
ttft_ms: float | None
|
||||
max_inter_token_ms: float | None
|
||||
raw: dict[str, Any]
|
||||
# —— 可观测字段(issue #3;带默认值,非 OpenAI 兼容的 transport 可不填)——
|
||||
cached_prompt_tokens: int | None = None
|
||||
model_reported: str | None = None
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
@@ -101,11 +179,26 @@ class SourceConfig:
|
||||
enable_thinking: bool | None = None
|
||||
missing_done: str = "retry"
|
||||
trust_env: bool = True
|
||||
extra_body: Mapping[str, Any] = field(default_factory=dict)
|
||||
"""本源恒定的采样参数(如 `temperature=0`),并入请求体(issue #4)。
|
||||
|
||||
优先级低于调用级 overlay。注: 本字段令 `SourceConfig` 不再 hashable
|
||||
(加任何 mapping 字段的固有代价,裸 dict 亦然),库内无以源作 key 的写法;
|
||||
要可变副本用 `dict(source.extra_body)`,要改字段用 `dataclasses.replace`。"""
|
||||
|
||||
def __post_init__(self) -> None:
|
||||
self._validate_identity()
|
||||
self._validate_gates()
|
||||
self._validate_watchdog()
|
||||
self._freeze_extra_body()
|
||||
|
||||
def effective_est_tokens(self) -> int:
|
||||
"""TPM 入场预扣量: 显式配置优先,否则按 tpm 派生(设计 §2.2)。"""
|
||||
if self.est_tokens > 0:
|
||||
return self.est_tokens
|
||||
if self.tpm > 0:
|
||||
return max(1, self.tpm // _EST_TOKENS_QUOTA_DIVISOR)
|
||||
return 0
|
||||
|
||||
def _validate_identity(self) -> None:
|
||||
for attr in ("name", "provider", "base_url", "api_key", "model"):
|
||||
@@ -122,8 +215,8 @@ class SourceConfig:
|
||||
for attr in ("max_concurrency", "rpm", "tpm", "est_tokens"):
|
||||
if getattr(self, attr) < 0:
|
||||
raise ValueError(f"SourceConfig.{attr} 不能为负(0 表示不启用)")
|
||||
if self.tpm > 0 and self.est_tokens <= 0:
|
||||
raise ValueError("启用 TPM 闸时 est_tokens 必须 > 0(入场预扣依据)")
|
||||
# 注: 不再强制 `tpm > 0 ⇒ est_tokens > 0`——预扣量由 effective_est_tokens()
|
||||
# 自 tpm 派生,运维只需填供应商配额页上抄得到的 tpm(设计 §3.2 #1)
|
||||
|
||||
def _validate_watchdog(self) -> None:
|
||||
# CHS config.py:66-82: 流式看门狗成对配置且 0 < inter < ttft < timeout_s
|
||||
@@ -134,6 +227,39 @@ class SourceConfig:
|
||||
):
|
||||
raise ValueError("看门狗不变式要求 0 < inter_token < ttft < timeout_s")
|
||||
|
||||
def _freeze_extra_body(self) -> None:
|
||||
"""校验后转只读视图: 装配完成的源不应再被就地改采样参数(设计决策 E)。"""
|
||||
validated = validate_request_overlay(
|
||||
self.extra_body, origin=f"SourceConfig({self.name}).extra_body"
|
||||
)
|
||||
object.__setattr__(self, "extra_body", MappingProxyType(validated))
|
||||
|
||||
|
||||
def strip_unsupported_extra_body(sources: list[SourceConfig], *, path: str) -> list[SourceConfig]:
|
||||
"""剥离非 chat 路径不消费的 `extra_body` 并 warning(issue #4 决策 G)。
|
||||
|
||||
剥离是必需的而非顺手清理: embedding 的 payload 硬编码 `{model, input}`、
|
||||
MonkeyOCR 只发 multipart 表单,两者都不会把 `extra_body` 发出去;但遥测的
|
||||
`sampling` 列会并上 `source.extra_body`,不剥离就等于**记录一个从未发出的
|
||||
参数**——那是数据造假,污染的恰是事后复现的唯一依据。
|
||||
|
||||
选择 warning 放行而非报错: 这两条路径本无采样语义,配错的后果远轻于 chat
|
||||
路径,不值得让下游整个装配起不来(2026-07-31 人类拍板)。
|
||||
"""
|
||||
stripped = []
|
||||
for source in sources:
|
||||
if source.extra_body:
|
||||
logger.warning(
|
||||
"{} 路径暂不支持 extra_body,源 {} 的该配置已被忽略"
|
||||
"(需要 dimensions 等参数请提 issue): {}",
|
||||
path,
|
||||
source.name,
|
||||
dict(source.extra_body),
|
||||
)
|
||||
source = dataclasses.replace(source, extra_body={})
|
||||
stripped.append(source)
|
||||
return stripped
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class RetryPolicy:
|
||||
@@ -270,7 +396,7 @@ class EmbeddingTransportResult:
|
||||
vectors: list[list[float]]
|
||||
dim: int
|
||||
prompt_tokens: int
|
||||
usage_source: str # measured | estimated
|
||||
usage_source: str # measured | estimated | unavailable
|
||||
raw: dict[str, Any]
|
||||
|
||||
|
||||
|
||||
@@ -84,6 +84,23 @@ class TestTpmGate:
|
||||
await p2.release()
|
||||
assert (await limiter.source_stats("s1")).tpm_used == 600
|
||||
|
||||
async def test_settle_equal_to_prededuct_keeps_deposit(self, limiter_factory):
|
||||
"""预扣量与结算量同为派生值时,双后端都必须留存押金(delta==0)。
|
||||
|
||||
这里只锁**后端算术**: 相等的两个数进出,窗口残留量恰为该值。
|
||||
"调用点是否真的取了派生值"是编排行为,由 tests/unit/test_retry.py
|
||||
经 RetryMW 端到端覆盖,不在本契约文件重复(否则只是自证同一个入参)。
|
||||
"""
|
||||
src = make_source(tpm=6000, est_tokens=0) # 派生值 = max(1, 6000 // 60) = 100
|
||||
derived = src.effective_est_tokens()
|
||||
assert derived == 100
|
||||
limiter = limiter_factory([src], _NO_GLOBAL)
|
||||
permit = await limiter.try_acquire("s1", derived)
|
||||
assert permit is not None
|
||||
await permit.settle(derived)
|
||||
await permit.release()
|
||||
assert (await limiter.source_stats("s1")).tpm_used == derived
|
||||
|
||||
async def test_failed_acquire_leaves_no_tpm_trace(self, limiter_factory):
|
||||
src = make_source(tpm=500, est_tokens=400)
|
||||
limiter = limiter_factory([src], _NO_GLOBAL)
|
||||
|
||||
@@ -4,6 +4,7 @@
|
||||
"""
|
||||
|
||||
import asyncio
|
||||
import dataclasses
|
||||
import json
|
||||
import sqlite3
|
||||
|
||||
@@ -204,3 +205,74 @@ class TestTransientErrorExport:
|
||||
|
||||
assert isinstance(ei.value, polygateway.AllSourcesExhausted)
|
||||
assert isinstance(ei.value.__cause__, TransientError)
|
||||
|
||||
|
||||
class TestSamplingThroughStack:
|
||||
"""issue #4: 采样参数经完整洋葱到达请求体,且缓存/遥测口径一致。"""
|
||||
|
||||
async def test_reaches_wire_and_lands_in_telemetry(self, tmp_path):
|
||||
seen = []
|
||||
|
||||
def handler(request):
|
||||
seen.append(json.loads(request.content))
|
||||
return _sse()
|
||||
|
||||
db = tmp_path / "t.db"
|
||||
recorder = SQLiteRecorder(db)
|
||||
client = _full_client(handler, telemetry=recorder)
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 42})
|
||||
recorder.close()
|
||||
|
||||
assert seen[0]["seed"] == 42 # 穿过全栈到达线上
|
||||
rows = sqlite3.connect(db).execute("SELECT sampling FROM llm_calls").fetchall()
|
||||
assert json.loads(rows[0][0]) == {"seed": 42}
|
||||
|
||||
async def test_config_level_merges_and_records(self, tmp_path):
|
||||
"""源级 extra_body 只有 emit_attempt 记得到(唯一有生效源的入口)。"""
|
||||
seen = []
|
||||
|
||||
def handler(request):
|
||||
seen.append(json.loads(request.content))
|
||||
return _sse()
|
||||
|
||||
src = dataclasses.replace(_source(), extra_body={"temperature": 0})
|
||||
db = tmp_path / "t.db"
|
||||
recorder = SQLiteRecorder(db)
|
||||
client = GatewayClient(
|
||||
scope="llm",
|
||||
sources=[src],
|
||||
selector=RoundRobinSelector(),
|
||||
limiter=InMemoryLimiter(
|
||||
scope="llm", sources={src.name: src}, global_limits=GlobalLimits(0, 0, 0)
|
||||
),
|
||||
breaker=InMemoryGate(config=_BREAKER),
|
||||
transport=OpenAICompatTransport(
|
||||
client_factory=lambda s: httpx.AsyncClient(transport=httpx.MockTransport(handler))
|
||||
),
|
||||
retry=RetryPolicy(2, 2.0, 30.0),
|
||||
backpressure=BackpressurePolicy(300.0, 0.01),
|
||||
telemetry=recorder,
|
||||
structured_strategy=JsonRepairStrategy(),
|
||||
sleep=_noop_sleep,
|
||||
)
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 1})
|
||||
recorder.close()
|
||||
|
||||
assert seen[0]["temperature"] == 0 and seen[0]["seed"] == 1
|
||||
rows = sqlite3.connect(db).execute("SELECT sampling FROM llm_calls").fetchall()
|
||||
assert json.loads(rows[0][0]) == {"seed": 1, "temperature": 0}
|
||||
|
||||
async def test_differing_seed_bypasses_cache_end_to_end(self):
|
||||
"""issue 场景全栈回归: 逐 rollout 变 seed 必须真的回源。"""
|
||||
calls = []
|
||||
|
||||
def handler(request):
|
||||
calls.append(json.loads(request.content)["seed"])
|
||||
return _sse()
|
||||
|
||||
client = _full_client(handler, cache=InMemoryCache())
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 1})
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 2})
|
||||
second_same = await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 1})
|
||||
assert calls == [1, 2] # 两个不同 seed 各自回源
|
||||
assert second_same.cache_hit is True # 同 seed 才命中
|
||||
|
||||
@@ -46,6 +46,10 @@ def _env(sources: dict[int, str], **extra: str) -> dict[str, str]:
|
||||
f"OCR__MONKEY__{n}__API_KEY": "none",
|
||||
f"OCR__MONKEY__{n}__MODEL": "monkey-ocr",
|
||||
f"OCR__MONKEY__{n}__TIMEOUT_S": "300",
|
||||
# 服务在 LAN,开发机若开着系统代理(httpx trust_env 读 macOS 系统配置,
|
||||
# 不是环境变量),代理会对内网地址回 403 —— 与本文件 raw httpx 用例
|
||||
# 显式传 trust_env=False 同因
|
||||
f"OCR__MONKEY__{n}__TRUST_ENV": "false",
|
||||
}
|
||||
env.update(extra)
|
||||
return env
|
||||
|
||||
@@ -11,6 +11,7 @@ DSN 走 .env `PGW_TELEMETRY_PG_DSN`,缺则 skip。该实例上有 app/chs_prod
|
||||
from __future__ import annotations
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import os
|
||||
from uuid import uuid4
|
||||
|
||||
@@ -39,6 +40,9 @@ _EXPECTED_COLUMNS = [
|
||||
"error",
|
||||
"cost",
|
||||
"created_at",
|
||||
"cached_prompt_tokens",
|
||||
"model_reported",
|
||||
"sampling",
|
||||
]
|
||||
|
||||
# run 级前缀: 同库并存的其他运行(迁移批跑/另一开发机)互不可见
|
||||
@@ -100,6 +104,9 @@ async def _record_minimal(
|
||||
"cache_hit": False,
|
||||
"error": None,
|
||||
"cost": None,
|
||||
"cached_prompt_tokens": None,
|
||||
"model_reported": None,
|
||||
"sampling": None,
|
||||
}
|
||||
fields.update(overrides)
|
||||
await recorder.record_llm_call(**fields)
|
||||
@@ -115,6 +122,111 @@ async def _fetch(dsn: str, sql: str, *args):
|
||||
await conn.close()
|
||||
|
||||
|
||||
_LEGACY_DDL = """
|
||||
CREATE TABLE {schema}.llm_calls (
|
||||
call_id TEXT PRIMARY KEY,
|
||||
parent_call_id TEXT,
|
||||
session_id TEXT,
|
||||
model TEXT NOT NULL,
|
||||
provider TEXT NOT NULL,
|
||||
source_name TEXT NOT NULL,
|
||||
messages TEXT NOT NULL,
|
||||
response TEXT NOT NULL,
|
||||
thinking TEXT NOT NULL DEFAULT '',
|
||||
prompt_tokens INTEGER NOT NULL,
|
||||
completion_tokens INTEGER NOT NULL,
|
||||
usage_source TEXT NOT NULL,
|
||||
latency_ms INTEGER NOT NULL,
|
||||
ttft_ms DOUBLE PRECISION,
|
||||
max_inter_token_ms DOUBLE PRECISION,
|
||||
cache_hit BOOLEAN NOT NULL DEFAULT FALSE,
|
||||
error TEXT,
|
||||
cost DOUBLE PRECISION,
|
||||
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
|
||||
)
|
||||
"""
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
async def legacy_schema(dsn):
|
||||
"""在**自建的临时 schema** 里造一张 18 列旧表,验证补列(issue #3)。
|
||||
|
||||
绝不碰共享的 public.llm_calls: 用 search_path 把 recorder 指向临时 schema,
|
||||
teardown 只 DROP 自己建的 schema。
|
||||
"""
|
||||
import asyncpg
|
||||
|
||||
name = f"pgwtest_{uuid4().hex[:8]}"
|
||||
conn = await asyncpg.connect(dsn, timeout=10)
|
||||
try:
|
||||
await conn.execute(f"CREATE SCHEMA {name}")
|
||||
await conn.execute(_LEGACY_DDL.format(schema=name))
|
||||
finally:
|
||||
await conn.close()
|
||||
sep = "&" if "?" in dsn else "?"
|
||||
yield f"{dsn}{sep}options=-csearch_path%3D{name}", name
|
||||
conn = await asyncpg.connect(dsn, timeout=10)
|
||||
try:
|
||||
await conn.execute(f"DROP SCHEMA {name} CASCADE")
|
||||
finally:
|
||||
await conn.close()
|
||||
|
||||
|
||||
class TestObservabilityColumns:
|
||||
"""issue #3: 两列写入可回读,且已存在的 18 列旧表会被自动补列。"""
|
||||
|
||||
async def test_values_round_trip(self, dsn):
|
||||
recorder = PostgresRecorder(dsn)
|
||||
try:
|
||||
await _record_minimal(recorder, call_id=_cid("hit"), cached_prompt_tokens=64)
|
||||
await _record_minimal(recorder, call_id=_cid("zero"), cached_prompt_tokens=0)
|
||||
await _record_minimal(recorder, call_id=_cid("model"), model_reported="MiniMax-01")
|
||||
await _record_minimal(
|
||||
recorder, call_id=_cid("samp"), sampling='{"seed": 42, "temperature": 0}'
|
||||
)
|
||||
rows = await _fetch(
|
||||
dsn,
|
||||
"SELECT call_id, cached_prompt_tokens, model_reported, sampling FROM llm_calls "
|
||||
"WHERE call_id LIKE $1",
|
||||
f"{_RUN_PREFIX}-%",
|
||||
)
|
||||
by_id = {r["call_id"]: r for r in rows}
|
||||
assert by_id[_cid("hit")]["cached_prompt_tokens"] == 64
|
||||
assert by_id[_cid("zero")]["cached_prompt_tokens"] == 0 # 真实零命中 ≠ NULL
|
||||
assert by_id[_cid("model")]["cached_prompt_tokens"] is None
|
||||
assert by_id[_cid("model")]["model_reported"] == "MiniMax-01"
|
||||
# issue #4: PG 侧也须验非空 sampling 能读回原值(不只是列存在)
|
||||
assert json.loads(by_id[_cid("samp")]["sampling"]) == {"seed": 42, "temperature": 0}
|
||||
assert by_id[_cid("hit")]["sampling"] is None
|
||||
finally:
|
||||
await recorder.aclose()
|
||||
|
||||
async def test_legacy_table_is_upgraded_in_place(self, legacy_schema):
|
||||
"""18 列旧表不补列的话,每行写入都会被逐行 warning 丢弃(遥测静默全失)。"""
|
||||
schema_dsn, schema = legacy_schema
|
||||
recorder = PostgresRecorder(schema_dsn)
|
||||
try:
|
||||
await _record_minimal(
|
||||
recorder, call_id=_cid("legacy"), cached_prompt_tokens=7, model_reported="m-real"
|
||||
)
|
||||
cols = await _fetch(
|
||||
schema_dsn,
|
||||
"SELECT column_name FROM information_schema.columns "
|
||||
"WHERE table_schema = $1 AND table_name = 'llm_calls' ORDER BY ordinal_position",
|
||||
schema,
|
||||
)
|
||||
# ALTER 只能追加到末尾: 与新建库的列序一致才不会分叉
|
||||
assert [r["column_name"] for r in cols] == _EXPECTED_COLUMNS
|
||||
rows = await _fetch(
|
||||
schema_dsn,
|
||||
"SELECT cached_prompt_tokens, model_reported FROM llm_calls WHERE call_id = $1",
|
||||
_cid("legacy"),
|
||||
)
|
||||
assert (rows[0]["cached_prompt_tokens"], rows[0]["model_reported"]) == (7, "m-real")
|
||||
finally:
|
||||
await recorder.aclose()
|
||||
|
||||
|
||||
class TestSchema:
|
||||
async def test_schema_has_frozen_columns_in_order(self, dsn):
|
||||
recorder = PostgresRecorder(dsn)
|
||||
|
||||
@@ -53,6 +53,35 @@ class TestKeyFormula:
|
||||
def test_any_dimension_change_changes_key(self, a, b):
|
||||
assert build_cache_key(*a) != build_cache_key(*b)
|
||||
|
||||
def test_empty_sampling_keeps_legacy_key(self):
|
||||
"""空采样参数时键形逐字不变,存量缓存不被全量作废(issue #4 决策 C)。
|
||||
|
||||
golden 值取自加 sampling 维度之前的实现,不得随实现漂移。
|
||||
"""
|
||||
assert build_cache_key("qwen-max", [{"role": "user", "content": "hi"}], "proj", None) == (
|
||||
"pgw:cache:c54544e8672f4c91373b4a72716a88497445b440b89445aa5379b356b228f58b"
|
||||
)
|
||||
assert build_cache_key("qwen-max", [{"role": "user", "content": "hi"}], "proj", "s1") == (
|
||||
"pgw:cache:eed9cd9cc06acc0dedf4f337b74e06ed3482afdc30fa2acedd194f6cc1df33bf"
|
||||
)
|
||||
|
||||
def test_differing_seed_changes_key(self):
|
||||
"""issue #4 的直接回归: 5 个 seed 若共用一个 key,标准差会恒为 0。"""
|
||||
k1 = build_cache_key("m", _MSGS, "proj", None, sampling={"seed": 1})
|
||||
k2 = build_cache_key("m", _MSGS, "proj", None, sampling={"seed": 2})
|
||||
assert k1 != k2
|
||||
|
||||
def test_sampling_key_order_irrelevant(self):
|
||||
k1 = build_cache_key("m", _MSGS, "proj", None, sampling={"seed": 1, "temperature": 0})
|
||||
k2 = build_cache_key("m", _MSGS, "proj", None, sampling={"temperature": 0, "seed": 1})
|
||||
assert k1 == k2
|
||||
|
||||
def test_empty_sampling_equals_omitted(self):
|
||||
"""空 dict 与不传须同键,否则升级后存量缓存全部 miss。"""
|
||||
assert build_cache_key("m", _MSGS, "proj", None, sampling={}) == build_cache_key(
|
||||
"m", _MSGS, "proj", None
|
||||
)
|
||||
|
||||
def test_multimodal_part_digested_not_inlined(self):
|
||||
big_b64 = "data:image/png;base64," + "A" * 1_000_000
|
||||
messages = [
|
||||
@@ -120,6 +149,32 @@ class TestCacheFlow:
|
||||
assert second.call_id != first.call_id # 命中生成独立 cache_call_id
|
||||
assert terminal.calls == 1 # 未再触达内层
|
||||
|
||||
async def test_differing_sampling_does_not_hit(self):
|
||||
"""issue #4 的中间件层回归: 逐 rollout 变 seed 必须回源,不得复用响应。"""
|
||||
backend = InMemoryCache()
|
||||
mw = _mw(backend)
|
||||
terminal = _Terminal(_resp())
|
||||
await mw(ChatRequest(messages=_MSGS, sampling={"seed": 1}), terminal)
|
||||
await mw(ChatRequest(messages=_MSGS, sampling={"seed": 2}), terminal)
|
||||
assert terminal.calls == 2 # 两次都回源
|
||||
# 同 seed 才命中
|
||||
third = await mw(ChatRequest(messages=_MSGS, sampling={"seed": 1}), terminal)
|
||||
assert third.cache_hit is True and terminal.calls == 2
|
||||
|
||||
async def test_structured_injection_does_not_pollute_key(self):
|
||||
"""CacheMW 读 sampling 而非 overlay: 结构化注入不该改变缓存身份。"""
|
||||
backend = InMemoryCache()
|
||||
mw = _mw(backend)
|
||||
terminal = _Terminal(_resp())
|
||||
await mw(ChatRequest(messages=_MSGS, sampling={"seed": 1}), terminal)
|
||||
polluted = ChatRequest(
|
||||
messages=_MSGS,
|
||||
sampling={"seed": 1},
|
||||
overlay={"seed": 1, "response_format": {"type": "json_object"}},
|
||||
)
|
||||
assert (await mw(polluted, terminal)).cache_hit is True
|
||||
assert terminal.calls == 1
|
||||
|
||||
async def test_per_call_namespace_overrides_default(self):
|
||||
backend = InMemoryCache()
|
||||
mw = _mw(backend)
|
||||
@@ -150,6 +205,49 @@ class TestCacheFlow:
|
||||
assert terminal.calls == 2
|
||||
|
||||
|
||||
class TestObservabilityFieldsOnHit:
|
||||
"""issue #3 决策 B1: 命中行原样回放,与 model/prompt_tokens 同一口径。"""
|
||||
|
||||
async def test_fields_replayed_on_hit(self):
|
||||
backend = InMemoryCache()
|
||||
mw = _mw(backend)
|
||||
terminal = _Terminal(
|
||||
_resp(cached_prompt_tokens=64, model_reported="MiniMax-Text-01-250321")
|
||||
)
|
||||
await mw(ChatRequest(messages=_MSGS), terminal)
|
||||
hit = await mw(ChatRequest(messages=_MSGS), terminal)
|
||||
assert hit.cache_hit is True
|
||||
assert hit.cached_prompt_tokens == 64
|
||||
assert hit.model_reported == "MiniMax-Text-01-250321"
|
||||
|
||||
async def test_legacy_cache_entry_without_new_keys_rehydrates(self):
|
||||
"""旧格式条目(无这两个键)必须照常重建为 None,不得抛异常回源。"""
|
||||
backend = InMemoryCache()
|
||||
mw = _mw(backend)
|
||||
key = build_cache_key("m", _MSGS, "proj", None)
|
||||
legacy = {
|
||||
"content": "legacy",
|
||||
"thinking": "",
|
||||
"model": "m",
|
||||
"provider": "p",
|
||||
"prompt_tokens": 1,
|
||||
"completion_tokens": 2,
|
||||
"latency_ms": 30,
|
||||
"ttft_ms": 5.0,
|
||||
"max_inter_token_ms": 2.0,
|
||||
"cache_hit": False,
|
||||
"call_id": "orig",
|
||||
"source_name": "s1",
|
||||
"cost": None,
|
||||
"usage_source": "measured",
|
||||
}
|
||||
await backend.set(key, json.dumps(legacy), ttl_s=100)
|
||||
terminal = _Terminal(_resp())
|
||||
hit = await mw(ChatRequest(messages=_MSGS), terminal)
|
||||
assert hit.content == "legacy" and terminal.calls == 0 # 真的走了缓存
|
||||
assert hit.cached_prompt_tokens is None and hit.model_reported is None
|
||||
|
||||
|
||||
class _BrokenBackend:
|
||||
async def get(self, key):
|
||||
raise ConnectionError("redis down")
|
||||
|
||||
@@ -19,6 +19,7 @@ from polygateway.backends.memory.cache import InMemoryCache
|
||||
from polygateway.backends.memory.limiter import InMemoryLimiter
|
||||
from polygateway.sources import RoundRobinSelector
|
||||
from polygateway.structured.json_repair import JsonRepairStrategy
|
||||
from polygateway.structured.native_schema import NativeSchemaStrategy
|
||||
from polygateway.transports.openai_compat import OpenAICompatTransport
|
||||
from polygateway.types import (
|
||||
BackpressurePolicy,
|
||||
@@ -117,6 +118,106 @@ class TestChatEndToEnd:
|
||||
await client.chat([{"role": "user", "content": "hi"}], structured="json")
|
||||
|
||||
|
||||
class TestSamplingOverlay:
|
||||
"""调用级采样参数入口(issue #4 Task 3)。"""
|
||||
|
||||
def _capturing_client(self, captured, **overrides):
|
||||
def handler(request):
|
||||
captured.append(json.loads(request.content))
|
||||
return _sse()
|
||||
|
||||
return _client(handler=handler, **overrides)
|
||||
|
||||
async def test_overlay_reaches_request_body(self):
|
||||
captured = []
|
||||
async with self._capturing_client(captured) as client:
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay={"seed": 42})
|
||||
assert captured[0]["seed"] == 42
|
||||
|
||||
async def test_call_level_beats_config_level(self):
|
||||
"""优先级: 调用级 > 配置级(设计决策 A)。"""
|
||||
captured = []
|
||||
source = _source(extra_body={"temperature": 0, "top_p": 0.9})
|
||||
async with self._capturing_client(captured, sources=[source]) as client:
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay={"temperature": 1})
|
||||
assert captured[0]["temperature"] == 1 # 调用级覆盖
|
||||
assert captured[0]["top_p"] == 0.9 # 配置级未被顶掉的键保留
|
||||
|
||||
async def test_structured_injection_beats_call_level(self):
|
||||
"""结构化注入优先级最高: 它关系到响应能否被解析(设计决策 A)。"""
|
||||
captured = []
|
||||
client = self._capturing_client(captured, structured_strategy=NativeSchemaStrategy())
|
||||
async with client:
|
||||
await client.chat(
|
||||
[{"role": "user", "content": "hi"}],
|
||||
structured="json",
|
||||
overlay={"response_format": {"type": "text"}},
|
||||
)
|
||||
assert captured[0]["response_format"] != {"type": "text"}
|
||||
|
||||
async def test_protected_key_rejected_before_onion(self):
|
||||
"""保护键在进洋葱之前就报错,transport 一次都不该被碰到。"""
|
||||
captured = []
|
||||
async with self._capturing_client(captured) as client:
|
||||
with pytest.raises(ValueError, match="stream"):
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay={"stream": False})
|
||||
assert captured == []
|
||||
|
||||
async def test_unserializable_value_rejected_before_onion(self):
|
||||
"""裸 TypeError 会在 CacheMW 的降级 try 之外炸且无遥测(设计决策 B)。"""
|
||||
captured = []
|
||||
async with self._capturing_client(captured) as client:
|
||||
with pytest.raises(ValueError, match="JSON"):
|
||||
await client.chat(
|
||||
[{"role": "user", "content": "hi"}], overlay={"temperature": object()}
|
||||
)
|
||||
assert captured == []
|
||||
|
||||
async def test_caller_dict_mutation_does_not_leak(self):
|
||||
"""调用方逐次改 seed 复用同一 dict 是预期模式(设计决策 E)。"""
|
||||
captured = []
|
||||
caller_overlay = {"seed": 1}
|
||||
async with self._capturing_client(captured) as client:
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay=caller_overlay)
|
||||
caller_overlay["seed"] = 2
|
||||
await client.chat([{"role": "user", "content": "hi"}], overlay=caller_overlay)
|
||||
assert [c["seed"] for c in captured] == [1, 2]
|
||||
|
||||
|
||||
class TestModelFingerprint:
|
||||
"""配置级采样参数须进缓存身份,否则改 temperature 后仍读旧缓存(决策 C)。"""
|
||||
|
||||
def test_empty_extra_body_keeps_legacy_fingerprint(self):
|
||||
"""全源无 extra_body 时字面量与旧实现逐字相同,不触发存量缓存冷启动。"""
|
||||
from polygateway.client import build_model_fingerprint
|
||||
|
||||
sources = [_source(), _source(name="qwen_2", model="qwen-plus")]
|
||||
assert build_model_fingerprint(sources) == "qwen-max,qwen-plus"
|
||||
|
||||
def test_extra_body_changes_fingerprint(self):
|
||||
from polygateway.client import build_model_fingerprint
|
||||
|
||||
plain = build_model_fingerprint([_source()])
|
||||
tuned = build_model_fingerprint([_source(extra_body={"temperature": 0})])
|
||||
assert plain != tuned
|
||||
assert tuned.startswith("qwen-max|") # 旧字面量仍是前缀,便于人眼辨认
|
||||
|
||||
def test_source_rename_does_not_change_fingerprint(self):
|
||||
"""指纹按 (model, extra_body) 而非源名: 改名不该误触全量冷启动。"""
|
||||
from polygateway.client import build_model_fingerprint
|
||||
|
||||
a = build_model_fingerprint([_source(name="qwen_1", extra_body={"temperature": 0})])
|
||||
b = build_model_fingerprint([_source(name="renamed", extra_body={"temperature": 0})])
|
||||
assert a == b
|
||||
|
||||
def test_differing_extra_body_across_sources_is_distinguished(self):
|
||||
from polygateway.client import build_model_fingerprint
|
||||
|
||||
a = build_model_fingerprint([_source(extra_body={"temperature": 0})])
|
||||
b = build_model_fingerprint([_source(extra_body={"temperature": 1})])
|
||||
assert a != b
|
||||
|
||||
|
||||
class TestFactories:
|
||||
def test_from_env_assembles(self):
|
||||
client = GatewayClient.from_env("LLM", env=_ENV)
|
||||
|
||||
+332
-2
@@ -1,8 +1,13 @@
|
||||
"""config.py 配置聚合测试(设计 §8): 多源命名、键优先级、缺失报错。"""
|
||||
|
||||
import pytest
|
||||
import contextlib
|
||||
import dataclasses
|
||||
|
||||
from polygateway.config import GatewaySettings
|
||||
import pytest
|
||||
from loguru import logger
|
||||
|
||||
from polygateway.client import GatewayClient
|
||||
from polygateway.config import EmbeddingSettings, GatewaySettings, OcrSettings
|
||||
|
||||
_BASE_ENV = {
|
||||
"LLM__QWEN__1__BASE_URL": "https://gw-a.example/v1",
|
||||
@@ -19,6 +24,17 @@ _BASE_ENV = {
|
||||
}
|
||||
|
||||
|
||||
@contextlib.contextmanager
|
||||
def _captured_warnings():
|
||||
"""捕获库发出的 WARNING;loguru 不经标准 logging,pytest 的 caplog 抓不到。"""
|
||||
messages: list[str] = []
|
||||
sink_id = logger.add(messages.append, level="WARNING")
|
||||
try:
|
||||
yield messages
|
||||
finally:
|
||||
logger.remove(sink_id)
|
||||
|
||||
|
||||
def _env(**overrides):
|
||||
env = dict(_BASE_ENV)
|
||||
env.update({k: v for k, v in overrides.items() if v is not None})
|
||||
@@ -87,6 +103,35 @@ class TestSourceAggregation:
|
||||
GatewaySettings.from_env("LLM", env=env)
|
||||
|
||||
|
||||
class TestExtraBodyParsing:
|
||||
"""配置级采样参数的 env 解析(issue #4 Task 2)。"""
|
||||
|
||||
def test_json_object_parsed(self):
|
||||
env = _env(**{"LLM__QWEN__1__EXTRA_BODY": '{"temperature": 0, "seed": 42}'})
|
||||
s = GatewaySettings.from_env("LLM", env=env)
|
||||
assert s.sources[0].extra_body == {"temperature": 0, "seed": 42}
|
||||
|
||||
def test_absent_defaults_to_empty(self):
|
||||
assert GatewaySettings.from_env("LLM", env=_env()).sources[0].extra_body == {}
|
||||
|
||||
def test_invalid_json_fails_loudly(self):
|
||||
env = _env(**{"LLM__QWEN__1__EXTRA_BODY": "{invalid"})
|
||||
with pytest.raises(ValueError, match="EXTRA_BODY"):
|
||||
GatewaySettings.from_env("LLM", env=env)
|
||||
|
||||
def test_non_object_json_fails(self):
|
||||
"""数组/标量都不是请求体片段,静默接受会让参数悄悄不生效。"""
|
||||
env = _env(**{"LLM__QWEN__1__EXTRA_BODY": "[1, 2]"})
|
||||
with pytest.raises(ValueError, match="JSON 对象"):
|
||||
GatewaySettings.from_env("LLM", env=env)
|
||||
|
||||
def test_protected_key_rejected_through_assembly(self):
|
||||
"""校验确实挂在装配路径上(而非只在 types.py 里孤立存在)。"""
|
||||
env = _env(**{"LLM__QWEN__1__EXTRA_BODY": '{"model": "sneaky"}'})
|
||||
with pytest.raises(ValueError, match="model"):
|
||||
GatewaySettings.from_env("LLM", env=env)
|
||||
|
||||
|
||||
class TestResilienceKeys:
|
||||
def test_flat_legacy_keys(self):
|
||||
s = GatewaySettings.from_env("LLM", env=_env())
|
||||
@@ -155,6 +200,11 @@ class TestAssemblyGuards:
|
||||
s2 = GatewaySettings.from_env("LLM", env=_env(PGW_STRUCTURED_MAX_RETRIES="0"))
|
||||
assert s2.structured_max_retries == 0
|
||||
|
||||
def test_negative_structured_retries_rejected_with_env_key(self):
|
||||
"""env 层的检查保留是为了报错能点出键名(构造期那道点的是字段名)。"""
|
||||
with pytest.raises(ValueError, match="PGW_STRUCTURED_MAX_RETRIES"):
|
||||
GatewaySettings.from_env("LLM", env=_env(PGW_STRUCTURED_MAX_RETRIES="-1"))
|
||||
|
||||
def test_cache_requires_namespace_and_ttl(self):
|
||||
env = _env(PGW_CACHE_BACKEND="memory")
|
||||
with pytest.raises(ValueError, match="NAMESPACE"):
|
||||
@@ -255,6 +305,11 @@ class TestAssemblyGuards:
|
||||
with pytest.raises(ValueError, match="TELEMETRY_BACKEND"):
|
||||
GatewaySettings.from_env("LLM", env=_env(PGW_TELEMETRY_BACKEND="mysql"))
|
||||
|
||||
def test_cache_backend_whitelist(self):
|
||||
"""对称于上一条: env 层的域检查保留是为了报错能点出键名,得有测试守着。"""
|
||||
with pytest.raises(ValueError, match="CACHE_BACKEND"):
|
||||
GatewaySettings.from_env("LLM", env=_env(PGW_CACHE_BACKEND="rediss"))
|
||||
|
||||
def test_pricing_path_optional(self):
|
||||
assert GatewaySettings.from_env("LLM", env=_env()).pricing_path is None
|
||||
s = GatewaySettings.from_env("LLM", env=_env(PGW_PRICING_PATH="conf/prices.json"))
|
||||
@@ -319,3 +374,278 @@ class TestOcrSettings:
|
||||
env = {k: v for k, v in self._OCR_ENV.items() if k != "OCR__MONKEY__1__BASE_URL"}
|
||||
with pytest.raises(ValueError):
|
||||
OcrSettings.from_env("OCR", env=env)
|
||||
|
||||
|
||||
class TestCrossFieldInvariants:
|
||||
"""四条跨字段不变量必须在**任何**构造路径上生效(设计 2026-07-29)。
|
||||
|
||||
这些约束单看一个字段都合法,组合起来才非法,因此 types.py 各子配置的
|
||||
__post_init__ 看不见——只能由聚合层 GatewaySettings 把关。守卫若只挂在
|
||||
from_env 上,from_settings 这条同等官方的装配路(CLAUDE.md §4.5)就能
|
||||
装出违反不变量的配置,类会存在于自己 docstring 声称不可能的状态。
|
||||
|
||||
每条不变量测两侧: 越界必拒、边界值(恰好相等)必过——收紧的是错的组合,
|
||||
不是所有直接构造。
|
||||
"""
|
||||
|
||||
def _base(self, **overrides) -> GatewaySettings:
|
||||
return GatewaySettings.from_env("LLM", env=_env(**overrides))
|
||||
|
||||
def _with_watchdog(self) -> GatewaySettings:
|
||||
"""带看门狗的基准: TTFT/inter-token 成对配置才满足 SourceConfig 不变式。"""
|
||||
return self._base(
|
||||
**{
|
||||
"LLM__QWEN__1__TTFT_TIMEOUT_S": "30",
|
||||
"LLM__QWEN__1__INTER_TOKEN_TIMEOUT_S": "15",
|
||||
"LLM__BACKPRESSURE__STALL_WINDOW_S": "300",
|
||||
}
|
||||
)
|
||||
|
||||
# —— 源超时 ≤ permit 租约 TTL(ARCH §7.3: 防租约先于请求过期,并发悄悄超配额)——
|
||||
|
||||
def test_lease_rejects_timeout_above_ttl_on_direct_construction(self):
|
||||
base = self._base() # 源 timeout_s=120
|
||||
with pytest.raises(ValueError, match="lease_ttl_s"):
|
||||
dataclasses.replace(base, lease_ttl_s=1.0)
|
||||
|
||||
def test_lease_accepts_timeout_equal_to_ttl(self):
|
||||
base = self._base()
|
||||
assert dataclasses.replace(base, lease_ttl_s=120.0).lease_ttl_s == 120.0
|
||||
|
||||
# —— stall 窗口 ≥ 最大源 TTFT(ARCH §7.3: 防正常慢首包被误判卡死掐断)——
|
||||
|
||||
def test_stall_rejects_window_below_max_ttft_on_direct_construction(self):
|
||||
base = self._with_watchdog() # 源 ttft_timeout_s=30
|
||||
narrowed = dataclasses.replace(base.backpressure, stall_window_s=20.0)
|
||||
with pytest.raises(ValueError, match="stall_window_s"):
|
||||
dataclasses.replace(base, backpressure=narrowed)
|
||||
|
||||
def test_stall_accepts_window_equal_to_max_ttft(self):
|
||||
base = self._with_watchdog()
|
||||
exact = dataclasses.replace(base.backpressure, stall_window_s=30.0)
|
||||
assert dataclasses.replace(base, backpressure=exact).backpressure.stall_window_s == 30.0
|
||||
|
||||
# —— 探针租约 ≥ 最慢源超时 + 5(M2 设计 §3: 防半开探针在途即被接管)——
|
||||
|
||||
def test_probe_rejects_ttl_below_floor_on_direct_construction(self):
|
||||
base = self._base() # 最慢 timeout_s=120,故下限 125
|
||||
shortened = dataclasses.replace(base.breaker, probe_ttl_s=100.0)
|
||||
with pytest.raises(ValueError, match="probe_ttl_s"):
|
||||
dataclasses.replace(base, breaker=shortened)
|
||||
|
||||
def test_probe_accepts_ttl_at_floor(self):
|
||||
base = self._base()
|
||||
at_floor = dataclasses.replace(base.breaker, probe_ttl_s=125.0)
|
||||
assert dataclasses.replace(base, breaker=at_floor).breaker.probe_ttl_s == 125.0
|
||||
|
||||
# —— sources 非空 ——
|
||||
|
||||
def test_empty_sources_rejected_with_actionable_message(self):
|
||||
"""零源装出来的 client 选源必然失败;消息须点明原因,不能泄漏 max() 的内置异常。"""
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError) as exc:
|
||||
dataclasses.replace(base, sources=())
|
||||
assert "至少一个源" in str(exc.value)
|
||||
assert "empty sequence" not in str(exc.value)
|
||||
|
||||
# —— 装配路径覆盖 ——
|
||||
|
||||
def test_factory_cannot_receive_invalid_settings(self):
|
||||
"""from_settings 这条路吃不到非法配置。
|
||||
|
||||
异常实际抛在实参求值(构造 settings)那一刻,而不是工厂内部——这正是
|
||||
把守卫放构造期换来的性质: 非法实例根本不存在,无需每个工厂各自设防。
|
||||
"""
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="lease_ttl_s"):
|
||||
GatewayClient.from_settings(dataclasses.replace(base, lease_ttl_s=1.0))
|
||||
|
||||
def test_ocr_settings_cannot_wrap_invalid_gateway(self):
|
||||
"""OcrSettings/EmbeddingSettings 只是包一层 GatewaySettings,自动继承同一把关。"""
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="lease_ttl_s"):
|
||||
OcrSettings(gateway=dataclasses.replace(base, lease_ttl_s=1.0))
|
||||
|
||||
# —— 第二轮(设计 2026-07-30): 后端枚举合法域 ——
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("field", "bad_value"),
|
||||
[
|
||||
("limiter_backend", "rediss"),
|
||||
("breaker_backend", "sqlite"),
|
||||
("cache_backend", "postgres"),
|
||||
("telemetry_backend", "redis"),
|
||||
("selector", "random"),
|
||||
("quota_full", "block"),
|
||||
],
|
||||
)
|
||||
def test_enum_field_rejects_value_outside_domain(self, field, bad_value):
|
||||
"""域外取值此前只有 from_env 拦得住,直接构造会落进 _build_* 的 else 分支。"""
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match=field):
|
||||
dataclasses.replace(base, **{field: bad_value})
|
||||
|
||||
# —— 条件必填: 取 redis 的后端必须有 redis_url ——
|
||||
|
||||
@pytest.mark.parametrize("field", ["limiter_backend", "breaker_backend"])
|
||||
def test_redis_backend_requires_redis_url(self, field):
|
||||
"""client.py 的 assert settings.redis_url is not None 依赖的正是这条。"""
|
||||
base = self._base() # redis_url=None
|
||||
with pytest.raises(ValueError, match="redis_url"):
|
||||
dataclasses.replace(base, **{field: "redis"})
|
||||
|
||||
def test_redis_cache_requires_redis_url(self):
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="redis_url"):
|
||||
dataclasses.replace(base, cache_backend="redis", cache_namespace="ns", cache_ttl_s=60)
|
||||
|
||||
# —— 条件必填: 启用缓存必须有命名空间与正 TTL ——
|
||||
|
||||
def test_cache_requires_namespace(self):
|
||||
"""缺命名空间即失去租户隔离,踩"无缓存毒化"铁律。"""
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="cache_namespace"):
|
||||
dataclasses.replace(base, cache_backend="memory", cache_ttl_s=60)
|
||||
|
||||
def test_cache_ttl_must_be_positive(self):
|
||||
"""from_env 明令禁止的"永不过期"不能从另一条路进来。"""
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="cache_ttl_s"):
|
||||
dataclasses.replace(base, cache_backend="memory", cache_namespace="ns", cache_ttl_s=0)
|
||||
|
||||
# —— 条件必填: 遥测后端各自的落点 ——
|
||||
|
||||
def test_sqlite_telemetry_requires_path(self):
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="telemetry_sqlite_path"):
|
||||
dataclasses.replace(base, telemetry_backend="sqlite")
|
||||
|
||||
def test_postgres_telemetry_requires_dsn(self):
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="telemetry_pg_dsn"):
|
||||
dataclasses.replace(base, telemetry_backend="postgres")
|
||||
|
||||
# —— 标量域 ——
|
||||
|
||||
def test_negative_structured_retries_rejected(self):
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="structured_max_retries"):
|
||||
dataclasses.replace(base, structured_max_retries=-1)
|
||||
|
||||
def test_blank_scope_rejected(self):
|
||||
"""空 scope 会污染遥测与缓存命名空间。"""
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="scope"):
|
||||
dataclasses.replace(base, scope=" ")
|
||||
|
||||
# —— 合法组合仍可构造(收紧的是错的那些)——
|
||||
|
||||
def test_full_redis_stack_constructible(self):
|
||||
base = self._base()
|
||||
settings = dataclasses.replace(
|
||||
base,
|
||||
limiter_backend="redis",
|
||||
breaker_backend="redis",
|
||||
cache_backend="redis",
|
||||
cache_namespace="ns",
|
||||
cache_ttl_s=60,
|
||||
redis_url="redis://127.0.0.1:6379/3",
|
||||
)
|
||||
assert settings.cache_ttl_s == 60 and settings.redis_url is not None
|
||||
|
||||
# —— Postgres DSN: 剥 SQLAlchemy 驱动后缀并出声(设计 §5 方案 C)——
|
||||
|
||||
def test_sqlalchemy_dsn_suffix_stripped_with_warning(self):
|
||||
"""asyncpg 不认 `+driver`;库替调用方剥掉,但不静默——日志里看得见。"""
|
||||
base = self._base()
|
||||
with _captured_warnings() as warnings:
|
||||
settings = dataclasses.replace(
|
||||
base,
|
||||
telemetry_backend="postgres",
|
||||
telemetry_pg_dsn="postgresql+asyncpg://u:s3cret@h/db",
|
||||
)
|
||||
assert settings.telemetry_pg_dsn == "postgresql://u:s3cret@h/db"
|
||||
assert any("asyncpg" in m for m in warnings)
|
||||
|
||||
def test_dsn_warning_does_not_leak_credentials(self):
|
||||
"""DSN 带密码,日志只能出现 scheme 段(P5: 敏感信息只走 .env)。"""
|
||||
base = self._base()
|
||||
with _captured_warnings() as warnings:
|
||||
dataclasses.replace(
|
||||
base,
|
||||
telemetry_backend="postgres",
|
||||
telemetry_pg_dsn="postgresql+asyncpg://u:s3cret@h/db",
|
||||
)
|
||||
assert warnings and not any("s3cret" in m or "@h/db" in m for m in warnings)
|
||||
|
||||
def test_env_path_strips_dsn_without_warning(self):
|
||||
"""env 路已在 _load_pg_dsn 剥过,不该给三项目的历史 DSN 写法刷噪音。"""
|
||||
with _captured_warnings() as warnings:
|
||||
settings = GatewaySettings.from_env(
|
||||
"LLM",
|
||||
env=_env(
|
||||
PGW_TELEMETRY_BACKEND="postgres",
|
||||
PGW_TELEMETRY_PG_DSN="postgresql+asyncpg://u@h/db",
|
||||
),
|
||||
)
|
||||
assert settings.telemetry_pg_dsn == "postgresql://u@h/db"
|
||||
assert not warnings
|
||||
|
||||
# —— 构造期规范化: env 路一直在做的,构造路也要做(否则两条路产出不同的值)——
|
||||
|
||||
@pytest.mark.parametrize("raw", ["LLM", " llm ", " LLM "])
|
||||
def test_scope_normalized_on_direct_construction(self, raw):
|
||||
"""scope 直接进 Redis key(pgw:limit:{scope}:…)。
|
||||
|
||||
大小写不一致会让同一逻辑 scope 的限流/熔断状态分裂到两套命名空间——
|
||||
两边各记各的配额与熔断状态,分布式治理静默失效且不报错。
|
||||
"""
|
||||
base = self._base()
|
||||
assert dataclasses.replace(base, scope=raw).scope == "llm"
|
||||
|
||||
def test_blank_redis_url_normalized_to_none(self):
|
||||
"""空串此前只有 env 路归 None,构造路留着它骗过 `is None` 判断。"""
|
||||
base = self._base()
|
||||
assert dataclasses.replace(base, redis_url="").redis_url is None
|
||||
|
||||
def test_blank_redis_url_still_blocks_redis_backend(self):
|
||||
"""归 None 后必须落进条件必填,而不是放行到 redis 库去抛连接串天书。"""
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="redis_url"):
|
||||
dataclasses.replace(base, limiter_backend="redis", redis_url="")
|
||||
|
||||
def test_blank_pricing_path_normalized_to_none(self):
|
||||
base = self._base()
|
||||
assert dataclasses.replace(base, pricing_path="").pricing_path is None
|
||||
|
||||
# —— EmbeddingSettings 自身的字段域(此前只有 from_env 校验)——
|
||||
|
||||
@pytest.mark.parametrize("bad", [0, -3])
|
||||
def test_embedding_settings_rejects_non_positive_batch_size(self, bad):
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="batch_size"):
|
||||
EmbeddingSettings(gateway=base, batch_size=bad)
|
||||
|
||||
def test_embedding_settings_rejects_non_positive_expected_dim(self):
|
||||
base = self._base()
|
||||
with pytest.raises(ValueError, match="expected_dim"):
|
||||
EmbeddingSettings(gateway=base, batch_size=8, expected_dim=0)
|
||||
|
||||
def test_embedding_settings_accepts_valid_values(self):
|
||||
base = self._base()
|
||||
settings = EmbeddingSettings(gateway=base, batch_size=8, expected_dim=1024)
|
||||
assert settings.batch_size == 8 and settings.expected_dim == 1024
|
||||
|
||||
# —— 回归护栏: client.py 的 assert 前提确实被保证了 ——
|
||||
|
||||
def test_factory_accepts_valid_redis_stack(self):
|
||||
"""补齐校验后,client.py:262/282/302 的 assert 退回成纯内部不变量声明。"""
|
||||
base = self._base()
|
||||
settings = dataclasses.replace(
|
||||
base,
|
||||
limiter_backend="redis",
|
||||
breaker_backend="redis",
|
||||
redis_url="redis://127.0.0.1:6379/3",
|
||||
)
|
||||
client = GatewayClient.from_settings(settings)
|
||||
assert client is not None
|
||||
|
||||
@@ -4,11 +4,13 @@
|
||||
VT adapters/embedding.py(归一化);库裁决见设计 §7.3 表。
|
||||
"""
|
||||
|
||||
import contextlib
|
||||
import dataclasses
|
||||
import json
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from loguru import logger
|
||||
|
||||
from polygateway.errors import (
|
||||
RequestRejectedError,
|
||||
@@ -102,12 +104,14 @@ class TestEmbedTransport:
|
||||
assert result.dim == 2
|
||||
assert result.prompt_tokens == 5 and result.usage_source == "measured"
|
||||
|
||||
async def test_missing_usage_falls_back_estimated(self):
|
||||
async def test_missing_usage_is_unavailable(self):
|
||||
"""usage 缺失不再退到 `est_tokens`(夹具填 7),与 chat 同口径记 0 + unavailable。"""
|
||||
|
||||
def handler(request):
|
||||
return httpx.Response(200, json=_ok_body([[1.0]]))
|
||||
|
||||
result = await _transport_with(handler).embed(texts=["a"], source=_src(), call_id="c")
|
||||
assert result.prompt_tokens == 7 and result.usage_source == "estimated" # est_tokens
|
||||
assert result.prompt_tokens == 0 and result.usage_source == "unavailable"
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
("status", "exc_type"),
|
||||
@@ -161,6 +165,7 @@ from polygateway.backends.memory.breaker import InMemoryGate # noqa: E402
|
||||
from polygateway.backends.memory.limiter import InMemoryLimiter # noqa: E402
|
||||
from polygateway.config import EmbeddingSettings # noqa: E402
|
||||
from polygateway.embedding import EmbeddingClient # noqa: E402
|
||||
from polygateway.pricing import ModelPrice, PricingTable # noqa: E402
|
||||
from polygateway.sources import RoundRobinSelector # noqa: E402
|
||||
from polygateway.types import ( # noqa: E402
|
||||
BackpressurePolicy,
|
||||
@@ -260,6 +265,28 @@ class TestEmbedBatching:
|
||||
assert resp.usage_source == "estimated" # 任一批 estimated 则整体 estimated
|
||||
assert resp.prompt_tokens == 2 + 9
|
||||
|
||||
async def test_unavailable_batch_dominates_and_voids_cost(self):
|
||||
"""三态合并优先级(设计 §3.2 #10/#11): 任一批不可得 → 整体不可得且 cost NULL。
|
||||
|
||||
改前二值合并只看 `estimated`,measured+unavailable 会误标 measured;
|
||||
`_total_cost` 逐批求和还会给出一个偏低却看似有效的金额。
|
||||
"""
|
||||
estimated = EmbeddingTransportResult(
|
||||
vectors=[[1.0], [1.0]], dim=1, prompt_tokens=9, usage_source="estimated", raw={}
|
||||
)
|
||||
unavailable = EmbeddingTransportResult(
|
||||
vectors=[[1.0], [1.0]], dim=1, prompt_tokens=0, usage_source="unavailable", raw={}
|
||||
)
|
||||
client, _ = _embed_client(
|
||||
[_src()],
|
||||
["ok", estimated, unavailable],
|
||||
batch_size=2,
|
||||
pricing=PricingTable({"embed-1": ModelPrice(input_per_1m=1.0, output_per_1m=0.0)}),
|
||||
)
|
||||
resp = await client.embed(["a", "b", "c", "d", "e", "f"])
|
||||
assert resp.usage_source == "unavailable" # unavailable 压过 estimated 与 measured
|
||||
assert resp.cost is None
|
||||
|
||||
|
||||
class TestEmbedPostProcess:
|
||||
async def test_normalize_l2(self):
|
||||
@@ -325,6 +352,46 @@ class TestEmbedTelemetry:
|
||||
assert len(rec.rows[1]["messages"]) < 1000 # 长文本截断后入库
|
||||
|
||||
|
||||
@contextlib.contextmanager
|
||||
def _captured_warnings():
|
||||
"""捕获库发出的 WARNING;loguru 不经标准 logging,pytest 的 caplog 抓不到。"""
|
||||
messages: list[str] = []
|
||||
sink_id = logger.add(messages.append, level="WARNING")
|
||||
try:
|
||||
yield messages
|
||||
finally:
|
||||
logger.remove(sink_id)
|
||||
|
||||
|
||||
class TestExtraBodyStripped:
|
||||
"""issue #4 决策 G: embedding 路径不消费 extra_body,剥离并 warning。"""
|
||||
|
||||
async def test_stripped_with_warning_but_assembly_succeeds(self):
|
||||
"""报错会让下游整个装配起不来,而这条路径本无采样语义(人类拍板)。"""
|
||||
with _captured_warnings() as warnings:
|
||||
client, _ = _embed_client([_src(extra_body={"temperature": 0})], ["ok"])
|
||||
assert client._sources[0].extra_body == {}
|
||||
assert any("extra_body" in m for m in warnings)
|
||||
assert any("dimensions" in m for m in warnings) # 文案须指路,不能只说不支持
|
||||
await client.embed(["hi"]) # 装配后可正常工作
|
||||
|
||||
async def test_telemetry_never_records_a_parameter_that_was_not_sent(self):
|
||||
"""剥离的真正理由: embed payload 硬编码 {model, input},不剥离则审计表
|
||||
|
||||
会显示这次调用带了 temperature=0——那是数据造假,比参数失效更坏。
|
||||
"""
|
||||
rec = _MemoryRecorder()
|
||||
client, _ = _embed_client([_src(extra_body={"temperature": 0})], ["ok"], telemetry=rec)
|
||||
await client.embed(["hi"])
|
||||
assert rec.rows[0]["sampling"] is None
|
||||
|
||||
async def test_no_warning_without_extra_body(self):
|
||||
with _captured_warnings() as warnings:
|
||||
client, _ = _embed_client([_src()], ["ok"])
|
||||
assert client._sources[0].extra_body == {}
|
||||
assert not [m for m in warnings if "extra_body" in m]
|
||||
|
||||
|
||||
class TestEmbeddingSettings:
|
||||
_ENV = {
|
||||
"EMBED__QWEN__1__BASE_URL": "https://gw.example/v1",
|
||||
@@ -358,6 +425,11 @@ class TestEmbeddingSettings:
|
||||
s = EmbeddingSettings.from_env("EMBED", env=env)
|
||||
assert s.normalize is True and s.expected_dim == 768
|
||||
|
||||
def test_expected_dim_must_be_positive(self):
|
||||
"""env 层的检查保留是为了报错能点出键名(构造期那道点的是字段名)。"""
|
||||
with pytest.raises(ValueError, match="EXPECTED_DIM"):
|
||||
EmbeddingSettings.from_env("EMBED", env={**self._ENV, "EMBED__EXPECTED_DIM": "0"})
|
||||
|
||||
def test_from_settings_assembles_client(self):
|
||||
s = EmbeddingSettings.from_env("EMBED", env=self._ENV)
|
||||
client = EmbeddingClient.from_settings(s)
|
||||
|
||||
@@ -7,6 +7,7 @@ retry_exhausted/circuit_open/stalled 三组断言即设计 §6 ③ 的 G1 契约
|
||||
import asyncio
|
||||
|
||||
import pytest
|
||||
from loguru import logger
|
||||
|
||||
from polygateway.backends.memory.breaker import InMemoryGate
|
||||
from polygateway.backends.memory.limiter import InMemoryLimiter
|
||||
@@ -377,6 +378,27 @@ class TestCheckHealth:
|
||||
await task
|
||||
|
||||
|
||||
class TestExtraBodyStripped:
|
||||
"""issue #4 决策 G: OCR 路径只发 multipart 表单,剥离 extra_body 并 warning。"""
|
||||
|
||||
def test_stripped_with_warning_but_assembly_succeeds(self):
|
||||
messages: list[str] = []
|
||||
sink_id = logger.add(messages.append, level="WARNING")
|
||||
try:
|
||||
client, _, _ = _client([_src(extra_body={"temperature": 0})], ["text"])
|
||||
finally:
|
||||
logger.remove(sink_id)
|
||||
assert client._sources[0].extra_body == {}
|
||||
assert any("extra_body" in m for m in messages)
|
||||
|
||||
async def test_telemetry_never_records_a_parameter_that_was_not_sent(self):
|
||||
"""不剥离则审计表会显示这次 OCR 带了 temperature=0——数据造假。"""
|
||||
recorder = _MemoryRecorder()
|
||||
client, _, _ = _client([_src(extra_body={"temperature": 0})], ["text"], telemetry=recorder)
|
||||
await client.recognize_text(b"IMG")
|
||||
assert recorder.rows[0]["sampling"] is None
|
||||
|
||||
|
||||
class TestTelemetry:
|
||||
async def test_success_and_failure_recorded_without_image_bytes(self):
|
||||
recorder = _MemoryRecorder()
|
||||
@@ -394,6 +416,21 @@ class TestTelemetry:
|
||||
assert recorder.rows[1]["error"] is None
|
||||
assert recorder.rows[1]["prompt_tokens"] == 0
|
||||
|
||||
async def test_success_row_stays_measured_and_settles_zero(self):
|
||||
"""OCR 的 0 token 是**事实**而非未知(est_tokens 解耦设计 §3.3 剔出决定)。
|
||||
|
||||
三态化不得把 OCR 成功行改成 `unavailable`——那会灌水缺口度量
|
||||
`COUNT(*) WHERE usage_source='unavailable'`;settle 恒 0 的差异①同样不动。
|
||||
"""
|
||||
recorder = _MemoryRecorder()
|
||||
client, limiter, _ = _client([_src(tpm=1000, est_tokens=400)], ["text"], telemetry=recorder)
|
||||
await client.recognize_text(b"jpg")
|
||||
row = recorder.rows[0]
|
||||
assert row["usage_source"] == "measured"
|
||||
assert row["prompt_tokens"] == 0 and row["completion_tokens"] == 0
|
||||
assert row["error"] is None
|
||||
assert (await limiter.source_stats("m1")).tpm_used == 0 # settle(0) 全额退回预扣
|
||||
|
||||
|
||||
class TestAssembly:
|
||||
_ENV = {
|
||||
|
||||
@@ -13,12 +13,14 @@ from polygateway.errors import (
|
||||
SourceDeadError,
|
||||
TransientError,
|
||||
)
|
||||
from polygateway.middleware.telemetry import TelemetryEmitter
|
||||
from polygateway.pricing import ModelPrice, PricingTable
|
||||
from polygateway.transports.openai_compat import (
|
||||
OpenAICompatTransport,
|
||||
_iter_sse_deltas,
|
||||
_sse_data_payload,
|
||||
)
|
||||
from polygateway.types import SourceConfig
|
||||
from polygateway.types import ChatRequest, LLMResponse, SourceConfig
|
||||
|
||||
|
||||
def _source(**overrides):
|
||||
@@ -34,7 +36,7 @@ def _source(**overrides):
|
||||
return SourceConfig(**base)
|
||||
|
||||
|
||||
def _chunk(content=None, reasoning=None, usage=None):
|
||||
def _chunk(content=None, reasoning=None, usage=None, model=None):
|
||||
delta = {}
|
||||
if content is not None:
|
||||
delta["content"] = content
|
||||
@@ -43,6 +45,8 @@ def _chunk(content=None, reasoning=None, usage=None):
|
||||
body = {"choices": [{"delta": delta}]} if (delta or usage is None) else {"choices": []}
|
||||
if usage is not None:
|
||||
body["usage"] = usage
|
||||
if model is not None:
|
||||
body["model"] = model
|
||||
return f"data: {json.dumps(body)}\n\n"
|
||||
|
||||
|
||||
@@ -71,6 +75,53 @@ async def _complete(transport, source, *, stream=True, overlay=None):
|
||||
)
|
||||
|
||||
|
||||
# 单价刻意取"输出贵于输入"的真实形态: est_tokens 兜底把整估值塞进 completion
|
||||
# 时,虚高才显形(设计 §1 的 26 倍算例即此单价)。
|
||||
_PRICING = PricingTable({"qwen-max": ModelPrice(input_per_1m=1.0, output_per_1m=8.0)})
|
||||
|
||||
|
||||
class _MemoryRecorder:
|
||||
def __init__(self):
|
||||
self.rows = []
|
||||
|
||||
async def record_llm_call(self, **fields):
|
||||
self.rows.append(fields)
|
||||
|
||||
|
||||
async def _recorded_cost(result, source):
|
||||
"""把 transport 产物走一遍真实计费路径,返回落库的 cost。
|
||||
|
||||
`unavailable` → cost=None 的判定在 `TelemetryEmitter` 里(设计 §3.2 #5),
|
||||
直接调 `PricingTable.cost` 对 `0/0` 只会得到 `0.0`——那正是本组用例要防的
|
||||
假金额,故断言必须穿过 emitter 而不是单测 pricing。
|
||||
"""
|
||||
recorder = _MemoryRecorder()
|
||||
response = LLMResponse(
|
||||
content=result.content,
|
||||
thinking=result.thinking,
|
||||
model=source.model,
|
||||
provider=source.provider,
|
||||
prompt_tokens=result.prompt_tokens,
|
||||
completion_tokens=result.completion_tokens,
|
||||
latency_ms=1,
|
||||
ttft_ms=result.ttft_ms,
|
||||
max_inter_token_ms=result.max_inter_token_ms,
|
||||
cache_hit=False,
|
||||
call_id="cid-1",
|
||||
source_name=source.name,
|
||||
usage_source=result.usage_source,
|
||||
)
|
||||
await TelemetryEmitter(recorder, pricing=_PRICING).emit_attempt(
|
||||
request=ChatRequest(messages=[{"role": "user", "content": "hi"}]),
|
||||
source=source,
|
||||
call_id="cid-1",
|
||||
latency_ms=1,
|
||||
response=response,
|
||||
error=None,
|
||||
)
|
||||
return recorder.rows[0]["cost"]
|
||||
|
||||
|
||||
class TestSsePureFunctions:
|
||||
def test_data_payload_filters_noise(self):
|
||||
assert _sse_data_payload("") is None
|
||||
@@ -129,13 +180,21 @@ class TestStreamHappyPath:
|
||||
assert result.content == "answer"
|
||||
assert result.thinking == "hmm"
|
||||
|
||||
async def test_usage_missing_falls_back_to_est(self):
|
||||
async def test_usage_missing_is_unavailable_with_null_cost(self):
|
||||
"""usage 帧缺失 → 0/0 + unavailable + cost NULL(设计 §3.2 #3)。
|
||||
|
||||
改前拿 `est_tokens` 当实测并整估值塞 completion,同一条调用记成
|
||||
`0/4000` → cost 0.032(设计 §1 的 26 倍虚高)。
|
||||
"""
|
||||
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="ok"))
|
||||
|
||||
result = await _complete(_transport_for(handler), _source(tpm=1000, est_tokens=333))
|
||||
assert result.usage_source == "estimated"
|
||||
assert result.prompt_tokens == 0 and result.completion_tokens == 333
|
||||
source = _source(tpm=1000, est_tokens=4000)
|
||||
result = await _complete(_transport_for(handler), source)
|
||||
assert result.usage_source == "unavailable"
|
||||
assert result.prompt_tokens == 0 and result.completion_tokens == 0
|
||||
assert await _recorded_cost(result, source) is None
|
||||
|
||||
|
||||
class TestMissingDoneSemantics:
|
||||
@@ -146,12 +205,28 @@ class TestMissingDoneSemantics:
|
||||
with pytest.raises(TransientError, match="missing_done|truncated"):
|
||||
await _complete(_transport_for(self._no_done_handler), _source())
|
||||
|
||||
async def test_salvage_policy_keeps_content_as_estimated(self):
|
||||
result = await _complete(
|
||||
_transport_for(self._no_done_handler), _source(missing_done="salvage")
|
||||
)
|
||||
async def test_salvage_with_usage_frame_degrades_to_estimated(self):
|
||||
"""打捞且收到 usage 帧: 数字真实、可信度降级 → estimated 且照常计费。"""
|
||||
source = _source(missing_done="salvage", tpm=1000, est_tokens=4000)
|
||||
result = await _complete(_transport_for(self._no_done_handler), source)
|
||||
assert result.content == "partial"
|
||||
assert result.usage_source == "estimated" # 打捞路径强制 estimated
|
||||
assert result.usage_source == "estimated"
|
||||
assert result.prompt_tokens == 11 and result.completion_tokens == 7
|
||||
assert await _recorded_cost(result, source) == pytest.approx(
|
||||
11 / 1_000_000 * 1.0 + 7 / 1_000_000 * 8.0
|
||||
)
|
||||
|
||||
async def test_salvage_without_usage_frame_stays_unavailable(self):
|
||||
"""打捞且 usage 帧缺失: 0/0 不得被洗成 estimated,否则算出假的 0.0(设计 §3.2 #4)。"""
|
||||
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="partial"), done=False)
|
||||
|
||||
source = _source(missing_done="salvage", tpm=1000, est_tokens=4000)
|
||||
result = await _complete(_transport_for(handler), source)
|
||||
assert result.content == "partial"
|
||||
assert result.usage_source == "unavailable"
|
||||
assert await _recorded_cost(result, source) is None
|
||||
|
||||
async def test_early_eof_always_transient_even_under_salvage(self):
|
||||
def handler(request):
|
||||
@@ -188,6 +263,137 @@ class TestEmptyCompletion:
|
||||
await _complete(_transport_for(handler), _source())
|
||||
|
||||
|
||||
class TestObservabilityFields:
|
||||
"""issue #3: 供应商 prompt cache 命中数与 API 实际返回的模型版本串。
|
||||
|
||||
网关报文一律不可信: 形态异常只归 None,绝不因一个可观测字段打断调用。
|
||||
"""
|
||||
|
||||
def _cached_usage(self, cached):
|
||||
return {**_USAGE, "prompt_tokens_details": {"cached_tokens": cached}}
|
||||
|
||||
async def test_stream_reads_cached_tokens(self):
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._cached_usage(128)))
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert result.cached_prompt_tokens == 128
|
||||
|
||||
async def test_non_stream_reads_cached_tokens(self):
|
||||
def handler(request):
|
||||
return httpx.Response(
|
||||
200,
|
||||
json={
|
||||
"choices": [{"message": {"content": "42"}}],
|
||||
"usage": self._cached_usage(128),
|
||||
},
|
||||
)
|
||||
|
||||
result = await _complete(_transport_for(handler), _source(), stream=False)
|
||||
assert result.cached_prompt_tokens == 128
|
||||
|
||||
async def test_zero_cached_tokens_is_a_real_zero(self):
|
||||
"""0(真实零命中)与 None(该源未上报)必须可区分——issue #3 的核心诉求。"""
|
||||
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._cached_usage(0)))
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert result.cached_prompt_tokens == 0
|
||||
|
||||
async def test_usage_without_details_is_none(self):
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="ok"), _chunk(usage=_USAGE))
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert result.cached_prompt_tokens is None
|
||||
|
||||
async def test_missing_usage_frame_is_none(self):
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="ok"))
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert result.cached_prompt_tokens is None
|
||||
|
||||
@pytest.mark.parametrize("bad", ["abc", -1, True, 1.5, None, [], {"x": 1}])
|
||||
async def test_malformed_cached_tokens_degrade_to_none(self, bad):
|
||||
"""`True` 必须排除: Python 里 isinstance(True, int) 为真。"""
|
||||
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._cached_usage(bad)))
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert result.cached_prompt_tokens is None
|
||||
|
||||
async def test_details_not_a_dict_is_none(self):
|
||||
def handler(request):
|
||||
usage = {**_USAGE, "prompt_tokens_details": "oops"}
|
||||
return _sse_stream(_chunk(content="ok"), _chunk(usage=usage))
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert result.cached_prompt_tokens is None
|
||||
|
||||
async def test_non_stream_reads_reported_model(self):
|
||||
def handler(request):
|
||||
return httpx.Response(
|
||||
200,
|
||||
json={
|
||||
"choices": [{"message": {"content": "42"}}],
|
||||
"usage": _USAGE,
|
||||
"model": "MiniMax-Text-01-250321",
|
||||
},
|
||||
)
|
||||
|
||||
result = await _complete(_transport_for(handler), _source(), stream=False)
|
||||
assert result.model_reported == "MiniMax-Text-01-250321"
|
||||
|
||||
async def test_stream_keeps_the_first_reported_model(self):
|
||||
"""末帧异常值不得覆盖首帧: 首次写入即固定。"""
|
||||
|
||||
def handler(request):
|
||||
return _sse_stream(
|
||||
_chunk(content="a", model="MiniMax-Text-01-250321"),
|
||||
_chunk(content="b", model="something-else"),
|
||||
_chunk(usage=_USAGE),
|
||||
)
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert result.model_reported == "MiniMax-Text-01-250321"
|
||||
|
||||
async def test_empty_first_model_does_not_block_a_later_real_one(self):
|
||||
"""首帧报空串不得锁死 sink: 守卫按"有效值"判断,否则真实版本会丢。"""
|
||||
|
||||
def handler(request):
|
||||
return _sse_stream(
|
||||
_chunk(content="a", model=""),
|
||||
_chunk(content="b", model="MiniMax-Text-01-250321"),
|
||||
_chunk(usage=_USAGE),
|
||||
)
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert result.model_reported == "MiniMax-Text-01-250321"
|
||||
|
||||
@pytest.mark.parametrize("bad", [None, "", " ", 123, {}])
|
||||
async def test_missing_or_malformed_model_is_none(self, bad):
|
||||
def handler(request):
|
||||
body = {"choices": [{"message": {"content": "42"}}], "usage": _USAGE}
|
||||
if bad is not None:
|
||||
body["model"] = bad
|
||||
return httpx.Response(200, json=body)
|
||||
|
||||
result = await _complete(_transport_for(handler), _source(), stream=False)
|
||||
assert result.model_reported is None
|
||||
|
||||
async def test_raw_payload_is_unchanged(self):
|
||||
"""新字段是独立格子,不改动 raw 的既有内容。"""
|
||||
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="ok"), _chunk(usage=self._cached_usage(5)))
|
||||
|
||||
result = await _complete(_transport_for(handler), _source())
|
||||
assert set(result.raw) == {"usage"}
|
||||
|
||||
|
||||
class TestNonStreamFastPath:
|
||||
async def test_non_stream_parses_message(self):
|
||||
def handler(request):
|
||||
@@ -239,6 +445,27 @@ class TestRequestShaping:
|
||||
)
|
||||
assert seen["response_format"] == {"type": "json_object"}
|
||||
|
||||
async def test_extra_body_merged_and_outranked_by_overlay(self):
|
||||
"""顺序即优先级: thinking profile → extra_body → overlay(issue #4)。"""
|
||||
seen = {}
|
||||
|
||||
def handler(request):
|
||||
seen.update(json.loads(request.content))
|
||||
return _sse_stream(_chunk(content="x"), _chunk(usage=_USAGE))
|
||||
|
||||
await _complete(
|
||||
_transport_for(handler),
|
||||
_source(extra_body={"temperature": 0, "top_p": 0.9}),
|
||||
overlay={"temperature": 1},
|
||||
)
|
||||
assert seen["temperature"] == 1 # 调用级覆盖配置级
|
||||
assert seen["top_p"] == 0.9 # 未被顶掉的配置级键保留
|
||||
|
||||
async def test_extra_body_cannot_break_governed_keys(self):
|
||||
"""治理键由 payload 骨架拥有;extra_body 的保护键在构造期已被拦下。"""
|
||||
with pytest.raises(ValueError, match="stream"):
|
||||
_source(extra_body={"stream": False})
|
||||
|
||||
|
||||
class TestErrorTranslation:
|
||||
@pytest.mark.parametrize(
|
||||
|
||||
@@ -114,6 +114,9 @@ class _DummyRecorder:
|
||||
cache_hit,
|
||||
error,
|
||||
cost,
|
||||
cached_prompt_tokens,
|
||||
model_reported,
|
||||
sampling,
|
||||
) -> None: ...
|
||||
|
||||
|
||||
|
||||
@@ -56,6 +56,74 @@ class TestPricingTable:
|
||||
ModelPrice(input_per_1m=-1.0, output_per_1m=0.0)
|
||||
|
||||
|
||||
class TestCachedInputTier:
|
||||
"""issue #3: 供应商 prompt cache 命中部分按更低单价计费,不配则不猜折扣。"""
|
||||
|
||||
_CACHED = PricingTable(
|
||||
{"m": ModelPrice(input_per_1m=10.0, output_per_1m=20.0, cached_input_per_1m=2.0)}
|
||||
)
|
||||
_PLAIN = PricingTable({"m": ModelPrice(input_per_1m=10.0, output_per_1m=20.0)})
|
||||
|
||||
def test_hit_is_billed_at_the_cached_rate(self):
|
||||
# 1M prompt 中 600k 命中: 400k×10 + 600k×2 = 4.0 + 1.2
|
||||
assert self._CACHED.cost("m", 1_000_000, 0, 600_000) == pytest.approx(5.2)
|
||||
|
||||
def test_without_the_tier_the_result_is_unchanged(self):
|
||||
"""未配缓存档 = 退化为现状全额计价,绝不按经验折扣率猜(P5)。"""
|
||||
full = self._PLAIN.cost("m", 1_000_000, 0)
|
||||
assert self._PLAIN.cost("m", 1_000_000, 0, 600_000) == full == pytest.approx(10.0)
|
||||
|
||||
@pytest.mark.parametrize("cached", [None, 0])
|
||||
def test_no_hit_is_billed_in_full(self, cached):
|
||||
assert self._CACHED.cost("m", 1_000_000, 0, cached) == pytest.approx(10.0)
|
||||
|
||||
def test_negative_cached_is_billed_in_full(self):
|
||||
"""负数命中数不得抬高成本: cost() 是公共方法,外部输入须校验后使用(P5)。"""
|
||||
assert self._CACHED.cost("m", 1_000_000, 0, -500_000) == pytest.approx(10.0)
|
||||
|
||||
def test_cached_over_prompt_is_clamped_and_never_negative(self):
|
||||
"""网关口径异常时按输入总数夹取: 全部按缓存价,不得算出负成本。"""
|
||||
clamped = self._CACHED.cost("m", 1_000_000, 0, 5_000_000)
|
||||
assert clamped == pytest.approx(2.0) and clamped >= 0
|
||||
|
||||
def test_legacy_three_arg_call_still_works(self):
|
||||
"""embedding.py 的三参调用形态必须零改动可用。"""
|
||||
assert self._CACHED.cost("m", 1_000_000, 0) == pytest.approx(10.0)
|
||||
|
||||
def test_from_file_accepts_and_validates_the_tier(self, tmp_path):
|
||||
path = tmp_path / "p.json"
|
||||
path.write_text(
|
||||
json.dumps(
|
||||
{"m": {"input_per_1m": 10.0, "output_per_1m": 20.0, "cached_input_per_1m": 2.0}}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
assert PricingTable.from_file(path).cost("m", 1_000_000, 0, 1_000_000) == pytest.approx(2.0)
|
||||
|
||||
@pytest.mark.parametrize("bad", [-1.0, "x"])
|
||||
def test_from_file_rejects_a_bad_tier(self, tmp_path, bad):
|
||||
path = tmp_path / "bad.json"
|
||||
path.write_text(
|
||||
json.dumps(
|
||||
{"m": {"input_per_1m": 1.0, "output_per_1m": 2.0, "cached_input_per_1m": bad}}
|
||||
),
|
||||
encoding="utf-8",
|
||||
)
|
||||
with pytest.raises(ValueError, match="cached_input_per_1m"):
|
||||
PricingTable.from_file(path)
|
||||
|
||||
def test_legacy_price_file_without_the_tier_still_loads(self, tmp_path):
|
||||
path = tmp_path / "old.json"
|
||||
path.write_text(
|
||||
json.dumps({"m": {"input_per_1m": 1.0, "output_per_1m": 2.0}}), encoding="utf-8"
|
||||
)
|
||||
assert PricingTable.from_file(path).cost("m", 1_000_000, 0, 500_000) == pytest.approx(1.0)
|
||||
|
||||
def test_negative_tier_rejected_on_construction(self):
|
||||
with pytest.raises(ValueError):
|
||||
ModelPrice(input_per_1m=1.0, output_per_1m=1.0, cached_input_per_1m=-0.1)
|
||||
|
||||
|
||||
class _MemoryRecorder:
|
||||
def __init__(self):
|
||||
self.rows = []
|
||||
|
||||
@@ -152,6 +152,76 @@ class TestSuccessPath:
|
||||
# 预扣 400,实际 15 → settle 后窗口只记 15
|
||||
assert (await limiter.source_stats("a")).tpm_used == 15
|
||||
|
||||
@pytest.mark.parametrize("usage_source", ["measured", "estimated"])
|
||||
async def test_settle_uses_measured_sum_when_usage_available(self, usage_source):
|
||||
"""用量可得(含打捞降级的 estimated)时结算恒取实测之和,不落派生兜底分支。"""
|
||||
src = _src("a", tpm=1000, est_tokens=400)
|
||||
result = TransportResult(
|
||||
content="ok",
|
||||
thinking="",
|
||||
prompt_tokens=40,
|
||||
completion_tokens=60,
|
||||
usage_source=usage_source,
|
||||
ttft_ms=12.0,
|
||||
max_inter_token_ms=3.0,
|
||||
raw={},
|
||||
)
|
||||
mw, limiter, *_ = _harness([src], [result])
|
||||
await mw(_REQ)
|
||||
# 预扣 400,实测 40+60 → settle 后窗口记 100(而非派生兜底的 400)
|
||||
assert (await limiter.source_stats("a")).tpm_used == 100
|
||||
|
||||
async def test_settle_keeps_derived_deposit_when_usage_unavailable(self):
|
||||
"""未填 est_tokens + usage 帧缺失的**成功**调用: 押金留存而非整笔退回。
|
||||
|
||||
入场预扣与结算须同取 `effective_est_tokens()`(delta==0),否则对
|
||||
"从不返回 usage 帧"的源等于 TPM 闸进门即放行、出门即清账(设计 §3.2 #9)。
|
||||
"""
|
||||
src = _src("a", tpm=1000, est_tokens=0) # 派生预扣量 = max(1, 1000 // 60) = 16
|
||||
result = TransportResult(
|
||||
content="ok",
|
||||
thinking="",
|
||||
prompt_tokens=0,
|
||||
completion_tokens=0,
|
||||
usage_source="unavailable",
|
||||
ttft_ms=12.0,
|
||||
max_inter_token_ms=3.0,
|
||||
raw={},
|
||||
)
|
||||
mw, limiter, *_ = _harness([src], [result])
|
||||
await mw(_REQ)
|
||||
assert src.effective_est_tokens() == 16
|
||||
assert (await limiter.source_stats("a")).tpm_used == 16
|
||||
|
||||
|
||||
class TestObservabilityPassthrough:
|
||||
"""issue #3: transport 采到的两个可观测字段必须原样上浮到 LLMResponse。"""
|
||||
|
||||
async def test_fields_reach_the_response(self):
|
||||
result = TransportResult(
|
||||
content="ok",
|
||||
thinking="",
|
||||
prompt_tokens=10,
|
||||
completion_tokens=5,
|
||||
usage_source="measured",
|
||||
ttft_ms=12.0,
|
||||
max_inter_token_ms=3.0,
|
||||
raw={},
|
||||
cached_prompt_tokens=64,
|
||||
model_reported="MiniMax-Text-01-250321",
|
||||
)
|
||||
mw, *_ = _harness([_src("a")], [result])
|
||||
resp = await mw(_REQ)
|
||||
assert resp.cached_prompt_tokens == 64
|
||||
assert resp.model_reported == "MiniMax-Text-01-250321"
|
||||
# model 仍是配置别名: 真实版本是旁证,不顶替溯源主字段
|
||||
assert resp.model == "m"
|
||||
|
||||
async def test_absent_fields_stay_none(self):
|
||||
mw, *_ = _harness([_src("a")], [_ok()])
|
||||
resp = await mw(_REQ)
|
||||
assert resp.cached_prompt_tokens is None and resp.model_reported is None
|
||||
|
||||
|
||||
class TestRetryAndFailover:
|
||||
async def test_transient_switches_source_then_succeeds(self):
|
||||
@@ -202,6 +272,18 @@ class TestRetryAndFailover:
|
||||
assert sleep.delays == [] # 源死亡不退避
|
||||
assert not (await gate.try_enter("a", "w")).allowed # a 已 force_open
|
||||
|
||||
async def test_transient_failure_keeps_derived_deposit(self):
|
||||
"""未填 est_tokens 的**非 dead 瞬时失败**同样按派生预扣量保守结算。
|
||||
|
||||
失败请求可能已被网关计费,退掉押金会低估用量(设计 §3.2 #8);
|
||||
max_attempts=1 保证恰一次尝试,窗口残留量即单次预扣量。
|
||||
"""
|
||||
src = _src("a", tpm=1000, est_tokens=0) # 派生预扣量 = 16
|
||||
mw, limiter, *_ = _harness([src], [TransientError("boom")], max_attempts=1)
|
||||
with pytest.raises(AllSourcesExhausted):
|
||||
await mw(_REQ)
|
||||
assert (await limiter.source_stats("a")).tpm_used == 16
|
||||
|
||||
|
||||
class TestNonRetryableOutcomes:
|
||||
async def test_request_rejected_propagates_without_retry(self):
|
||||
|
||||
@@ -177,3 +177,55 @@ class TestNativeOverlayFirstAttempt:
|
||||
resp = await _mw()(ChatRequest(messages=_MSGS, structured=Verdict), terminal)
|
||||
with pytest.raises(dataclasses.FrozenInstanceError):
|
||||
resp.structured_data = None
|
||||
|
||||
|
||||
class TestSamplingSnapshotInvariant:
|
||||
"""地基不变式: `sampling` 跨洋葱层恒定,`overlay` 会被结构化注入(issue #4)。
|
||||
|
||||
缓存 key(决策 C)与三个遥测入口(决策 D)都建立在这条之上,而它此前只靠
|
||||
"dataclasses.replace 恰好保留未提及字段"的约定成立,无任何机械执法。
|
||||
这个测试是那份执法——它红了就意味着两个决策同时失效。
|
||||
"""
|
||||
|
||||
async def test_sampling_survives_feedback_ladder_while_overlay_diverges(self):
|
||||
caller_sampling = {"temperature": 0, "seed": 42}
|
||||
# 先坏后好,强制走一次带反馈重问(重问会 replace messages)
|
||||
terminal = ScriptedTerminal(["not json at all", '{"answer": 1, "reason": "r"}'])
|
||||
mw = _mw(strategy=NativeSchemaStrategy(), max_retries=1)
|
||||
await mw(
|
||||
# overlay 与 sampling 传**同一个对象**,复现 client.py 的别名关系
|
||||
# ——否则中间件就地改写 overlay 时不会波及 sampling,这条执法就是空的
|
||||
ChatRequest(
|
||||
messages=_MSGS,
|
||||
structured=Verdict,
|
||||
overlay=caller_sampling,
|
||||
sampling=caller_sampling,
|
||||
),
|
||||
terminal,
|
||||
)
|
||||
assert len(terminal.requests) == 2 # 确实重问过
|
||||
for seen in terminal.requests:
|
||||
# ① 跨层恒定: 每次尝试看到的 sampling 与调用方传入的逐字相同
|
||||
assert seen.sampling == caller_sampling
|
||||
# ② 确实分叉: 同一时刻 overlay 已被注入 response_format
|
||||
assert seen.overlay["response_format"]["type"] == "json_schema"
|
||||
assert "response_format" not in seen.sampling
|
||||
|
||||
async def test_middleware_does_not_mutate_caller_mapping(self):
|
||||
"""决策 E 的第二条约束: 中间件只能 replace 派生,不得就地改这两个 dict。
|
||||
|
||||
同样传同一对象: 生产中 overlay 与 sampling 是别名,任何对 overlay 的
|
||||
就地改写都会同步毒化缓存 key 与遥测列。
|
||||
"""
|
||||
caller_sampling = {"seed": 7}
|
||||
terminal = ScriptedTerminal(['{"answer": 1, "reason": "r"}'])
|
||||
await _mw(strategy=NativeSchemaStrategy())(
|
||||
ChatRequest(
|
||||
messages=_MSGS,
|
||||
structured=Verdict,
|
||||
overlay=caller_sampling,
|
||||
sampling=caller_sampling,
|
||||
),
|
||||
terminal,
|
||||
)
|
||||
assert caller_sampling == {"seed": 7} # 调用方的对象未被污染
|
||||
|
||||
+473
-11
@@ -1,6 +1,7 @@
|
||||
"""遥测子系统测试: SQLiteRecorder(18 列)+ TelemetryEmitter 单一 helper + TelemetryMW。"""
|
||||
"""遥测子系统测试: SQLiteRecorder(21 列)+ TelemetryEmitter 单一 helper + TelemetryMW。"""
|
||||
|
||||
import asyncio
|
||||
import json
|
||||
import sqlite3
|
||||
import subprocess
|
||||
from pathlib import Path
|
||||
@@ -9,6 +10,7 @@ import pytest
|
||||
|
||||
from polygateway.errors import CircuitOpenError, RequestRejectedError
|
||||
from polygateway.middleware.telemetry import TelemetryEmitter, TelemetryMW
|
||||
from polygateway.pricing import ModelPrice, PricingTable
|
||||
from polygateway.telemetry.sqlite import SQLiteRecorder
|
||||
from polygateway.types import ChatRequest, LLMResponse, SourceConfig
|
||||
|
||||
@@ -34,6 +36,9 @@ _EXPECTED_COLUMNS = [
|
||||
"error",
|
||||
"cost",
|
||||
"created_at",
|
||||
"cached_prompt_tokens",
|
||||
"model_reported",
|
||||
"sampling",
|
||||
]
|
||||
|
||||
|
||||
@@ -57,15 +62,21 @@ def _resp(**overrides):
|
||||
return LLMResponse(**base)
|
||||
|
||||
|
||||
def _source():
|
||||
return SourceConfig(
|
||||
name="s1",
|
||||
provider="p",
|
||||
base_url="https://gw.example/v1",
|
||||
api_key="sk",
|
||||
model="m",
|
||||
timeout_s=10.0,
|
||||
)
|
||||
def _source(**overrides):
|
||||
base = {
|
||||
"name": "s1",
|
||||
"provider": "p",
|
||||
"base_url": "https://gw.example/v1",
|
||||
"api_key": "sk",
|
||||
"model": "m",
|
||||
"timeout_s": 10.0,
|
||||
}
|
||||
base.update(overrides)
|
||||
return SourceConfig(**base)
|
||||
|
||||
|
||||
# 输出单价 8 元/百万: 改前 `unavailable` 行按兜底的 0/4000 换算恰好是 0.032
|
||||
_PRICING = PricingTable({"m": ModelPrice(input_per_1m=1.0, output_per_1m=8.0)})
|
||||
|
||||
|
||||
async def _record_minimal(recorder, call_id="c1", **overrides):
|
||||
@@ -88,6 +99,9 @@ async def _record_minimal(recorder, call_id="c1", **overrides):
|
||||
"cache_hit": False,
|
||||
"error": None,
|
||||
"cost": None,
|
||||
"cached_prompt_tokens": None,
|
||||
"model_reported": None,
|
||||
"sampling": None,
|
||||
}
|
||||
fields.update(overrides)
|
||||
await recorder.record_llm_call(**fields)
|
||||
@@ -129,6 +143,181 @@ class TestSQLiteRecorder:
|
||||
await _record_minimal(recorder) # 不抛
|
||||
recorder.close()
|
||||
|
||||
async def test_observability_columns_round_trip(self, tmp_path):
|
||||
recorder = SQLiteRecorder(tmp_path / "t.db")
|
||||
await _record_minimal(recorder, call_id="c-hit", cached_prompt_tokens=64)
|
||||
await _record_minimal(recorder, call_id="c-zero", cached_prompt_tokens=0)
|
||||
await _record_minimal(recorder, call_id="c-none", model_reported="MiniMax-Text-01")
|
||||
recorder.close()
|
||||
rows = dict(
|
||||
sqlite3.connect(tmp_path / "t.db")
|
||||
.execute("SELECT call_id, cached_prompt_tokens FROM llm_calls")
|
||||
.fetchall()
|
||||
)
|
||||
assert rows["c-hit"] == 64
|
||||
assert rows["c-zero"] == 0 # 真实零命中,读回仍是 0 而非 NULL
|
||||
assert rows["c-none"] is None
|
||||
|
||||
async def test_sampling_column_round_trips(self, tmp_path):
|
||||
"""issue #4: 采样参数落库,否则事后无法证明某批数据跑在什么温度下。"""
|
||||
recorder = SQLiteRecorder(tmp_path / "t.db")
|
||||
await _record_minimal(recorder, call_id="c-s", sampling='{"seed": 42, "temperature": 0}')
|
||||
await _record_minimal(recorder, call_id="c-plain")
|
||||
recorder.close()
|
||||
rows = dict(
|
||||
sqlite3.connect(tmp_path / "t.db")
|
||||
.execute("SELECT call_id, sampling FROM llm_calls")
|
||||
.fetchall()
|
||||
)
|
||||
assert json.loads(rows["c-s"]) == {"seed": 42, "temperature": 0}
|
||||
assert rows["c-plain"] is None # 无采样参数为 NULL,便于 SQL 过滤
|
||||
|
||||
|
||||
class TestSQLiteColumnBackfill:
|
||||
"""issue #3: 已存在的 18 列旧表必须自动补列,否则每行写入都被丢弃。"""
|
||||
|
||||
_LEGACY_DDL = """
|
||||
CREATE TABLE llm_calls (
|
||||
call_id TEXT PRIMARY KEY,
|
||||
parent_call_id TEXT,
|
||||
session_id TEXT,
|
||||
model TEXT NOT NULL,
|
||||
provider TEXT NOT NULL,
|
||||
source_name TEXT NOT NULL,
|
||||
messages TEXT NOT NULL,
|
||||
response TEXT NOT NULL,
|
||||
thinking TEXT NOT NULL DEFAULT '',
|
||||
prompt_tokens INTEGER NOT NULL,
|
||||
completion_tokens INTEGER NOT NULL,
|
||||
usage_source TEXT NOT NULL,
|
||||
latency_ms INTEGER NOT NULL,
|
||||
ttft_ms REAL,
|
||||
max_inter_token_ms REAL,
|
||||
cache_hit INTEGER NOT NULL DEFAULT 0,
|
||||
error TEXT,
|
||||
cost REAL,
|
||||
created_at TEXT NOT NULL DEFAULT (datetime('now'))
|
||||
);
|
||||
"""
|
||||
|
||||
async def test_legacy_table_is_upgraded_in_place(self, tmp_path):
|
||||
db = tmp_path / "legacy.db"
|
||||
legacy = sqlite3.connect(db)
|
||||
legacy.execute(self._LEGACY_DDL)
|
||||
legacy.commit()
|
||||
legacy.close()
|
||||
|
||||
recorder = SQLiteRecorder(db)
|
||||
await _record_minimal(recorder, cached_prompt_tokens=7, model_reported="m-real")
|
||||
recorder.close()
|
||||
|
||||
conn = sqlite3.connect(db)
|
||||
cols = [r[1] for r in conn.execute("PRAGMA table_info(llm_calls)")]
|
||||
assert cols == _EXPECTED_COLUMNS # ALTER 追加到末尾,与新建库列序一致
|
||||
assert conn.execute(
|
||||
"SELECT cached_prompt_tokens, model_reported FROM llm_calls"
|
||||
).fetchone() == (7, "m-real")
|
||||
|
||||
async def test_backfill_failure_keeps_the_recorder_usable(self, tmp_path):
|
||||
"""补列失败只能逐行降级,绝不能把 recorder 整体变成 no-op(设计 D1 纪律)。
|
||||
|
||||
把 llm_calls 建成同名 view: `CREATE TABLE IF NOT EXISTS` 遇 view 静默
|
||||
no-op(不抛),随后的 ALTER 才抛 "Cannot add a column to a view"——正是
|
||||
补列失败这条分支。`_conn` 必须保持非 None,否则整个 recorder 永久失能。
|
||||
"""
|
||||
db = tmp_path / "view.db"
|
||||
conn = sqlite3.connect(db)
|
||||
conn.execute("CREATE TABLE real_rows (call_id TEXT)")
|
||||
conn.execute("CREATE VIEW llm_calls AS SELECT call_id FROM real_rows")
|
||||
conn.commit()
|
||||
conn.close()
|
||||
|
||||
recorder = SQLiteRecorder(db) # 不得抛
|
||||
assert recorder._conn is not None # 补列失败 ≠ recorder 失能(D1 纪律)
|
||||
await _record_minimal(recorder) # 不得抛
|
||||
recorder.close()
|
||||
|
||||
|
||||
class _FakePgConn:
|
||||
"""记录执行过的语句;可让 ALTER 抛错以模拟权限不足。"""
|
||||
|
||||
def __init__(self, existing: list[str], *, fail_alter: bool = False):
|
||||
self.existing = existing
|
||||
self.fail_alter = fail_alter
|
||||
self.statements: list[str] = []
|
||||
|
||||
async def execute(self, sql, *args):
|
||||
self.statements.append(sql)
|
||||
if sql.startswith("ALTER TABLE") and self.fail_alter:
|
||||
raise RuntimeError("must be owner of table llm_calls")
|
||||
|
||||
async def fetch(self, sql, *args):
|
||||
self.statements.append(sql)
|
||||
return [{"attname": name} for name in self.existing]
|
||||
|
||||
|
||||
class _FakePgPool:
|
||||
def __init__(self, conn):
|
||||
self._conn = conn
|
||||
|
||||
def acquire(self):
|
||||
conn = self._conn
|
||||
|
||||
class _Ctx:
|
||||
async def __aenter__(self):
|
||||
return conn
|
||||
|
||||
async def __aexit__(self, *exc):
|
||||
return False
|
||||
|
||||
return _Ctx()
|
||||
|
||||
|
||||
class TestPostgresBackfillDiscipline:
|
||||
"""PG 补列必须与 SQLite 侧对称: 失败只逐行降级,且稳态不抢排他锁(issue #3)。"""
|
||||
|
||||
_LEGACY = ["call_id", "cost", "created_at"]
|
||||
_CURRENT = [
|
||||
"call_id",
|
||||
"cost",
|
||||
"created_at",
|
||||
"cached_prompt_tokens",
|
||||
"model_reported",
|
||||
"sampling",
|
||||
]
|
||||
|
||||
def _recorder(self, conn):
|
||||
from polygateway.telemetry.postgres import PostgresRecorder
|
||||
|
||||
return PostgresRecorder("postgresql://u:p@h:5432/polygateway", pool=_FakePgPool(conn))
|
||||
|
||||
async def test_alter_failure_does_not_disable_the_recorder(self):
|
||||
"""ALTER 失败(如账号只有 INSERT 权限)不得置 _failed —— 那会让遥测全灭。"""
|
||||
conn = _FakePgConn(self._LEGACY, fail_alter=True)
|
||||
recorder = self._recorder(conn)
|
||||
await _record_minimal(recorder) # 不得抛
|
||||
assert recorder._failed is False
|
||||
assert any(s.startswith("INSERT INTO llm_calls") for s in conn.statements)
|
||||
|
||||
async def test_no_alter_when_columns_already_exist(self):
|
||||
"""ADD COLUMN IF NOT EXISTS 即使列已存在也会先抢 ACCESS EXCLUSIVE 锁,
|
||||
|
||||
而遥测是内联 await——稳态下必须一条 ALTER 都不发,否则每个进程的首次
|
||||
写入都会去锁共享审计表。
|
||||
"""
|
||||
conn = _FakePgConn(self._CURRENT)
|
||||
await _record_minimal(self._recorder(conn))
|
||||
assert not [s for s in conn.statements if s.startswith("ALTER TABLE")]
|
||||
|
||||
async def test_missing_columns_are_added_once(self):
|
||||
conn = _FakePgConn(self._LEGACY)
|
||||
await _record_minimal(self._recorder(conn))
|
||||
from polygateway.telemetry.postgres import _BACKFILL
|
||||
|
||||
altered = [s for s in conn.statements if s.startswith("ALTER TABLE")]
|
||||
assert len(altered) == len(_BACKFILL) # 旧表缺全部补列,故一列一条 ALTER
|
||||
assert all("IF NOT EXISTS" not in s for s in altered) # 探测已确认缺列,无需再判
|
||||
|
||||
|
||||
class _MemoryRecorder:
|
||||
def __init__(self):
|
||||
@@ -138,6 +327,228 @@ class _MemoryRecorder:
|
||||
self.rows.append(fields)
|
||||
|
||||
|
||||
class TestEmitterRecorderContract:
|
||||
"""emitter 的实参键集合必须与两个后端的 _COLUMNS 完全一致(issue #3)。
|
||||
|
||||
两个后端的 `row = tuple(fields[col] for col in _COLUMNS)` 都在 try **之外**,
|
||||
emitter 漏传一个键就抛 KeyError,被 `_record` 的 except Exception 吞成 warning
|
||||
→ 遥测静默全丢。而 8 个 `**fields` 形态的 fake 一个都拦不住,故显式断言。
|
||||
"""
|
||||
|
||||
async def test_emitter_supplies_exactly_the_backend_columns(self):
|
||||
from polygateway.telemetry.postgres import _COLUMNS as PG_COLUMNS
|
||||
from polygateway.telemetry.sqlite import _COLUMNS as SQLITE_COLUMNS
|
||||
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="cid-1",
|
||||
latency_ms=42,
|
||||
response=_resp(),
|
||||
error=None,
|
||||
)
|
||||
assert set(rec.rows[0]) == set(SQLITE_COLUMNS) == set(PG_COLUMNS)
|
||||
|
||||
@pytest.mark.parametrize("emit", ["attempt", "cache_hit", "terminal_failure"])
|
||||
async def test_every_entry_point_supplies_the_same_keys(self, emit):
|
||||
from polygateway.telemetry.sqlite import _COLUMNS as SQLITE_COLUMNS
|
||||
|
||||
rec = _MemoryRecorder()
|
||||
emitter = TelemetryEmitter(rec)
|
||||
if emit == "attempt":
|
||||
await emitter.emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="c",
|
||||
latency_ms=1,
|
||||
response=None,
|
||||
error="boom",
|
||||
)
|
||||
elif emit == "cache_hit":
|
||||
await emitter.emit_cache_hit(request=_REQ, response=_resp())
|
||||
else:
|
||||
await emitter.emit_terminal_failure(
|
||||
request=_REQ, call_id="c", latency_ms=1, error="dead"
|
||||
)
|
||||
assert set(rec.rows[0]) == set(SQLITE_COLUMNS)
|
||||
|
||||
|
||||
class TestEmitterObservabilityFields:
|
||||
"""issue #3: 三个入口各自的取值口径(设计 §5 表)。"""
|
||||
|
||||
async def test_attempt_carries_the_response_values(self):
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="cid-1",
|
||||
latency_ms=42,
|
||||
response=_resp(cached_prompt_tokens=64, model_reported="m-real"),
|
||||
error=None,
|
||||
)
|
||||
assert rec.rows[0]["cached_prompt_tokens"] == 64
|
||||
assert rec.rows[0]["model_reported"] == "m-real"
|
||||
|
||||
async def test_failed_attempt_has_no_provider_facts(self):
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="cid-2",
|
||||
latency_ms=7,
|
||||
response=None,
|
||||
error="boom",
|
||||
)
|
||||
assert rec.rows[0]["cached_prompt_tokens"] is None
|
||||
assert rec.rows[0]["model_reported"] is None
|
||||
|
||||
async def test_cache_hit_replays_the_recorded_values(self):
|
||||
"""决策 B1: 命中行原样回放,故命中率统计必须带 WHERE cache_hit = false。"""
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec).emit_cache_hit(
|
||||
request=_REQ, response=_resp(cached_prompt_tokens=64, model_reported="m-real")
|
||||
)
|
||||
row = rec.rows[0]
|
||||
assert row["cache_hit"] is True
|
||||
assert row["cached_prompt_tokens"] == 64 and row["model_reported"] == "m-real"
|
||||
|
||||
async def test_terminal_failure_records_none(self):
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec).emit_terminal_failure(
|
||||
request=_REQ, call_id="c", latency_ms=1, error="dead"
|
||||
)
|
||||
assert rec.rows[0]["cached_prompt_tokens"] is None
|
||||
assert rec.rows[0]["model_reported"] is None
|
||||
|
||||
|
||||
class TestEmitterSamplingColumn:
|
||||
"""issue #4: sampling 列在三个入口的口径(设计决策 D 表格)。
|
||||
|
||||
列语义 = 「调用方采样意图 ⊎ 生效源 extra_body」,**不含**结构化注入的
|
||||
response_format(列名是采样参数,schema 不是;且数 KB schema 逐行落库会让
|
||||
审计表无谓膨胀)。三入口若各读各的层,同一列在不同行含义就不同。
|
||||
"""
|
||||
|
||||
_SAMPLED = ChatRequest(
|
||||
messages=[{"role": "user", "content": "hi"}],
|
||||
sampling={"seed": 42},
|
||||
overlay={"seed": 42, "response_format": {"type": "json_object"}},
|
||||
)
|
||||
|
||||
async def test_attempt_merges_source_extra_body(self):
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec).emit_attempt(
|
||||
request=self._SAMPLED,
|
||||
source=_source(extra_body={"temperature": 0}),
|
||||
call_id="c",
|
||||
latency_ms=1,
|
||||
response=_resp(),
|
||||
error=None,
|
||||
)
|
||||
assert json.loads(rec.rows[0]["sampling"]) == {"seed": 42, "temperature": 0}
|
||||
|
||||
async def test_response_format_never_leaks_into_the_column(self):
|
||||
"""三行都不得出现 response_format——它不是采样参数。"""
|
||||
rec = _MemoryRecorder()
|
||||
emitter = TelemetryEmitter(rec)
|
||||
await emitter.emit_attempt(
|
||||
request=self._SAMPLED,
|
||||
source=_source(),
|
||||
call_id="c",
|
||||
latency_ms=1,
|
||||
response=_resp(),
|
||||
error=None,
|
||||
)
|
||||
await emitter.emit_cache_hit(request=self._SAMPLED, response=_resp())
|
||||
await emitter.emit_terminal_failure(
|
||||
request=self._SAMPLED, call_id="c", latency_ms=1, error="dead"
|
||||
)
|
||||
assert len(rec.rows) == 3
|
||||
for row in rec.rows:
|
||||
assert "response_format" not in row["sampling"]
|
||||
|
||||
@pytest.mark.parametrize("emit", ["cache_hit", "terminal_failure"])
|
||||
async def test_sourceless_entries_record_call_level_only(self, emit):
|
||||
"""两个最外层入口没有"生效源"可言,与 model/source_name 置空同一先例。"""
|
||||
rec = _MemoryRecorder()
|
||||
emitter = TelemetryEmitter(rec)
|
||||
if emit == "cache_hit":
|
||||
await emitter.emit_cache_hit(request=self._SAMPLED, response=_resp())
|
||||
else:
|
||||
await emitter.emit_terminal_failure(
|
||||
request=self._SAMPLED, call_id="c", latency_ms=1, error="dead"
|
||||
)
|
||||
assert json.loads(rec.rows[0]["sampling"]) == {"seed": 42}
|
||||
|
||||
async def test_absent_sampling_is_null(self):
|
||||
"""无采样参数时为 NULL,而非空字符串或 "{}"——便于 SQL 过滤。"""
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="c",
|
||||
latency_ms=1,
|
||||
response=_resp(),
|
||||
error=None,
|
||||
)
|
||||
assert rec.rows[0]["sampling"] is None
|
||||
|
||||
|
||||
class TestCostWithCachedTier:
|
||||
"""issue #3: 命中部分按缓存单价计费,避免 cost 系统性高估。"""
|
||||
|
||||
_TABLE = PricingTable(
|
||||
{"m": ModelPrice(input_per_1m=10.0, output_per_1m=20.0, cached_input_per_1m=2.0)}
|
||||
)
|
||||
|
||||
async def test_cached_hit_lowers_the_recorded_cost(self):
|
||||
rec = _MemoryRecorder()
|
||||
emitter = TelemetryEmitter(rec, pricing=self._TABLE)
|
||||
full = _resp(prompt_tokens=1_000_000, completion_tokens=0)
|
||||
await emitter.emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="c1",
|
||||
latency_ms=1,
|
||||
response=full,
|
||||
error=None,
|
||||
)
|
||||
await emitter.emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="c2",
|
||||
latency_ms=1,
|
||||
response=_resp(
|
||||
prompt_tokens=1_000_000, completion_tokens=0, cached_prompt_tokens=600_000
|
||||
),
|
||||
error=None,
|
||||
)
|
||||
assert rec.rows[0]["cost"] == pytest.approx(10.0)
|
||||
assert rec.rows[1]["cost"] == pytest.approx(5.2) # 400k×10 + 600k×2
|
||||
|
||||
async def test_cache_hit_row_still_costs_zero(self):
|
||||
"""缓存命中未产生新调用 → cost 恒 0.0,该短路必须排在任何换算之前。"""
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec, pricing=self._TABLE).emit_cache_hit(
|
||||
request=_REQ,
|
||||
response=_resp(prompt_tokens=1_000_000, cached_prompt_tokens=600_000),
|
||||
)
|
||||
assert rec.rows[0]["cost"] == 0.0
|
||||
|
||||
async def test_unavailable_usage_still_costs_none(self):
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec, pricing=self._TABLE).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="c",
|
||||
latency_ms=1,
|
||||
response=_resp(usage_source="unavailable", cached_prompt_tokens=5),
|
||||
error=None,
|
||||
)
|
||||
assert rec.rows[0]["cost"] is None
|
||||
|
||||
|
||||
class TestEmitter:
|
||||
async def test_attempt_success_row(self):
|
||||
rec = _MemoryRecorder()
|
||||
@@ -168,7 +579,58 @@ class TestEmitter:
|
||||
)
|
||||
row = rec.rows[0]
|
||||
assert row["error"].startswith("TransientError")
|
||||
assert row["response"] == "" and row["usage_source"] == "estimated"
|
||||
# 失败尝试没有任何用量信息可言 → unavailable(设计 §3.2 #6)
|
||||
assert row["response"] == "" and row["usage_source"] == "unavailable"
|
||||
assert row["cost"] is None
|
||||
|
||||
async def test_terminal_failure_row_is_unavailable(self):
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec, pricing=_PRICING).emit_terminal_failure(
|
||||
request=_REQ, call_id="cid-t", latency_ms=5, error="cancelled"
|
||||
)
|
||||
row = rec.rows[0]
|
||||
assert row["usage_source"] == "unavailable" and row["cost"] is None
|
||||
|
||||
@pytest.mark.parametrize(("prompt", "completion"), [(0, 0), (0, 4000)])
|
||||
async def test_unavailable_success_row_has_null_cost(self, prompt, completion):
|
||||
"""产生了真实调用但用量不可得 → cost 记 NULL(设计 §3.1 不变式)。
|
||||
|
||||
参数第二组是改前兜底写出的 `0/4000` 形态: 那时换算出 0.032 的假金额。
|
||||
"""
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec, pricing=_PRICING).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="cid-u",
|
||||
latency_ms=42,
|
||||
response=_resp(
|
||||
usage_source="unavailable", prompt_tokens=prompt, completion_tokens=completion
|
||||
),
|
||||
error=None,
|
||||
)
|
||||
assert rec.rows[0]["cost"] is None
|
||||
|
||||
async def test_measured_row_still_priced(self):
|
||||
"""对照组: 同一价格表下 measured 行照常换算,证明 None 不是价格表没接上。"""
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec, pricing=_PRICING).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_source(),
|
||||
call_id="cid-m",
|
||||
latency_ms=42,
|
||||
response=_resp(prompt_tokens=0, completion_tokens=4000),
|
||||
error=None,
|
||||
)
|
||||
assert rec.rows[0]["cost"] == pytest.approx(0.032)
|
||||
|
||||
async def test_cache_hit_keeps_zero_cost_even_when_unavailable(self):
|
||||
"""缓存命中未产生新调用,0.0 是事实而非未知 → 短路必须排在 cache_hit 之后。"""
|
||||
rec = _MemoryRecorder()
|
||||
await TelemetryEmitter(rec, pricing=_PRICING).emit_cache_hit(
|
||||
request=_REQ,
|
||||
response=_resp(cache_hit=True, usage_source="unavailable", completion_tokens=4000),
|
||||
)
|
||||
assert rec.rows[0]["cache_hit"] is True and rec.rows[0]["cost"] == 0.0
|
||||
|
||||
async def test_multimodal_messages_digested_before_storage(self):
|
||||
rec = _MemoryRecorder()
|
||||
|
||||
+185
-3
@@ -1,10 +1,12 @@
|
||||
"""types.py 冻结签名的行为测试(M1 设计 §2)。"""
|
||||
|
||||
import dataclasses
|
||||
import inspect
|
||||
|
||||
import pytest
|
||||
|
||||
from polygateway.types import (
|
||||
USAGE_SOURCES,
|
||||
BackpressurePolicy,
|
||||
BreakerConfig,
|
||||
ChatRequest,
|
||||
@@ -47,6 +49,29 @@ class TestLLMResponse:
|
||||
assert resp.usage_source == "measured"
|
||||
assert resp.structured_data is None
|
||||
|
||||
def test_observability_fields_default_to_none(self):
|
||||
"""issue #3: None = 该源未上报,与"上报了但是 0"区分(0 是真实零命中)。"""
|
||||
resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid")
|
||||
assert resp.cached_prompt_tokens is None
|
||||
assert resp.model_reported is None
|
||||
filled = LLMResponse(
|
||||
"c",
|
||||
"t",
|
||||
"m",
|
||||
"p",
|
||||
1,
|
||||
2,
|
||||
3,
|
||||
None,
|
||||
None,
|
||||
False,
|
||||
"cid",
|
||||
cached_prompt_tokens=0,
|
||||
model_reported="MiniMax-Text-01-250321",
|
||||
)
|
||||
assert filled.cached_prompt_tokens == 0 # 真实零命中,不得与 None 混同
|
||||
assert filled.model_reported == "MiniMax-Text-01-250321"
|
||||
|
||||
def test_frozen(self):
|
||||
resp = LLMResponse("c", "t", "m", "p", 1, 2, 3, None, None, False, "cid")
|
||||
with pytest.raises(dataclasses.FrozenInstanceError):
|
||||
@@ -91,9 +116,11 @@ class TestSourceConfig:
|
||||
with pytest.raises(ValueError):
|
||||
_make_source(timeout_s=0)
|
||||
|
||||
def test_tpm_requires_est_tokens(self):
|
||||
with pytest.raises(ValueError):
|
||||
_make_source(tpm=10000, est_tokens=0)
|
||||
def test_tpm_does_not_require_est_tokens(self):
|
||||
"""`tpm > 0 ⇒ est_tokens > 0` 已解绑: 供应商配额可独立于库实现细节填写。"""
|
||||
derived = _make_source(tpm=10000, est_tokens=0)
|
||||
assert derived.est_tokens == 0
|
||||
assert derived.effective_est_tokens() == 166 # max(1, 10000 // 60)
|
||||
assert _make_source(tpm=10000, est_tokens=800).est_tokens == 800
|
||||
|
||||
def test_negative_gate_rejected(self):
|
||||
@@ -117,6 +144,70 @@ class TestSourceConfig:
|
||||
_make_source(missing_done="ignore")
|
||||
|
||||
|
||||
class TestEffectiveEstTokens:
|
||||
"""TPM 入场预扣量的派生(est_tokens 解耦设计 §2.2)。"""
|
||||
|
||||
def test_derives_from_tpm_scale_free(self):
|
||||
"""派生量随配额同比缩放: 两种配额规模的在途上限同为 60 个调用。"""
|
||||
assert _make_source(tpm=6000).effective_est_tokens() == 100
|
||||
assert _make_source(tpm=600000).effective_est_tokens() == 10000
|
||||
|
||||
def test_derived_floor_is_one(self):
|
||||
"""极小配额下派生量不得塌到 0——0 预扣等于 TPM 闸不设防(设计 §2.2)。"""
|
||||
assert _make_source(tpm=30).effective_est_tokens() == 1
|
||||
|
||||
def test_zero_when_tpm_gate_disabled(self):
|
||||
"""tpm=0 即 TPM 闸未启用,无需预扣。"""
|
||||
assert _make_source().effective_est_tokens() == 0
|
||||
|
||||
def test_explicit_value_wins(self):
|
||||
"""显式配置是调优覆盖,优先于派生。"""
|
||||
assert _make_source(tpm=6000, est_tokens=4000).effective_est_tokens() == 4000
|
||||
|
||||
def test_is_pure_sync_function(self):
|
||||
"""纯方法: 非协程、可重复调用、不改动自身字段(设计 §5 并发前提)。"""
|
||||
assert not inspect.iscoroutinefunction(SourceConfig.effective_est_tokens)
|
||||
src = _make_source(tpm=6000)
|
||||
assert src.effective_est_tokens() == src.effective_est_tokens() == 100
|
||||
assert src.est_tokens == 0 # 派生不回写字段
|
||||
|
||||
|
||||
class TestUsageSourceDomain:
|
||||
"""`usage_source` 三态值域常量(设计 §3.1)。"""
|
||||
|
||||
def test_domain_is_exactly_three_values(self):
|
||||
assert set(USAGE_SOURCES) == {"measured", "estimated", "unavailable"}
|
||||
assert isinstance(USAGE_SOURCES, frozenset) # 不可变: 调用方无法就地扩张值域
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"build",
|
||||
[
|
||||
lambda v: LLMResponse(
|
||||
"c", "t", "m", "p", 1, 2, 3, None, None, False, "cid", usage_source=v
|
||||
),
|
||||
lambda v: Usage(prompt_tokens=1, completion_tokens=2, usage_source=v),
|
||||
lambda v: TransportResult(
|
||||
content="c",
|
||||
thinking="",
|
||||
prompt_tokens=1,
|
||||
completion_tokens=2,
|
||||
usage_source=v,
|
||||
ttft_ms=None,
|
||||
max_inter_token_ms=None,
|
||||
raw={},
|
||||
),
|
||||
],
|
||||
)
|
||||
def test_no_runtime_validation_on_public_dataclasses(self, build):
|
||||
"""越界值构造**不得**抛异常(锁定设计 §3.1 的落点裁决)。
|
||||
|
||||
这些是运行时构造点(如 retry.py:418),裸 `ValueError` 不属 errors.py
|
||||
四分类、`RetryMW` 不捕它,会直接逃出 `chat()`——故值域只约束生产侧,
|
||||
不落在公共 frozen dataclass 的 `__post_init__` 上。
|
||||
"""
|
||||
assert build("garbage").usage_source == "garbage"
|
||||
|
||||
|
||||
class TestResilienceConfigs:
|
||||
def test_retry_policy_validation(self):
|
||||
assert RetryPolicy(max_attempts=3, backoff_base_s=2.0, backoff_max_s=30.0)
|
||||
@@ -154,6 +245,8 @@ class TestAuxTypes:
|
||||
raw={"id": "x"},
|
||||
)
|
||||
assert s.raw["id"] == "x"
|
||||
# issue #3: 新字段带默认值,不填也能构造(OCR 等其他 transport 零改动)
|
||||
assert s.cached_prompt_tokens is None and s.model_reported is None
|
||||
|
||||
|
||||
class TestOcrTypes:
|
||||
@@ -212,3 +305,92 @@ class TestOcrTypes:
|
||||
|
||||
with pytest.raises(TypeError):
|
||||
OcrTextResult(text="x") # 溯源件不可省略
|
||||
|
||||
|
||||
class TestSamplingValidation:
|
||||
"""采样参数覆盖层的构造期校验(issue #4 设计决策 B)。"""
|
||||
|
||||
@pytest.mark.parametrize("key", ["model", "messages", "stream", "stream_options"])
|
||||
def test_protected_keys_rejected(self, key):
|
||||
"""保护键会击穿治理: 成本算错/口径失真/绕过看门狗与 usage 帧。"""
|
||||
from polygateway.types import validate_request_overlay
|
||||
|
||||
with pytest.raises(ValueError) as exc:
|
||||
validate_request_overlay({key: "x"}, origin="chat(overlay=...)")
|
||||
assert key in str(exc.value)
|
||||
assert "chat(overlay=...)" in str(exc.value) # 信息须能定位来源
|
||||
|
||||
def test_non_str_key_reports_key_problem(self):
|
||||
"""非 str 键须报"键必须是 str",不能被 sort_keys 的比较错误误报成不可序列化。"""
|
||||
from polygateway.types import validate_request_overlay
|
||||
|
||||
with pytest.raises(ValueError, match="str"):
|
||||
validate_request_overlay({1: "a", "b": 2}, origin="test")
|
||||
|
||||
def test_unserializable_value_becomes_value_error(self):
|
||||
"""裸 TypeError 会逃出 CacheMW 的降级 try 且一行遥测都没有(设计决策 B)。"""
|
||||
from polygateway.types import validate_request_overlay
|
||||
|
||||
with pytest.raises(ValueError, match="JSON"):
|
||||
validate_request_overlay({"temperature": object()}, origin="test")
|
||||
|
||||
def test_returns_independent_copy(self):
|
||||
"""调用方逐次改 seed 复用同一 dict 是预期模式,不拷贝会有竞态(决策 E)。"""
|
||||
from polygateway.types import validate_request_overlay
|
||||
|
||||
caller_dict = {"temperature": 0, "seed": 42}
|
||||
validated = validate_request_overlay(caller_dict, origin="test")
|
||||
caller_dict["seed"] = 43
|
||||
assert validated == {"temperature": 0, "seed": 42}
|
||||
|
||||
def test_merge_prefers_call_level(self):
|
||||
"""优先级: 调用级 > 配置级(设计决策 A)。"""
|
||||
from polygateway.types import merge_sampling
|
||||
|
||||
merged = merge_sampling({"temperature": 0, "top_p": 1}, {"temperature": 1})
|
||||
assert merged == {"temperature": 1, "top_p": 1}
|
||||
|
||||
def test_canonical_json_is_key_order_stable(self):
|
||||
"""缓存 key 与遥测列共用同一序列化口径,键序不得影响结果。"""
|
||||
from polygateway.types import canonical_sampling_json
|
||||
|
||||
assert canonical_sampling_json({"b": 1, "a": 2}) == canonical_sampling_json(
|
||||
{"a": 2, "b": 1}
|
||||
)
|
||||
assert canonical_sampling_json({}) is None
|
||||
|
||||
|
||||
class TestSourceConfigExtraBody:
|
||||
"""配置级采样参数(issue #4 设计决策 A/E)。"""
|
||||
|
||||
def test_defaults_to_empty_and_is_read_only(self):
|
||||
source = _make_source()
|
||||
assert source.extra_body == {}
|
||||
with pytest.raises(TypeError):
|
||||
source.extra_body["temperature"] = 0 # MappingProxyType 只读
|
||||
|
||||
def test_protected_key_rejected_at_construction(self):
|
||||
"""装配期报错,不放到运行时才炸(CLAUDE.md §4.5)。"""
|
||||
with pytest.raises(ValueError, match="model"):
|
||||
_make_source(extra_body={"model": "sneaky"})
|
||||
|
||||
def test_accepts_sampling_params(self):
|
||||
source = _make_source(extra_body={"temperature": 0})
|
||||
assert source.extra_body["temperature"] == 0
|
||||
|
||||
def test_replace_rebuilds_proxy(self):
|
||||
"""决策 G 的剥离依赖 replace 能重跑 __post_init__ 且不递归。"""
|
||||
source = _make_source(extra_body={"temperature": 0})
|
||||
stripped = dataclasses.replace(source, extra_body={})
|
||||
assert stripped.extra_body == {}
|
||||
with pytest.raises(TypeError):
|
||||
stripped.extra_body["x"] = 1
|
||||
|
||||
def test_no_longer_hashable_is_intentional(self):
|
||||
"""加 mapping 字段的固有代价(裸 dict 亦然),库内无调用点会踩。
|
||||
|
||||
锁定为有意行为: 将来踩到的人不应把它当 bug"修"回去——要可变副本用
|
||||
dict(source.extra_body),要改字段用 dataclasses.replace(设计 Task 1)。
|
||||
"""
|
||||
with pytest.raises(TypeError):
|
||||
hash(_make_source())
|
||||
|
||||
@@ -0,0 +1,295 @@
|
||||
"""`usage_source` 值域封闭: 库内所有生产点的产出恒落在 `USAGE_SOURCES` 内。
|
||||
|
||||
设计 §3.1 裁定值域**只约束生产侧**——公共 frozen dataclass 不加运行时校验
|
||||
(裸 `ValueError` 不属四分类,会逃出 `chat()`;该裁决的锁定断言在
|
||||
`test_types.py::TestUsageSourceDomain`)。因此封闭性只能由"逐个驱动生产点、
|
||||
断言其产出在三态内"来保证,本文件即该断言的载体。
|
||||
|
||||
独立成文件而非并入 `test_types.py`: 断言横跨 transports / embedding /
|
||||
telemetry 三层,放进最内层内核的类型测试会让它反向依赖具体实现。
|
||||
|
||||
覆盖的生产点(设计 §3.2 逐处改动表的字面量产出方):
|
||||
`_resolve_usage`、`_resolve_embedding_usage`、`_resolve_stream_usage`(打捞覆盖)、
|
||||
`EmbeddingClient._merge`、`EmbeddingClient.embed` 空输入短路、
|
||||
`OcrClient._emit`、`TelemetryEmitter.emit_attempt/emit_cache_hit/emit_terminal_failure`。
|
||||
"""
|
||||
|
||||
import itertools
|
||||
import json
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
|
||||
from polygateway.backends.memory.breaker import InMemoryGate
|
||||
from polygateway.backends.memory.limiter import InMemoryLimiter
|
||||
from polygateway.embedding import EmbeddingClient, _BatchOutcome
|
||||
from polygateway.middleware.telemetry import TelemetryEmitter
|
||||
from polygateway.ocr import OcrClient
|
||||
from polygateway.sources import RoundRobinSelector
|
||||
from polygateway.transports.openai_compat import (
|
||||
OpenAICompatTransport,
|
||||
_resolve_embedding_usage,
|
||||
_resolve_usage,
|
||||
)
|
||||
from polygateway.types import (
|
||||
USAGE_SOURCES,
|
||||
BackpressurePolicy,
|
||||
BreakerConfig,
|
||||
ChatRequest,
|
||||
EmbeddingTransportResult,
|
||||
GlobalLimits,
|
||||
LLMResponse,
|
||||
OcrTextTransportResult,
|
||||
RetryPolicy,
|
||||
SourceConfig,
|
||||
)
|
||||
|
||||
_REQ = ChatRequest(messages=[{"role": "user", "content": "hi"}])
|
||||
_DOMAIN = sorted(USAGE_SOURCES)
|
||||
|
||||
|
||||
def _src():
|
||||
return SourceConfig(
|
||||
name="s1",
|
||||
provider="p",
|
||||
base_url="https://gw.example/v1",
|
||||
api_key="sk",
|
||||
model="m",
|
||||
timeout_s=10.0,
|
||||
est_tokens=4000, # 兜底口径的历史来源: 生产点不得因它落到三态之外
|
||||
)
|
||||
|
||||
|
||||
class _MemoryRecorder:
|
||||
def __init__(self):
|
||||
self.rows = []
|
||||
|
||||
async def record_llm_call(self, **fields):
|
||||
self.rows.append(fields)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"usage",
|
||||
[
|
||||
{"prompt_tokens": 12, "completion_tokens": 34}, # 完整可信
|
||||
{}, # 整帧缺失
|
||||
{"prompt_tokens": 0, "completion_tokens": 0}, # 全 0(和不为正)
|
||||
{"prompt_tokens": "12", "completion_tokens": 34}, # 类型非法
|
||||
{"prompt_tokens": None, "completion_tokens": None},
|
||||
{"prompt_tokens": 12}, # 半帧
|
||||
],
|
||||
)
|
||||
def test_resolve_usage_stays_in_domain(usage):
|
||||
assert _resolve_usage(usage)[2] in USAGE_SOURCES
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"data",
|
||||
[
|
||||
{"usage": {"prompt_tokens": 12}},
|
||||
{},
|
||||
{"usage": None},
|
||||
{"usage": {}},
|
||||
{"usage": {"prompt_tokens": 0}},
|
||||
{"usage": {"prompt_tokens": "12"}},
|
||||
],
|
||||
)
|
||||
def test_resolve_embedding_usage_stays_in_domain(data):
|
||||
assert _resolve_embedding_usage(data)[1] in USAGE_SOURCES
|
||||
|
||||
|
||||
def _sse(*frames, done):
|
||||
"""构造 SSE 响应;done=False 触发打捞路径(`_complete_stream` 的覆盖分支)。"""
|
||||
text = "".join(f"data: {json.dumps(f)}\n\n" for f in frames) + (
|
||||
"data: [DONE]\n\n" if done else ""
|
||||
)
|
||||
return httpx.Response(200, content=text.encode(), headers={"content-type": "text/event-stream"})
|
||||
|
||||
|
||||
@pytest.mark.parametrize("usage", [{"prompt_tokens": 11, "completion_tokens": 7}, None])
|
||||
async def test_salvage_override_stays_in_domain(usage):
|
||||
"""打捞覆盖(`openai_compat._complete_stream`)是第三个字面量产出方。"""
|
||||
frames = [{"choices": [{"delta": {"content": "partial"}}]}]
|
||||
if usage is not None:
|
||||
frames.append({"choices": [], "usage": usage})
|
||||
transport = OpenAICompatTransport(
|
||||
client_factory=lambda src: httpx.AsyncClient(
|
||||
base_url=src.base_url,
|
||||
transport=httpx.MockTransport(lambda request: _sse(*frames, done=False)),
|
||||
)
|
||||
)
|
||||
source = SourceConfig(
|
||||
name="s1",
|
||||
provider="qwen",
|
||||
base_url="https://gw.example/v1",
|
||||
api_key="sk",
|
||||
model="m",
|
||||
timeout_s=10.0,
|
||||
est_tokens=4000,
|
||||
missing_done="salvage",
|
||||
)
|
||||
result = await transport.complete(
|
||||
messages=[{"role": "user", "content": "hi"}],
|
||||
source=source,
|
||||
stream=True,
|
||||
overlay={},
|
||||
call_id="cid",
|
||||
)
|
||||
assert result.usage_source in USAGE_SOURCES
|
||||
|
||||
|
||||
class _ScriptedOcrTransport:
|
||||
async def recognize_text(self, *, image, source, call_id):
|
||||
return OcrTextTransportResult(text="LINE-1", raw={"task_type": "text"})
|
||||
|
||||
async def parse_layout(self, *, image, source, call_id):
|
||||
raise NotImplementedError
|
||||
|
||||
async def check_health(self, *, source):
|
||||
raise NotImplementedError
|
||||
|
||||
|
||||
async def test_ocr_emit_stays_in_domain():
|
||||
"""`OcrClient._emit` 的字面量(ocr.py:411)同样纳入封闭性断言。
|
||||
|
||||
值取 `measured` 是设计 §3.3 的裁决(OCR 的 0 token 属事实);此处只断言
|
||||
落在三态内,精确取值的防回归钉在 `test_ocr_client.py`。
|
||||
"""
|
||||
source = SourceConfig(
|
||||
name="m1",
|
||||
provider="monkey",
|
||||
base_url="http://gw.example",
|
||||
api_key="none",
|
||||
model="monkey-ocr",
|
||||
timeout_s=10.0,
|
||||
)
|
||||
recorder = _MemoryRecorder()
|
||||
client = OcrClient(
|
||||
scope="ocr",
|
||||
sources=[source],
|
||||
selector=RoundRobinSelector(),
|
||||
limiter=InMemoryLimiter(
|
||||
scope="ocr",
|
||||
sources={source.name: source},
|
||||
global_limits=GlobalLimits(max_concurrency=0, rpm=0, tpm=0),
|
||||
lease_ttl_s=100.0,
|
||||
),
|
||||
breaker=InMemoryGate(
|
||||
config=BreakerConfig(fail_threshold=3, cooldown_s=60.0, probe_ttl_s=120.0)
|
||||
),
|
||||
transport=_ScriptedOcrTransport(),
|
||||
retry=RetryPolicy(max_attempts=1, backoff_base_s=0.001, backoff_max_s=0.01),
|
||||
backpressure=BackpressurePolicy(stall_window_s=300.0, poll_interval_s=0.001),
|
||||
telemetry=recorder,
|
||||
)
|
||||
await client.recognize_text(b"jpg")
|
||||
assert recorder.rows[0]["usage_source"] in USAGE_SOURCES
|
||||
|
||||
|
||||
def _merge_client():
|
||||
"""构造仅用于调用 `_merge` 的最小 EmbeddingClient(不发起任何调用)。"""
|
||||
source = _src()
|
||||
return EmbeddingClient(
|
||||
scope="embed",
|
||||
sources=[source],
|
||||
selector=RoundRobinSelector(),
|
||||
limiter=InMemoryLimiter(
|
||||
scope="embed",
|
||||
sources={source.name: source},
|
||||
global_limits=GlobalLimits(max_concurrency=0, rpm=0, tpm=0),
|
||||
lease_ttl_s=100.0,
|
||||
),
|
||||
breaker=InMemoryGate(
|
||||
config=BreakerConfig(fail_threshold=3, cooldown_s=60.0, probe_ttl_s=120.0)
|
||||
),
|
||||
transport=object(),
|
||||
retry=RetryPolicy(max_attempts=1, backoff_base_s=0.001, backoff_max_s=0.01),
|
||||
backpressure=BackpressurePolicy(stall_window_s=300.0, poll_interval_s=0.001),
|
||||
batch_size=2,
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(("first", "second"), list(itertools.product(_DOMAIN, repeat=2)))
|
||||
def test_merge_stays_in_domain(first, second):
|
||||
"""任意两批 usage_source 组合(含尚无生产者的 unavailable)合并后仍在三态内。"""
|
||||
source = _src()
|
||||
outcomes = [
|
||||
_BatchOutcome(
|
||||
result=EmbeddingTransportResult(
|
||||
vectors=[[1.0]], dim=1, prompt_tokens=1, usage_source=value, raw={}
|
||||
),
|
||||
source=source,
|
||||
call_id="c",
|
||||
latency_ms=1,
|
||||
)
|
||||
for value in (first, second)
|
||||
]
|
||||
assert _merge_client()._merge(outcomes).usage_source in USAGE_SOURCES
|
||||
|
||||
|
||||
async def test_empty_input_short_circuit_stays_in_domain():
|
||||
"""空输入短路自造响应(embedding.py:151),不经 transport 也须落在三态内。"""
|
||||
resp = await _merge_client().embed([])
|
||||
assert resp.usage_source in USAGE_SOURCES
|
||||
|
||||
|
||||
def _resp(usage_source):
|
||||
return LLMResponse(
|
||||
content="ok",
|
||||
thinking="",
|
||||
model="m",
|
||||
provider="p",
|
||||
prompt_tokens=1,
|
||||
completion_tokens=2,
|
||||
latency_ms=30,
|
||||
ttft_ms=None,
|
||||
max_inter_token_ms=None,
|
||||
cache_hit=False,
|
||||
call_id="cid",
|
||||
source_name="s1",
|
||||
usage_source=usage_source,
|
||||
)
|
||||
|
||||
|
||||
@pytest.mark.parametrize("emitted", _DOMAIN)
|
||||
async def test_emit_attempt_success_stays_in_domain(emitted):
|
||||
recorder = _MemoryRecorder()
|
||||
await TelemetryEmitter(recorder).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_src(),
|
||||
call_id="cid",
|
||||
latency_ms=10,
|
||||
response=_resp(emitted),
|
||||
error=None,
|
||||
)
|
||||
assert recorder.rows[0]["usage_source"] in USAGE_SOURCES
|
||||
|
||||
|
||||
async def test_emit_attempt_failed_attempt_stays_in_domain():
|
||||
"""失败尝试无 response,`usage_source` 取 emitter 自己的字面量。"""
|
||||
recorder = _MemoryRecorder()
|
||||
await TelemetryEmitter(recorder).emit_attempt(
|
||||
request=_REQ,
|
||||
source=_src(),
|
||||
call_id="cid",
|
||||
latency_ms=10,
|
||||
response=None,
|
||||
error="boom",
|
||||
)
|
||||
assert recorder.rows[0]["usage_source"] in USAGE_SOURCES
|
||||
|
||||
|
||||
@pytest.mark.parametrize("emitted", _DOMAIN)
|
||||
async def test_emit_cache_hit_stays_in_domain(emitted):
|
||||
recorder = _MemoryRecorder()
|
||||
await TelemetryEmitter(recorder).emit_cache_hit(request=_REQ, response=_resp(emitted))
|
||||
assert recorder.rows[0]["usage_source"] in USAGE_SOURCES
|
||||
|
||||
|
||||
async def test_emit_terminal_failure_stays_in_domain():
|
||||
"""终态失败无具体源,`usage_source` 同样取 emitter 字面量。"""
|
||||
recorder = _MemoryRecorder()
|
||||
await TelemetryEmitter(recorder).emit_terminal_failure(
|
||||
request=_REQ, call_id="cid", latency_ms=10, error="cancelled"
|
||||
)
|
||||
assert recorder.rows[0]["usage_source"] in USAGE_SOURCES
|
||||
Reference in New Issue
Block a user