feat: collect reasoning_tokens from the provider usage payload (issue #6)
Reasoning tokens are already counted inside completion_tokens, so the cost total was never wrong -- what was missing is the attribution: how much of a call was spent thinking rather than answering. LLMResponse and TransportResult each gain a trailing reasoning_tokens field, and the telemetry port grows from 21 to 22 columns with the new column appended in both backends so fresh and migrated schemas keep the same physical order. None means this particular call did not report the field, not that the source never reports it: a relay that falls back to a local tokenizer replaces the whole usage object and drops completion_tokens_details. Downstream checks must therefore read "in (None, 0)"; no provider was observed reporting a literal zero.
This commit is contained in:
@@ -99,6 +99,15 @@ class LLMResponse:
|
||||
model_reported: str | None = None
|
||||
"""API 响应体里的 model 字段;None = 未上报。与 `model`(配置别名)可能
|
||||
分叉——供应商把别名指向新权重时,实验复现必须认这个串。"""
|
||||
reasoning_tokens: int | None = None
|
||||
"""推理消耗的输出 token 数(含在 `completion_tokens` 内,故不影响成本总额,
|
||||
只补归因;issue #6)。
|
||||
|
||||
`None` = **本次调用**未上报,**不是**"该源不上报"——中转网关在上游不返回
|
||||
usage 时会用本地 tokenizer 补算并整体替换 usage 对象,把
|
||||
`completion_tokens_details` 一并吃掉(findings §4c 实测同一请求 10 轮呈
|
||||
6:4 双峰)。实测三家供应商在未推理时都是整个 details 缺失、无人上报 `0`,
|
||||
故下游判据须为 `in (None, 0)`,写 `== 0` 的条件永远不成立。"""
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
@@ -151,9 +160,10 @@ class TransportResult:
|
||||
ttft_ms: float | None
|
||||
max_inter_token_ms: float | None
|
||||
raw: dict[str, Any]
|
||||
# —— 可观测字段(issue #3;带默认值,非 OpenAI 兼容的 transport 可不填)——
|
||||
# —— 可观测字段(issue #3/#6;带默认值,非 OpenAI 兼容的 transport 可不填)——
|
||||
cached_prompt_tokens: int | None = None
|
||||
model_reported: str | None = None
|
||||
reasoning_tokens: int | None = None
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
|
||||
Reference in New Issue
Block a user