A sixteen-row matrix over 127 real calls: disable and enable on MiniMax-M3 in both streaming and non-streaming mode, extra_body winning over the profile slot, qwen and deepseek still disabling correctly, a drift sentinel that re-derives every registered capability from live behaviour, and the assembly guard refusing the models that cannot comply. Two judgement criteria had to be corrected by the data they were meant to judge. Output length cannot separate the two regimes at all -- the disabled runs reach 46 tokens when the model narrates its working in the visible answer, and the enabled runs drop to 13 when medium effort barely thinks. reasoning_tokens separates them cleanly in both directions, which is precisely what issue #6 was collected for. A second anchor compares prompt_tokens between the two regimes: the vendor injects a reasoning instruction when thinking is on, so the input side grows, and comparing the two runs relatively avoids hardcoding any vendor number. Provider names are mapped explicitly rather than guessed from the model string; guessing had silently skipped the qwen row behind a "source unavailable" reason that was not true.
This commit is contained in:
@@ -1,11 +1,11 @@
|
||||
---
|
||||
type: schema
|
||||
node_id: schema:llm-calls
|
||||
title: "表结构: llm_calls(遥测 21 字段)"
|
||||
title: "表结构: llm_calls(遥测 22 字段)"
|
||||
date: 2026-07-20
|
||||
---
|
||||
|
||||
# 表结构: llm_calls(遥测 21 字段)
|
||||
# 表结构: llm_calls(遥测 22 字段)
|
||||
|
||||
|
||||
## 列定义(冻结,M1 设计 §4.4 / ARCH §7.8)
|
||||
@@ -27,6 +27,7 @@ date: 2026-07-20
|
||||
| cached_prompt_tokens | INTEGER | 供应商 prompt cache 命中的输入 token(2026-07-31,issue #3);NULL = 该源未上报,`0` = 上报了真实零命中,两者不可混同 |
|
||||
| model_reported | TEXT | API 响应体实际返回的 model;NULL = 未上报。与 `model`(配置别名)可能分叉 |
|
||||
| sampling | TEXT | 本次调用的采样参数 canonical JSON(2026-07-31,issue #4);NULL = 未传。见下方口径 |
|
||||
| reasoning_tokens | INTEGER | 推理消耗的输出 token(2026-08-02,issue #6);**含在 completion_tokens 内**,不影响成本总额,只补归因。NULL = **本次调用**未上报 |
|
||||
|
||||
## usage/成本口径(2026-07-30,est_tokens 解耦)
|
||||
|
||||
@@ -53,6 +54,8 @@ FROM llm_calls WHERE cache_hit = false AND cached_prompt_tokens IS NOT NULL;
|
||||
|
||||
## 采样参数口径(2026-07-31,issue #4)
|
||||
|
||||
`reasoning_tokens` 的 NULL 语义与 `cached_prompt_tokens` **不同**: 后者的 NULL 是"该源不报这个数",前者只能读作"**本次调用**未上报"——中转在上游不返回 usage 时会用本地 tokenizer 补算并整体替换 usage 对象,把 `completion_tokens_details` 一并吃掉(实测同一请求 10 轮呈 6:4 双峰)。故统计口径须为 `IS NULL OR = 0` 才算"未推理",写 `= 0` 的条件永远不成立——实测三家供应商在未推理时都是整个 details 缺失,无人上报字面 `0`。**不可用 `completion_tokens` 反推是否推理**: 两档的输出长度分布重叠(关闭档实测最高 46,开启档最低 13)。
|
||||
|
||||
`sampling` 列 = 「调用方采样意图 ⊎ 生效源 `extra_body`」的 canonical JSON,空则 NULL。**不含**结构化输出注入的 `response_format`——列名是采样参数,schema 不是,且数 KB schema 逐行落库会让审计表无谓膨胀。补列纪律与 issue #3 两列逐字相同(排在末尾、先探测再 ALTER、失败只逐行降级)。
|
||||
|
||||
三个 emit 入口的取值必须各自定死,否则同一列在不同行含义不同:
|
||||
|
||||
Reference in New Issue
Block a user