A sixteen-row matrix over 127 real calls: disable and enable on MiniMax-M3 in both streaming and non-streaming mode, extra_body winning over the profile slot, qwen and deepseek still disabling correctly, a drift sentinel that re-derives every registered capability from live behaviour, and the assembly guard refusing the models that cannot comply. Two judgement criteria had to be corrected by the data they were meant to judge. Output length cannot separate the two regimes at all -- the disabled runs reach 46 tokens when the model narrates its working in the visible answer, and the enabled runs drop to 13 when medium effort barely thinks. reasoning_tokens separates them cleanly in both directions, which is precisely what issue #6 was collected for. A second anchor compares prompt_tokens between the two regimes: the vendor injects a reasoning instruction when thinking is on, so the input side grows, and comparing the two runs relatively avoids hardcoding any vendor number. Provider names are mapped explicitly rather than guessed from the model string; guessing had silently skipped the qwen row behind a "source unavailable" reason that was not true.
This commit is contained in:
@@ -82,3 +82,5 @@
|
||||
- [2026-08-02 09:49 UTC] 新增边: plan:2026-08-02-thinking-capability --implements--> design:2026-08-02-thinking-capability-design
|
||||
- [2026-08-02 09:49 UTC] 新增 plan: 推理开关能力建模与 reasoning_tokens 采集实施计划 (plan:2026-08-02-thinking-capability)
|
||||
- [2026-08-02 09:49 UTC] 重建索引: 53 篇页面
|
||||
- [2026-08-02 10:55 UTC] 重建索引: 53 篇页面
|
||||
- [2026-08-02 10:55 UTC] 更新 finding: 补 §2.5 输出长度不是有效判别量(e2e 各 15 轮实测)
|
||||
|
||||
Reference in New Issue
Block a user