37b4a557c2
The four red e2e cases were blamed on MiniMax-M3 no longer reasoning. Raw gateway probes show the opposite: M3 reasons fine (124 chars of reasoning_content, prompt 194 to 216, completion 3 to 60). What changed is that the MiniMax route stopped returning completion_tokens_details, while qwen and deepseek still do on the same gateway and key. The library already holds 185 chars of proof in LLMResponse.thinking and never feeds it into any verdict. The design turns that verdict into a first-class return value judged from multiple signals, says UNKNOWN when a single response cannot tell, and reconciles it against the capability table so a stale declaration becomes a warning instead of a silent illusion.