a1c4273a8b
resolve_thinking now takes an Effort instead of a tri-state bool, and the four gates become five. The new one sits ahead of the generic tier check on purpose: asking for `none` on GLM-5.3 used to fall through to "none is not supported, pick low/high/max", which loses both the fact that the model cannot stop reasoning and the one tier the caller could switch to right now. Without that alternative, downstream goes looking for extra_body — which is how issue #20 happened in the first place. The return type is a ThinkingResolution rather than the payload alone. Under fallback="nearest" the tier that goes out is not the tier that was asked for, and telemetry has to record the one that ran, or task 10 files a call under a tier it never used. Ties in that mapping go to the weaker side: a silent medium -> max is a multiple of the bill, and the library does not raise a caller's price on its own. Two readings the design left implicit, both settled the way its own compatibility promise requires: - `auto` is exempt from the tier list. It means "on, no tier named", which in the body is the absence of the effort key, not a value of it. Checking it against the list would break every existing source that sets ENABLE_THINKING=true against deepseek-v4 or glm-5.3. - `none` is never a mapping target. Turning "think less" into "do not think" reverses the decision instead of cheapening it; a switch-only model maps to `auto` and a model that only has `none` still errors. Both call sites convert enable_thinking in place for now; task 5 folds that into effective_effort along with the source- and call-level tiers.