The verifier caught that the disable-direction evidence only proved "no regression", not "actually took effect": on M3 the disabled runs and the no-opinion baseline are identically distributed, because that model does not reason by default anyway. So the disable runs alone cannot rule out the very failure mode issue #5 is about -- the parameter being silently dropped upstream. The bogus-value experiment that does rule it out was sitting in the findings document instead of the test suite; it is now case L3b, and the L3 assertion that could never fail is gone. Also from the review: the e2e helper caught bare Exception, which would have disguised a library bug as an unavailable source, exactly the silence the reporting discipline exists to prevent; the unregistered model warning fired on every request instead of once per source; and the transport caught ValueError broadly enough to mislabel unrelated errors, now narrowed to a dedicated ThinkingUnsupportedError. The design and plan still described the original judgement criteria, which the measurements had already overturned. Both now match what the tests actually do, and the design no longer claims the only new failure surface is the openai one -- dissect configures MiniMax-M2.7 with ENABLE_THINKING=false and will fail at assembly, which has to be coordinated before this merges.
This commit is contained in:
@@ -7,6 +7,7 @@ import json
|
||||
|
||||
import httpx
|
||||
import pytest
|
||||
from loguru import logger
|
||||
|
||||
from polygateway.errors import (
|
||||
RequestRejectedError,
|
||||
@@ -557,6 +558,45 @@ class TestRequestShaping:
|
||||
with pytest.raises(RequestRejectedError, match="MiniMax-M2.7"):
|
||||
await _complete(_transport_for(handler), source)
|
||||
|
||||
async def test_unregistered_model_warns_only_once_per_source(self):
|
||||
"""未登记模型的告警不能打在请求热路径上: 装配期已喊过,逐次再喊是刷屏。"""
|
||||
|
||||
def handler(request):
|
||||
return _sse_stream(_chunk(content="x"), _chunk(usage=_USAGE))
|
||||
|
||||
source = _source(name="mm", provider="minimax", model="MiniMax-M99", enable_thinking=False)
|
||||
transport = _transport_for(handler)
|
||||
messages: list[str] = []
|
||||
sink_id = logger.add(messages.append, level="WARNING")
|
||||
try:
|
||||
await _complete(transport, source)
|
||||
await _complete(transport, source)
|
||||
await _complete(transport, source)
|
||||
finally:
|
||||
logger.remove(sink_id)
|
||||
hits = [m for m in messages if "MiniMax-M99" in m]
|
||||
assert len(hits) == 1, f"三次调用应只告警一次,实得 {len(hits)} 次"
|
||||
|
||||
async def test_unrelated_value_error_is_not_mislabelled(self, monkeypatch):
|
||||
"""只捕 ThinkingUnsupportedError: 无关的 ValueError 不该被贴成推理开关的错。
|
||||
|
||||
今天 `_build_payload` 里只有 resolve_thinking 会抛 ValueError,所以这条
|
||||
是防御未来 —— 但正因如此才要钉住: 将来谁在那里加一处校验,宽 catch 会
|
||||
把它的错误信息盖掉,而这个用例会先红。
|
||||
"""
|
||||
|
||||
def handler(request): # pragma: no cover - 不该走到发请求
|
||||
raise AssertionError("请求不该发出")
|
||||
|
||||
def _boom(*args, **kwargs):
|
||||
raise ValueError("故意的无关错误")
|
||||
|
||||
monkeypatch.setattr("polygateway.transports.openai_compat.resolve_thinking", _boom)
|
||||
with pytest.raises(ValueError, match="故意的无关错误") as exc:
|
||||
await _complete(_transport_for(handler), _source(enable_thinking=False))
|
||||
assert "推理开关" not in str(exc.value)
|
||||
assert not isinstance(exc.value, RequestRejectedError)
|
||||
|
||||
async def test_unknown_shape_is_rejected(self):
|
||||
def handler(request): # pragma: no cover - 不该走到发请求
|
||||
raise AssertionError("请求不该发出")
|
||||
|
||||
Reference in New Issue
Block a user