a0a5cf7ecc
Independent verification found the first cut had swapped one bug for a worse one. The budgets were split by "did we send a request", so a 429 attempt counted as productive — but 429 is exempt from the retry budget, so its time burned neither budget. Against a queueing gateway that holds the request for the full timeout before answering 429, a call could hang for 301 attempts / 25.2 hours, measured, versus 301 seconds before the change. The split is now by which budget the time consumes: time that burns max_attempts is excluded from stall, time that does not (429 attempts included) belongs to stall. Measured again: back to one attempt / 301s. Only the chat loop needs this — embedding and ocr count 429 against max_attempts unconditionally, so the gap never existed there. The stall verdict moved into _stalled(), which both call sites had duplicated, to keep __call__ under the complexity gate.