602 lines
30 KiB
Markdown
602 lines
30 KiB
Markdown
# WP3 训练循环与进化引擎 Implementation Plan
|
||
|
||
> **For agentic workers:** REQUIRED SUB-SKILL: Use subagent-driven-development to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
|
||
> **前置依赖:WP1(模板已迁移,进化引擎可运行)+ WP4(cache_salt 能力已就绪)必须先完成。**
|
||
|
||
**Goal:** 修复训练循环与进化引擎的 9 处缺陷,使诊断拿到真实轨迹、早停按 epoch 语义、微型题型不崩、崩溃可幂等续跑、进化 patch 不破坏冻结区、降级信号不驱动错误进化、跨 epoch 评估真实重采样。
|
||
|
||
**Architecture:** 诊断经 `StepsJsonRunLog` 从 steps_json 重建轨迹(算法 #7 恢复);早停计数单位 step→epoch;可训练性预检在 gate 建立前剔除微型题型;`_run_step` 幂等(先 DELETE 再写);patch 冻结区检查整个 target 跨度(算法 #8 加固);降级/未判定题按 lapse 保守分流;训练推理用 run_id 作 cache_salt(run_id 已含 epoch,天然跨 epoch 重采样)。
|
||
|
||
**Tech Stack:** Python 3.11、asyncio、SQLite、pytest。
|
||
|
||
**设计源**:`research-wiki/designs/2026-07-16-preflight-fixes-design.md §6-7`
|
||
|
||
---
|
||
|
||
## 关键锚点(实现前必读)
|
||
|
||
| 用途 | 位置 |
|
||
|------|------|
|
||
| 训练主流程 | `app/harness/runner.py:789-859` `train`;`_setup_train_run:865-879`(预检插入点 L874 前);`_run_step:1001-1026`(run_id L1011、诊断 L1019、无 DELETE) |
|
||
| 早停 | `app/harness/runner.py:291-316` `_should_early_stop`(L315 `+= steps_this_epoch`);`_TrainState:95-121`(`steps_since_best_improved` L118);`_maybe_promote_best:1587`(置 0) |
|
||
| 诊断调用 | `app/harness/runner.py:2163-2196` `_run_diagnosis`(`RunLogImpl` L2173 未包 StepsJsonRunLog);DiagnosisResult `degraded_count` 未被引用 |
|
||
| gate 刷新 | `app/harness/runner.py:1833-1879` `_refresh_gate_ladder`(save L1878 → set observed L1879);checkpoint 落盘晚在 train L837 |
|
||
| holdout | `app/harness/runner.py:1881-1921` `_holdout_four_way`;`_pick_mixed_best:1923-1963`;`_eval_version_on_pool:2148-2161` |
|
||
| 推理落库 | `app/harness/inference.py:363-447` `_run_single_question`(prediction L422 未归一、insert L446 try 外);`_to_text_field:147-162`;`run_inference:473` |
|
||
| Agent Loop | `core/agent/loop.py:103` `run`(session_id L110);`_call_llm` chat 调用 `:329`(`self._llm.chat(messages, session_id=session_id)`) |
|
||
| batching | `app/harness/batching.py:186-202` `_classify_unit`(L200 缺 correctness→None) |
|
||
| checkpoint | `app/harness/checkpoint.py:37-44` `_STRUCTURAL_KEYS`;`serialize_state:76-106`;`write_checkpoint:212-260`(原子写) |
|
||
| 诊断分流 | `core/evolution/diagnose.py:1485-1519` `_build_skill_case_packs`(lapse 分流 L1492);`_process_question:2139-2206`(except L2194 cause_category 留 None) |
|
||
| traces 适配 | `app/harness/baseline_run_log.py:13` `StepsJsonRunLog`;`steps_json_traces.py:13` |
|
||
| patch | `core/evolution/patch.py:285-287` `_in_ranges`;`_do_insert_after:309-326`(L320);`_do_replace_delete:329-347`(L343);markers L11-18;`validate_skill` in `evolve.py:296` |
|
||
|
||
## 核心算法保真校验
|
||
|
||
触及算法 #7(诊断瀑布,Task 1 恢复轨迹)、#8(patch 引擎,Task 5 冻结区加固)、#10(Agent Loop,Task 9 透传 cache_salt)、#5/#12(信息阶梯/训练编排,Task 7 checkpoint 时序)。**均为恢复/加固/透传,不改算法逻辑**:Task 1 让诊断拿到本就该有的轨迹;Task 5 把"只查起点"补成"查整跨度"(保护方向不变);Task 9 只加透传参数;Task 7 只调 checkpoint 落盘时机。每个相关 Task 设保真检查点。
|
||
|
||
---
|
||
|
||
## Task 1: traces 适配(诊断拿到真实轨迹,算法 #7)
|
||
|
||
**Files:**
|
||
- Modify: `app/harness/runner.py:2163-2196`(_run_diagnosis)
|
||
- Test: `tests/unit/test_runner_diag_tree_inject.py` 或 `test_harness_runner.py`
|
||
|
||
- [ ] **Step 1: 写失败测试**
|
||
|
||
在 `tests/unit/test_runner_diag_tree_inject.py` 追加(构造只写 steps_json 不写 traces 表的 run,断言诊断能拿到轨迹):
|
||
```python
|
||
@pytest.mark.asyncio
|
||
async def test_diagnosis_reads_traces_from_steps_json(...):
|
||
"""traces 表为空但 predictions.steps_json 有轨迹时,诊断仍拿到非空 traces。"""
|
||
# 依现有 runner 测试 fixture 造一个 run:predictions 有 steps_json,traces 表空;
|
||
# 调 _run_diagnosis 后断言 diagnose 收到的 traces 非空(可 patch run_diagnosis 捕获入参)
|
||
...
|
||
```
|
||
> 依 `test_runner_diag_tree_inject.py` 现有 fixture;实现前读对齐 runner 构造与 patch 点。
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_runner_diag_tree_inject.py -k reads_traces_from_steps_json -v`
|
||
Expected: FAIL(当前 RunLogImpl 直读空 traces 表)。
|
||
|
||
- [ ] **Step 3: 包 StepsJsonRunLog**
|
||
|
||
`app/harness/runner.py` `_run_diagnosis`,把传给 `run_diagnosis` 的 `run_log`(当前 `RunLogImpl(...)`,L2173 附近)包一层:
|
||
```python
|
||
from app.harness.baseline_run_log import StepsJsonRunLog
|
||
from app.harness.log import RunLogImpl
|
||
|
||
run_log = StepsJsonRunLog(RunLogImpl(str(self._paths.db_path)))
|
||
```
|
||
(`StepsJsonRunLog.get_traces` 在底层 traces 空时从 predictions.steps_json 经 `steps_json_to_trace_rows` 重建;`get_predictions` 透传。)
|
||
|
||
- [ ] **Step 4: 保真检查点 + 测试通过**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_runner_diag_tree_inject.py tests/unit/test_baseline_run_log.py -q`
|
||
Expected: 全 PASS。保真:确认诊断瀑布拿到的是逐 step `{tool_name,tool_args,tool_output,thought}` 行(对齐 TRM4 诊断输入)。
|
||
|
||
- [ ] **Step 5: 提交**
|
||
|
||
```bash
|
||
git add app/harness/runner.py tests/unit/test_runner_diag_tree_inject.py
|
||
git commit -m "fix: wrap diagnosis run_log with StepsJsonRunLog (restore algo #7 traces)"
|
||
```
|
||
|
||
---
|
||
|
||
## Task 2: prediction 归一化 + 落库加固
|
||
|
||
**Files:**
|
||
- Modify: `app/harness/inference.py:417-447`
|
||
- Test: `tests/unit/test_harness_inference.py`
|
||
|
||
- [ ] **Step 1: 写失败测试**
|
||
|
||
在 `tests/unit/test_harness_inference.py` 追加(LLM 提交非标量 answer 不崩 gather):
|
||
```python
|
||
@pytest.mark.asyncio
|
||
async def test_nonscalar_prediction_does_not_crash(...):
|
||
"""submit_answer 返回 {'answer': ['B']} 等非标量时归一化落库,不抛 sqlite 绑定异常。"""
|
||
# 依现有 inference 测试 fixture,让 AgentLoop 返回 result={'answer': ['B']};
|
||
# run_inference 应正常完成、predictions 行 prediction 为字符串(如 '["B"]'),不崩
|
||
...
|
||
```
|
||
> 依 `test_harness_inference.py` 现有 fake loop/dispatch fixture;实现前读对齐。
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_inference.py -k nonscalar_prediction -v`
|
||
Expected: FAIL(sqlite `InterfaceError: Error binding parameter`)。
|
||
|
||
- [ ] **Step 3: 归一化 prediction + insert 加固**
|
||
|
||
`app/harness/inference.py`:新增归一化 helper(None 保留、str 原样、其余 `_to_text_field`):
|
||
```python
|
||
def _normalize_prediction(answer: object) -> str | None:
|
||
"""归一化 prediction:None 保留(INFRA 空预测语义),str 原样,其余 JSON 序列化。"""
|
||
if answer is None or isinstance(answer, str):
|
||
return answer
|
||
return _to_text_field(answer)
|
||
```
|
||
L422 `"prediction": result_dict.get("answer"),` 改为 `"prediction": _normalize_prediction(result_dict.get("answer")),`。
|
||
L446 的 `await asyncio.to_thread(log.insert, "predictions", record)` 包 try,绑定异常降级为最小 error 行不击穿 gather:
|
||
```python
|
||
try:
|
||
await asyncio.to_thread(log.insert, "predictions", record)
|
||
except (sqlite3.InterfaceError, sqlite3.ProgrammingError):
|
||
logger.exception("[{}] QA {} 落库绑定异常,降级为 error 行", qa.video_id, qa.question_id)
|
||
record["prediction"] = None
|
||
record["stop_reason"] = "error"
|
||
await asyncio.to_thread(
|
||
log.insert, "predictions",
|
||
{k: v for k, v in record.items() if isinstance(v, (str, int, float, type(None)))},
|
||
)
|
||
```
|
||
确认 `import sqlite3` 在文件顶部(无则加)。
|
||
|
||
- [ ] **Step 4: 测试通过 + 回归**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_inference.py -q`
|
||
Expected: 全 PASS。
|
||
|
||
- [ ] **Step 5: 提交**
|
||
|
||
```bash
|
||
git add app/harness/inference.py tests/unit/test_harness_inference.py
|
||
git commit -m "fix: normalize non-scalar prediction; harden predictions insert"
|
||
```
|
||
|
||
---
|
||
|
||
## Task 3: early_stop 改 epoch 计数
|
||
|
||
**Files:**
|
||
- Modify: `app/harness/runner.py:291-316,118,1587,849-857`
|
||
- Modify: `app/harness/checkpoint.py`(若字段入 state)
|
||
- Test: `tests/unit/test_harness_runner.py`
|
||
|
||
- [ ] **Step 1: 写失败测试**
|
||
|
||
在 `tests/unit/test_harness_runner.py` 追加:
|
||
```python
|
||
def test_early_stop_counts_epochs_not_steps(tmp_path):
|
||
"""patience=2 表示连续 2 个 epoch 无 best 刷新才停(不是步数)。"""
|
||
from app.harness.runner import _should_early_stop, _TrainState
|
||
# 造 state + workspace,best 停在 epoch 1;
|
||
# epoch 2 无刷新 → epochs_since_best_improved=1 → 不停;
|
||
# epoch 3 无刷新 → =2 → 停
|
||
...
|
||
```
|
||
> 依现有 `_TrainState`/`read_best` fixture;实现前读对齐 workspace best 写入。
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_runner.py -k early_stop_counts_epochs -v`
|
||
Expected: FAIL(当前累加 steps_this_epoch)。
|
||
|
||
- [ ] **Step 3: 字段改名 + 计数改 epoch**
|
||
|
||
全局把 `steps_since_best_improved` 改名 `epochs_since_best_improved`(`grep -rn steps_since_best_improved app/`:`_TrainState:118`、`_should_early_stop:313,315`、`_maybe_promote_best:1587`,以及 checkpoint serialize/deserialize 若含此字段)。
|
||
`_should_early_stop`(L315)`state.steps_since_best_improved += steps_this_epoch` 改为 `state.epochs_since_best_improved += 1`;签名删除 `steps_this_epoch` 参数(改为 `_should_early_stop(workspace_dir, epoch, state, patience)`),train 调用点(L849-857)同步去掉 `len(batches)` 实参。docstring 改为"epoch 粒度"。
|
||
|
||
- [ ] **Step 4: checkpoint 兼容**
|
||
|
||
若 `epochs_since_best_improved` 入 checkpoint state(`grep -n steps_since_best_improved app/harness/checkpoint.py`),同步改名。本轮为 fresh 训练无旧 checkpoint,无迁移负担。
|
||
|
||
- [ ] **Step 5: 测试通过 + 回归**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_runner.py tests/unit/test_harness_checkpoint.py -q`
|
||
Expected: 全 PASS。
|
||
|
||
- [ ] **Step 6: 提交**
|
||
|
||
```bash
|
||
git add app/harness/runner.py app/harness/checkpoint.py tests/unit/test_harness_runner.py
|
||
git commit -m "fix: early_stop patience counts epochs not steps"
|
||
```
|
||
|
||
---
|
||
|
||
## Task 4: 可训练性预检
|
||
|
||
**Files:**
|
||
- Modify: `app/harness/runner.py:865-879`(_setup_train_run)
|
||
- Modify: `app/harness/config.py`(RunConfig 加 `trainable_min_units`)+ `config/train_videomme.yaml`(WP 收尾建,先加字段默认)
|
||
- Test: `tests/unit/test_harness_runner.py`
|
||
|
||
- [ ] **Step 1: 写失败测试**
|
||
|
||
```python
|
||
def test_untrainable_types_filtered_before_gate():
|
||
"""val<eval_min_per_class 或 非test单元<trainable_min_units 的题型从 diag/val/task_types 剔除。"""
|
||
from app.harness.runner import _filter_untrainable_types
|
||
# 构造 pools:题型 A(val=5, units=40)可训;B(val=0, units=2)不可训
|
||
# 调 _filter_untrainable_types(pools, task_types=[A,B], eval_min_per_class=2, trainable_min_units=8)
|
||
# 断言返回 pools 不含 B、task_types 不含 B
|
||
...
|
||
```
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_runner.py -k untrainable_types_filtered -v`
|
||
Expected: FAIL(`_filter_untrainable_types` 不存在)。
|
||
|
||
- [ ] **Step 3: 实现纯函数预检**
|
||
|
||
`app/harness/runner.py` 新增模块级纯函数:
|
||
```python
|
||
def _filter_untrainable_types(
|
||
pools: Pools,
|
||
task_types: list[str] | None,
|
||
eval_min_per_class: int,
|
||
trainable_min_units: int,
|
||
) -> tuple[Pools, list[str] | None]:
|
||
"""剔除不可训练题型(val<eval_min_per_class 或 非test单元<trainable_min_units)。
|
||
|
||
非test单元数 = 该题型 diag+val 题数(single 题 unit==题;等于 gate 阶梯该类候选数)。
|
||
test 池不过滤(继续报告全题型准确率)。返回过滤后 (pools, task_types)。
|
||
"""
|
||
from collections import Counter
|
||
|
||
diag_by_type = Counter(q.task_type for q in pools.diagnosis)
|
||
val_by_type = Counter(q.task_type for q in pools.validation)
|
||
keep: set[str] = set()
|
||
dropped: list[tuple[str, str]] = []
|
||
for tt in set(diag_by_type) | set(val_by_type):
|
||
n_val = val_by_type.get(tt, 0)
|
||
n_units = diag_by_type.get(tt, 0) + n_val
|
||
if n_val < eval_min_per_class:
|
||
dropped.append((tt, f"val={n_val}<{eval_min_per_class}"))
|
||
elif n_units < trainable_min_units:
|
||
dropped.append((tt, f"units={n_units}<{trainable_min_units}"))
|
||
else:
|
||
keep.add(tt)
|
||
for tt, why in sorted(dropped):
|
||
logger.warning("可训练性预检剔除题型 {}({})", tt, why)
|
||
new_pools = replace(
|
||
pools,
|
||
diagnosis=[q for q in pools.diagnosis if q.task_type in keep],
|
||
validation=[q for q in pools.validation if q.task_type in keep],
|
||
)
|
||
new_types = [t for t in task_types if t in keep] if task_types is not None else sorted(keep)
|
||
return new_pools, new_types
|
||
```
|
||
(确认 `from dataclasses import replace` 已 import;`Pools` 是否 frozen dataclass 支持 `replace`——若非,按其构造方式重建。)
|
||
在 `_setup_train_run`(pools 加载后、`_init_gate_pools` L874 前)调用:
|
||
```python
|
||
pools, self._config.task_types = _filter_untrainable_types(
|
||
pools, self._config.task_types,
|
||
self._config.eval_min_per_class, self._config.trainable_min_units,
|
||
)
|
||
```
|
||
(`RunConfig` 若 frozen,改用局部变量传递而非改 self._config;实现前读 config.py 确认可变性。)
|
||
`app/harness/config.py` `RunConfig` 加字段 `trainable_min_units: int`(无默认,显式配置;train yaml 提供)。
|
||
|
||
- [ ] **Step 4: 测试通过 + 回归**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_runner.py -q`
|
||
Expected: 全 PASS。
|
||
|
||
- [ ] **Step 5: 提交**
|
||
|
||
```bash
|
||
git add app/harness/runner.py app/harness/config.py tests/unit/test_harness_runner.py
|
||
git commit -m "feat: pre-flight filter of untrainable task types before gate"
|
||
```
|
||
|
||
---
|
||
|
||
## Task 5: patch 冻结区跨度 + 注入 + marker 校验(算法 #8)
|
||
|
||
**Files:**
|
||
- Modify: `core/evolution/patch.py:285-347`
|
||
- Modify: `core/evolution/evolve.py:296`(validate_skill)
|
||
- Test: `tests/unit/test_patch.py`
|
||
|
||
- [ ] **Step 1: 写失败测试**
|
||
|
||
在 `tests/unit/test_patch.py` 追加:
|
||
```python
|
||
def test_replace_spanning_into_protected_is_skipped():
|
||
"""target 起点在正文、末端伸入冻结区的 replace 被跳过(不破坏 marker)。"""
|
||
from core.evolution.patch import apply_patch_with_report, APPENDIX_START, APPENDIX_END
|
||
|
||
body = "正文最后一段。"
|
||
appendix = f"{APPENDIX_START}\n## 执行提醒\n- 规则A\n{APPENDIX_END}"
|
||
content = body + "\n\n" + appendix
|
||
# target 从正文末尾跨入 APPENDIX_START
|
||
target = "正文最后一段。\n\n" + APPENDIX_START
|
||
edits = [{"op": "delete", "target": target, "content": ""}]
|
||
new_content, report = apply_patch_with_report(content, edits, protected_spans=[appendix])
|
||
assert APPENDIX_START in new_content and APPENDIX_END in new_content # marker 未被破坏
|
||
|
||
|
||
def test_edit_payload_with_marker_literal_rejected():
|
||
"""edit payload/target 含 marker 字面量 → 拒绝该 edit。"""
|
||
from core.evolution.patch import apply_patch_with_report, APPENDIX_START
|
||
edits = [{"op": "append", "target": "", "content": f"注入 {APPENDIX_START} 破坏"}]
|
||
_, report = apply_patch_with_report("正文", edits, protected_spans=[])
|
||
assert any("marker" in str(s).lower() or "reject" in str(s).lower() for s in report)
|
||
```
|
||
> marker 常量名以 `core/evolution/patch.py:11-18` 为准(`APPENDIX_START` 等),实现前读对齐 import 名。
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_patch.py -k "spanning_into_protected or marker_literal" -v`
|
||
Expected: FAIL(当前只查起点 pos,跨入未拦;无注入检查)。
|
||
|
||
- [ ] **Step 3: 跨度检查 + 注入检查**
|
||
|
||
`core/evolution/patch.py` 新增跨度 helper:
|
||
```python
|
||
def _span_overlaps_ranges(pos: int, length: int, ranges: list[tuple[int, int]]) -> bool:
|
||
"""判断 [pos, pos+length) 是否与任一冻结区间相交(不止起点)。"""
|
||
end = pos + length
|
||
return any(start < end and pos < r_end for start, r_end in ranges)
|
||
```
|
||
`_do_insert_after` L320 `if _in_ranges(pos, ranges):` 改为 `if _span_overlaps_ranges(pos, len(target), ranges):`。
|
||
`_do_replace_delete` L343 `if _in_ranges(pos, ranges):` 改为 `if _span_overlaps_ranges(pos, len(target), ranges):`。
|
||
在 `apply_patch_with_report`(L387)应用每个 edit 前加注入检查:payload/target 含任一 marker 字面量(`APPENDIX_START/END`、`MOMENTUM_START/END`)→ 跳过该 edit 并记 `skipped_marker_injection`。
|
||
|
||
- [ ] **Step 4: validate_skill 加 marker 完整性校验**
|
||
|
||
`core/evolution/evolve.py:296` `validate_skill`:在现有 frontmatter/长度/代码块校验后加——统计 evolved 中 `APPENDIX_START/END`、`MOMENTUM_START/END` 出现次数,要求成对(START 数==END 数)、各至多一对、START 在 END 前;违反则返回校验失败(该候选整体 reject)。
|
||
|
||
- [ ] **Step 5: 保真检查点 + 测试通过**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_patch.py tests/unit/test_evolve.py -q`
|
||
Expected: 全 PASS。保真:确认"保护跨度"方向未变(仍是保护 appendix/momentum 不被误改),只是从"查起点"补成"查整跨度"+ 注入/完整性双防线。
|
||
|
||
- [ ] **Step 6: 提交**
|
||
|
||
```bash
|
||
git add core/evolution/patch.py core/evolution/evolve.py tests/unit/test_patch.py
|
||
git commit -m "fix: patch checks full target span + marker injection/integrity (algo #8)"
|
||
```
|
||
|
||
---
|
||
|
||
## Task 6: 诊断降级分流 + 占比中止
|
||
|
||
**Files:**
|
||
- Modify: `core/evolution/diagnose.py:1485-1497`
|
||
- Modify: `app/harness/runner.py:1019`(诊断后 degraded 占比检查)
|
||
- Test: `tests/unit/test_diagnose.py`、`tests/unit/test_harness_runner.py`
|
||
|
||
- [ ] **Step 1: 写失败测试(分流)**
|
||
|
||
在 `tests/unit/test_diagnose.py` 追加:
|
||
```python
|
||
def test_none_cause_and_degraded_route_to_lapse():
|
||
"""cause_category=None(判别失败)与 degraded 题按 lapse 处置,不进 defect 正文路径。"""
|
||
from core.evolution.diagnose import _build_skill_case_packs
|
||
# 构造 metrics_group:一题 attr.cause_category=None(非 degraded)、一题 qm.degraded=True;
|
||
# 断言二者都不出现在 failure_cases(wrong_by_error),只可能进 lapse_notes
|
||
...
|
||
```
|
||
> 依 `test_diagnose.py` 现有 QuestionMetrics/ErrorAttribution 构造(`:753` 附近);实现前读对齐字段。
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_diagnose.py -k none_cause_and_degraded_route -v`
|
||
Expected: FAIL(当前 None → wrong_by_error 走 defect 正文)。
|
||
|
||
- [ ] **Step 3: 改分流逻辑**
|
||
|
||
`core/evolution/diagnose.py` `_build_skill_case_packs` 的分流循环(L1488-1497)改为"仅明确 defect 且非 degraded 才进正文路径":
|
||
```python
|
||
for qm in metrics_group:
|
||
if qm.correct:
|
||
continue
|
||
attr = attribution_map.get(qm.question_id)
|
||
is_defect = (
|
||
attr is not None
|
||
and attr.cause_category == "defect"
|
||
and not qm.degraded
|
||
)
|
||
if not is_defect:
|
||
# lapse / None(判别失败)/ degraded → 保守,不驱动正文进化
|
||
if attr is not None and attr.lapse_note and attr.lapse_note.strip():
|
||
lapse_notes.append(attr.lapse_note)
|
||
continue
|
||
wrong_by_error[attr.error_type].append(qm)
|
||
```
|
||
|
||
- [ ] **Step 4: 写失败测试(占比中止)+ 实现**
|
||
|
||
`tests/unit/test_harness_runner.py` 追加:诊断结果 degraded_count/总题数 > 0.5 时 `_run_step` 后应 raise(疑似基础设施故障)。
|
||
`app/harness/runner.py` `_run_step`(L1019 拿到 `diagnosis` 后)加:
|
||
```python
|
||
n_wrong = sum(1 for q in batch if not state.correctness.get(q.question_id, True))
|
||
if n_wrong > 0 and diagnosis.degraded_count / n_wrong > 0.5:
|
||
raise RuntimeError(
|
||
f"本 step 诊断降级占比 {diagnosis.degraded_count}/{n_wrong} > 50%,"
|
||
"疑似 judge 基础设施故障,中止训练(不以降级信号驱动进化)。"
|
||
)
|
||
```
|
||
(确认 `DiagnosisResult.degraded_count` 字段可用,L2311。)
|
||
|
||
- [ ] **Step 5: 测试通过 + 回归**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_diagnose.py tests/unit/test_harness_runner.py -q`
|
||
Expected: 全 PASS。
|
||
|
||
- [ ] **Step 6: 提交**
|
||
|
||
```bash
|
||
git add core/evolution/diagnose.py app/harness/runner.py tests/unit/test_diagnose.py tests/unit/test_harness_runner.py
|
||
git commit -m "fix: route None/degraded diagnoses to lapse; abort on high degrade rate"
|
||
```
|
||
|
||
---
|
||
|
||
## Task 7: step 幂等(DELETE)+ gate_epoch_observed 立即落盘
|
||
|
||
**Files:**
|
||
- Modify: `app/harness/runner.py:1001-1026`(_run_step 开头 DELETE)
|
||
- Modify: `app/harness/runner.py:1378-1498,1833-1879`(checkpoint 提到 gate save 之后)
|
||
- Test: `tests/unit/test_harness_runner.py`
|
||
|
||
- [ ] **Step 1: 写失败测试(幂等)**
|
||
|
||
```python
|
||
@pytest.mark.asyncio
|
||
async def test_run_step_deletes_stale_rows_before_rerun(...):
|
||
"""同 run_id 重跑前先清 predictions/traces,避免重复行双计。"""
|
||
# 预置该 step run_id 的旧 predictions 行;调 _run_step;断言旧行被清、只剩本次
|
||
...
|
||
```
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_runner.py -k deletes_stale_rows -v`
|
||
Expected: FAIL(当前 append,无 DELETE)。
|
||
|
||
- [ ] **Step 3: _run_step 开头 DELETE**
|
||
|
||
`app/harness/runner.py` `_run_step`,在 rollout(L1012)之前加:
|
||
```python
|
||
with HarnessLog(str(self._paths.db_path), run_id, register_run=False) as log:
|
||
log.execute("DELETE FROM predictions WHERE run_id=?", (run_id,))
|
||
log.execute("DELETE FROM traces WHERE run_id=?", (run_id,))
|
||
```
|
||
(`register_run=False` 来自 WP4 Task 8;若 HarnessLog 无 `execute` 便捷方法,用其现有 `query`/连接接口,实现前读 log.py 对齐。)
|
||
|
||
- [ ] **Step 4: checkpoint 提到 gate save 之后(消除双计窗口)**
|
||
|
||
目标:`_refresh_gate_ladder` 内 `gate_pools.save`(L1878)+ `gate_epoch_observed=True`(L1879)之后,**立即** `write_checkpoint`(phase="epoch_done"),不等到 `train` L837。实现:把 `_slow_update_cycle` Phase 10(调 `_refresh_gate_ladder` L1496-1498)之后的 checkpoint 落盘从 `train`(L837)移入 `_slow_update_cycle` 末尾,或让 `_refresh_gate_ladder` 接收 checkpoint 所需上下文(epoch/progress/batches)并在 save 后落盘。实现前读 `write_checkpoint` 签名(checkpoint.py:212)与 `train` L833-848 对齐参数,确保 gate_pools.json 与 checkpoint 的 `gate_epoch_observed` 同一时刻一致。
|
||
|
||
- [ ] **Step 5: 保真检查点 + 测试通过**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_runner.py tests/unit/test_harness_checkpoint.py -q`
|
||
Expected: 全 PASS。保真(算法 #5/#12):确认 γ-EMA 更新(`update_probs`)仍每 epoch 一次、checkpoint 落盘不改变慢更新十步序的语义顺序。
|
||
|
||
- [ ] **Step 6: 提交**
|
||
|
||
```bash
|
||
git add app/harness/runner.py tests/unit/test_harness_runner.py
|
||
git commit -m "fix: idempotent _run_step (DELETE stale) + checkpoint after gate save"
|
||
```
|
||
|
||
---
|
||
|
||
## Task 8: holdout 四向去重
|
||
|
||
**Files:**
|
||
- Modify: `app/harness/runner.py:1881-1963`(_holdout_four_way / _pick_mixed_best)
|
||
- Test: `tests/unit/test_harness_runner.py`
|
||
|
||
> **说明**:目标是每 epoch 的四向 test 评估从"4×600 全跑"降为"仅 final 必跑 + best_hard 未评过才跑 + baseline 从基线预测推导 + best_mixed 引用赢家"。去重做在 harness 逻辑层(配合 WP4 epoch 盐,同版本不重采样)。
|
||
|
||
- [ ] **Step 1: 写失败测试**
|
||
|
||
```python
|
||
@pytest.mark.asyncio
|
||
async def test_holdout_dedup_skips_reevaluated_versions(...):
|
||
"""baseline 不跑推理(从基线预测推导);best_hard==final 时不重复评估。"""
|
||
# 统计 _eval_version_on_pool 被调次数:baseline=0,best_hard==final 时该向复用不重跑
|
||
...
|
||
```
|
||
> 依现有 runner holdout fixture;实现前读 `_holdout_four_way`/`write_holdout_eval` 对齐。
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_runner.py -k holdout_dedup -v`
|
||
Expected: FAIL(当前四向各跑一次)。
|
||
|
||
- [ ] **Step 3: 实现去重**
|
||
|
||
`app/harness/runner.py` `_holdout_four_way`:
|
||
- **baseline 向**:不调 `_eval_version_on_pool`,改从基线 predictions(baseline_run_id)读 test 题对错算 acc(test 题在 infer_adhoc 已推理);epoch>1 复用首次结果。
|
||
- **final 向**:真评 600(不变)。
|
||
- **best_hard 向**:维护"已评版本→test_acc"备忘录,resume 时从 `holdout_eval` 表 hydrate(查已落库 `(skills_v,prompts_v)→acc`);版本对未评过且 ≠ final 才真评,否则引用。
|
||
- **best_mixed 向**:引用赢家(best_hard 或 final 之一,`_pick_mixed_best` 已定)的 test 成绩,写 holdout_eval 时标指针,不推理。
|
||
|
||
实现前完整读 `_holdout_four_way`(1881-1921)、`write_holdout_eval`、`_eval_version_on_pool`、`read`/hydrate 路径对齐。
|
||
|
||
- [ ] **Step 4: 测试通过 + 回归**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_runner.py -q`
|
||
Expected: 全 PASS。
|
||
|
||
- [ ] **Step 5: 提交**
|
||
|
||
```bash
|
||
git add app/harness/runner.py tests/unit/test_harness_runner.py
|
||
git commit -m "perf: dedup holdout four-way eval (baseline derive, best_hard memo)"
|
||
```
|
||
|
||
---
|
||
|
||
## Task 9: 训练推理注入 epoch 盐(cache_salt=run_id,算法 #10 透传)
|
||
|
||
**Files:**
|
||
- Modify: `core/agent/loop.py:103-141,271-329`
|
||
- Modify: `app/harness/inference.py:408-415,473`
|
||
- Test: `tests/unit/test_agent_loop`(或现有 loop 测试)、`tests/unit/test_harness_inference.py`
|
||
|
||
> **原理**:训练/val/test/holdout 推理的 run_id 已含 `_e{epoch}`(如 `{base}_e{epoch}_s{step}`、`{run_id}_holdout_{kind}_e{epoch}`),用 run_id 作 cache_salt 即天然跨 epoch 重采样、同 epoch 续跑仍命中。judge/evolve 不经此路径(默认 salt=None)。gate 基线臂走 BaselineCache 不受影响;候选臂 messages 含 skill 版本天然区分。
|
||
|
||
- [ ] **Step 1: 写失败测试**
|
||
|
||
```python
|
||
@pytest.mark.asyncio
|
||
async def test_agent_loop_forwards_cache_salt():
|
||
"""AgentLoop.run(cache_salt=...) 透传到 llm.chat。"""
|
||
from core.agent.loop import AgentLoop
|
||
# fake llm 记录 chat 收到的 cache_salt kwarg;loop.run(..., cache_salt='run:e2')
|
||
# 断言 fake_llm.chat 收到 cache_salt='run:e2'
|
||
...
|
||
```
|
||
|
||
- [ ] **Step 2: 运行确认失败**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest -k agent_loop_forwards_cache_salt -v`
|
||
Expected: FAIL(`run()` 无 cache_salt 参数)。
|
||
|
||
- [ ] **Step 3: AgentLoop 透传 cache_salt**
|
||
|
||
`core/agent/loop.py`:`run`(L103)、`_step`/`_call_llm`(L271,317)签名加 `cache_salt: str | None = None`(keyword,随 session_id 透传);L329 `self._llm.chat(messages, session_id=session_id)` 改为 `self._llm.chat(messages, session_id=session_id, cache_salt=cache_salt)`。
|
||
|
||
- [ ] **Step 4: inference 用 run_id 作 salt**
|
||
|
||
`app/harness/inference.py` `_run_single_question`:`loop.run(...)`(L409)加 `cache_salt=run_id`(run_id 从 run_inference 透传到每题;`_run_single_question` 已有 run_id 上下文——若无则从 run_inference 参数透传)。确认 run_inference→_run_single_question 的 run_id 传递链完整。
|
||
|
||
- [ ] **Step 5: 保真检查点 + 测试通过**
|
||
|
||
Run: `conda run -n Video-Tree-TRM python -m pytest tests/unit/test_harness_inference.py tests/integration/test_agent_governed_e2e.py -q`
|
||
Expected: 全 PASS。保真(算法 #10):确认只加透传参数,Thinking+JSON/json_repair/pluggy hook 逻辑不变。
|
||
|
||
- [ ] **Step 6: 提交**
|
||
|
||
```bash
|
||
git add core/agent/loop.py app/harness/inference.py tests/unit/
|
||
git commit -m "feat: inject run_id as cache_salt for per-epoch resampling (algo #10)"
|
||
```
|
||
|
||
---
|
||
|
||
## Self-Review(作者自查,执行者复核)
|
||
|
||
- [ ] traces 适配后诊断拿到真实轨迹(算法 #7 恢复)。
|
||
- [ ] prediction 归一化 None 保留、非标量序列化;insert 绑定异常不击穿 gather。
|
||
- [ ] early_stop 字段全局改名一致、计数改 epoch。
|
||
- [ ] 预检剔除不可训练题型(test 池不动)。
|
||
- [ ] patch 查整跨度 + 注入 + marker 完整性三防线(算法 #8 保护方向不变)。
|
||
- [ ] None/degraded 按 lapse 保守分流;降级占比>50% 中止。
|
||
- [ ] _run_step 幂等;gate_pools 与 checkpoint 的 gate_epoch_observed 同刻一致。
|
||
- [ ] holdout 去重后 baseline 0 推理、best_hard 备忘录 resume 可 hydrate。
|
||
- [ ] cache_salt=run_id 贯穿 AgentLoop,run_id 含 epoch 保证跨 epoch 重采样。
|
||
|
||
## 核心算法保真校验结论
|
||
|
||
触及算法 #5/#7/#8/#10/#12,均为恢复(#7 轨迹)/加固(#8 跨度)/透传(#10 salt)/时序(#5/#12 checkpoint),各 Task 已设保真检查点,不改算法核心逻辑。Task 5/7/9 需在实现时对照 TRM4 参考确认无行为漂移。
|
||
|
||
## 验收标准
|
||
|
||
1. `pytest tests/unit/test_harness_runner.py tests/unit/test_harness_inference.py tests/unit/test_diagnose.py tests/unit/test_patch.py tests/unit/test_harness_checkpoint.py tests/unit/test_baseline_run_log.py tests/unit/test_runner_diag_tree_inject.py` 全绿。
|
||
2. 诊断拿到非空轨迹;非标量 prediction 不崩;early_stop 按 epoch。
|
||
3. 微型题型被预检剔除;patch 不破坏 marker;降级题不驱动进化。
|
||
4. _run_step 幂等;cache_salt=run_id 贯穿。
|