iomgaa
0d0f275134
feat(core): add PoolConfig dataclass for pool strategy configuration
2026-07-12 22:33:49 -04:00
iomgaa
4fb7a61f8b
fix(question_gen): resolve pipeline integration issues from final review
...
1. Apply postprocess shuffle result (pp.options, pp.answer) to final
GeneratedQuestion output instead of using original candidate values.
2. Record dedup rejection in store via new mark_item_rejected() method,
preventing items from staying as 'accepted' after dedup rejects them.
3. Add .flatten() to embed_fn outputs in _is_duplicate and embed_pool
append to handle 2D (1,D) arrays from embedding implementations.
4. Validate exactly 4 options in _validate_parsed_fields (was >= 2),
matching the A-D answer constraint.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 00:14:12 -04:00
iomgaa
eecb86e27a
feat(question_gen): add generate-v2 CLI subcommand and experiment script
...
- Add generate-v2 subparser with --config, --store-dir, --db-path,
--seed, and --dry-run arguments to tools/generate_questions.py
- Implement _run_generate_v2 async handler: config loading, video
discovery, DI client construction, TreeIndex loading, pipeline
invocation, and result persistence
- Add scripts/generate_questions_v2.sh following build_trees.sh
conventions (source .env, conda run python path, MODE=mock support)
- Update app/question_gen/__init__.py to export full v2 public API:
run_pipeline_v2, PipelineConfig, PipelineResult, QuestionFamilySpec,
ALL_FAMILIES, CandidateQuestion, generate_one_v2, GateReport, run_gates
- Add QuestionGenStore.load_progress() for pipeline resumption
- Add integration tests for CLI help and dry-run behavior
- Update test_question_gen_api to match expanded __all__
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 00:04:42 -04:00
iomgaa
206c553143
feat(question_gen): add v2 pipeline with retry loop and heavy check
...
- PipelineConfig: YAML-driven configuration with family_ratios, retry,
concurrency, dedup threshold, and heavy sampling rate
- _assign_slots: deterministic round-robin slot assignment across videos
with per-family weighted random selection
- _process_one_slot: full retry loop (generate → postprocess → gates →
dedup) with reject-reason feedback to VLM on retry
- _heavy_check_one: blind LLM agent trial-answer for difficulty_steps
- run_pipeline_v2: orchestration with semaphore-bounded concurrency,
progress/resume support, and store integration
- is_duplicate: cosine similarity dedup against embedding pool
Tests: 11 integration tests covering slot assignment, retry behavior,
max-retries exhaustion, full pipeline flow, progress resume, heavy
sampling, and store record completeness.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:55:15 -04:00
iomgaa
cf51d2de9d
feat(question_gen): add v2 generator with per-family prompt templates
...
Implement generator_v2.py with:
- CandidateQuestion dataclass (canonical location)
- _load_prompt_template: loads per-family .md from store/prompts/
- _build_v2_prompt: constructs system+user messages with material context
- _parse_v2_response: JSON extraction, json_repair, field validation
- generate_one_v2: async VLM call orchestration with reject_reason support
Add 5 family-specific prompt templates:
- retrieval.md: factual recall from visible content
- reasoning.md: multi-hop inference across segments
- enumeration.md: counting/listing entities and actions
- visual.md: visual details requiring frame observation
- spatial.md: spatial relationships between objects/people
Tests: 11 unit tests covering prompt build, parse, and e2e generation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:41:36 -04:00
iomgaa
9053233f99
fix(question_gen): gates reject_reason returns raw reason; short-circuit skips LLM
...
- GateReport.reject_reason now returns gate.reason directly (no [name] prefix)
- verbatim short-circuit sets other 3 gates to SKIP without calling LLM
- test_high_verbatim_shortcircuits asserts zero LLM calls and SKIP verdicts
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:33:57 -04:00
iomgaa
f74711cd11
feat(question_gen): add v2 material sampler with family constraints
...
Implement sample_material_v2 module that samples tree nodes with
QuestionFamilySpec-aware constraint validation, providing richer
MaterialContext output (subtitles, cross-L2 context, frame paths).
Key components:
- AnchorContext/MaterialContext frozen dataclasses
- _validate_sampling_constraints: multi-level constraint checking
- _collect_subtitle_sentences: subtree subtitle extraction
- _collect_cross_l2_context: peer L2 event descriptions
- sample_material_v2: main entry with retry-on-constraint-violation
Tests: 11 unit tests covering normal sampling, used-node exclusion,
constraint violation retries, cross-L2 population, and subtitle
collection.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:30:56 -04:00
iomgaa
271d1682c9
feat(question_gen): add lightweight 4-gate quality check
...
Implement 4 concurrent LLM-based quality gates for generated questions:
- key_verify: validates answer evidence in source material
- blind_answer: rejects questions answerable without video context
- multi_true: detects ambiguous multi-correct options
- leak_test: per-family shortcut detection (5 probe templates)
Includes run_gates orchestrator with verbatim_ratio short-circuit,
JSON response parsing with fallback, and 9 unit tests (all passing).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:30:13 -04:00
iomgaa
9627ac9cf9
fix(question_gen): raise ValueError on UPDATE of missing rows
...
record_run_end, update_gates, and update_difficulty now check
cursor.rowcount after UPDATE+commit and raise ValueError if 0 rows
were affected. Prevents silent telemetry loss.
Adds three negative-path tests:
- test_record_run_end_missing_run_raises
- test_update_gates_missing_item_raises
- test_update_difficulty_missing_item_raises
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:23:48 -04:00
iomgaa
7abe92eb1c
fix(question_gen): check_verbatim covers question_text + add missing blacklist patterns
...
- check_verbatim now computes n-gram overlap for BOTH question_text and
correct_option vs source texts, returning max(question_ratio, option_ratio).
Extracted _ngram_overlap_ratio helper for reuse.
- Added 4 missing blacklist patterns: 'this segment', 'this frame',
'the current frame', 'frame summary'.
- Added 5 new test cases covering the above changes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:23:37 -04:00
iomgaa
6a5424a618
feat(question_gen): add SQLite run store for generation telemetry
...
Implements QuestionGenStore with:
- Idempotent schema initialization (question_gen_runs + question_gen_items)
- Run lifecycle: record_run_start / record_run_end / get_run_stats
- Per-item recording: record_item / update_gates / update_difficulty
- GateReportLike Protocol for duck-type gate report compatibility
- WAL mode + foreign keys + check_same_thread=False
DDL aligns with research-wiki/schemas/question-gen-{runs,items}.md.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:15:53 -04:00
iomgaa
c83d771923
feat(question_gen): add deterministic postprocess layer
...
Add app/question_gen/postprocess.py with zero-LLM deterministic
post-processing for generated questions:
- shuffle_options: deterministic option permutation with answer remapping
- check_referent_blacklist: detect self-referential language (this clip, etc.)
- check_verbatim: word-level n-gram overlap ratio measurement
- has_time_anchor: timestamp and temporal phrase detection
- check_forbidden_material: T1/T7 source material validation
- run_postprocess: orchestration returning PostprocessResult
Tests: 31 unit tests covering all functions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:14:43 -04:00
iomgaa
75e6d8c550
feat(question_gen): add 5 question family specs with sampling constraints
...
Define QuestionFamilySpec, LeakTestProfile, SamplingConstraint dataclasses
and instantiate 5 families (RETRIEVAL/REASONING/ENUMERATION/VISUAL/SPATIAL)
targeting failure mechanisms M1-M5. Implement get_family_for_slot with
legal-type filtering + weighted random selection.
13 unit tests cover: full task-type coverage, skill_target uniqueness,
deterministic seeding, invalid input errors, and chi-square distribution.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:14:22 -04:00
iomgaa
811ffa648b
feat(types): extend GeneratedQuestion with skill_target & difficulty_steps
...
- Add skill_target (str | None) and difficulty_steps (int | None) fields
to GeneratedQuestion dataclass with field(default=None)
- Update loader.py to pass new fields from JSON (backward-compatible)
- Update pools.py _q_to_dict/_dict_to_q for serialization compat
- Add question_gen_v2 config section to default.yaml
- Add comprehensive test coverage (7 tests)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:05:38 -04:00
iomgaa
56fe138a46
feat(tools): batch tree build orchestration with shared API semaphore
2026-07-11 11:56:36 -04:00
iomgaa
978ddef91b
feat(tools): add build_trees skeleton with discovery helpers
2026-07-11 11:41:22 -04:00
iomgaa
e9073bfdc2
feat(tree): expose build_async and accept injected API semaphore
...
Core algorithms #1/#2/#3 unchanged: only entry wrapping and semaphore
source switch (injected vs self-built); build logic untouched.
- rename _build_async to public build_async (body unchanged)
- __init__ accepts keyword-only api_semaphore for cross-video sharing
- default path (no injection) is verbatim-equivalent to previous code
2026-07-11 11:14:45 -04:00
iomgaa
1c21e215e2
test(tree): cover L1 entity fields and block ordering in view_node
2026-07-11 09:02:03 -04:00
iomgaa
4e0e05210d
feat(search): append raw entity fields after view_node summary
...
view_node 按题两轮摘要(summarize_node)会吞掉 entities/visible_text
字段信号,Agent 站在证据节点上仍漏读实体(benchmark 错题 M1,案例
786-2、872-3、750-1)。dispatcher 侧在摘要后确定性追加 [实体]/[画面
文字] 原文区块,LLM 无法吞掉。
- TreeEnvironment.node_entity_fields:按层级取 card 实体字段原文,
去空白、去重、分号拼接;空字段省键;未知节点抛 KeyError
- _handle_view_node Phase 2.5:摘要后、子节点概览前追加实体区块
- 附带 ruff format 修正 test_tree_environment.py 两处既有格式
算法 #11(树环境语义搜索)数据访问层扩展,不改搜索算法本身。
2026-07-11 08:53:56 -04:00
iomgaa
291a8108e1
test(agent): tighten type annotations in step-retry tests
2026-07-11 08:48:04 -04:00
iomgaa
badfcce4cb
fix(agent): validate non-empty retry delays; pin exhaustion semantics
...
核心算法 #10(Agent Loop):Codex 质量审查跟进,仅加固防御与
可观测,不改变循环语义。
- __init__ 校验 step_retry_delays 非空,空序列直接抛 ValueError
(P5 fail-fast,避免首次可重试异常时 IndexError 掩盖原始 LLM
异常、违背方法契约)
- 补耗尽语义窄测试(耗尽后原样抛出最后一次原始异常)与空序列
构造校验测试
- run() 最终失败日志补异常类型名,便于回溯归因
2026-07-11 08:44:47 -04:00
iomgaa
e3184c11f9
feat(agent): step-level retry for transient LLM errors (20s/40s backoff)
...
核心算法 #10(Agent Loop):仅加固异常路径的韧性兜底,不改变
解析协议、hook 时序与步数语义。benchmark 错题 796-3 显示一次
SSL BAD_RECORD_MAC 穿透 GovernedLLMClient 重试栈后废掉 13 步
已积累上下文;本次在 run() Phase 1 增加步级重试(默认 2 次,
20s/40s 退避),可重试异常限定 (TimeoutError, OSError),非可
重试异常照旧 fail-fast 整题终止,行为与现状一致(P5 显式异常)。
为满足 radon C 级复杂度约束,重试循环抽取为私有方法
_call_llm_with_step_retry,行为不变。
2026-07-11 08:37:35 -04:00
iomgaa
cf529f2c8f
test(agent): add docstring to empty-content boundary test
2026-07-11 08:33:36 -04:00
iomgaa
439dc29b3b
fix(agent): reject argless action in normalization; add boundary test
...
核心算法 #10(Agent Loop):修复 Codex 质量审查 Critical——
_normalize_action 仅在除 tool 外至少存在一个平铺参数键时才收拢,
{"tool": "x"} 无参结构不再被静默升级为空 args 合法结构,照旧
返回 None 走 retry 追问路径。补边界测试 + 测试辅助方法类型注解
与中文 docstring。
2026-07-11 08:28:59 -04:00
iomgaa
6034d4172d
test(question_gen): update __all__ assertion to include synthesizer exports
2026-07-11 08:21:06 -04:00
iomgaa
a31b1fbf37
test(question_gen): drop tests for removed calibrate pairing validator
2026-07-11 08:19:45 -04:00
iomgaa
8d84d5e236
fix(agent): normalize fenced and flat-args LLM outputs in parser
...
核心算法 #10(Agent Loop):仅加固 _parse_response 解析路径——
剥除 ```json 围栏 + 收拢 action 平铺参数(deepseek 稳定输出变体,
案例 637-3/615-3 三连拒 0 步阵亡)。Thinking+JSON 协议、json_repair
兜底链、hook 时序与步数语义均未改动。
2026-07-11 08:16:28 -04:00
iomgaa
ba400417e2
config: concurrency=24, max_steps=40, breaker_threshold=48
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 12:27:05 -04:00
iomgaa
a7ca6d15ed
feat(harness): InferenceDepsRouter per-video 路由器
...
按 video_id 懒加载 InferenceDeps 并缓存,路由 dispatch/prompt_builder:
- create_dispatch: 按 session_id 路由到对应视频的工具调度
- create_prompt_builder: 自动注册 qid→vid 映射并路由 prompt 构建
- 三元组 (video_id, skills_dir, prompts_dir) 缓存键
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 12:20:32 -04:00
iomgaa
f21bf345a6
feat(runner): 注入 tool_dispatch_factory/prompt_builder_factory + fail-fast
...
Runner.__init__ 新增 2 个可选参数:
- tool_dispatch_factory: 工具调度工厂
- prompt_builder_factory: prompt 构建工厂
infer/eval/train 模式缺少工厂时 fail-fast 抛 ValueError。
_make_tool_dispatch_fn/_make_prompt_builder 优先使用注入工厂。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 12:20:25 -04:00
iomgaa
924160c779
feat(ports): 新增 ToolDispatchFactory/PromptBuilderFactory Protocol
...
4 个新 Protocol 类型:
- ToolDispatchFn: 工具调度函数签名
- ToolDispatchFactory: per-version 工具调度工厂
- PromptBuilderFn: Prompt 构建函数签名
- PromptBuilderFactory: per-version prompt 构建工厂
含 runtime_checkable isinstance 测试。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 12:20:16 -04:00
iomgaa
d3be9b1322
refactor(tree): subtitle 迁入 L3Card/L2Card + 建树管线修正
...
- L3Card/L2Card 新增 subtitle: str 字段(L1Card 不加)
- L3Node 移除 subtitle 字段(数据迁入 Card)
- assign_subtitles_voronoi 改写 Card.subtitle + L2 聚合
- _collect_card_strings 增加 skip_fields 排除 subtitle
- _node_full_text/_node_anchored_text 保持 字幕:/[sN] 语义
- get_subtitle 读 Card.subtitle(L2/L3)
- verify.py/synthesizer.py: l3.subtitle → l3.card.subtitle
- 迁移脚本 tools/migrate_subtitle_to_card.py(幂等,300 棵树已迁移)
- 9→6 个测试文件适配(3 个无需改动)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 11:57:41 -04:00
iomgaa
45403b23b4
feat(repair): regenerator + supplement 防御修复 + 迁移脚本
...
- 新增 app/tree/repair/regenerator.py(VLM 重生成 + 级联修复)
- supplement.py: deduplicate_field str() 防御 + inject_value strip
- patch.py: ruff format 格式化
- repair_trees.sh: conda source 激活修复
- 新增 migrate_from_trm4.sh 迁移工具
- enhance/__init__.py → repair/__init__.py 重命名
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 07:53:59 -04:00
iomgaa
f57ee45dc0
fix: Codex 全量审查修正
...
Critical:
- C1: assemble_mode 'plain' → 'ids'(合法枚举值)
- C2: question_id 加入 task_type slug 避免跨题型冲突
Important/Minor:
- generate_one 移除未用的 embed_fn/similarity_threshold 参数
- config.py 注释 11→12 同步
- 测试 question_id 断言更新
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 07:40:27 -04:00
iomgaa
93c9be8bfa
feat(tools): generate_questions.py calibrate 子命令
...
- Fisher exact test + effect size 组合判定(PASS/WARN/FAIL)
- 按 video_id 分组推理,避免跨视频树错用
- baseline 支持从 DB 读取或自动跑推理
- 对比表输出 + 退出码控制
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:47:52 -04:00
iomgaa
11f3c90200
feat(tools): generate_questions.py generate 子命令
...
- VLM 出题 + embedding 去重 + 断点续跑 + 并发控制
- 单线程汇总点保证去重原子性
- 18 个单元测试覆盖 progress/exemplar/pool rebuild/JSON append
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:42:50 -04:00
iomgaa
6e46d184b8
feat(harness): add factory.py — InferenceDeps dataclass + build_inference_deps
...
组装一次推理所需的全套依赖的工厂函数:
- TreeIndex 加载(FileNotFoundError if missing)
- TreeEnvironment 构建
- SkillRegistry 按需发现
- SearchToolDispatcher 装配
- PromptManager + prompt_builder 闭包
测试覆盖:正常路径、缺失树文件、skills 注入、frozen 不可变性。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:36:07 -04:00
iomgaa
5aa7cc48c5
feat(question_gen): is_duplicate + generate_one — 去重判定与单题生成编排
...
- is_duplicate: 余弦相似度去重,空池短路
- generate_one: 异步重试循环,不含去重(由调用方汇总点原子执行)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:32:39 -04:00
iomgaa
90f17e330e
feat(question_gen): build_generation_prompt + parse_vlm_response
...
- prompt 组装:system(角色+题型+约束+few-shot) + user(card+字幕+干扰项)
- VLM 响应解析:JSON 直接 + markdown code block 回退,四选一 schema 校验
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:28:25 -04:00
iomgaa
40b04f886e
fix(question_gen): sample_anchor Codex 审查修正
...
- C1/C2: Temporal Reasoning ≥3 L2 + Object Reasoning ≥2 L2 下限检查
- I1: L2 题型子帧不足时 ValueError
- I2: L3 过滤无 frame_path 的节点
- I3/I4: distractor_texts 扩展到整棵树范围
- I5-I7: 测试补强 Information Synopsis/Temporal Reasoning/Object Reasoning
- M1: Spatial Reasoning 测试断言 spatial_layout
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:24:30 -04:00
iomgaa
a597a9f901
feat(question_gen): sample_anchor — 按题型层级采样锚节点
...
含 6 种层级分支:L3 单帧、L2 多帧、Temporal Perception 特例、
L1 全量/采样 L2、L1-L2 混合。时间排序 + used_node_ids 排除。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:18:25 -04:00
iomgaa
9eb9b86954
feat(question_gen): AnchorContext + 12 题型-层级映射常量
...
- AnchorContext frozen dataclass: 锚节点生成上下文(node_id, card_text, frame_paths, subtitle, distractor_texts)
- TaskTypeSpec frozen dataclass: 题型生成规格(level, needs_frames, frame_count, context_fields)
- TASK_TYPE_LEVEL_MAP: 12 种 Video-MME 题型 → 树层级 + 生成规格映射
- 11 项单元测试覆盖:映射完整性、值类型、层级合法性、frozen 不变性
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:10:58 -04:00
iomgaa
eb15ab315e
fix(config): _VIDEO_MME_TASK_TYPE_COUNT 11→12,Video-MME 实际有 12 种题型
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 05:08:05 -04:00
iomgaa
8c383354c5
fix(detector): 移除 ongoing_actions/visible_entities 空字段检测 + 新增 repair sh
...
抽检确认这两个字段为空是合法内容状态(静物/黑帧/模糊帧),
VLM 重修也修不好,保留会导致断点续跑死循环。
L3 空字段检测缩减为 frame_summary + spatial_layout。
2026-07-09 00:54:27 -04:00
iomgaa
afe80a8b32
feat(repair): 断点续跑 progress 文件管理
...
load_progress / save_progress(asyncio.Lock + os.replace 原子写入)
/ should_skip_video。支持并发安全的读改写和 --reaggregate-all 兜底。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 00:23:04 -04:00
iomgaa
8182cb86b1
feat(detector): 扩展空字段检测到 L2 event_description / L1 scene_summary
...
断点续跑判据需要 L2/L1 层的 empty_field 检测。零 LLM 成本。
2026-07-09 00:12:55 -04:00
iomgaa
f733c13dd1
fix(llm): call_id 移入重试循环,每 attempt 独立
...
消除重试时遥测主键冲突的根因。每次 attempt 独立记录,
parent_call_id 不受影响(循环外固定),更利于事后诊断重试轨迹。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 00:12:38 -04:00
iomgaa
5a91f392f0
fix(telemetry): INSERT OR IGNORE + WAL + try/except 三层防御加固
...
根治遥测写入主键冲突(UNIQUE constraint)和并发写锁(database is locked)
导致的异常冒泡,遥测侧信道错误不再污染 LLM 重试链。
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-09 00:11:54 -04:00
iomgaa
d6e74f2734
feat(harness): runner.py — train loop orchestrator ( #13 algorithm fidelity)
...
Three-level nesting (epoch -> step -> per-skill), slow update 10-step
sequence, checkpoint/resume, early stop, probation accept/reject/rollback.
Key TRM4->TRM5 changes:
- sync -> async (all inference/diagnosis/evolve/validate awaited)
- LLMClient.from_env -> injected LLMProvider (DI via constructor)
- Direct DB/file access -> module functions (workspace/store/log)
- _TrainState as train() local, explicit param passing to helpers
Module-level pure functions extracted for testability:
resume_plan, _guard_infra_failures, _apply_batch_correctness,
_compute_total_steps, _should_early_stop, _format_applied_edits,
_fallback_summary, _write_skip_report, _outcome_to_quadrant_pairs,
_build_comparison_pairs, _batch_from_ids, _snapshot_current_skills.
Tests: 34 unit tests covering 13a-13e sub-tasks.
Radon: all functions Grade B or better.
2026-07-07 13:43:20 -04:00
iomgaa
6baddcc17d
feat(harness): validate.py — async 块序贯验证编排 + Probation 统一定义
2026-07-07 13:20:43 -04:00