iomgaa
1468a53b7a
fix: read-only baseline queries skip _runs upsert (register_run flag)
2026-07-16 05:36:47 -04:00
iomgaa
a0c7e043e8
fix: validate global frozen pools baseline_run_id + sha256 on load
2026-07-16 05:05:15 -04:00
iomgaa
5bb8319220
feat: tier-aware diag/val split with val-power repair (design 5.1)
2026-07-16 04:58:41 -04:00
iomgaa
6a21d80313
feat: add video-split config knobs and reproducible script
2026-07-15 12:54:25 -04:00
iomgaa
aa10485b9f
fix: make pools.json freeze atomic + add split manifest
2026-07-15 12:28:48 -04:00
iomgaa
fd907aab46
fix: extend correctness fail-fast to test-side pool questions (P5)
2026-07-15 12:22:11 -04:00
iomgaa
20eea98cdd
refactor: add video-atomic pool split (algo #5 gate input preserved)
2026-07-15 12:14:36 -04:00
iomgaa
bd1f7a22a2
feat(harness): pools.json 序列化 pair 四字段防孤儿 single
...
_q_to_dict 写出 pair_id/question_role/flip_axis/unit_id,_dict_to_q 用 .get
兼容旧 workspace 的 pools.json 读回并回填(unit_id 缺省交 __post_init__)。
pools.json 是训练主回路读回题目处,此前漏写会让孪生对解冻后退化成孤儿
single、配对指标失真。categories 块沿用 per-qid 记录,Task 3 的 unit 原子
切分已保证两 pair 成员同池同 key,序列化不破坏该原子性。
2026-07-15 08:05:15 -04:00
iomgaa
d6a3107e4e
feat(question_gen): loader 按 unit 分层采样 + load_benchmark 读回 pair 字段
...
stratified_sample 先 build_units 聚合,以 QuestionUnit 为采样原子做
分层/去重/补足/rng.sample,返回前 flatten_units 展开为逐题列表;
size/correct_ratio/min_per_class 均按 unit 计数,单元正确性走成员 AND,
孪生对两题永不被劈开。纯 single 输入下 build_units 1:1 折叠、顺序不变,
rng 消耗与旧逐题实现字节级一致(新增回归测试守护)。
_backfill_per_class candidates 改按 unit 枚举去重;build_units/flatten_units
函数内延迟导入以规避 question_gen<->harness 循环依赖(沿用 adversarial_filter)。
load_benchmark 反序列化补 pair_id/question_role/flip_axis/unit_id 四字段,
用 .get 兼容旧 JSON(缺失退化为 single,unit_id 由 __post_init__ 回填)。
pools._sample_excluding 随之改为透传 flatten_units(candidates) 给已单元化的
stratified_sample(不再用 lone pair-original 代表),行为对 single-only 保持等价。
2026-07-15 06:27:41 -04:00
iomgaa
ddb9a44f75
feat(pools): 三池切分以 unit 为原子,孪生对同池不被拆散
...
build_pools/_sample_excluding 与 PerCategoryPoolStrategy._split_one_category/
build_incremental 两条切分路径均改为以 QuestionUnit 为采样原子:progressive
exclusion 互斥集合与 train/val 分层划分都按 unit_id 计数(pair 计 1 个 unit),
命中单元整体展开,AR 孪生对两题永不落入不同池/split。
复用 app.harness.question_units 的 build_units/flatten_units,不重写分组逻辑。
single-only 输入下 unit 与 question 一一对应、rng 消耗量不变,采样与划分结果
与逐题口径完全一致;抽出 _unit_correct/_assert_correctness_complete 两个 helper
将 _split_one_category 复杂度压回基线以下。
新增 tests/unit/test_pools_pair_atomic.py 覆盖两条路径的 pair 原子性回归。
2026-07-15 06:09:14 -04:00
iomgaa
84b52a0311
feat(pools): auto-supplement maintenance correct questions in PerCategoryPoolStrategy
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 10:41:46 -04:00
iomgaa
72befa2bd4
feat(pools): add batch_correct_ratio field to PoolConfig
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 10:34:28 -04:00
iomgaa
dec7346da3
feat(harness): add Action Recognition training experiment
...
- PerCategoryPoolStrategy: filter test pool by task_types
- RunConfig: add run_holdout_eval toggle (default true)
- load_config: fix YAML task_types list-to-tuple conversion
- Runner: conditionally skip _holdout_four_way when disabled
- CLI: add --no-run-holdout-eval flag
- New config/train_action_recognition.yaml (3 epochs, per_category)
- New scripts/train_action_recognition.sh (baseline + seed + train)
2026-07-14 00:58:54 -04:00
iomgaa
37d4519905
chore: lint and format per-category pool strategy implementation
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 22:58:10 -04:00
iomgaa
c66a00c924
feat(harness): refactor build_or_load_pools to accept PoolStrategy + per_category freeze format
...
- save_pools: extended with split_mode and config params; per_category
mode writes categories metadata (seed, train_ratio, test_source) for
incremental append and consistency validation
- load_pools: compatible with both old format (no split_mode) and new
format; extra metadata fields ignored during load
- build_or_load_pools: signature changed to (config, strategy, db_path);
baseline_run_id read from seed.json (not config.run_id); per_category
mode does consistency check on reload and supports incremental category
append via strategy.build_incremental
- Added _to_pool_config, _read_baseline_run_id,
_validate_per_category_consistency helpers
- Tests: TestPerCategorySaveLoad with 5 test cases covering roundtrip,
missing config error, global split_mode field, legacy format compat,
multi-type categories
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 22:53:19 -04:00
iomgaa
21c6a53aed
feat(harness): add PerCategoryPoolStrategy with correctness-stratified 2:1 split
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 22:41:48 -04:00
iomgaa
811ffa648b
feat(types): extend GeneratedQuestion with skill_target & difficulty_steps
...
- Add skill_target (str | None) and difficulty_steps (int | None) fields
to GeneratedQuestion dataclass with field(default=None)
- Update loader.py to pass new fields from JSON (backward-compatible)
- Update pools.py _q_to_dict/_dict_to_q for serialization compat
- Add question_gen_v2 config section to default.yaml
- Add comprehensive test coverage (7 tests)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:05:38 -04:00
iomgaa
48b423ef35
feat(harness): pools.py — 三池切分(test→validation→diagnosis)
2026-07-07 12:48:54 -04:00