iomgaa
d4e9852864
feat: add cheater gate with resume-safe survivor recovery
2026-07-14 16:00:14 -04:00
iomgaa
24ed7ca322
fix: harden canonical_answer_text and align decision-core signatures to spec
2026-07-14 15:55:36 -04:00
iomgaa
c109f2257a
feat: add pure decision core for adversarial filter (hash/fingerprint/canonical/flip)
2026-07-14 15:51:28 -04:00
iomgaa
334fbbc94d
feat: add AdversarialFilterConfig for Phase B post-hoc filter layer
2026-07-14 15:47:21 -04:00
iomgaa
8b9e8aa19f
feat: add optional backfill params to run_pipeline_v2
2026-07-14 15:39:57 -04:00
iomgaa
d77cbc95eb
feat: add adversarial_verdicts table with resume and terminal-verdict query
...
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-14 15:35:11 -04:00
iomgaa
e41a2b0d08
test: guard non-AR flip isolation and AR flip-capable set
2026-07-14 15:30:18 -04:00
iomgaa
f12dd7e559
feat: add supports_flip/flip_axis to SubPattern for Phase B flip gate
2026-07-14 15:27:21 -04:00
iomgaa
96e314c3a0
fix: mark item rejected on selector_error for consistent bookkeeping
...
selector_error 分支(捕获 ValueError/FileNotFoundError)此前只返回 reason,
未 mark_item_rejected,导致 Phase 3 已 record 的 pending attempt 行永远停在
pending;而 hard-fail 分支会标记 rejected。两条失败路径落库风格现统一为
mark_item_rejected(异常路径无 outcome/observation,故不写 selector_scores)。
补单测 test_apply_grounded_selector_marks_rejected_on_error 守卫该路径。
2026-07-14 14:49:41 -04:00
iomgaa
46ac848176
fix: converge non-numeric selector scores to ValueError
2026-07-14 14:47:21 -04:00
iomgaa
76f719018c
feat: loosen multi_true gate to qualifier-scoped correctness
2026-07-14 14:33:22 -04:00
iomgaa
b1f15ddb3a
fix: keep cross_segment rule single-dimension, guard no-absent-events
2026-07-14 14:31:30 -04:00
iomgaa
58278c6de4
feat: enforce single-dimension counterfactual in AR distractor rules
2026-07-14 14:28:32 -04:00
iomgaa
b13eab0659
feat: wire grounded selector into AR slot processing
...
将 Task 5 的 grounded selector 织入 AR 出题路径(Phase 3.5,位于
record_item 与 postprocess 之间),仅在 strategy.uses_grounded_selector
为真时进入。observation 始终落库(含 hard-fail),硬失败走重出。
PipelineConfig 新增 candidate_pool_size/selector_delta_low/
selector_delta_high 三参,YAML 与 CLI seed override 同步。
2026-07-14 14:15:50 -04:00
iomgaa
d0194f5840
fix: degrade distractor pool gracefully on malformed VLM response
2026-07-14 14:11:13 -04:00
iomgaa
8a54055d02
feat: add grounded distractor selector with visual scoring
2026-07-14 14:06:07 -04:00
iomgaa
207e834f30
feat: add selector_scores observation column to question_gen_items
2026-07-14 13:59:29 -04:00
iomgaa
3f984acc18
feat: add uses_grounded_selector strategy switch (AR only)
2026-07-14 13:53:17 -04:00
iomgaa
2608a3841f
style: reformat sub_pattern round-trip tests
2026-07-14 13:51:51 -04:00
iomgaa
e68e4b7d57
fix: restore sub_pattern when loading benchmark JSON
2026-07-14 13:51:05 -04:00
iomgaa
111c88488f
style: format sub_pattern test file
2026-07-14 13:46:25 -04:00
iomgaa
ae0a718f67
feat: thread and persist sub_pattern into accepted questions
2026-07-14 13:45:31 -04:00
iomgaa
25f2a845ff
fix: run_gates must use current_tree after video resample
2026-07-14 13:40:46 -04:00
iomgaa
d7d7ce5bdc
feat(question_gen): register ActionRecognitionStrategy, replace temp VISUAL binding
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 06:51:50 -04:00
iomgaa
d0a8019fe1
feat(question_gen): add ActionRecognitionStrategy with 6 SubPatterns
...
Self-contained strategy targeting 6 Agent failure modes in Action
Recognition: premature_evidence_anchoring, temporal_reasoning_failure,
semantic_rigidity, fine_grained_visual_action,
cross_segment_entity_tracking, evidence_gap_confabulation.
- L2 default sampling (upgrade from L3) with 3 patterns overriding to L1
- Weighted random SubPattern selection (0.20/0.20/0.15/0.15/0.15/0.15)
- Each SubPattern includes instruction, examples, distractor rules
- Satisfies TaskTypeStrategy Protocol without extending BaseTaskTypeStrategy
- 32 unit tests covering all properties, definitions, and selection behavior
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 06:49:03 -04:00
iomgaa
b9616de21e
fix(tests): update test_generator_v2 to use new generator signatures
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 05:55:08 -04:00
iomgaa
afa77173e3
refactor(question_gen): adapt generator/gates/store signatures for strategy
...
- generator_v2: _load_prompt_template takes template_name str instead of
QuestionFamilySpec; _build_v2_prompt takes prompt_template + strategy_name
+ sub_pattern_instruction; generate_one_v2 takes discrete params
(prompt_template, strategy_name, skill_target, sub_pattern_instruction)
- gates: _gate_leak_test and run_gates take leak_probe_template str
instead of QuestionFamilySpec
- run_store: add sub_pattern column to DDL + idempotent migration;
record_item accepts optional sub_pattern param
- Remove QuestionFamilySpec imports from generator_v2 and gates modules
- Update test call sites accordingly
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 05:50:47 -04:00
iomgaa
c49d0ff12f
fix(sampler): validate level param rejects invalid values
...
Add ValueError guard at the top of sample_material_v2 for level not in
{1, 2, 3}, preventing silent fallthrough to L1 sampling. Add unit test
for the new validation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 05:43:28 -04:00
iomgaa
e2325b6535
refactor(sampler): replace family_spec param with level+constraint
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 05:38:14 -04:00
iomgaa
b6b6a48503
feat(question_gen): add TaskTypeStrategy Protocol and BaseTaskTypeStrategy
...
- TaskTypeStrategy Protocol: pipeline 的唯一接口,定义 task_type、
sampling_level、sampling_constraint、prompt_template 等属性
- SubPattern frozen dataclass: 出题子模式,靶向特定失败机制
- BaseTaskTypeStrategy: 封装现有 QuestionFamilySpec 行为的默认策略,
所有属性委托给绑定的 family
- _TASK_TYPE_TO_FAMILY: 消歧绑定表,12 个题型确定性绑定到 1 个 family
- register_strategy/get_strategy: 注册表 API,未注册题型自动创建
BaseTaskTypeStrategy
- 13 个单元测试全部通过
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-14 05:24:29 -04:00
iomgaa
dec7346da3
feat(harness): add Action Recognition training experiment
...
- PerCategoryPoolStrategy: filter test pool by task_types
- RunConfig: add run_holdout_eval toggle (default true)
- load_config: fix YAML task_types list-to-tuple conversion
- Runner: conditionally skip _holdout_four_way when disabled
- CLI: add --no-run-holdout-eval flag
- New config/train_action_recognition.yaml (3 epochs, per_category)
- New scripts/train_action_recognition.sh (baseline + seed + train)
2026-07-14 00:58:54 -04:00
iomgaa
c66a00c924
feat(harness): refactor build_or_load_pools to accept PoolStrategy + per_category freeze format
...
- save_pools: extended with split_mode and config params; per_category
mode writes categories metadata (seed, train_ratio, test_source) for
incremental append and consistency validation
- load_pools: compatible with both old format (no split_mode) and new
format; extra metadata fields ignored during load
- build_or_load_pools: signature changed to (config, strategy, db_path);
baseline_run_id read from seed.json (not config.run_id); per_category
mode does consistency check on reload and supports incremental category
append via strategy.build_incremental
- Added _to_pool_config, _read_baseline_run_id,
_validate_per_category_consistency helpers
- Tests: TestPerCategorySaveLoad with 5 test cases covering roundtrip,
missing config error, global split_mode field, legacy format compat,
multi-type categories
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 22:53:19 -04:00
iomgaa
e5b07ac974
feat(harness): add task_types, pool_split_mode, train_ratio, test_questions to RunConfig
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 22:46:28 -04:00
iomgaa
73ae1f7143
fix(harness): change _runs INSERT OR IGNORE to ON CONFLICT DO UPDATE for incremental infer
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 22:45:18 -04:00
iomgaa
21c6a53aed
feat(harness): add PerCategoryPoolStrategy with correctness-stratified 2:1 split
...
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 22:41:48 -04:00
iomgaa
cd5c9c01fb
feat(app): add PoolStrategy Protocol to application ports
2026-07-12 22:34:47 -04:00
iomgaa
0d0f275134
feat(core): add PoolConfig dataclass for pool strategy configuration
2026-07-12 22:33:49 -04:00
iomgaa
4fb7a61f8b
fix(question_gen): resolve pipeline integration issues from final review
...
1. Apply postprocess shuffle result (pp.options, pp.answer) to final
GeneratedQuestion output instead of using original candidate values.
2. Record dedup rejection in store via new mark_item_rejected() method,
preventing items from staying as 'accepted' after dedup rejects them.
3. Add .flatten() to embed_fn outputs in _is_duplicate and embed_pool
append to handle 2D (1,D) arrays from embedding implementations.
4. Validate exactly 4 options in _validate_parsed_fields (was >= 2),
matching the A-D answer constraint.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 00:14:12 -04:00
iomgaa
eecb86e27a
feat(question_gen): add generate-v2 CLI subcommand and experiment script
...
- Add generate-v2 subparser with --config, --store-dir, --db-path,
--seed, and --dry-run arguments to tools/generate_questions.py
- Implement _run_generate_v2 async handler: config loading, video
discovery, DI client construction, TreeIndex loading, pipeline
invocation, and result persistence
- Add scripts/generate_questions_v2.sh following build_trees.sh
conventions (source .env, conda run python path, MODE=mock support)
- Update app/question_gen/__init__.py to export full v2 public API:
run_pipeline_v2, PipelineConfig, PipelineResult, QuestionFamilySpec,
ALL_FAMILIES, CandidateQuestion, generate_one_v2, GateReport, run_gates
- Add QuestionGenStore.load_progress() for pipeline resumption
- Add integration tests for CLI help and dry-run behavior
- Update test_question_gen_api to match expanded __all__
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-12 00:04:42 -04:00
iomgaa
cf51d2de9d
feat(question_gen): add v2 generator with per-family prompt templates
...
Implement generator_v2.py with:
- CandidateQuestion dataclass (canonical location)
- _load_prompt_template: loads per-family .md from store/prompts/
- _build_v2_prompt: constructs system+user messages with material context
- _parse_v2_response: JSON extraction, json_repair, field validation
- generate_one_v2: async VLM call orchestration with reject_reason support
Add 5 family-specific prompt templates:
- retrieval.md: factual recall from visible content
- reasoning.md: multi-hop inference across segments
- enumeration.md: counting/listing entities and actions
- visual.md: visual details requiring frame observation
- spatial.md: spatial relationships between objects/people
Tests: 11 unit tests covering prompt build, parse, and e2e generation.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:41:36 -04:00
iomgaa
9053233f99
fix(question_gen): gates reject_reason returns raw reason; short-circuit skips LLM
...
- GateReport.reject_reason now returns gate.reason directly (no [name] prefix)
- verbatim short-circuit sets other 3 gates to SKIP without calling LLM
- test_high_verbatim_shortcircuits asserts zero LLM calls and SKIP verdicts
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:33:57 -04:00
iomgaa
f74711cd11
feat(question_gen): add v2 material sampler with family constraints
...
Implement sample_material_v2 module that samples tree nodes with
QuestionFamilySpec-aware constraint validation, providing richer
MaterialContext output (subtitles, cross-L2 context, frame paths).
Key components:
- AnchorContext/MaterialContext frozen dataclasses
- _validate_sampling_constraints: multi-level constraint checking
- _collect_subtitle_sentences: subtree subtitle extraction
- _collect_cross_l2_context: peer L2 event descriptions
- sample_material_v2: main entry with retry-on-constraint-violation
Tests: 11 unit tests covering normal sampling, used-node exclusion,
constraint violation retries, cross-L2 population, and subtitle
collection.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:30:56 -04:00
iomgaa
271d1682c9
feat(question_gen): add lightweight 4-gate quality check
...
Implement 4 concurrent LLM-based quality gates for generated questions:
- key_verify: validates answer evidence in source material
- blind_answer: rejects questions answerable without video context
- multi_true: detects ambiguous multi-correct options
- leak_test: per-family shortcut detection (5 probe templates)
Includes run_gates orchestrator with verbatim_ratio short-circuit,
JSON response parsing with fallback, and 9 unit tests (all passing).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:30:13 -04:00
iomgaa
9627ac9cf9
fix(question_gen): raise ValueError on UPDATE of missing rows
...
record_run_end, update_gates, and update_difficulty now check
cursor.rowcount after UPDATE+commit and raise ValueError if 0 rows
were affected. Prevents silent telemetry loss.
Adds three negative-path tests:
- test_record_run_end_missing_run_raises
- test_update_gates_missing_item_raises
- test_update_difficulty_missing_item_raises
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:23:48 -04:00
iomgaa
7abe92eb1c
fix(question_gen): check_verbatim covers question_text + add missing blacklist patterns
...
- check_verbatim now computes n-gram overlap for BOTH question_text and
correct_option vs source texts, returning max(question_ratio, option_ratio).
Extracted _ngram_overlap_ratio helper for reuse.
- Added 4 missing blacklist patterns: 'this segment', 'this frame',
'the current frame', 'frame summary'.
- Added 5 new test cases covering the above changes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:23:37 -04:00
iomgaa
6a5424a618
feat(question_gen): add SQLite run store for generation telemetry
...
Implements QuestionGenStore with:
- Idempotent schema initialization (question_gen_runs + question_gen_items)
- Run lifecycle: record_run_start / record_run_end / get_run_stats
- Per-item recording: record_item / update_gates / update_difficulty
- GateReportLike Protocol for duck-type gate report compatibility
- WAL mode + foreign keys + check_same_thread=False
DDL aligns with research-wiki/schemas/question-gen-{runs,items}.md.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:15:53 -04:00
iomgaa
c83d771923
feat(question_gen): add deterministic postprocess layer
...
Add app/question_gen/postprocess.py with zero-LLM deterministic
post-processing for generated questions:
- shuffle_options: deterministic option permutation with answer remapping
- check_referent_blacklist: detect self-referential language (this clip, etc.)
- check_verbatim: word-level n-gram overlap ratio measurement
- has_time_anchor: timestamp and temporal phrase detection
- check_forbidden_material: T1/T7 source material validation
- run_postprocess: orchestration returning PostprocessResult
Tests: 31 unit tests covering all functions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:14:43 -04:00
iomgaa
75e6d8c550
feat(question_gen): add 5 question family specs with sampling constraints
...
Define QuestionFamilySpec, LeakTestProfile, SamplingConstraint dataclasses
and instantiate 5 families (RETRIEVAL/REASONING/ENUMERATION/VISUAL/SPATIAL)
targeting failure mechanisms M1-M5. Implement get_family_for_slot with
legal-type filtering + weighted random selection.
13 unit tests cover: full task-type coverage, skill_target uniqueness,
deterministic seeding, invalid input errors, and chi-square distribution.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:14:22 -04:00
iomgaa
811ffa648b
feat(types): extend GeneratedQuestion with skill_target & difficulty_steps
...
- Add skill_target (str | None) and difficulty_steps (int | None) fields
to GeneratedQuestion dataclass with field(default=None)
- Update loader.py to pass new fields from JSON (backward-compatible)
- Update pools.py _q_to_dict/_dict_to_q for serialization compat
- Add question_gen_v2 config section to default.yaml
- Add comprehensive test coverage (7 tests)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com >
2026-07-11 23:05:38 -04:00
iomgaa
56fe138a46
feat(tools): batch tree build orchestration with shared API semaphore
2026-07-11 11:56:36 -04:00