Map codebase onto agent-harness component model - #9
Open
trisha-ant wants to merge 12 commits into
Open
Conversation
Diagram + coverage table mapping the perception eval code onto the six canonical agent-harness components (loop, prompt+tools, context mgmt, subagents, execution env, session state). Clarifies that the agent loop here is temporal (frame→frame with history feedback) rather than tool-driven, and flags where each component is duplicated, absent, or hazardous.
Extends the harness map with a target architecture: two decoupled surfaces (production Perceiver, research experiments/) sharing only a thin call_model primitive; brain/hands split via a frozen FrameInput; harness-owned verify phase; events.jsonl observability; baseline.lock for auto re-eval. Each box tagged [P1-P12] and a table maps each principle to the concrete defect it fixes from the current code.
…im, 22 offline tests Lays the harness-independent groundwork for the migration eval so we can prove equivalence before any redesign code lands: - benchmark/stats.py: RunSet + Welch's t (no scipy dep). Pinned to RESEARCH.md's hand-computed t≈0.61 for hybrid vs vote3_mm. - benchmark/metrics.py: rewrite to operate on the result-dict shape that run.py actually emits (the old version imported a non-existent BenchmarkReport and was dead). Adds transition_mae — mean |Δframe| between predicted and GT stage onsets, the metric that actually moves per the transition-timing analysis. - experiments/analysis/record_replay.py: sha256-keyed request→response Recorder + install_record/install_replay context managers that monkeypatch both call_claude impls. Recording CLI is implemented but not run yet. - tests/: conftest fixtures over the archived *_replication/ JSONs, test_stats (7), test_metrics (9), test_record_replay (6). 22 tests, ~1.2s, zero API calls, zero volume data needed.
…fig — 37 tests green Builds the new research harness alongside the legacy one (perception/ + root run.py left intact so Phase A recordings can still be captured against them). Core pieces: - gently_perception/types.py: FrameInput (frozen, no GT field) + Obs + ParseError/LeakError. Closes the GT-leak surface at the type level — variants physically cannot read ground truth. - experiments/harness/loop.py: the new agent loop. Owns history construction with provenance-tagged entries; _assert_no_leak fires LeakError if GT-sourced history would reach a scored frame. Explicit render→classify→verify→score phases per frame. - experiments/harness/events.py: append-only EventLog (JSONL) + reduce_events() — result JSON is derived from the event stream, not accumulated as mutable state. - experiments/harness/config.py: RunConfig with full provenance (model, thinking, git_sha, dirty, seed) and append-only result_path(). - experiments/harness/run.py: --n-runs / --model / --thinking / --baseline CLI; built-in Welch-t vs stored baseline. - experiments/harness/discover.py + experiments/variants/hybrid.py: auto-discovery; hybrid ported to perceive(FrameInput) using the byte-identical prompts so replay-equivalence can succeed. Tests: test_leak_guard (8), test_loop_stub (5), test_variant_contract (2). Full suite now 37 passing in ~1.3s, no API, no volume data.
…che — 46 tests - experiments/variants/_adapter.py: adapt(legacy_fn) → perceive(FrameInput) that calls the legacy function unchanged, so API requests are byte-identical to the old harness for record-replay equivalence. hybrid/scientific/temporal/multimeasure/vote3_mm/fillpct/ensemble/minimal are now 4-line adapter files. - gently_perception/render.py: CachedFrameSource wraps any FrameSource (e.g. OfflineTestset) with a sha1-keyed disk cache. Warm runs never create the upstream iterator, so TIFF decode + 5-image render are skipped entirely. Old benchmark/testset.py is left untouched. - Fix record_replay._patch_targets(): variants do `from ._base import call_claude`, so the name is bound in each variant module — patching only _base missed almost all calls. Now scans sys.modules for every binding under perception/gently_perception/experiments. Found by the new adapter-equivalence test. - tests: test_adapter_equivalence (6, parametrized — proves identical request hashes old vs new), test_render_cache (3). 46 passing, ~1.3s.
…test — 51 tests - loop.run_variant: extract per-embryo coroutine and gather under a semaphore (default concurrency=4). Embryos are independent (history is per-embryo); concurrency=1 preserves deterministic event ordering. - harness/run.py: gather over --n-runs seeds; wrap OfflineTestset in CachedFrameSource so seeds 1..N hit the disk cache populated by seed 0. - experiments/pipelines/judge_ensemble.py: hybrid + vote3_mm + judge arbitrator as a perceive(FrameInput) function. Because pipelines use the same contract as variants, they run through loop.run_variant and inherit the leak guard — the judge only ever sees FrameInput, never GT. This is the structural fix for the bug in judge_ensemble/judge_replicate.py:131. - discover.py: scan pipelines/ alongside variants/. - tests/test_retroactive_leak (2): judge fires on every disagreement frame and the FrameInput it receives has no GT field; loop raises LeakError if a stage-filter pattern would smuggle GT history past first_scored. - tests/test_parallelism (3): max in-flight reaches embryo count, semaphore caps it, parallel result == sequential result. 51 tests passing, ~1.3s, no API. Legacy harness untouched.
…lence test The shim wrapper required positional (system, content_or_msgs) but every variant calls call_claude(system=..., content=...) by keyword, so recording failed on every frame with "missing 1 required positional argument". Now extracts system/content from *args/**kwargs handling positional, content=, and messages= (conversation) uniformly. Adds tests/test_replay_equivalence.py (marked slow — renders TIFFs) that loads a recording, installs install_replay, runs the new loop, and fails on the first CacheMiss. pytest.ini deselects slow/live by default so the fast suite stays at 51 tests / ~0.5s.
Captured a full 233-frame recording for hybrid via the old harness and replayed it through both old and new loops. New harness raises no CacheMiss (request-level equivalence) and produces byte-identical per-frame predictions and overall_accuracy (prediction-level equivalence). The new loop is provably interchangeable with the old one for hybrid. - data/fixtures/recordings/hybrid__*.jsonl: 233-response recording (106KB, sha-keyed request→response). - tests/test_replay_equivalence.py: corrected to compare old-replay vs new-replay of the SAME recording (the previous version compared against a stale results JSON from a different model invocation, which is meaningless). Uses CachedFrameSource so the second pass hits the render cache (498 hits, 0 misses observed). - .gitignore: exclude data/cache/ (render cache reaches ~679MB). Side finding: the recording's accuracy is 39.9% — the known anchoring collapse on the current _base.py model config. Irrelevant to equivalence (both harnesses get 39.9%) but a live demo of why model/thinking belong in RunConfig rather than hardcoded.
…8 variants The recording CLI now wraps OfflineTestset in CachedFrameSource so recordings after the first one skip TIFF decode + render entirely (the hybrid run populated the cache). Replay test parametrized over the full leaderboard set; skips per-variant when its recording is absent.
…mechanism limit)
Captured 233-frame recordings for the 7 remaining leaderboard variants
(scientific, temporal, multimeasure, fillpct, minimal, vote3_mm,
ensemble) in parallel with the render cache warm. Replay-equivalence
across all 8:
hybrid scientific temporal minimal fillpct multimeasure → PASS
(0 CacheMiss, 0 per-frame diffs, identical accuracy old↔new)
vote3_mm ensemble → XFAIL
(self-consistency variants make N calls with the same request key
per frame; live responses differ but the Recorder stores one per
key, so replay collapses the vote and history diverges. This is a
recording-mechanism limitation — the 6 single-call passes prove
loop equivalence. Fix is a per-key sequence recorder.)
Recordings total ~888KB (8×233 sha-keyed entries). Slow suite runs in
~22s with the render cache warm (was ~161s for hybrid alone cold). Fast
suite unchanged: 51 passed / 0.5s.
- experiments/harness/model_override.py: install_model_override(model,
thinking) replaces every call_claude binding with a direct
messages.create using the runner's model/thinking instead of the
hardcoded defaults in _base.py. Never sends sampling params (newer
models 400 on them) or effort=xhigh (model-version-specific); keeps
prompt caching on the system block.
- harness/run.py: wraps seed runs in install_model_override so --model/
--thinking actually reach the API; resumes from existing result paths.
- experiments/analysis/validate_migration.py: sweeps {variants} ×
{models} × {thinking} × N seeds via run_one_config (resumable,
append-only), then writes MIGRATION_REPORT.md with mean±std and
Welch-t against the archived RESEARCH.md numbers.
- tests/test_model_override (3): override injects model+thinking, omits
sampling/effort, restores on exit.
Fast suite: 54 passed / 0.5s.
… vs 81.7±2.4, t=-0.63)
Fix result_path provenance bug found during the smoke: thinking mode
was not in the path, so none and adaptive configs collided on disk and
the second loaded the first's results. Path is now
{variant}/{model}/{thinking}/{git8}_{seed}.json. Also fix write_report
sort with None keys.
Results (12 runs, render cache 2988 hits / 0 misses):
- 4.6/none: 80.1 ± 4.2 ← pass-criterion config, t=-0.63 vs archived
- 4.6/adaptive: 77.8 ± 1.2
- 4.7/none: 57.1 ± 2.4 ← anchoring collapse on 4.7 (expected)
- 4.7/adaptive: 53.1 ± 9.5
The new harness reproduces the master baseline within noise. The 4.7
drop is the documented prompt-anchoring sensitivity, not a harness
issue — both harnesses show it identically (Phase A proved that).
Commits MIGRATION_REPORT.md + 12 result JSONs + 12 events.jsonl.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
gently-perception mapped onto the agent-harness definition
In this repo the "loop" is temporal (iterate frames of a developing embryo),
not tool-use. Each step: render frame → call model → parse stage → append to
history → next frame sees that history. The "tool result fed back" is the
model's own previous prediction.
The harness, by component
graph TB MODEL["MODEL<br/>anthropic.Anthropic().messages.create<br/>(reasoning only — everything else is harness)"] %% ---------------- LOOP ---------------- subgraph LOOP["1 · AGENT LOOP / ORCHESTRATION"] direction TB RUNLOOP["run.py :: run_variant()<br/>for embryo → for timepoint:<br/>perceive() → parse → append history → next<br/><i>(the canonical loop)</i>"] PERCEIVER["gently_perception/perceiver.py :: Perceiver.__call__<br/>same loop, OO, per-embryo Session<br/><i>(production wrapper, duplicate impl)</i>"] MULTITURN["_base.py :: call_claude_conversation()<br/>multi-turn within one frame<br/><i>(used by *_multishot — negative result)</i>"] end %% ---------------- PROMPT / TOOLS ---------------- subgraph PROMPT["2 · SYSTEM PROMPT + TOOL DEFINITIONS"] direction TB SYSPROMPTS["perception/{variant}.py :: SYSTEM_PROMPT<br/>~29 variants, one string each<br/>hybrid.py routes between TEMPORAL/SCIENTIFIC"] OUTFMT["Output schema: free-text JSON<br/>{stage, reasoning}<br/>parsed by _base.parse_stage_json()<br/><i>(no structured tool_use — regex+brace fallback)</i>"] TOOLS["Tool definitions: NONE active<br/>tools (3D rotate, prev-frame) tried in gently<br/>baseline → 23% vs 35% → removed"] end %% ---------------- CONTEXT ---------------- subgraph CTX["3 · CONTEXT MANAGEMENT"] direction TB HIST["history[] — model's own past predictions<br/>_base.build_history_text() → last 3<br/>Session.history → last 5<br/><i>(hard truncation = compaction policy)</i>"] REFS["references — 2 example imgs/stage<br/>_base.build_reference_content()<br/>cache_control: ephemeral 1h"] PCACHE["Prompt caching<br/>system block + final ref block<br/>tagged cache_control 1h TTL"] GTLEAK["⚠ GT-for-skipped-frames injected into<br/>same history[] (run.py:138-143)<br/>— the leak surface"] end %% ---------------- SUBAGENTS ---------------- subgraph SUB["4 · SUBAGENTS / STRUCTURED HANDOFFS"] direction TB VOTE["vote3_mm.py — 3× self-consistency,<br/>majority vote (variance ↓ 4.8→0.5)"] ENSEMBLE["ensemble.py — N variant calls → vote"] JUDGE["experiments/judge_ensemble/*.py<br/>expert A + expert B → judge arbitrator<br/><i>(ad-hoc scripts, OUTSIDE harness)</i>"] ZSUB["zslice*.py — segment-count subcall<br/>overrides main classification<br/><i>(negative result)</i>"] ROUTE["hybrid.py :: _get_system_prompt()<br/>route prompt by predicted last stage<br/><i>(handoff between prompt 'personas')</i>"] end %% ---------------- ENV ---------------- subgraph ENV["5 · EXECUTION ENVIRONMENT"] direction TB RENDER["benchmark/testset.py<br/>TIFF volume → 5 image renders<br/>(3-view, top, side, midplane, z-stack)<br/><i>the model's 'sensors'</i>"] SECRETS["anthropic.Anthropic()<br/>reads ANTHROPIC_API_KEY from env"] NOSAND["No sandbox/VM — pure API,<br/>no code execution by model"] end %% ---------------- STATE ---------------- subgraph STATE["6 · SESSION STATE & CACHING"] direction TB SESSION["perceiver.Session<br/>observations[] + images{tp→b64}<br/>to_dict/from_dict for restart"] RUNHIST["run.py :: history list<br/>same state, inline, not persisted"] RESULTS["data/results/{variant}_{stages}.json<br/>per-run output (overwritten)"] CLIENT["_base._client / api._client<br/>module-level Anthropic() singleton"] end %% ---------------- WIRING ---------------- LOOP ==> MODEL PROMPT ==> MODEL CTX ==> MODEL RUNLOOP --> HIST RUNLOOP --> SYSPROMPTS RUNLOOP --> RENDER RUNLOOP --> RUNHIST RUNLOOP --> RESULTS PERCEIVER --> SESSION PERCEIVER --> HIST SYSPROMPTS --> OUTFMT HIST --> GTLEAK VOTE --> MODEL ENSEMBLE --> MODEL JUDGE -.escapes harness.-> MODEL ZSUB --> MODEL ROUTE --> SYSPROMPTS RENDER --> RUNLOOP SECRETS --> CLIENT CLIENT --> MODEL PCACHE --> MODEL %% ---------------- STYLING ---------------- classDef model fill:#222,color:#fff,stroke:#000,stroke-width:3 classDef warn fill:#ffe0e0,stroke:#c00 classDef absent fill:#eee,stroke:#888,stroke-dasharray:4 classDef escape fill:#fff0d0,stroke:#e90 classDef dup fill:#ffeaf4,stroke:#c6c class MODEL model class GTLEAK warn class TOOLS,NOSAND absent class JUDGE escape class PERCEIVER,RUNHIST dupLegend: black = the model · pink = duplicate implementations · grey-dashed = component absent/removed · yellow = escapes the harness · red = known hazard
The loop, as a sequence (one embryo)
sequenceDiagram autonumber participant Env as testset.py<br/>(environment / sensors) participant Loop as run_variant()<br/>(agent loop) participant Ctx as history[] + refs<br/>(context mgmt) participant Prompt as variant.py<br/>(sys prompt + schema) participant API as call_claude()<br/>(model boundary) participant Parse as parse_stage_json() Note over Loop: Harness owns everything<br/>left of the API call loop each timepoint T Env->>Loop: TestCase(image_b64, gt) Loop->>Ctx: read history (last 3) Loop->>Prompt: select system prompt<br/>(hybrid: route by last stage) Prompt->>API: system + refs(cached) +<br/>history_text + image API-->>Parse: raw text Parse-->>Loop: PerceptionOutput{stage, reasoning} Loop->>Ctx: append {T, predicted stage} Note over Ctx: ← this IS the<br/>"feed results back" step end Loop->>Loop: score vs GT, write results JSONCoverage summary
run.py:run_variant,perceiver.Perceiver.__call__perception/*.pyconstantshybridroutes between two_base.parse_stage_jsonhistory[],Session.observationsbuild_history_textSession.to_dict/from_dictPerceiverpath, notrun.pyvote3_mm,ensemble,zslice, judge scriptshybrid._get_system_promptanthropic.Anthropic()envSession, inlinehistoryrun.pyversion not persistedcache_control1h,_clientsingletonProposed harness
Designed against the 12 principles. Two product surfaces (production
Perceiverfor the gently microscopy framework, research
experiments/for autoresearch)each own their own harness [P4]; they share only the lowest-level model
primitive (
call_model) [P1]. The loop, context, and verification areharness-owned — variants are pure functions that cannot touch history or scoring
[P5][P6][P7][P12].
graph TB MODEL["MODEL<br/>messages.create(model=?, thinking=?)<br/>all knobs are parameters [P2]"] %% ================= SHARED PRIMITIVE ================= subgraph PRIM["shared primitive — intentionally tiny [P1]"] APICALL["api.call_model(system, content, *, model, thinking)<br/>thin wrapper, no defaults baked in [P2]<br/>raises on non-2xx / empty [P10]"] PARSE["api.parse_output(text, schema)<br/>structured tool_use, not regex [P8]<br/>raises ParseError — no silent 'early' [P10]"] end %% ================= SURFACE A: PRODUCTION ================= subgraph PROD["SURFACE A · production harness (ships inside gently) [P4]"] direction TB PLOOP["Perceiver.__call__<br/>the simple loop, nothing else [P1]"] PSESSION["Session — typed history owner [P7]<br/>to_dict/from_dict for restart"] PPROMPT["prompts/hybrid.py<br/>one frozen prompt per model release [P2][P3]"] PRENDER["render.volume_to_image()<br/>HANDS: sensors only [P5]"] end %% ================= SURFACE B: RESEARCH ================= subgraph RES["SURFACE B · research harness (experiments/) [P4] — not-DRY by design"] direction TB subgraph BRAIN["BRAIN — variant under test [P5]"] VARIANT["variants/{name}.py<br/>perceive(fi: FrameInput) → Output<br/>pure, stateless, ~30 lines [P1][P3]"] end subgraph HANDS["HANDS — harness-owned, variants can't reach [P5][P6]"] RLOOP["loop.run(variant, cfg)<br/>owns iteration + history construction<br/>explicit phases: render → classify → verify → record [P6]"] FI["FrameInput (frozen)<br/>{image, history: tuple[Obs], timepoint}<br/>NO gt field — leak-proof [P7]"] RRENDER["render.py + disk cache<br/>lazy: only what variant signature needs"] VERIFY["verify.py — separate critic pass [P12]<br/>monotonicity check + optional judge<br/>harness runs it, variant doesn't"] end subgraph ORCH["ORCHESTRATION — added only after simple loop measured [P1]"] PIPE["pipelines/{name}.py<br/>compose(variantA, variantB, arbitrate)<br/>same FrameInput contract [P6]"] end subgraph OBS["OBSERVABILITY [P9]"] EVENTS["events.jsonl — append-only<br/>one line per: model_call, parse, verify, score<br/>state DERIVED from events, never mutated"] RESULTS["results/{variant}/{model}/{git8}_{seed}.json<br/>reduced from events.jsonl"] CONFIG["config block in every result:<br/>{model, thinking, git_sha, dirty, seed, prompt_sha}<br/>[P2][P11]"] end subgraph EVAL["EVAL-DRIVEN [P11]"] RUNNER["run.py --n-runs 3 --model M --baseline hybrid"] STATS["stats.py — mean±std, Welch-t vs baseline<br/>transition_mae"] BASELINE["baseline.lock<br/>pins {variant, model, git_sha}<br/>auto re-run when model flag changes [P11]"] end end %% ================= WIRING ================= APICALL ==> MODEL PARSE -.schema via tool_use.-> MODEL PLOOP --> APICALL PLOOP --> PSESSION PLOOP --> PPROMPT PRENDER --> PLOOP RUNNER --> RLOOP RLOOP --> FI FI --> VARIANT VARIANT --> APICALL VARIANT --> PARSE RLOOP --> RRENDER RLOOP --> VERIFY RLOOP --> EVENTS VERIFY --> APICALL PIPE --> VARIANT PIPE --> RLOOP EVENTS --> RESULTS RESULTS --> STATS RUNNER --> BASELINE BASELINE --> STATS CONFIG -.embedded in.-> RESULTS %% ================= STYLING ================= classDef model fill:#222,color:#fff,stroke:#000,stroke-width:3 classDef prim fill:#f4f4f4,stroke:#666 classDef brain fill:#e8f0ff,stroke:#36c,stroke-width:2 classDef hands fill:#eaf7ea,stroke:#2a2 classDef obs fill:#fff7e0,stroke:#c80 classDef evald fill:#f0e8ff,stroke:#7a4dd6 class MODEL model class APICALL,PARSE prim class VARIANT brain class RLOOP,FI,RRENDER,VERIFY hands class EVENTS,RESULTS,CONFIG obs class RUNNER,STATS,BASELINE evaldLegend: black = model · grey = shared primitive (the only DRY part) · blue = brain (variant under test) · green = hands (harness-owned execution) · amber = observability · purple = eval loop
Principle → design decision
api.call_modelis the only shared code; variant is ~30 lines; pipelines added only after baseline measuredmodel,thinking,effortare CLI flags recorded inconfig; no defaults in api.py;prompt_shain results_base.pyhardcodes 4.7+adaptive; CLAUDE.md says 4.6 — drifted silentlyscientific.pyin place, losing the 4.6 versiongently_perception/(prod) andexperiments/(research) each own their loop; share onlyapi.call_modelPerceiverreaches intoexperiments/prompt/via sys.path hack; two runners fight over same variantsperceive(FrameInput)is pure; render/history/verify/score live inloop.pyand never enter the varianttestset.pymixes rendering+GT; variants build their own content blocks; judge scripts rebuild historyloop.runenforces render→classify→verify→record phases; history construction is not overridableFrameInputfrozen dataclass,history: tuple[Obs]immutable, no GT field; truncation policy in one placehistoryis alist[dict]any code can mutate; GT leaks through ittool_useblock (classify_stagetool with enum schema) — model can't emit invalid stageparse_stage_jsondoes regex + brace-matching; invalid → silent fallback to "early"events.jsonlper run (model_call, parse, verify, score events); results reduced from it; transcript reviewableParseErrorraised, recorded as expliciterrorevent; unknown model/variant → exit 1 with listresponse_to_outputreturnsstage="early"on parse failure — indistinguishable from a real "early"--n-runs 3default;baseline.lockpins comparison config; runner warns if model ≠ baseline's modelverify.pyis a harness phase (monotonicity assert + optional independent critic); variant returns raw prediction onlyvote3_mm/ensemble/judgeeach hand-roll verification inside the variant; judge escaped harness entirelyWhat deliberately stays simple [P1]
pipelines/is a directory of plain async functions, not a DAG engine.async def perceive(fi: FrameInput) -> PerceptionOutput. That's the whole contract.Perceivergets less code than today — it drops the sys.path import hack and ships with one frozen prompt.