Skip to content

Codex/control 001 seed - #3

Open
samrishtt wants to merge 65 commits into
arcprize:mainfrom
samrishtt:codex/control-001-seed
Open

samrishtt wants to merge 65 commits into
arcprize:mainfrom
samrishtt:codex/control-001-seed

Conversation

@samrishtt

Copy link
Copy Markdown

No description provided.

samrishtt added 30 commits May 31, 2026 18:00
- Fixed PicklingError: register game_mod in sys.modules
- Fixed GameAction Enum unpickling: custom copyreg reducer
- Fixed TypeError in _sample_novelty_guided: dict comparison
- Added encoding=utf-8 to build_notebook.py for Windows
- Added reference notebooks: Forge v16 (0.35), Forge v18 (~0.39), Arc26 v15, Ash agent
- Added extraction/injection utility scripts
- Score: 0.23 on Kaggle (up from 0.18 baseline)
- FIX arcprize#1: Hidden field probing BEFORE Phase 1 (was Phase 3)
- FIX arcprize#2: Clock field elimination prevents timer state explosion
- FIX arcprize#3: CBAM removed from ForgeNet (pretrained weight mismatch)
- FIX arcprize#4: MCTS f-string double-brace bug (log signals restored)
- FIX arcprize#5: _state_hash selective mode (hidden_fields used properly)
- NEW: ACMD counter search as Phase 3 (threshold-gated games)
- NEW: Professional README with architecture diagrams
- All search phases now use trigger-aware hashing consistently
…schema_helpers breakthrough, Exp G prepared (45K context)
…K (1.06) is optimal; transition to building state_dedup.py custom graft
…e backtracking; recommendation to restore Exp C 1.06 baseline
…NFIRMED POSITIVE with full ranked leaderboard
…ind_objects, death_memory) + Level 3 search loop + post-death cleanup
samrishtt and others added 30 commits August 10, 2026 19:51
…anup

arc3x/ (twin.py, explore.py, verify_twin.py) was built last session but never
committed and never run. Committing it as a safety net before deleting ~700MB
of spent experiment artifacts.

Also records the 2.14 notebook: stock TAAF solver + Qwen3.8-27B-FP8, with an
empty customization hook. The +0.81 over 1.33 came from the model swap alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…now cap at 115

Bug 1 (fatal, algorithmic): the archive key hashed the raw 64x64 frame and was
effectively bijective - tn36 60 steps -> 60 distinct keys. Every step looked
novel, so the archive became a log of every state visited, novelty carried no
signal, and Go-Explore degenerated into a random walk. It also silently killed
compress()'s loop-splice pass, which can only fire when a state repeats.

Root cause: only 2-3% of the 4096 pixels ever move, and tn36 row 1 is a
49-pixel HUD timer draining by exactly 6 per action (441,435,429,417...). That
one row makes every frame globally unique.

General fix in cell.py: a pixel that moves monotonically with the action count
is a clock, not state. State revisits values; a clock never does. Calibrated
per game from ~600 free simulated actions. No per-game knowledge.
Collapse measured: sk48 91->28 keys, m0r0 77->29, tn36 60->43.

Bug 2 (performance): snapshotting every new archive cell cost a 20ms deepcopy
per step -> 19 steps/sec. Now snapshots only on restore (~48x rarer) and skips
nodes shallower than 16 actions, since deterministic replay is cheaper there.
19 -> 546 steps/sec.

Bug 3 (coverage): nothing tracked which actions had been tried from a cell, so
the rollout re-picked the same few forever. Added Node.tried plus a selection
weight favouring cells with untried actions. Stickiness now applies only to
non-click actions.

Measured, on the two games a month of LLM work never cracked:
  sk48 L0: never completed (781-940 actions) -> 29 actions (baseline 61), 115.0
  m0r0 L0: 477 actions, score 0.02        -> 15 actions (baseline 30), 115.0
Game scores 2.78 and 4.76 from level 0 alone, vs a 2.14 leaderboard baseline.

Also corrects the docstring's 6,811 steps/sec claim: that was an artifact of
benchmarking an already-dead game. Real sustained ceiling is ~700 steps/sec.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…dent policy

Three changes, each measured.

1. Per-level recalibration (the depth bug). calibrate() probes one root, so the
   mask 'varies & ~clock' describes only that level's screen. Measured wrong in
   both directions on later levels: cd82 L2 has 139 informative pixels the L0
   mask excludes (14% of state invisible), and vc33 L1 has 1,779 clock-like
   pixels vs 494 at L0, so ~1,285 timer pixels get hashed into the archive key
   and it turns bijective again - the exact fatal bug the mask was added to fix,
   reappearing at every level past the first. This is why games clear L0 in ~10
   actions then stall. arc3x/diag_mask.py reproduces the measurement.

2. Restarts on failure. solve_level is randomised search with heavy-tailed
   runtime; a bad opening traps a rollout distribution for the whole budget.
   Also, the 25% reserved for compression was silently discarded whenever a
   level failed. Now up to 3 re-seeded attempts share that budget.

3. Student policy (arc3x/student.py, train_student.py). 4096 -> 256 -> 261 MLP
   in pure numpy, no torch: one-hot of the 4x4-max-pooled frame -> 5 simple
   actions + 256 coarse click cells, softmax masked to the legal action set.
   Trained by imitating the searcher's own compressed plans. Measured 0.424
   held-out vs 0.194 random-choice = 2.18x on 223 examples. Numpy-only so it
   cannot fail the way vLLM did (Exp 11: 2.68 local -> 0.60 Kaggle on prefill
   timeout).

25-game sweep: mean 2.978 -> 4.695 after the snapshot-memory fix. lp85 reached
5/8 levels for 41.67, confirming depth is the lever.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The imagination was never in the agent. Dream and Progress were only ever
exercised by measure_dream.py's 240-step harness, so every number measured
about them had no path to a score. Wired into agent.py: Dream is constructed
alongside Mechanics, fed every transition, cut on level change, and imagine()
re-plans after each action.

Scored with run_suite over all 25 dev games: 0.111 -> 0.142, 2 levels
completed, and every game still burns the full 3000-action cap. The Kaggle
submission scores 2.14, so arc3x does not ship as the agent - it is a
component source. lp85 clears level 0 in 6 actions against a baseline of 17,
which is the one place the objective demonstrably works.

Also classified every declared button as MOVE / ACT / DEAD instead of only
asking whether the avatar translated (probe_actions.py). Mechanics.moves
filters d != (0,0), which discarded a whole class of button:

  12 of 25 games have an ACT button - use, grab, select, rotate
  cd82 has five working buttons and the planner saw none of them
  tr87 has four and the planner saw none

Exposed as Mechanics.acts, built from new changes/shifts counters. cd82 and
tr87 were filed as 'no steerable avatar'; they are not click games at all.

settle() also loses real MOVE buttons: wa30 and sc25 each have four, the
planner sees [2, 4]. Added _fill_deltas to accept single-observation deltas
for the believed avatar once its step sizes are known - identifying who I am
needs strong evidence, using a button does not. It did not fix those two:
there are zero votes for the missing axes, so the cause is upstream in
moved_objects, where a sprite that rotates as it turns never matches as a
rigid translation. Recorded rather than papered over.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… them

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e it loses

Three things the suite could not do: run in parallel, say why a game failed, or
tell a fitted result from a general one.

PARALLEL (graded.run_suite workers=N). Every game is fully independent - its own
Twin, its own engine, no shared state, no writes - so the parallel result is
exact, not approximate. The naive version regressed badly: 6 workers each sized
a BLAS thread pool to all 12 cores, 72 threads thrashed over 12, and ls20 managed
32 actions in 60s where serial did 401. Thread limits have to be set in the
parent before the pool exists, because 'spawn' re-imports numpy in the child and
sizes the pool at import time. With that fixed, parallel reproduces serial to the
action (ls20 401 actions, 1 level, 0.78 both ways) and runs faster.

Also made the wall clock loud. Kaggle bills actions, not seconds, so a run cut
short by the deadline is a machine-dependent experiment rather than a
measurement, and two of them are not comparable. Agent.timed_out records it and
stall_summary shouts about it instead of averaging it in.

WHY IT LOST (GradedRun.note). agent_fn returns nothing, so the score was the only
channel out, and a score cannot distinguish a game the agent never understood
from one it understood and could not execute - which need opposite fixes. Notes
are keyed by the level the attempt *started* on: the attempt that clears a level
advances the counter first, and filing 'cover completed a level' under the level
it moved us to blames the wrong one.

First run already names the biggest leak: click:change=386 of all stall reasons.
click_round returns success on any pixel change whatsoever, so on cd82 the agent
clicks 259 times, something changes every time, and it never once asks whether
the change was progress.

OVERFIT CHECK (suite.py --split both). The evaluation set is 110 private games;
these 25 are development. The risk is not a rule that fails here, it is a rule
that succeeds here for a reason that does not exist there. Fixed alphabetical
every-third split: 17 tune, 8 hold. Fixed and never re-drawn, because a holdout
re-rolled after a disappointing result is a lottery.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
States the 3.52 / 10.57 / 21.14 ceilings, the five components and what each is
for, the two load-bearing priors (reversibility for avatar identification, the
ratchet rule for a label-free objective), and the measurement discipline -
including the three powers the local engine offers that the graded gateway does
not, and which every earlier local measurement here was inflated by.

Keeps the negative results, because they are the useful part: the mirror-panel
hypothesis refuted by its own control, graft tuning 0-for-11, prediction accuracy
proved not to be the bottleneck, and compute proved not to be the bottleneck.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The notebook is the deliverable. It is the 2.14 notebook plus one cell inserted
at its own documented customization hook, carrying two grafts and nothing else.

Both are expressible as inline notebook code, which is the binding constraint:
the solver is imported from a read-only Kaggle dataset this repo cannot change,
so a graft is either a wrapper or a string. That is what rules the arc3x modules
out and these two in.

  * Level probe (inert). taaf.game.Game.finish_game already computes
    actions_per_level and base_actions_per_level and prints them (game.py:636-650);
    TAAF_MINIMAL_DIAGNOSTICS=1 on a true submission is why no artifact of them has
    ever survived. Appends one JSON line per finished game. This is the
    measurement the plan has been blocked on: 2.14 is equally consistent with
    "level 0 on most games" and "two levels on a few", which need opposite fixes.

  * Family priors (behavioural). Appends one addendum to the agent's system
    prompt: (1) the win condition is usually a cover predicate and is therefore
    drawn on frame 0 as a set of small identical static objects, groupable by the
    segmentation hash the agent already has; (2) level score is
    (baseline/actions)^2 and only for levels you clear, so exploring level 0 is
    cheap, later levels must be crisp, and giving up early never pays.

Deliberately excluded, each for a measured reason: markers.py (its own grader
returned 0/1 on the only row with source-read ground truth — it proposed floor
tiling instead of the exit, because ranking candidates by repeat-count promotes
scenery); the ACTION1-4=N/S/W/E convention (measured at exactly 0.00 as a planner
seed); any change to GRAFT_FLAGS including context_window (untested, and changing
it alongside two new grafts would make the result unattributable again).

Verified rather than assumed. tools/verify_submission_notebook.py compiles every
cell (top-level await allowed), checks the graft lands after GRAFT_FLAGS and
before bm.run, and confirms the addendum survived two levels of escaping.
tools/smoke_grafts.py execs the cell lifted out of the generated notebook against
the real shipped taaf/inference source: all eleven GameRun field names resolve,
the wrapper is idempotent and writes one row per game rather than per call, the
missing-baseline branch records None, and the system prompt grows 12358 -> 15573
by exactly the addendum. That run also confirms the scoring weights are i+1 —
1 of 4 levels scored exactly 10.00, the completion cap 1/(1+2+3+4).

Suite at this commit: tune 0.167 / hold 0.001 over 25 games. arc3x remains a
component source, not the agent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The finished suite showed four holdout games clearing level 0 and still scoring
0.00: ft09 1633 actions vs baseline 43, cd82 2502 vs 55, vc33 289 vs 7, m0r0 1179
vs 30 - all about 40x, and (1/40)^2 is 0.06% of that level's points.

So 'exploring level 0 is cheap' is true only relative to the later levels. If
level 0 is the only level you clear, its efficiency IS your score. The addendum
now says that, and says to stop exploring once the objective is visible - which is
exactly the failure mode those four games display.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…he notebook

Sam asked why GRAFT_FLAGS arms only 3 of the 10 flags composite.install accepts.
docs/EXPERIMENT_LOG.md answers it: all seven others were tested on the real
leaderboard over Exp 3-11 and every combination scored below the three-flag
baseline (banking 1.10, helpers+void+transfer 1.06, state_dedup 0.77, schema_notes
0.47). The three are the survivors of a finished ablation, not an oversight.

Sam is also right that some of those experiments were good: Exp 11 set the highest
local score ever recorded (2.6848) and produced the first m0r0 level-0 clear, then
scored 0.60 on Kaggle. The log names the mechanism - 57K context plus verbose
per-turn notes caused vLLM prefill timeouts on the shared GPU, stranding games at
zero. Good locally, destroyed by latency remotely.

Two findings from reading the grafts change what ships:

recovery is separable. Both recovery regressions are attributed by the log to R2's
probe tax alone, and R2 spends up to 16 actions per stalled level. R1 (clear a
death-spiralled history, rewrite the world model) and R3 (carry mechanics across
the vendor level wipe via cross_level_notes) cost zero actions by construction.
build_probe_plan returns plan[:PROBE_MAX_ACTIONS] and _do_probe returns on an empty
plan, so PROBE_MAX_ACTIONS = 0 kills R2 through a guard that already exists. R3 is
the only mechanism in the bundle aimed at level 1+, and level-0 breadth is capped
at 3.52. The arming order is load-bearing and verified: kill R2, then arm recovery,
then install - a failure anywhere leaves the byte-exact 2.14 config.

The efficiency note is economically wrong. It opens every turn with "every wasted
action costs you quadratically" and reports "you are 38.0x over the target" on a
level the agent has not cleared, where the ratio costs nothing at all. The four
holdout games that do clear level 0 clear it at 38-45x for ~0.06 each, strictly
better than the 0.00 they score by giving up. One appended clause, in the most
salient channel there is.

Also folds in the two measured priors still missing from the prompt: RESET is a
one-action rewind of the current level on a deterministic engine, and the default
button convention with its avatar-less fallbacks.

Verified: 9 cells compile, arming order holds, addendum intact at 5799 chars.
Smoked against real shipped source: probe plan 8 -> 0 actions, rider present on a
loud note and absent on a silent one, prompt +5799, level probe writes 11 fields.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The mind (arc3x/mindgraft.py) induces a forward model from nothing but the
(action, frame) history the framework already writes, and predicts frames it has
never seen. First honest measurement, 25 dev games, 250 random actions each, last
30% held out, no engine consulted: place=47%, exact=4%.

Three defects found by measurement, each with a named cause and a counter-example.

1. Ordering: `Mechanics.observe` gates all passability learning on
   `if self.avatar >= 0`, but `settle()` is what establishes an avatar and it runs
   *after* a batch is folded. So on a first pass the avatar is -1 throughout and
   every collision observation is skipped: `shifts` and `blocking` came out zero
   on every button of all 25 games, `walk_mask` fell back to all-ones, `_free`
   was true everywhere, and the model could not predict a refusal at all. Held-out
   misses were 100% "walked through a wall in imagination", zero the other way.
   Fixed by splitting the avatar-gated half of `observe` into `_track` and adding
   `replay_geometry` — a second pass over the same history, recording a disjoint
   set of facts so replaying cannot double-count the votes that chose the avatar.
   place 47% -> 60%.

2. Units: `_blame_block` incremented `blocking[c]` once per swept *pixel* while
   `passable[c]` increments once per *move*, so a six-wide sprite cast six wall
   votes against one floor vote for the same colour and `blocked_set` condemned
   the floor it had been walking on. The walkable mask collapsed to 4% of the
   board (sp80, re86) and 1% (r11l), with 42/73/53 phantom refusals and no error
   the other way. Now one vote per colour per refusal. place 60% -> 72%,
   sp80 0% -> 100%, cn04 exact 3% -> 45%.

3. Measurement: `movecall` compared `pred.moved` against `(before != after).any()`
   — "did the frame change" — not "did the sprite displace". Nearly every board
   has a HUD that ticks on every action (ls20: 2 pixels in rows 61-62; s5i5: 1
   pixel on 175 of 175 actions), so that test was true even for refused moves and
   the reported 84% meant nothing. Judged against located displacement it is 80%.

Now: place=72% exact=14% movecall=80%, 18 of 25 games with a model. The 7 silent
games are the avatar-less ones, where a translation model correctly declines.

Remaining failures are modelling gaps rather than bugs, and now visible: ar25 and
tu93 refuse moves for non-spatial reasons (ghost destinations are floor with
passable=108, blocking=14), and r11l is a click game where the sprite moves to
where you clicked, so no per-button delta exists to learn.

tools/mind_why.py now reports the miss *direction* (ghost vs phantom) plus the
destination colours and their evidence, which is what made 2 and 3 findable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Prior-session engine work that was never committed. agent.py gains the Relive
fallback for when the imagination has no objective to descend (RESET-and-replay
instead of deepcopy, since the gateway has no snapshot), an idle-round
convergence test so a round that bills nothing three times in a row stops
grinding, and the ACT-button branches that make Mechanics.acts plannable.

dream/progress/runner carry the objective ratchet and the retire-on-floor /
retire-on-stale rules. The why_* modules are single-question diagnostics kept
next to the code they interrogate.

None of this reaches Kaggle yet: arc3x is a component source, not the agent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The click model was built on a 25-cell lattice, which reported tn36 as having
1.0 distinct outcomes. tools/click_targets.py replaces the sample with the
engine as oracle: freeze frame 0, deepcopy, click every one of the 4096 cells,
group by the exact resulting board. tn36 has ELEVEN outcomes, with 96% of cells
landing on one of them. The lattice had sampled only the boring 96%.

    m0r0 1  sp80 1  sc25 1  ka59 2  vc33 3  lf52 3  dc22 3  cn04 3
    s5i5 5  sb26 5  bp35 8  ft09 9  tn36 11  su15 144  r11l 3096

That inverts the advice v15 was about to ship. The first draft told the agent
that agreeing frames mean "the coordinate is decoration, stop searching" — but
agreeing samples are the EXPECTED observation when a handful of cells in 4096
do everything interesting, so they can never license stopping. Section 5 of the
prompt addendum now says so, and adds how to find those cells: they sit on
non-background colours at 2.75x chance and on minority colours (<5% of board)
at 7.83x, in a median of 2 compact clusters. Sampling the board evenly is the
one search guaranteed to miss them. r11l is the negative control.

Section 6 is new and comes from the harvested run being clock-bound, not
action-bound: all four games report wallclock_s 7920.2 and state "gave_up", so
actions were purchased with tokens at 394-661 tok/action. Batching is the lever,
and the two costs the model cannot see from inside a batch are read out of
Solver.step_env — only the last frame returns, and the loop does not break on an
unchanged board. Hence "batch what you can predict, single-step what you probe".

Also in v15: the bundle loader no longer takes the first rglob hit. That bug
attached the dataset WITHOUT src/taaf-grafts on 2026-08-23 and silently played
stock for 2h12m; rglob order is filesystem order, so a rerun could not have
reproduced it. verify_submission_notebook.py was itself pinned to v14, so its
checks had never run on the upload candidate — it now defaults to v15 and check
[5] rejects the first-hit loop outright.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… done item pending

42 __pycache__/*.pyc files were tracked from before .gitignore existed. The
ignore rules already cover them, so this only removes the stale bytes.

ARCHITECTURE.md:250, PLAN.md:186 and TOMORROW.md:136 all said Mechanics.acts was
exposed but that "nothing uses them". It has two consumers: act_round
(agent.py:580) presses each act button at most once, ordered by how often it has
been seen to do anything, and push_frontier (agent.py:718) tries them against a
blocking colour after shoving it fails — which is the "walk there, then press
use" capability the plan asked for. Found while auditing what actually ships; a
doc that under-reports progress costs a session re-deciding solved questions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant