Conversation
- Fixed PicklingError: register game_mod in sys.modules - Fixed GameAction Enum unpickling: custom copyreg reducer - Fixed TypeError in _sample_novelty_guided: dict comparison - Added encoding=utf-8 to build_notebook.py for Windows - Added reference notebooks: Forge v16 (0.35), Forge v18 (~0.39), Arc26 v15, Ash agent - Added extraction/injection utility scripts - Score: 0.23 on Kaggle (up from 0.18 baseline)
- FIX arcprize#1: Hidden field probing BEFORE Phase 1 (was Phase 3) - FIX arcprize#2: Clock field elimination prevents timer state explosion - FIX arcprize#3: CBAM removed from ForgeNet (pretrained weight mismatch) - FIX arcprize#4: MCTS f-string double-brace bug (log signals restored) - FIX arcprize#5: _state_hash selective mode (hidden_fields used properly) - NEW: ACMD counter search as Phase 3 (threshold-gated games) - NEW: Professional README with architecture diagrams - All search phases now use trigger-aware hashing consistently
…d reference files
… architecture documentation
…breakthrough, and Exp 6 recommendation
…schema_helpers breakthrough, Exp G prepared (45K context)
…K (1.06) is optimal; transition to building state_dedup.py custom graft
…e backtracking; recommendation to restore Exp C 1.06 baseline
…NFIRMED POSITIVE with full ranked leaderboard
…ind_objects, death_memory) + Level 3 search loop + post-death cleanup
…anup arc3x/ (twin.py, explore.py, verify_twin.py) was built last session but never committed and never run. Committing it as a safety net before deleting ~700MB of spent experiment artifacts. Also records the 2.14 notebook: stock TAAF solver + Qwen3.8-27B-FP8, with an empty customization hook. The +0.81 over 1.33 came from the model swap alone. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…now cap at 115 Bug 1 (fatal, algorithmic): the archive key hashed the raw 64x64 frame and was effectively bijective - tn36 60 steps -> 60 distinct keys. Every step looked novel, so the archive became a log of every state visited, novelty carried no signal, and Go-Explore degenerated into a random walk. It also silently killed compress()'s loop-splice pass, which can only fire when a state repeats. Root cause: only 2-3% of the 4096 pixels ever move, and tn36 row 1 is a 49-pixel HUD timer draining by exactly 6 per action (441,435,429,417...). That one row makes every frame globally unique. General fix in cell.py: a pixel that moves monotonically with the action count is a clock, not state. State revisits values; a clock never does. Calibrated per game from ~600 free simulated actions. No per-game knowledge. Collapse measured: sk48 91->28 keys, m0r0 77->29, tn36 60->43. Bug 2 (performance): snapshotting every new archive cell cost a 20ms deepcopy per step -> 19 steps/sec. Now snapshots only on restore (~48x rarer) and skips nodes shallower than 16 actions, since deterministic replay is cheaper there. 19 -> 546 steps/sec. Bug 3 (coverage): nothing tracked which actions had been tried from a cell, so the rollout re-picked the same few forever. Added Node.tried plus a selection weight favouring cells with untried actions. Stickiness now applies only to non-click actions. Measured, on the two games a month of LLM work never cracked: sk48 L0: never completed (781-940 actions) -> 29 actions (baseline 61), 115.0 m0r0 L0: 477 actions, score 0.02 -> 15 actions (baseline 30), 115.0 Game scores 2.78 and 4.76 from level 0 alone, vs a 2.14 leaderboard baseline. Also corrects the docstring's 6,811 steps/sec claim: that was an artifact of benchmarking an already-dead game. Real sustained ceiling is ~700 steps/sec. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…dent policy Three changes, each measured. 1. Per-level recalibration (the depth bug). calibrate() probes one root, so the mask 'varies & ~clock' describes only that level's screen. Measured wrong in both directions on later levels: cd82 L2 has 139 informative pixels the L0 mask excludes (14% of state invisible), and vc33 L1 has 1,779 clock-like pixels vs 494 at L0, so ~1,285 timer pixels get hashed into the archive key and it turns bijective again - the exact fatal bug the mask was added to fix, reappearing at every level past the first. This is why games clear L0 in ~10 actions then stall. arc3x/diag_mask.py reproduces the measurement. 2. Restarts on failure. solve_level is randomised search with heavy-tailed runtime; a bad opening traps a rollout distribution for the whole budget. Also, the 25% reserved for compression was silently discarded whenever a level failed. Now up to 3 re-seeded attempts share that budget. 3. Student policy (arc3x/student.py, train_student.py). 4096 -> 256 -> 261 MLP in pure numpy, no torch: one-hot of the 4x4-max-pooled frame -> 5 simple actions + 256 coarse click cells, softmax masked to the legal action set. Trained by imitating the searcher's own compressed plans. Measured 0.424 held-out vs 0.194 random-choice = 2.18x on 223 examples. Numpy-only so it cannot fail the way vLLM did (Exp 11: 2.68 local -> 0.60 Kaggle on prefill timeout). 25-game sweep: mean 2.978 -> 4.695 after the snapshot-memory fix. lp85 reached 5/8 levels for 41.67, confirming depth is the lever. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The imagination was never in the agent. Dream and Progress were only ever exercised by measure_dream.py's 240-step harness, so every number measured about them had no path to a score. Wired into agent.py: Dream is constructed alongside Mechanics, fed every transition, cut on level change, and imagine() re-plans after each action. Scored with run_suite over all 25 dev games: 0.111 -> 0.142, 2 levels completed, and every game still burns the full 3000-action cap. The Kaggle submission scores 2.14, so arc3x does not ship as the agent - it is a component source. lp85 clears level 0 in 6 actions against a baseline of 17, which is the one place the objective demonstrably works. Also classified every declared button as MOVE / ACT / DEAD instead of only asking whether the avatar translated (probe_actions.py). Mechanics.moves filters d != (0,0), which discarded a whole class of button: 12 of 25 games have an ACT button - use, grab, select, rotate cd82 has five working buttons and the planner saw none of them tr87 has four and the planner saw none Exposed as Mechanics.acts, built from new changes/shifts counters. cd82 and tr87 were filed as 'no steerable avatar'; they are not click games at all. settle() also loses real MOVE buttons: wa30 and sc25 each have four, the planner sees [2, 4]. Added _fill_deltas to accept single-observation deltas for the believed avatar once its step sizes are known - identifying who I am needs strong evidence, using a button does not. It did not fix those two: there are zero votes for the missing axes, so the cause is upstream in moved_objects, where a sprite that rotates as it turns never matches as a rigid translation. Recorded rather than papered over. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… them Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e it loses Three things the suite could not do: run in parallel, say why a game failed, or tell a fitted result from a general one. PARALLEL (graded.run_suite workers=N). Every game is fully independent - its own Twin, its own engine, no shared state, no writes - so the parallel result is exact, not approximate. The naive version regressed badly: 6 workers each sized a BLAS thread pool to all 12 cores, 72 threads thrashed over 12, and ls20 managed 32 actions in 60s where serial did 401. Thread limits have to be set in the parent before the pool exists, because 'spawn' re-imports numpy in the child and sizes the pool at import time. With that fixed, parallel reproduces serial to the action (ls20 401 actions, 1 level, 0.78 both ways) and runs faster. Also made the wall clock loud. Kaggle bills actions, not seconds, so a run cut short by the deadline is a machine-dependent experiment rather than a measurement, and two of them are not comparable. Agent.timed_out records it and stall_summary shouts about it instead of averaging it in. WHY IT LOST (GradedRun.note). agent_fn returns nothing, so the score was the only channel out, and a score cannot distinguish a game the agent never understood from one it understood and could not execute - which need opposite fixes. Notes are keyed by the level the attempt *started* on: the attempt that clears a level advances the counter first, and filing 'cover completed a level' under the level it moved us to blames the wrong one. First run already names the biggest leak: click:change=386 of all stall reasons. click_round returns success on any pixel change whatsoever, so on cd82 the agent clicks 259 times, something changes every time, and it never once asks whether the change was progress. OVERFIT CHECK (suite.py --split both). The evaluation set is 110 private games; these 25 are development. The risk is not a rule that fails here, it is a rule that succeeds here for a reason that does not exist there. Fixed alphabetical every-third split: 17 tune, 8 hold. Fixed and never re-drawn, because a holdout re-rolled after a disappointing result is a lottery. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
States the 3.52 / 10.57 / 21.14 ceilings, the five components and what each is for, the two load-bearing priors (reversibility for avatar identification, the ratchet rule for a label-free objective), and the measurement discipline - including the three powers the local engine offers that the graded gateway does not, and which every earlier local measurement here was inflated by. Keeps the negative results, because they are the useful part: the mirror-panel hypothesis refuted by its own control, graft tuning 0-for-11, prediction accuracy proved not to be the bottleneck, and compute proved not to be the bottleneck. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The notebook is the deliverable. It is the 2.14 notebook plus one cell inserted
at its own documented customization hook, carrying two grafts and nothing else.
Both are expressible as inline notebook code, which is the binding constraint:
the solver is imported from a read-only Kaggle dataset this repo cannot change,
so a graft is either a wrapper or a string. That is what rules the arc3x modules
out and these two in.
* Level probe (inert). taaf.game.Game.finish_game already computes
actions_per_level and base_actions_per_level and prints them (game.py:636-650);
TAAF_MINIMAL_DIAGNOSTICS=1 on a true submission is why no artifact of them has
ever survived. Appends one JSON line per finished game. This is the
measurement the plan has been blocked on: 2.14 is equally consistent with
"level 0 on most games" and "two levels on a few", which need opposite fixes.
* Family priors (behavioural). Appends one addendum to the agent's system
prompt: (1) the win condition is usually a cover predicate and is therefore
drawn on frame 0 as a set of small identical static objects, groupable by the
segmentation hash the agent already has; (2) level score is
(baseline/actions)^2 and only for levels you clear, so exploring level 0 is
cheap, later levels must be crisp, and giving up early never pays.
Deliberately excluded, each for a measured reason: markers.py (its own grader
returned 0/1 on the only row with source-read ground truth — it proposed floor
tiling instead of the exit, because ranking candidates by repeat-count promotes
scenery); the ACTION1-4=N/S/W/E convention (measured at exactly 0.00 as a planner
seed); any change to GRAFT_FLAGS including context_window (untested, and changing
it alongside two new grafts would make the result unattributable again).
Verified rather than assumed. tools/verify_submission_notebook.py compiles every
cell (top-level await allowed), checks the graft lands after GRAFT_FLAGS and
before bm.run, and confirms the addendum survived two levels of escaping.
tools/smoke_grafts.py execs the cell lifted out of the generated notebook against
the real shipped taaf/inference source: all eleven GameRun field names resolve,
the wrapper is idempotent and writes one row per game rather than per call, the
missing-baseline branch records None, and the system prompt grows 12358 -> 15573
by exactly the addendum. That run also confirms the scoring weights are i+1 —
1 of 4 levels scored exactly 10.00, the completion cap 1/(1+2+3+4).
Suite at this commit: tune 0.167 / hold 0.001 over 25 games. arc3x remains a
component source, not the agent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The finished suite showed four holdout games clearing level 0 and still scoring 0.00: ft09 1633 actions vs baseline 43, cd82 2502 vs 55, vc33 289 vs 7, m0r0 1179 vs 30 - all about 40x, and (1/40)^2 is 0.06% of that level's points. So 'exploring level 0 is cheap' is true only relative to the later levels. If level 0 is the only level you clear, its efficiency IS your score. The addendum now says that, and says to stop exploring once the objective is visible - which is exactly the failure mode those four games display. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…he notebook Sam asked why GRAFT_FLAGS arms only 3 of the 10 flags composite.install accepts. docs/EXPERIMENT_LOG.md answers it: all seven others were tested on the real leaderboard over Exp 3-11 and every combination scored below the three-flag baseline (banking 1.10, helpers+void+transfer 1.06, state_dedup 0.77, schema_notes 0.47). The three are the survivors of a finished ablation, not an oversight. Sam is also right that some of those experiments were good: Exp 11 set the highest local score ever recorded (2.6848) and produced the first m0r0 level-0 clear, then scored 0.60 on Kaggle. The log names the mechanism - 57K context plus verbose per-turn notes caused vLLM prefill timeouts on the shared GPU, stranding games at zero. Good locally, destroyed by latency remotely. Two findings from reading the grafts change what ships: recovery is separable. Both recovery regressions are attributed by the log to R2's probe tax alone, and R2 spends up to 16 actions per stalled level. R1 (clear a death-spiralled history, rewrite the world model) and R3 (carry mechanics across the vendor level wipe via cross_level_notes) cost zero actions by construction. build_probe_plan returns plan[:PROBE_MAX_ACTIONS] and _do_probe returns on an empty plan, so PROBE_MAX_ACTIONS = 0 kills R2 through a guard that already exists. R3 is the only mechanism in the bundle aimed at level 1+, and level-0 breadth is capped at 3.52. The arming order is load-bearing and verified: kill R2, then arm recovery, then install - a failure anywhere leaves the byte-exact 2.14 config. The efficiency note is economically wrong. It opens every turn with "every wasted action costs you quadratically" and reports "you are 38.0x over the target" on a level the agent has not cleared, where the ratio costs nothing at all. The four holdout games that do clear level 0 clear it at 38-45x for ~0.06 each, strictly better than the 0.00 they score by giving up. One appended clause, in the most salient channel there is. Also folds in the two measured priors still missing from the prompt: RESET is a one-action rewind of the current level on a deterministic engine, and the default button convention with its avatar-less fallbacks. Verified: 9 cells compile, arming order holds, addendum intact at 5799 chars. Smoked against real shipped source: probe plan 8 -> 0 actions, rider present on a loud note and absent on a silent one, prompt +5799, level probe writes 11 fields. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The mind (arc3x/mindgraft.py) induces a forward model from nothing but the (action, frame) history the framework already writes, and predicts frames it has never seen. First honest measurement, 25 dev games, 250 random actions each, last 30% held out, no engine consulted: place=47%, exact=4%. Three defects found by measurement, each with a named cause and a counter-example. 1. Ordering: `Mechanics.observe` gates all passability learning on `if self.avatar >= 0`, but `settle()` is what establishes an avatar and it runs *after* a batch is folded. So on a first pass the avatar is -1 throughout and every collision observation is skipped: `shifts` and `blocking` came out zero on every button of all 25 games, `walk_mask` fell back to all-ones, `_free` was true everywhere, and the model could not predict a refusal at all. Held-out misses were 100% "walked through a wall in imagination", zero the other way. Fixed by splitting the avatar-gated half of `observe` into `_track` and adding `replay_geometry` — a second pass over the same history, recording a disjoint set of facts so replaying cannot double-count the votes that chose the avatar. place 47% -> 60%. 2. Units: `_blame_block` incremented `blocking[c]` once per swept *pixel* while `passable[c]` increments once per *move*, so a six-wide sprite cast six wall votes against one floor vote for the same colour and `blocked_set` condemned the floor it had been walking on. The walkable mask collapsed to 4% of the board (sp80, re86) and 1% (r11l), with 42/73/53 phantom refusals and no error the other way. Now one vote per colour per refusal. place 60% -> 72%, sp80 0% -> 100%, cn04 exact 3% -> 45%. 3. Measurement: `movecall` compared `pred.moved` against `(before != after).any()` — "did the frame change" — not "did the sprite displace". Nearly every board has a HUD that ticks on every action (ls20: 2 pixels in rows 61-62; s5i5: 1 pixel on 175 of 175 actions), so that test was true even for refused moves and the reported 84% meant nothing. Judged against located displacement it is 80%. Now: place=72% exact=14% movecall=80%, 18 of 25 games with a model. The 7 silent games are the avatar-less ones, where a translation model correctly declines. Remaining failures are modelling gaps rather than bugs, and now visible: ar25 and tu93 refuse moves for non-spatial reasons (ghost destinations are floor with passable=108, blocking=14), and r11l is a click game where the sprite moves to where you clicked, so no per-button delta exists to learn. tools/mind_why.py now reports the miss *direction* (ghost vs phantom) plus the destination colours and their evidence, which is what made 2 and 3 findable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Prior-session engine work that was never committed. agent.py gains the Relive fallback for when the imagination has no objective to descend (RESET-and-replay instead of deepcopy, since the gateway has no snapshot), an idle-round convergence test so a round that bills nothing three times in a row stops grinding, and the ACT-button branches that make Mechanics.acts plannable. dream/progress/runner carry the objective ratchet and the retire-on-floor / retire-on-stale rules. The why_* modules are single-question diagnostics kept next to the code they interrogate. None of this reaches Kaggle yet: arc3x is a component source, not the agent. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The click model was built on a 25-cell lattice, which reported tn36 as having
1.0 distinct outcomes. tools/click_targets.py replaces the sample with the
engine as oracle: freeze frame 0, deepcopy, click every one of the 4096 cells,
group by the exact resulting board. tn36 has ELEVEN outcomes, with 96% of cells
landing on one of them. The lattice had sampled only the boring 96%.
m0r0 1 sp80 1 sc25 1 ka59 2 vc33 3 lf52 3 dc22 3 cn04 3
s5i5 5 sb26 5 bp35 8 ft09 9 tn36 11 su15 144 r11l 3096
That inverts the advice v15 was about to ship. The first draft told the agent
that agreeing frames mean "the coordinate is decoration, stop searching" — but
agreeing samples are the EXPECTED observation when a handful of cells in 4096
do everything interesting, so they can never license stopping. Section 5 of the
prompt addendum now says so, and adds how to find those cells: they sit on
non-background colours at 2.75x chance and on minority colours (<5% of board)
at 7.83x, in a median of 2 compact clusters. Sampling the board evenly is the
one search guaranteed to miss them. r11l is the negative control.
Section 6 is new and comes from the harvested run being clock-bound, not
action-bound: all four games report wallclock_s 7920.2 and state "gave_up", so
actions were purchased with tokens at 394-661 tok/action. Batching is the lever,
and the two costs the model cannot see from inside a batch are read out of
Solver.step_env — only the last frame returns, and the loop does not break on an
unchanged board. Hence "batch what you can predict, single-step what you probe".
Also in v15: the bundle loader no longer takes the first rglob hit. That bug
attached the dataset WITHOUT src/taaf-grafts on 2026-08-23 and silently played
stock for 2h12m; rglob order is filesystem order, so a rerun could not have
reproduced it. verify_submission_notebook.py was itself pinned to v14, so its
checks had never run on the upload candidate — it now defaults to v15 and check
[5] rejects the first-hit loop outright.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… done item pending 42 __pycache__/*.pyc files were tracked from before .gitignore existed. The ignore rules already cover them, so this only removes the stale bytes. ARCHITECTURE.md:250, PLAN.md:186 and TOMORROW.md:136 all said Mechanics.acts was exposed but that "nothing uses them". It has two consumers: act_round (agent.py:580) presses each act button at most once, ordered by how often it has been seen to do anything, and push_frontier (agent.py:718) tries them against a blocking colour after shoving it fails — which is the "walk there, then press use" capability the plan asked for. Found while auditing what actually ships; a doc that under-reports progress costs a session re-deciding solved questions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.