Skip to content
Open
Show file tree
Hide file tree
Changes from 1 commit
Commits
Show all changes
65 commits
Select commit Hold shift + click to select a range
daf55fc
feat: Implement Heuristic Explorer Agent for ARC-AGI-3
samrishtt May 31, 2026
3d2c722
Integrate Forge v3 BFS + CNN solver and update README
samrishtt May 31, 2026
48f0199
Update README with attributions and results
samrishtt May 31, 2026
f136a2c
Update README to clarify project description
samrishtt May 31, 2026
96b6776
MASTER BASELINE v10: 3 bug fixes + 5 reference agent analyses
samrishtt Jun 7, 2026
a9e01c5
v11: Trigger-Aware BFS + ACMD + 5 critical fixes
samrishtt Jun 15, 2026
90eb5f4
security: remove .env from tracking, add to gitignore
samrishtt Jun 15, 2026
3d253f2
Update MASTER BASELINE v10 version number to 0.29
samrishtt Jun 15, 2026
a24e8d6
v12: deploy Forge v20 Patched agent (0.33 score) and clean up unwante…
samrishtt Jun 21, 2026
48f1457
Merge branch 'main' of https://github.com/samrishtt/arc-agi-3-kaggle-…
samrishtt Jun 21, 2026
e28b6c2
docs: rewrite README to match Forge v20 Patched architecture
samrishtt Jun 21, 2026
636a064
docs: enrich README with detailed solved_path bug description and Go-…
samrishtt Jun 21, 2026
e1b4660
docs: add go_explore_memory infographic and update README reference
samrishtt Jun 21, 2026
783968e
hjfhj
samrishtt Jul 13, 2026
5d769f7
jttd
samrishtt Jul 13, 2026
a396e08
research: add analyzer seed control
samrishtt Jul 13, 2026
78b43ee
research: record baseline action evidence
samrishtt Jul 13, 2026
db8cca7
research: log failed seed control
samrishtt Jul 14, 2026
c149eae
research: record control submission score
samrishtt Jul 17, 2026
353d011
,,gjdjg
samrishtt Jul 18, 2026
3bd5207
jf
samrishtt Jul 18, 2026
4df7d57
docs: comprehensive ARC-AGI-3 ablation research, experiment logs, and…
samrishtt Jul 24, 2026
8cbb370
docs: document Experiment B (0.66 score), tn36 -54% action reduction …
samrishtt Jul 25, 2026
3c2aa86
docs: Experiment C scored 1.06 — tn36 cleared Level 1 (10.04 score), …
samrishtt Jul 26, 2026
0fa49d0
docs: Experiment D scored 0.95 — empirical confirmation that Stock 32…
samrishtt Jul 27, 2026
52d87df
docs: Experiment F scored 0.77 — false-positive trimming on legitimat…
samrishtt Jul 28, 2026
610be92
docs: Exp 9 schema_notes=0.47 NET NEGATIVE and Exp 10 banking=1.10 CO…
samrishtt Aug 8, 2026
5f44ea4
feat: Experiment 11 (Beat 1.33) — Level 2 sandbox tools (find_path, f…
samrishtt Aug 8, 2026
277e27d
feat: Add sam agi.ipynb notebook with Exp 11 (Level 2+3 tools and arc…
samrishtt Aug 8, 2026
cfd4bf0
fix: Guarantee submission.parquet creation at start of benchmark run …
samrishtt Aug 9, 2026
3cbd7ef
fix: Resolve Cell 13 docstring syntax error in sam agi.ipynb
samrishtt Aug 9, 2026
9d98f5d
docs: Exp 11 scored 0.60 Kaggle but 2.6848 local (HIGHEST EVER) - m0r…
samrishtt Aug 10, 2026
18d68ec
feat: Exp 12 — calibrated Level 2+3 with 32K context + compact prompt…
samrishtt Aug 10, 2026
825e231
fix: Re-add the banking flag to Exp 12 config
samrishtt Aug 10, 2026
7b4d41b
docs: Save multi-agent swarm architecture ideas for Kaggle constraints
samrishtt Aug 10, 2026
8d94db3
docs: Add Dynamic Agent Scaling idea
samrishtt Aug 10, 2026
1fb3957
fix: Correct indentation in Cell 13 try/except block
samrishtt Aug 10, 2026
4985cb4
docs: Clean ASCII comments and verify Exp 12 notebook readiness
samrishtt Aug 12, 2026
7548a0a
chore: track arc3x search engine + add .gitignore before artifact cle…
samrishtt Aug 22, 2026
1bb3859
fix(arc3x): three bugs that made Go-Explore score 0.00; sk48+m0r0 L0 …
samrishtt Aug 22, 2026
dfb73d0
feat(arc3x): per-level cell recalibration, search restarts, numpy stu…
samrishtt Aug 22, 2026
40a1d75
arc3x: measure the mind end to end, and classify every button
samrishtt Aug 22, 2026
1944e2a
docs: Queue the next session - four tasks and the number that governs…
samrishtt Aug 22, 2026
0c75ffa
arc3x: make the measurement fast, reproducible, and honest about wher…
samrishtt Aug 23, 2026
e688696
docs: architecture and method note, written for an outside researcher
samrishtt Aug 23, 2026
02eaa30
feat: submittable notebook (v13) — level probe + the two family priors
samrishtt Aug 23, 2026
02377cd
fix(notebook): state the level-0 exception in the efficiency prior
samrishtt Aug 23, 2026
60cd9d9
arc3x: the experiment log answers the flag question, and it changes t…
samrishtt Aug 23, 2026
ea74840
arc3x: the mind learns walls — place 47% -> 72% on held-out play
samrishtt Aug 23, 2026
7fe7a67
arc3x: Relive search, per-level recalibration, and the why-* diagnostics
samrishtt Aug 23, 2026
42d8135
v15: measure the click space exhaustively, and fix the prior it refuted
samrishtt Aug 23, 2026
adb7fe3
chore: untrack committed bytecode, and correct three docs that call a…
samrishtt Aug 23, 2026
ca0b99f
kghxyx
samrishtt Aug 27, 2026
83d3e27
jjgc
samrishtt Aug 27, 2026
1184d0c
test: document and guard marker proposer state
samrishtt Aug 30, 2026
3df7975
fix: keep level scene cuts out of mind learning
samrishtt Aug 30, 2026
30d3513
feat: package human-mind pilot notebook
samrishtt Aug 30, 2026
d37fdf6
feat: add history-based progress objectives to pilot
samrishtt Aug 30, 2026
4202350
feat: explore learned level maps before first objective
samrishtt Aug 31, 2026
241d7af
feat: learn click semantics across levels
samrishtt Aug 31, 2026
cb81a52
feat: preserve v12 baseline with cautious level sidecar
samrishtt Aug 31, 2026
a428e0a
fix: make v12 sidecar import path explicit
samrishtt Aug 31, 2026
e6eafbf
docs: document v12 sidecar architecture
samrishtt Aug 31, 2026
ce97c67
feat: add guarded mental simulation solver
samrishtt Aug 31, 2026
3784caa
docs: record mental solver checkpoint
samrishtt Aug 31, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Prev Previous commit
Next Next commit
fix(notebook): state the level-0 exception in the efficiency prior
The finished suite showed four holdout games clearing level 0 and still scoring
0.00: ft09 1633 actions vs baseline 43, cd82 2502 vs 55, vc33 289 vs 7, m0r0 1179
vs 30 - all about 40x, and (1/40)^2 is 0.06% of that level's points.

So 'exploring level 0 is cheap' is true only relative to the later levels. If
level 0 is the only level you clear, its efficiency IS your score. The addendum
now says that, and says to stop exploring once the objective is visible - which is
exactly the failure mode those four games display.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
  • Loading branch information
samrishtt and claude committed Aug 23, 2026
commit 02377cd8866d5e230195128180080c98194dead6
Original file line number Diff line number Diff line change
Expand Up @@ -632,7 +632,11 @@
" clearing. Spend actions to establish the mechanics: which button moves what, what\n",
" blocks movement, what the objective is, what an action does when nothing appears\n",
" to change. The first level also carries the smallest weight of any level in the\n",
" game.\n",
" game. But \"cheap\" here means cheap *relative to the later levels*: if the first\n",
" level turns out to be the only one you clear, then its efficiency is your entire\n",
" score for this game. So explore it freely while you still do not understand it,\n",
" and the moment you can see the objective, go and complete it rather than\n",
" continuing to explore for its own sake.\n",
"- Being crisp pays on later levels. Once you know the mechanics, execute. Later\n",
" levels carry more weight and charge every wasted action quadratically, and you\n",
" should re-derive nothing you already established earlier in the same game - the\n",
Expand Down
6 changes: 5 additions & 1 deletion tools/make_submission_notebook.py
Original file line number Diff line number Diff line change
Expand Up @@ -230,7 +230,11 @@ def finish_game(self, generated_tokens: int = 0, uncached_input_tokens: int = 0)
clearing. Spend actions to establish the mechanics: which button moves what, what
blocks movement, what the objective is, what an action does when nothing appears
to change. The first level also carries the smallest weight of any level in the
game.
game. But "cheap" here means cheap *relative to the later levels*: if the first
level turns out to be the only one you clear, then its efficiency is your entire
score for this game. So explore it freely while you still do not understand it,
and the moment you can see the objective, go and complete it rather than
continuing to explore for its own sake.
- Being crisp pays on later levels. Once you know the mechanics, execute. Later
levels carry more weight and charge every wasted action quadratically, and you
should re-derive nothing you already established earlier in the same game - the
Expand Down