Skip to content

[New Record] Kernel Optimizations by AI System Aster (-1.6s) - #217

Merged
ClassicLarry merged 3 commits into
KellerJordan:masterfrom
emmett-bicker:aster_record
Feb 16, 2026
Merged

ClassicLarry merged 3 commits into
KellerJordan:masterfrom
emmett-bicker:aster_record

Conversation

@emmett-bicker

@emmett-bicker emmett-bicker commented Feb 2, 2026 •

Copy link
Copy Markdown
Contributor

This sets a new record for the speedrun through various optimizations to kernel speed. The improvement was found by my AI system, Aster!

Here's a summary of the changes:

Kernel / Function Change Type Description
XXT_kernel Memory Switched to coalesced loads + in-register transpose (tl.trans).
ba_plus_cAA_kernel Memory Switched to coalesced loads + in-register transpose.
Host Configs (XXT, etc.) Tuning Increased num_warps from 4 to 8.
softcapped_entropy Math Replaced inner-loop divisions with pre-computed multiplications.

And some stats on the improvement:
Screenshot 2026-02-02 at 4 23 52 AM

>>> from scipy import stats
>>> losses = [3.2763,3.2777,3.2794,3.2773,3.2774,3.279,3.279,3.2791,3.2767]
>>> print("p=%.4f" % stats.ttest_1samp(losses, 3.28, alternative="less").pvalue)
p=0.0004

Some additional things I think I should share:

  • I ran all of these tests on a GPU but neglected to save the logs for them, and the GPU was provided through a cloud service where I do not have the ability to reclaim that specific GPU. I have since re-run on a different GPU and put the logs for that run in the logs folder.
  • I was unable to run the existing code without pinning torch to 2.10.0 in the dockerfile (RUN pip install --pre torch==2.10.0 --index-url https://download.pytorch.org/whl/nightly/cu126 --upgrade) because of a fa3 prebuilt kernel not found issue

@KellerJordan and @ClassicLarry , thank you so much for your work in organizing this speedrun competition!

@emmett-bicker emmett-bicker changed the title Merge in [New Record] Kernel Optimization Feb 2, 2026
@emmett-bicker emmett-bicker changed the title [New Record] Kernel Optimization [New Record] Various Optimizations by AI System Aster Feb 2, 2026
@emmett-bicker emmett-bicker changed the title [New Record] Various Optimizations by AI System Aster [New Record] Various Optimizations by AI System Aster (-1.6s) Feb 2, 2026
@emmett-bicker emmett-bicker changed the title [New Record] Various Optimizations by AI System Aster (-1.6s) [New Record] Kernel Optimizations by AI System Aster (-1.6s) Feb 2, 2026
@emmett-bicker

emmett-bicker commented Feb 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Just merged the changes from #207 into this one. I also removed the change of
increasing the BLOCK_SIZE from 1024 to 4096 in softcapped_entropy since it had merge conflicts with the most recent record

@ClassicLarry

Copy link
Copy Markdown
Collaborator

neat, will time later

@emmett-bicker

Copy link
Copy Markdown
Contributor Author

Thank you so much!

@ClassicLarry

ClassicLarry commented Feb 16, 2026 •

Copy link
Copy Markdown
Collaborator

Merging at 91.7. (-0.4s)

Im getting 0.4s improvement over current train.py. The other recent kernel improvements may be diluting this gain from the original 1.6s. Attribution is assigned based on PR submission order, and 0.4s is still significant.
sum([92.738, 92.740, 92.764])/3-sum([92.295, 92.385, 92.408])/3

If anything about the timing seems off, please let me know and include logs comparing the before and after (where before includes all prior records up to this PR), and I can update the timing retroactively.

@ClassicLarry
ClassicLarry merged commit 866e243 into KellerJordan:master Feb 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants