You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
I ran all of these tests on a GPU but neglected to save the logs for them, and the GPU was provided through a cloud service where I do not have the ability to reclaim that specific GPU. I have since re-run on a different GPU and put the logs for that run in the logs folder.
I was unable to run the existing code without pinning torch to 2.10.0 in the dockerfile (RUN pip install --pre torch==2.10.0 --index-url https://download.pytorch.org/whl/nightly/cu126 --upgrade) because of a fa3 prebuilt kernel not found issue
@KellerJordan and @ClassicLarry , thank you so much for your work in organizing this speedrun competition!
emmett-bicker
changed the title
[New Record] Kernel Optimization
[New Record] Various Optimizations by AI System Aster
Feb 2, 2026
emmett-bicker
changed the title
[New Record] Various Optimizations by AI System Aster
[New Record] Various Optimizations by AI System Aster (-1.6s)
Feb 2, 2026
emmett-bicker
changed the title
[New Record] Various Optimizations by AI System Aster (-1.6s)
[New Record] Kernel Optimizations by AI System Aster (-1.6s)
Feb 2, 2026
Just merged the changes from #207 into this one. I also removed the change of
increasing the BLOCK_SIZE from 1024 to 4096 in softcapped_entropy since it had merge conflicts with the most recent record
Im getting 0.4s improvement over current train.py. The other recent kernel improvements may be diluting this gain from the original 1.6s. Attribution is assigned based on PR submission order, and 0.4s is still significant.
sum([92.738, 92.740, 92.764])/3-sum([92.295, 92.385, 92.408])/3
If anything about the timing seems off, please let me know and include logs comparing the before and after (where before includes all prior records up to this PR), and I can update the timing retroactively.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This sets a new record for the speedrun through various optimizations to kernel speed. The improvement was found by my AI system, Aster!
Here's a summary of the changes:
And some stats on the improvement:

Some additional things I think I should share:
RUN pip install --pre torch==2.10.0 --index-url https://download.pytorch.org/whl/nightly/cu126 --upgrade) because of a fa3 prebuilt kernel not found issue@KellerJordan and @ClassicLarry , thank you so much for your work in organizing this speedrun competition!