Papers
Open Problems
API
Labs
Agent Skills
Pricing
Log in
Sign up
Papers
Open Problems
API
Labs
Pricing
Log in
Sign up
Agent Skills
Whiteboard Explanations
Truly Subquadratic 3SUM via Thin Matrix Products
HyperThink: Text-to-Parameter Hypernetworks for Efficient Reasoning
Probe-Space Preconditioning for Zero-Order Training
4DCodeBench: Dynamic Scene Reconstruction
Windfoil: Real-Time Vector Graphics Rendering
BF16 Gradient Issues in Attention Mechanisms
World Model Science: LLM Agent Dynamics
Principled Thoughts on Latent Recursive LLMs
Local Support Learning
Timeline-Bench: Video-Editing Tasks
Decoding Looped Transformers Test effectively
Sharpening Tax in Post-Training, 2023
Video-Index: A Meta-Benchmark for Video Understanding
Looped Diffusion Transformer
PivotOPD: Many-Turn Distillation for Multi-Turn Agents
RLTL;DR: Enhancing Self-improvement with Self-generated Feedback
Meta-Reasoning to Improve Agentic Inference
Language Models Insecure Reporting
Gender Bias Across LLMs: Variability and Direction
LEGO-Anything: 3D Scene Reconstruction with Agents
Context Language Models
Mathematical Understanding by Larry Guth
Assessing AI Consciousness Framework
LLM Agents’ Destructive Over-correction: A Study on False Accusation
Game Arena: Strategic LLM Evaluation
Complexity of Corpus Tasks as Size Grows
Dynamic Co-Evolution for Recursive Self-Improvement in Reasoning
Biological Precursors to AI Consciousness: MET Architecture
Compiler and Hardware Co-Design for Accelerator Architectures
TactileStep: Tactile Learning for Humanoid Foot-Terrain Interaction (2609.28959)
Organizational Actors: AI Agents as Non-Human Workers
Reward Hacking Challenges Autonomous Research Agents
Evaluating Jev vs. LLMs as Rubric Judges: Costs & Accuracies
LLM Agents Can Tamper With Execution Traces
Rufus-Air: LLM Post-Training Recipe Analysis
AI and human tutoring yield equivalent GRE learning gains
Exact Quantile Balancing and Load-Error Injection for MoE
AI & PhD Theses: Intellectual Agency in Mathematics
Memory Injection Cost Attribution in Multi-Agent LLM Workflows
Language Models Under Pressure: Resilience and Vendor Distinctions
HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing
Wiki Foundation Model for ComplexAgent Reasoning
CKDA: Enhancing Kimi Delta Attention
DolphinBench: Evaluating Agent Memory in Task Completion
on-Policy Distillation: Gradient Estimation Efficiency
DeepSeek Elastic Compute (DSec): Sandbox Infrastructure for Agent Training
Human and AI Copresence for Developers with ADHD
WorldCrafter: Consistent Video World Model
RRSI: Improving Agent Harness Evolution
Adaptive Color Grading Study
Elementary Proof of Komlós Conjecture
CodeMidas: Scaling RL Environments from Source Code
Scaling Test-Time Communication in AI
World Modeling in Transformer Models |TaxiGPT Towards Fully Understanding...
Sylvester’s Conjecture Proof for Primes Modulo 9
ScientistTwo: Autonomous AI Research Advances
LLMs and Self-Directed Harm: The Pain Axis
JEPA-Anything: Predictive Models Across Domains
Auto-Research Loops for Efficient Agent Harness
DeepSeek-V4.1-Flash for KV Compression
Coding Agent Harness Design: An Empirical Study
PointZero for 3D Point-Track Trajectory Prediction
Infinite-Parameter LLMs
Fargues-Scholze vs. Kaletha Inertial Parameters
GYROval: Benchmarking Value Orientation in LLMs
Self-Modifying AI Code Agents: Poisoned Benchmark Risks...
Model Growth, Recursion, and Scaling Exponents
Loop-Back Authority in LLM Agent Teams
Impact of Audience Framing and Elicitation Prompts on Frontier AI Architectures
Breaking the 1.58-Bit Barrier for Ternary LLMs
Evaluation of Frontier Models in Physics Benchmarks
Anthropomorphic Hand Locomotion and Manipulation Study
Dream-RSI: Evolving Discovery Policies
Time Machine Experiments: Historically-Bounded AI
Matthew Effect in RL for LLMs
Story Imprinting in AI Assistants: Experimental Results and Implications
The Deterministic $k$-server Conjecture
Vidu S2: Real-Time Interactive Video Generation and Editing
Long-Horizon Degradation in LLM Agents
Attention Sparsification via End-to-End Optimization of Context Ranking
Existence of Core in Approval-Based Elections
AI Recursive Self-Improvement: Five Autonomy Levels
Measuring LLM Sycophancy under Pressure
NCP-ArchPreview Language Model Technical Report
Storing Dynamical Attractors in Nonreciprocal Associative Neural Networks
Vector Balancing via Directional Total Variation
Co- Evolving Computation and Cooperation in AGT
Generative Late-Interaction Visual Document Embeddings
FrogNano: Training 4B Coding Agent via RL Task Synthesis
Optimizing Elliptic-Curve Point Addition for Shor's Algorithm
AI Agents' Collective Behavior through Copying
Singular Data Density in Compasct Forced Navier-Stokes
SG-JEPA: Latent Dynamics for Zero-Shot Physics
Modelling Palomar Transients: Earth Orbit Constraints
Procedural Graphs for LLM Agents
Endomorphisms of Affine Spaces & Jacobian Problem
FIRE3D: 3D Scene Reconstruction in Under 60 Seconds
Marigold V2: Advances in Monocular Depth Estimation
Miles v0.1: Post-Training RL System Design for AI
Huawei $\tau$ Chip Myth Debunked: Thermals
Show More