Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

LLMs Get Lost in Evolving User Intent

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales






























Improved Distribution Matching Distillation for Fast Image Synthesis
THREE-BODY SCATTERING FOR GENERATIVE MODELING
Scaling Native Multimodal Pre-Training From Scratch
DataPrep-Bench: Benchmarking LLMs as Training Data Preparators
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation
From Memory to Skills: Evidence-Grounded Co-Evolution Governance for Long-Horizon LLM Agents
Context-weighted Discrete Flow Matching
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
Show, Don’t Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
VISUAL CONTRASTIVE SELF-DISTILLATION
ReferTrack: Referring Then Tracking for Embodied Visual Tracking
K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
Scaling Laws for HyperNetwork-Based Knowledge Injection in Large Language Models
An Exam for Active Observers
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
SELF GRADIENT FORCING: NATIVE LONG VIDEO EX-TRAPOLATION
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Vera: A Layered Difusion Model for Content-Preserving Video Editing
Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
Automated Discovery Has No Universally Superior Harness
Towards a Science of Scaling Agent Systems
AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report
Mage-Flow: An Eficient Native-Resolution Foundation Model for Image Generation and Editing
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers