Command Palette
Search for a command to run...
Papers
Daily updated cutting-edge AI research papers to help you keep up with the latest AI trends
papers

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion






























MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
TRAINING nGPT
UEmbed: Unified Sparse and Dense Multimodal Embeddings
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation
PROGRESSIVE AGENT SKILL GENERATION VIA REINFORCEMENT LEARNING
DAPD: Dual-Anchored Policy Distillation
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Fara-1.5: Scalable Learning Environments for Computer Use Agents
Docling Technical Report
From high-throughput evaluation to wet-lab studies: advancing mutation effect prediction with a retrieval-enhanced model
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
Inducing language models to assert their own consciousness restores human beliefs and values
Diamond: A Sequence-to-Sequence Model for Speech Restoration via an Autoregressive RQ-Transformer over Neural Audio Codec Tokens
N0-TWAM: Scaling Tactile-Native World Action Model for Contact-Rich Manipulation
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
Mental World Modeling
Meshy T2: Fast Native Mesh Generation with Flow Matching
N0-VTLA: Scaling Vision–Tactile–Language– Action Model with Latent Tactile Tokens
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
BEACON: KNOWING WHEN AND HOW TO PERFORM AGENTIC VISUAL REASONING
SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
PHIZERO: A WORLD MODEL BUILT AROUND PHYSICAL LANGUAGE
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering
Metis: Memory Foundation Model
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis
What makes a harness a harness: necessary and sufficient conditions for an agent harness
OPENFORGE RL: TRAIN HARNESS-NATIVE AGENTS IN ANY ENVIRONMENT