AI 论文 · 2026-06
浏览 2026-06 发布的AI 论文内容,第 4 页,共 818 条。
- Learning User Simulators with Turing Rewards
- Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States
- Does VLA Even Know the Basics? Measuring Commonsense and World Knowledge Retention in Vision-Language-Action Models
- TurboServe: Serving Streaming Video Generation Efficiently and Economically
- STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
- Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance
- The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL
- Seeing Before Reasoning: Decoupling Perception and Reasoning for Shortcut-Resilient Multimodal On-Policy Self-Distillation
- RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents
- Sumi: Open Uniform Diffusion Language Model from Scratch
- EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts
- Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
- Physics-IQ Verified
- REVES: REvision and VErification--Augmented Training for Test-Time Scaling
- Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness
- WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
- Learning from Your Own Mistakes: Constructing Learnable Micro-Reflective Trajectories for Self-Distillation
- Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning
- GateMem: Benchmarking Memory Governance in Multi-Principal Shared-Memory Agents
- Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish
- MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction
- Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
- Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns
- SproutRAG: Attention-Guided Tree Search with Progressive Embeddings for Long-Document RAG
- CEO-Bench: Can Agents Play the Long Game?
- When Does Trajectory-Level Supervision Permit Efficient Offline Reinforcement Learning?
- MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval
- PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation
- Guava: An Effective and Universal Harness for Embodied Manipulation
- SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
- Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification
- Variable-Width Transformers
- Adaptive Volumetric Mechanical Property Fields Invariant to Resolution
- Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
- Looped World Models
- Learning from the Self-future: On-policy Self-distillation for dLLMs
- EgoCS-400K: An Egocentric Gameplay Dataset for World Models
- Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding
- LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
- LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI
- ChLogic: Evaluating Robustness of Logical Reasoning in Chinese Expressions
- GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine?
- MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
- ActWorld: From Explorable to Interactive World Model via Action-Aware Memory
- From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
- OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation
- Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
- GeneralVLA-2: Geometry-Aware Reconstruction and Governed Memory for Robot Planning
- Kairos: A Native World Model Stack for Physical AI
- Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence
- Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion
- ProCUA-SFT Technical Report
- ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining
- Verified Detection and Prevention of Concurrency Anomalies in Multi-Agent Large Language Model Systems
- RepSelect: Robust LLM Unlearning via Representation Selectivity
- MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision
- Human Universal Grasping
- Context-Aware RL for Agentic and Multimodal LLMs
- BRDFusion: Physics Meets Generation for Urban Scene Inverse Rendering
- Geometric Action Model for Robot Policy Learning
- Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes
- Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
- ExpRL: Exploratory RL for LLM Mid-Training
- TokenPilot: Cache-Efficient Context Management for LLM Agents
- TuneJury: An Open Metric for Improving Music Generation Preference Alignment
- DreamX-World 1.0: A General-Purpose Interactive World Model
- Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences
- The Ghosts of Polymarket: When Off-Chain Matches Meet On-Chain Reverts
- OneRank: Unified Transformer-Native Ranking Architecture for Multi-Task Recommendation
- No Resource, No Benchmarks, No Problem? Evaluating and Improving LLMs for Code Generation in No-Resource Languages
- Understanding the Behaviors of Environment-aware Information Retrieval
- GD^2PO: Mitigating Multi-Reward Conflicts via Group-Dynamic reward-Decoupled Policy Optimization
- Text-Vision Co-Instructed Image Editing
- MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents
- Multi-Turn Reflective Masking Elicits Reasoning in Mask Diffusion Models
- MMDiff: Extending Diffusion Transformers for Multi-Modal Generation
- CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
- BadWorld: Adversarial Attacks on World Models
- How Post-Training Shapes Biological Reasoning Models
- PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory
- Taylor-Calibrate: Principled Initialization for Hybrid Linear Attention Distillation
- LectūraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
- SP^3: Spherical Priors for Plug-and-Play Restoration
- RL-Index: Reinforcement Learning for Retrieval Index Reasoning
- VisualClaw: A Real-Time, Personalized Agent for the Physical World
- Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models
- UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion Transformer
- EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video
- A Gradient Perspective on RLVR Stability and Winner Advantage Policy Optimization
- VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
- Thinking with Visual Grounding
- Implicit Reasoning for Large Language Model-based Generative Recommendation
- Tangram: Unlocking Non-Uniform KV Cache Compression for Efficient Multi-turn LLM Serving
- SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
- Who Flips? Self- and Cross-Model Counterarguments Reveal Answer Instability in LLMs
- You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences
- The Data Manifold under the Microscope
- SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage Reconstruction
- Distilling Examples into Task Instructions: Enhanced In-Context Learning for Real-World B2B Conversations