AI 论文 · 2026-09
浏览 2026-09 发布的AI 论文内容,第 6 页,共 741 条。
- StepAudio 3 Gen Technical Report
- Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
- SteerDuplex: Steerable Duplex Speech Dialogue Models
- RelateAnything: Real-Time Open-Vocabulary Relation Prediction From Any Inputs
- Agent as Policy for Robotic Manipulation
- MInTRL: Off-policy Intervention can boost On-policy RL
- Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning
- Competence-Gated Pooling of Language Models and Priors for Event Forecasting
- Feature Recovery for Object Understanding After Irreversible Fire Damage
- Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence
- SenseNova-U1.5: Towards Native Unified Visual Intelligence
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement
- Generative Late-Interaction Embeddings For Visual Document Retrieval
- Negative Self-Distillation: Learning to Reason by Avoiding Flaws
- COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization
- Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
- Memory as Plans: World-Action Modeling with Memory-Grounded Planning
- World in World: Explore the World with World Models
- Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs
- FreeFlow: A Bias-free Hierarchical Transformer for Optical Flow Estimation
- X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
- Mi-Ripple: Restoring Images Degraded by Iterative AI Editing
- UniH^3: Unifying Hierarchical Homogeneity and Heterogeneity for All-in-One Medical Image Restoration
- DRG-MAPPO: Hierarchical Dynamic Role-Graph Multi-Agent Reinforcement Learning for Cooperative Air Combat
- Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation
- Beyond Solver Verdicts: Generative Reward Models for Autoformalization
- The information geometry of large language models is shared, learned, and controllable
- T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks
- Ambient @ EgoLongQA 2026: Distilling Long-Video perception into a Sub-2B Model
- Ambient @ EgoProactive 2026 : Proactive Egocentric Assistance with Visually Grounded Supervision
- LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents
- ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs
- Studying Without a Syllabus: Task-Agnostic Environment Preprocessing
- Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking
- Towards a Deterministic Math Solver for Clinical Language Models
- NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
- An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics
- Programmable World Model
- IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
- Show-Harness: Just a VLM Agent Can Play Robots
- Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
- Why Is Video Still So Expensive? A Survey of Inference-Efficiency Mechanisms in Video and Audiovisual LLMs
- TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI Agents
- The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
- Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them?
- Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States
- MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes
- RESCUE-BENCH: Towards Relation-Aware Multi-Party Emotional Support Conversation Systems
- Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction
- StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
- TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model
- SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
- Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
- Studying Image Tokenizers as Visual Languages in Unified Multimodal Models
- NOAH: Learning the Full Patient Journey. A Longitudinal Multimodal Time-Aware Model for Representation and Forecasting
- Co-Evolving Harnesses and Models: On-Policy Correction Helps Weaker Models Catch Up Where Imitation Fails
- Mask Forcing: Improving Autoregressive Video Diffusion Distillation via Dual-Noise Masking Rollout
- SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
- ActReview: Rebuttal-Guided Training Data and Rubric Rewards for Actionable Peer Review Generation
- Omni Interaction Agent Technical Report
- PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving
- AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing
- Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
- Difficulty-Adaptive Tree-Structured Policy Optimization for Expanding Reasoning Coverage in RLVR
- AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems
- Feyospace-v1: How the Cyber Mercury Seven Trained Frontier Cyber Models
- Environments as Scaffold: Enriching Feedback to Bootstrap Self-Evolving Agents in Long-Horizon Tasks
- Miles v0.1: Production-Level Post-Training
- ReMoMask-2: Latent Retrieval-Augmented Masked Motion Generation
- CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
- ActionSplice: In-Flight Action Editing for Interactive World Models
- NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
- SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation
- Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation
- The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements
- A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM
- LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay
- Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents
- ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding
- Kalman Delta Networks: Uncertainty-aware Associative Memory
- Harnessing CLIP and DINO: An Uncertainty-Aware Cascaded Fusion Network for Generalizable Deepfake Image Detection
- CosmoH2G: A Hand-to-Gripper Transfer Dataset and Baseline Method for Object Manipulation with Complex Spatial Movements
- Measuring Language Transfer in Robot Policies: Adding Greek to a Cosmos3 Vision-Language-Action Policy
- RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting
- OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
- DF26: We Cannot Tell Fake From Real Anymore
- Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise
- Online Draft Co-Training for Speculative Decoding in Large-Scale, Long-Context RL Post-Training
- Revisiting Complete Reasoning Traces for Post-Training
- SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem
- Continual Learning Mechanisms Compose for Long-Horizon Memorization
- MOLE: Detecting Insider Threats in AI Agents
- CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation
- SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
- Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations
- Train Smarter, Not Harder: Switching Signal-Guided Training in Active Learning
- Agentic Visual Generation: From Generative Models to Agentic Control
- Reason Through the Latent! Making Latent Visual Reasoning Necessary