AI 论文 · 2026-09
浏览 2026-09 发布的AI 论文内容,第 1 页,共 741 条。
- LOCI: Spatial Linear Memory for Streaming World Models
- Video Generation Models: A Survey of Post-Training and Alignment
- SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation
- Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
- PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
- Memorizon: Training World Models Beyond Their Context Window
- Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
- JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces
- Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution
- Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces
- Removing Timing Shortcuts Improves Non-Invasive Brain-to-Text
- I Have a Stream: Making Self-Supervised Learning Work on Continuous Video
- Scaling Laws for Looped Mixture of Experts
- How Much Is an AI Token Worth? Scaling Laws for Wild AI-Generated Web Text
- MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories
- Dream4ACT: A Shared Visual Action Interface for Multi-Embodiment Video-Action Modeling
- Learning Functional Subspaces for Neural Network Compression
- Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
- LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception
- OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
- DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes
- Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
- Learning from Runtime Feedback through Failure-Bank Self-Evolution for Vision-Language-Action Models
- Safety of Latent Communication in Multi-Agent Systems
- Better Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
- The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends
- OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning
- Can Computation from Earlier Problems Help LLMs Solve New Ones?
- EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos
- Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings
- Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds
- DAGent: Evaluate-then-Grow Planning for Deep Research Agents
- False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
- LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
- SparseEngine: Sparse-First Inference Engine
- Smaller Models, Better Rejects: Preference Distillation Scaling
- Benchmarking and Enhancing Skill-Level Memory for Partially Observable Robotic Manipulation
- Does Learning Protein Folding Generalize to Broader Reasoning?
- Mitigating the Length-Scaling Tax with Online Distillation
- Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
- Soft Spatial Reasoning
- SpatialCORE: Confidence-Aware Grounded Spatial Reasoning in Large Vision--Language Models
- ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images
- Honeycomb: Constant-Size Scene Memory Representation for Video World Models
- ATLAS: Aligned Transport of Latent Structure for Reliable World Model Planning
- Architect-Ant: Editable Automatic Furnishing of Architectural Floor Plans
- Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
- Tacit-TTS: From Autoregressive Decoding to Masked Prediction for Efficient Transcript-Free Voice Cloning
- TERRA: Terrain-Aware Reconstruction, Retargeting and Control for Musculoskeletal Locomotion
- PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation
- LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation
- Behavioral Persistence and Incomplete Functional Transfer of Co-evolved Communication in Evolutionary Robotics
- PAMI: Part Anchored Motion for Text to Human-Object Interaction Generation
- LoopVL: Recurrent Visual Intelligence
- MILO: Automated Harness Discovery via Orchestrated Multi-Agent Evolution
- EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
- Strike a Chord! Modal Kinetic Typography
- STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization
- EmoRES-TTS: Residual-Enhanced Vector Steering for Emotional Speech Generation
- LIFT: Layout-In-Future Video Generation under Large Viewpoint Change via On-Policy Self-Distillation
- Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE
- HelixWorld: A Real-time Interactive Audio-Visual World Model
- Explore Broadly, Reason Sharply: Push Small Models toward the Frontier via Sampling
- Tail-Influence Sampling for CVaR Policy Evaluation
- MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
- SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video
- Rollout-Marginal Distillation for Long-Horizon Autoregressive Video Generation
- EpiCon: Collective Agent Learning through Co-Evolving Multimodal Memory
- Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
- Retrieval Capacity of Self-Attention Under Competition
- It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
- The Geometry of Inference in Transformer Residual Streams
- Context Language Models
- WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation
- EngiWorld: What Can Frontier Agents Deliver in Professional Engineering Environments?
- Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
- APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants
- SkillGym: Training Skill-Use Agents with Automatic Verifiable Environment Generation
- E-MoE: Enhanced Mixture-of-Experts for Non-Factorized Diffusion Language Models
- Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing
- Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing
- Foundations of Proactive Agents: Principles, Technical Layers, and Proactivity-Gym
- Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents
- Follow the Entities: A Corpus Map for Agentic Search
- Improved Distributional Diffusion Models
- Selecting The Most Informative Tokens in Natural Language Autoencoders
- Salt++: Context-Aligned Post-Training for Few-Step Streaming Multimodal Generation
- Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
- Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation
- WEFT: Scaling Tool-Use Post-Training for General-Purpose Agents
- CorrGRPO: Correlation-Normalized GRPO for Multi-Reward Learning
- Can Agents Design Libraries for Agents?
- Video2Skill: From Streaming Experience to Reusable Embodied Skills
- ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context
- On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
- Scheduling Recursive Reasoning in Looped Transformers
- FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution
- What Makes Recurrence Effective in Looped Language Models?
- OTRetarget: Joint Robot and Object Motion Retargeting via Optimal Transport
- Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It