AI 论文 · 2026-09
浏览 2026-09 发布的AI 论文内容,第 2 页,共 741 条。
- Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling
- MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization
- ALICE: In-context, Zero-shot, Mutual Information Estimation
- LVMT: Video Mask Transformer for Long-term Video Segmentation
- RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers
- Prefill-Free Cross-Family KV Cache Transfer for Heterogeneous Multi-Agent LLMs
- Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning
- PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
- Evaluating Bounded Autonomy in Regulated Agentic AI: A Diagnostic Harness with Constitutional Rewards, Escalation Labels, and Runtime Governance
- Persona Dosing: Calibrated Activation Steering for Graded Trait Control
- StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
- Periodic Weak Spots: Phase Sensitivity from Chunked KV-Cache Compression
- Fractional State Space Transition for Long Sequence Modeling
- CheatBench: Measuring Reward Gaming in AI Agents
- Language Models Are "Insecure" Reporters
- When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs
- Data Unlearning via Inverse Distillation
- LongCat-DeepResearch Technical Report
- Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion
- Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection
- Prompted Identity Degrades Cooperation in Multi-Agent LLM Systems
- Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning
- How to Loop MoE: Flatten the Experts, Untie the Attention
- InfiniHand: Streaming World-Space Hand Motion Estimation from Egocentric Video
- GeoVerse: World-Consistent Novel View Synthesis in Geometric Latent Space
- Hard Vision, Easy Vision: What GPT-6 Astra Reveals Across Computer Vision
- Reinforcing Agentic Creativity in Scientific Ideation with Night Science
- Distillation Defenses Easily Break After Reinforcement Learning
- FlowTool: Controlling Tool Parameter in Image Retouching via Flow Matching
- MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining
- Rubric Rewards from Item Response Theory
- On-Policy Self-Distillation for Multi-Turn Image Editing
- FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
- AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation
- One Proposal for Every Margin: Zero-Shot Amortized Sequential Importance Sampling for Binary Matrices
- An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
- How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
- ReSPO: Reshaped Sequence Policy Optimization for Gradient Starvation in Off-Policy Learning
- LLMs are General Asynchronous Agents
- Multilinguality in Hybrid Attention LLMs
- Jev thinks "I don't know'', but doesn't say it: Introducing Sys1Cal-v1 Dataset for Probability Calibration
- Imprint Reader: From Weight-Update Readout to Behavioral Intervention
- On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics
- AnswerMap: Faithful Spatial Interpretability of VLMs from Answer Posteriors
- Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
- ColNanoVDR: Document-Free Query Distillation for Multi-Vector Visual Document Retrieval via Optimal Transport
- Can We Trust the Teacher? Decoupled Credit Direction-Magnitude for Self-Distillation
- WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning
- When Do Model Internals Help? Exploring the Role of Representation Engineering in LLM Safety
- SeLMRoute: Probabilistic Semantic Evidence for Large Language Model Routing
- Geometry as Address: Routing Attention to Visual Memory for Long-Horizon Camera-Controlled Video Generation
- PDE-JEPA: Predictive Representation Learning of Latent Dynamics Modeling for Parametric PDEs
- Nereus: Adaptive Parallelism for LLM Post-Training
- Does Native 3D Texture Generation Necessarily Require 3D Assets for Training?
- PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
- The Low-Rank Structure of VLA Reinforcement Learning
- Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence
- Preference-Guided Adaptation for Open-Vocabulary Semantic Segmentation via Prompt Disagreement
- SEAD: A State-Based Perspective on Attack and Defense in Tool-Using Agents
- SentZero: An Enhanced Sentence-Centric Vision-Language Pretraining for Multi-Task Zero-Shot Chest X-Ray Analysis
- Org-Agent: Beyond Personal Assistants Towards Organizational Agents
- Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
- NavHarness: Towards Lifelong Embodied Navigation
- BIABench: Evaluating AI agents on real-world bioimage analysis tasks
- When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model
- SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs
- SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving
- Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models
- SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models
- QuantWM: Temporally Consistent 2-Bit KV Cache Quantization for Video World Models
- The Endless Exam: Mathematical Constructions from Today's Models toward Superintelligence
- Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces
- Adaptive Fused Prior Transfer for Controllable Generative Image Compression
- QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents
- MinkowskiPE: Minkowski Positional Encoding for Spatiotemporal Perception
- Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?
- AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models
- Learning Multimodal Embeddings with Evidence-Aligned Readout
- Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
- StoryEngine: A State-Grounded Agentic Framework for Video Storytelling
- Learning Dynamics of Continual Learning: A Unified View of Data Attribution, Forgetting, and Plasticity Loss
- Pretraining Transformers with Quantized Softmax in Attention
- What masking geometry works best for EEG foundation models?
- DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation
- Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
- SMAT: Simple and Efficient Merge-Aware Training
- TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
- DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration
- WideSWE: Can Coding Agents Coordinate Changes Across Repositories?
- Recursive Harness Distillation across Agents for Robot Manipulation
- Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models
- VisionHOPE: Visual Backbones as Self-Modifying Learning Systems
- When Privacy Moves ML-Mediated Decisions On Device: Information and Incentive Misalignment in Auctions
- Structured Residual Connectivity Matters for Diffusion Transformers
- KernelZero: Co-Evolving Proposer and Coder for Continuously Improved GPU Kernel Generation
- WhiteMatter: All-to-All Cross-Layer Connections via KV Source Mixing
- GeoSET: Generalist Foundation Model for SAR-to-EO Image Translation
- X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization
- Relic: From Multi-Agent Collaboration to Persistent Organizational Capability
- Adaptive Latent Capacity for World Models