AI 论文 · 2026-10
浏览 2026-10 发布的AI 论文内容,第 2 页,共 207 条。
- VepAgent: Bridging Causal-Transition via Tool-Augmented Reinforcement Learning for Video Event Prediction
- Conditional Trajectory Peaks: Single-Pass Multimodal Policies over Action Chunks
- Execution-Aligned Progressive Noise for Consistent Asynchronous Replanning in Generative Robot Policies
- JLD: Perceptual Distance Through A Jacobian Lens
- MiniCorp: The Last Mile of the AI Agent Firm
- Collaborative Personalized Preference Alignment for LLMs under Data Deficiency
- Learning to Learn a Language
- HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing
- FairRSFM: A Biome-Aware Benchmark and Debiasing Framework for Remote Sensing Foundation Models
- Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks
- HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models
- DeCoPrune: Efficient KV-Cache Pruning for Autoregressive Video Diffusion via Denoising Consistency
- TIDES: Implicit Time-Awareness in Selective State Space Models
- UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents
- LiFT: Loop Flow Transformers
- SheetSage2: Coherent Lead-Sheet Transcription with Synthetic Supervision
- CoDance: Learning Reactive and Compliant Human-Humanoid Interaction from Video
- SearchJev: A Fast and Calibrated System-1 Model for Search Agents
- Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation
- How corner is a corner case? Percentile control for highway scenario generation
- Do LLMs Understand Sequential Structure? A Controlled Study of Inference and Generation
- RobotUse: Allocating Computation, Context, and Decisions
- LLM-as-Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
- Agentic discovery of blood biomarker from distilled private health records
- NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis
- Learning Discriminative Geometry for Drifting Models
- The Numerical Linear Algebra of Large Language Models
- PerturBot: Breaking Shortcut Priors in Vision-Language-Action Models with Perturbative Training
- ConEx: Human-Interpretable Saliency Maps via Concept-Aware Attribution
- DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation
- LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures
- CurveCodec 2: Skeleton-agnostic animation compression with a learned entropy model
- Magic-W0: A Structured World-Action Foundation Model for Physical Intelligence
- Toward Real-Time VLAs: Stage-Aware Two-Step Flow Denoising and System-Level Evaluation
- What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document
- Periscope: Extending Frozen Language Models Beyond Their Context Window
- Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models
- Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale
- Learning Latent Protein Languages for Autoregressive Generation
- 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
- EyeRobot 2.0: Active Gaze for Precise Manipulation without Wrist Cameras
- Language Models that Play Chess and Explain Their Moves
- FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution
- Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
- ProAR: Learning Prospective Reasoning with Autoregressive Video Models
- DEPICT: Scoring Text-to-Image Alignment by Answer Agreement
- HyperBrowseComp: A Multilingual and Multimodal Stress Test for Web-Browsing Agents
- Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability
- Native Action-Prior Learning from Videos for World Action Models
- Multilingual GSM-Symbolic: What determines capability transfer across languages?
- COSMI: COmpositional Synthesis of Multi-object Interactions
- Collective Bias Mitigation via Model Routing and Collaboration
- Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective
- Investigating the Role of Reasoning-Language Alignment in Monolingual Retrieval-Augmented Generation
- Foresight: planning future perception in streaming VLMs without retraining
- In-Distribution Forcing for Long Video Generation at Test Time
- Post-Training Frontier Text-to-Image Models by Composing Preference and Rubric Rewards
- TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows
- Harness-Aware Distillation for Small Language Model Agents
- FastOPD: On-Policy Distillation for Lightweight VLA Deployment
- OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation
- Improving Atomic-Fact Recall via Focused Views in Unstructured Knowledge Editing
- Self-Supervised Scaling of Terminal Environments for Scientific Domains
- Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
- Intent Interpretation at RIC Timescales: Jev Decision Models versus Large Language Models in 6G Open RAN
- StepCAD: Mesh-to-CAD Code Generation via LLM Policy and Geometry-Guided Search
- Labels Override Definitions in Jev-Style Typed Decision Models
- Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory
- World Action Modeling with Progressive Visual Planning
- From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders
- MEA: A Reward-Driven Multi-Agent System for Faithful Model Explanations
- OpenRUA: Robot-Use Agents Are Zero-Shot Visuomotor Policies
- Inherit-MAS: Test-Time Evolution of Multi-Agent Systems through Workflow and Execution Inheritance
- Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
- Rethinking World-Action Model for Compositional and In-Context Robotic Manipulation
- World Editing: Intervening on Executable Worlds at Increasing Depth
- DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents
- SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation
- KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards
- ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
- SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation
- InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
- DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
- Decoding Looped Transformers Better for (Almost) Free
- From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
- Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
- Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models
- Local Support Learning
- Argo-Bench: Evaluating Data Agents on Enterprise-Scale Workflows
- Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
- UniWAM: Unified World-Action Model
- Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models
- Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
- iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
- RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations
- OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
- Before It Fades: Reinforcing Temporal Representations at Inference Time in VideoLLMs
- VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation
- Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
- Science Utopia? Closed-Loop LLM Simulation of Academic Research Ecosystems