AI 论文 · 2026-06
浏览 2026-06 发布的AI 论文内容,第 5 页,共 818 条。
- Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time
- Track2View: 4D-Consistent Camera-Controlled Video Generation via Paired 3D Point Tracks
- Selective Synergistic Learning for Video Object-Centric Learning
- AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining
- Reinforcement Learning-Guided Retrieval with Soft Fusion for Robust Multimodal Imitation Learning under Missing Modalities
- Rethinking the Role of Efficient Attention in Hybrid Architectures
- Beyond Monolingual Deep Research: Evaluating Agents and Retrievers with Cross-Lingual BrowseComp-Plus
- CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
- Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
- RefGC-SR^2: Reference-guided Generated Content Super-Resolution and Refinement
- MotionVLA: Vision-Language-Action Model for Humanoid Motion
- Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings
- DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects
- Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
- ReSyn: A Generalized Recursive Regular Expression Synthesis Framework
- IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products
- Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- FastMix: Fast Data Mixture Optimization via Gradient Descent
- MVEB: Massive Video Embedding Benchmark
- Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion
- PhoneHarness: Harnessing Phone-Use Agents through Mixed GUI, CLI, and Tool Actions
- Selective Control under Noisy Perception: Governance Failures Hidden by Aggregate Metrics in Modular Networks
- OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains
- RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space
- ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning
- AdaSR: Adaptive Streaming Reasoning with Hierarchical Relative Policy Optimization
- Memento: Reconstruct to Remember for Consistent Long Video Generation
- LoSoNA: A Benchmark for Local Social Norm Adaptation in Group Conversations
- VISTA: View-Consistent Self-Verified Training for GUI Grounding
- From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
- Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
- Squeeze-Release: Iterative Pruning with Exact Structural Minimization
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties
- FastContext: Training Efficient Repository Explorer for Coding Agents
- LLM Agents Can See Code Repositories
- ViT-Up: Faithful Feature Upsampling for Vision Transformers
- The Price of Anarchy in Disaggregated Inference
- Self-Evolving Visual Questioner
- HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing
- Avatar V: Scaling Video-Reference Avatar Video Generation
- Aligning Quantum Operators with Large Language Models
- A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets
- μ_0: A Scalable 3D Interaction-Trace World Model
- EvoArena: Tracking Memory Evolution for Robust LLM Agents in Dynamic Environments
- InterleaveThinker: Reinforcing Agentic Interleaved Generation
- RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
- SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
- WEAVER, Better, Faster, Longer: An Effective World Model for Robotic Manipulation
- EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery
- Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
- World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible
- Surflo: Consistent 3D Surface Flow Model with Global State
- See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents
- LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
- ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages
- MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling
- OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data
- MiniMax Sparse Attention
- MoVerse: Real-Time Video World Modeling with Panoramic Gaussian Scaffold
- VideoMDM: Towards 3D Human Motion Generation From 2D Supervision
- HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers
- Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents
- Rethinking RAG in Long Videos: What to Retrieve and How to Use It?
- EvoBrowseComp: Benchmarking Search Agents on Evolving Knowledge
- Demystifying Hidden-State Recurrence: Switchable Latent Reasoning with On-Policy Reinforcement Learning
- No Hidden Prompts Needed! You Can Game AI Peer Review with Presentation-Only Revisions
- The Hidden Power of Scaling Factor in LoRA Optimization
- HarnessBridge: Learnable Bidirectional Controller for LLM Agent Harness
- LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
- APPO: Agentic Procedural Policy Optimization
- On Subquadratic Architectures: From Applications to Principles
- Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training
- Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics
- APEX: A Network-Native Time-Series Foundation Model for Forecasting and Anomaly Detection for Wireless Edge Operations
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
- Orchestra-o1: Omnimodal Agent Orchestration
- Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
- Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
- From AGI to ASI
- Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents
- Pythagoras-Prover: Advancing Efficient Formal Proving via Augmented Lean Formalisation
- High-Fidelity Two-Step Image Generation via Teacher-Aligned End-to-End Distillation
- A Stationary (and Therefore Compatible) Representation is All You Need
- Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
- World Pilot: Steering Vision-Language-Action Models with World-Action Priors
- Redesign Mixture-of-Experts Routers with Manifold Power Iteration
- Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs
- Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization
- Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- An Efficient Method for the Optimal Control of Microgrids Under Uncertainties using Local Reduction
- Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
- From 2D Grids to 1D Tokens: Reforming Shared Representations for Multimodal Image Fusion
- Measuring Epistemic Resilience of LLMs Under Misleading Medical Context
- PianoKontext: Expressive Performance Rendering from Deadpan Context
- VIA-SD: Verification via Intra-Model Routing for Speculative Decoding
- Adaptive Multi-Resolution Procedural Knowledge Compression for Large Language Models
- InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
- Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application