AI 论文 · 2026-06
浏览 2026-06 发布的AI 论文内容,第 8 页,共 818 条。
- Flash-WAM: Modality-Aware Distillation for World Action Models
- STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations
- GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
- Streaming Communication in Multi-Agent Reasoning
- Reinforcement Learning from Rich Feedback with Distributional DAgger
- Deep Embedded Multiplicative DMD for Algebra-Preserving Koopman Learning
- Self-Evaluation Is Already There: Eliciting Latent Judge Calibration in Base LLMs with Minimal Data
- Audio Interaction Model
- Evaluating Large Language Models in Dynamic Clinical Decision-Making with Standardized Patient Cases
- ZipSplat: Fewer Gaussians, Better Splats
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- CIPER: A Unified Framework for Cross-view Image-retrieval and Pose-estimation
- DAR: Deontic Reasoning with Agentic Harnesses
- M^3Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks
- Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game
- Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
- Measuring Model Robustness via Fisher Information: Spectral Bounds, Theoretical Guarantees, and Practical Algorithms
- TIDE: Proactive Multi-Problem Discovery via Template-Guided Iteration
- Rethinking Continual Experience Internalization for Self-Evolving LLM Agents
- MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation
- Why Muon Outperforms Adam: A Curvature Perspective
- Echo-Infinity: Learning Evolving Memory for Real-Time Infinite Video Generation
- GENEB: Why Genomic Models Are Hard to Compare
- MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation
- SparDA: Sparse Decoupled Attention for Efficient Long-Context LLM Inference
- SePO: Self-Evolving Prompt Agent for System Prompt Optimization
- Online Skill Learning for Web Agents via State-Grounded Dynamic Retrieval
- Video2LoRA: Parametric Video Internalization for Vision-Language Models
- Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
- Value-Aware Stochastic KV Cache Eviction for Reasoning Models
- KletterMix: Climbing Toward High-Quality German Pretraining Data
- World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning
- ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning
- WebRISE: Requirement-Induced State Evaluation for MLLM-Generated Web Artifacts
- Small RL Controller, Large Language Model: RL-Guided Adaptive Sampling for Test-Time Scaling
- Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates
- SkillHarness: Harnessing Safe Skills for Computer-Use Agents
- GridVQA-X: A Framework for Evaluating Multimodal Explainability Methods
- Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion
- Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory
- MAOAM: Unified Object and Material Selection with Vision-Language Models
- A Cookbook of 3D Vision: Data, Learning Paradigms, and Application
- Can Generalist Agents Automate Data Curation?
- Large Language Models Hack Rewards, and Society
- Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents, with an Affine-Typed Rust Mitigation as a Case Study
- Unlocking Feature Learning in Gated Delta Networks at Scale
- Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking
- Skill-RM: Unifying Heterogeneous Evaluation Criteria via Agent Skill
- Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
- AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video Generation
- Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning
- SEAOTTER: Sensor Embedded Autoencoding with One-Time Transcode for Efficient Reconstruction
- Benchmarking Visual State Tracking in Multimodal Video Understanding
- Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching
- Agent libOS: A Library-OS-Inspired Runtime for Long-Running, Capability-Controlled LLM Agents
- OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs
- EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management
- Training-Free Multi-Concept LoRA Composition with Prompt-Aware Weighting
- Reasoning over Grammar: Can Synthetic Linguistic Reasoning Traces Enhance Low-Resource Machine Translation?
- Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models
- Qwen-Image-Flash: Beyond Objective Design
- Text-to-Image Models Need Less from Text Encoders Than You Think
- When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models
- KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks
- Large Language Models Are Overconfident in Their Own Responses
- RobotValues: Evaluating Household Robots When Human Values Conflict
- BA-T: An Iterative Transformer for Two-View Bundle Adjustment
- PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training
- MemTrain: Self-Supervised Context Memory Training
- NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation
- EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
- The Shadow Price of Reasoning: Economic Perspective on Optimal Budget Allocation for LLMs
- AUDITFLOW: Executable Symbolic Environments for Structured Financial Reporting Verification
- Neural Networks Provably Learn Spectral Representations for Group Composition
- BraveGuard: From Open-World Threats to Safer Computer-Use Agents
- Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
- Prompt-Level Distillation: A Non-Parametric Alternative to Model Fine-Tuning for Efficient Reasoning
- AURA: Action-Gated Memory for Robot Policies at Constant VRAM
- Ψ-Bench: Evaluating Persona-Sensitive Influencing in Persuasive Dialogues
- MMG2Skill: Can Agents Distill In-the-Wild Guides into Self-Evolving Skills?
- HarnessForge: Joint Harness and Policy Evolution for Adaptive Agent Systems
- Adaptive Auto-Harness: Sustained Self-Improvement for Agentic System Deployment on Open-Ended Task Streams
- TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL
- The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset
- Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions
- Cosmos 3: Omnimodal World Models for Physical AI
- Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation
- Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models
- Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling
- AdaCodec: A Predictive Visual Code for Video MLLMs
- LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation
- AFUN: Towards an Affordance Foundation Model for Functionality Understanding
- SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction
- MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation
- A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL
- Policy and World Modeling Co-Training for Language Agents
- Harness-1: Reinforcement Learning for Search Agents with State-Externalizing Harnesses
- TVIR: Building Deep Research Agents Towards Text--Visual Interleaved Report Generation
- Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
- Geometric Latent Reasoning Induces Shorter Generations in LLMs