AI 论文 · 2026-06
浏览 2026-06 发布的AI 论文内容,第 3 页,共 818 条。
- Vera: A Layered Diffusion Model for Content-Preserving Video Editing
- Causal Discovery in the Era of Agents
- Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views
- VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct
- Self-Compacting Language Model Agents
- Arbor: Explicit Geometric Conditioning for Controllable 3D Asset Generation
- UniverSat: Resolution- and Modality-Agnostic Transformers for Earth Observation
- TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization
- MeshFlow: Mesh Generation with Equivariant Flow Matching
- AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
- ReasoningLens: Hierarchical Visualization and Diagnostic Auditing for Large Reasoning Models
- Tmax: A simple recipe for terminal agents
- Safe Few-Step Generation via Velocity Editing
- Capable but Careless: Do Computer-Use Agents Follow Contextual Integrity?
- ShotcreteDepth: A Bi-modal Dataset for Robust Robotic Depth Perception in Shotcrete Construction Environments
- Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation
- ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation
- Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents
- Unlimited OCR Works
- Training Open Models for Agentic Phone Use
- Plans Don't Persist: Why Context Management Is Load Bearing for LLM Agents
- When Agents Commit Too Soon: Diagnosing Premature Commitment in LLM Agents
- CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents
- FedOT: Ownership Verification and Leakage Tracing via Watermarks for Federated LDMs
- SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
- KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking
- HAKARI-Bench: A Lightweight Benchmark for Comparing Retrieval Architectures and Efficiency Settings under Unified Conditions
- RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
- EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos
- Libretto: Giving LLM Agents a Sense of Musical Structure
- Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do
- PolicyTrim: Boosting Intrinsic Policy Efficiency of Vision-Language-Action Models
- Interleaved Speech Language Models Latently Work In Text
- PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
- Lexical Consensus: Grounded Word Learning and Shared Meaning in Artificial Agents
- Multi4D: High-Fidelity Dynamic Gaussian Splatting via Multi-Level Competitive Allocation
- BioMatrix: Towards a Comprehensive Biological Foundation Model Spanning the Modality Matrix of Sequences, Structures, and Language
- OpenBioRQ: Unsolved Biomedical Research Questions for Agents
- Deeper is Not Always Better: Mitigating the Alignment Tax via Confident Layer Decoding
- A Verifiable Search Is Not a Learnable Chain-of-Thought
- EBench: Elemental Diagnosis of Generalist Mobile Manipulation Policies
- Discretizing Reward Models
- PrivacyAlign: Contextual Privacy Alignment for LLM Agents
- Improving Text-to-Music Generation with Human Preference Rewards
- UnityShots: Memory-Driven Multi-Shot Audio-Video Generation with Boundary-Aware Gating
- Toward Open Weight Models Without Risks: Separating Public and Private Capabilities in LLMs
- Counsel: A Meta-Evaluation Dataset for Agentic Tasks
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach
- An Exploratory Case Study of LLM-Assisted Refactoring and Gameplay Feature Generation in an Endless Runner Game
- PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning
- BioInsight: Multi-Agent Orchestration for Interactive Biomedical Knowledge Discovery
- Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
- Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
- CogniRoute: Learning to Route Social Evidence in Omni-Modal Models
- Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention
- Vesta: A Generalist Embodied Reasoning Model
- Go-with-the-Track: Video Compositing and Motion Control with Point Tracking
- Toward Parking Spot Occupancy Recognition: A Self-Supervised Approach
- World Action Models: A Survey
- JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising
- Current World Models Lack a Persistent State Core
- The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation
- LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
- StylisticBias: A Few Human Visual Cues Drive Most Social Biases in MLLMs
- HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
- Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
- S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence
- FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining
- FlowBender: Feedback-Aware Training for Self-Correcting Conditional Flows
- Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe
- Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation
- HydraHead: From Head-Level Functional Heterogeneity to Specialized Attention Hybridization
- EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies
- Holo-World: Unified Camera, Object and Weather Control for Video World Model
- QG-MIL: A Gated Transformer Aggregator for Domain-Agnostic Multiple Instance Learning in Medical Imaging
- When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
- Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
- ENPIRE: Agentic Robot Policy Self-Improvement in the Real World
- MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization
- MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management
- JAMER: Project-Level Code Framework Dataset and Benchmark on Professional Game Engines
- When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning
- Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning
- Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models
- Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
- Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
- DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
- BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation
- FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines
- Comparing Linear Probes with Mahalanobis Cosine Similarity
- Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why
- PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
- ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
- LooseControlVideo: Directorial Video Control using Spatial Blocking
- Characterizing Narrative Content in Web-scale LLM Pretraining Data
- Playful Agentic Robot Learning
- OpenRath: Session-Centered Runtime State for Agent Systems
- Native Active Perception as Reasoning for Omni-Modal Understanding
- Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games