AI 论文 · 2026-06
浏览 2026-06 发布的AI 论文内容,第 7 页,共 818 条。
- Chiaroscuro Attention: Spending Compute in the Dark
- Light-WAM: Efficient World Action Models with State-Fusion Action Decoding
- Revisiting Articulated Parts Perception in Robot Manipulation
- Robust-U1: Can MLLMs Self-Recover Corrupted Visual Content for Robust Understanding?
- When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
- MuJoCo-Drones-Gym: A GPU-Accelerated Multi-Drone Simulator for Control and Reinforcement Learning
- DyCo-RL: Dynamic Cross-Modal Coordination for Visual Reasoning
- POISE: Position-Aware Undetectable Skill Injection on LLM Agents
- Skill-3D: Evolving Scene-Aware Skills for Agentic 3D Spatial Reasoning
- Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path
- SigmaScale: LLM Compression with SVD-based Low-Rank Decomposition and Learned Scaling Matrices
- Breaking the Bubble: Asynchronous Pipeline Parallel Training with Bounded Weight Inconsistency
- The Cold-Start Safety Gap in LLM Agents
- VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation
- Struct-Searcher: Agentic Structural Thinking Advances Multimodal Deep Information Seeking
- UniSHARP: Universal Sharp Monocular View Synthesis
- MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism
- Streaming Video Generation with Streaming Force Control
- Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
- Whisper Hallucination Detection and Mitigation via Hidden Representation Steering and Sparse AutoEncoders
- PaperFlow: Profiling, Recommending, and Adapting Across Daily Paper Streams
- Watch, Remember, Reason: Human-View Video Understanding with MLLMs
- Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
- Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
- How Far Can Chord-Symbol Time-Series Adaptation Carry Genre Identity? Capabilities and Boundaries in Multi-Genre Chord-Symbol Modeling
- AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
- DuMate-DeepResearch: An Auditable Multi-Agent System with Recursive Search and Rubric-Grounded Reasoning
- MMAE: A Massive Multitask Audio Editing Benchmark
- Robotic Policy Adaptation via Weight-Space Meta-Learning
- Entropy as a Structural Prior: How a Log-Barrier on DiT Belief Space Drives Musical Diversity and Development
- dots.tts Technical Report
- SlimSearcher: Training Efficiency-Aware Web Agents via Adaptive Reward Gating
- LIMMT: Less is More for Motion Tracking
- Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors
- Towards Retrieving Interaction Spaces for Agentic Search
- Empirical Study on the Characteristics and Evolution of AI-usage in GitHub Repositories: Evidence from Code Comments
- ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction
- StepPO: Step-Aligned Policy Optimization for Agentic Reinforcement Learning
- ECI_{sem}: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives
- In-Context Multiple Instance Learning
- IR3DE: A Linear Router for Large Language Models
- Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems
- Retrospective Harness Optimization: Improving LLM Agents via Self-Preference over Trajectory Rollouts
- Answer Presence Drives RAG Rewriting Gains
- ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs
- OpenSkill: Open-World Self-Evolution for LLM Agents
- A Geometric Account of Activation Steering through Angle-Norm Decomposition
- Data-Efficient Autoregressive-to-Diffusion Language Models via On-Policy Distillation
- UnpredictaBench: A Benchmark for Evaluating Distributional Randomness in LLMs
- Re-Centering Humans in LLM Personalization
- Direct 3D-Aware Object Insertion via Decomposed Visual Proxies
- Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
- Robots Need More than VLA and World Models
- WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark
- Code2LoRA: Hypernetwork-Generated Adapters for Code Language Models under Software Evolution
- Regret Minimization with Adaptive Opponents in Repeated Games
- Complexity-Balanced Diffusion Splitting
- Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators
- MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery
- Benchmark Everything Everywhere All at Once
- Latent Reasoning with Normalizing Flows
- Revising Context, Shifting Simulated Stance: Auditing LLM-Based Stance Simulation in Online Discussions
- Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation
- Unsupervised Skill Discovery for Agentic Data Analysis
- Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them
- RhymeFlow: Training-Free Acceleration for Video Generation with Asynchronous Denoising Flow Scheduling
- Towards One-to-Many Temporal Grounding
- LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs
- ActiveMimic: Egocentric Video Pretraining with Active Perception
- AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
- Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback
- LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
- LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing
- Memory is Reconstructed, Not Retrieved: Graph Memory for LLM Agents
- OPRD: On-Policy Representation Distillation
- Compress-Distill: Reasoning Trace Compression for Efficient Knowledge Distillation
- World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis
- LLM Explainability with Counterfactual Chains and Causal Graphs
- Towards Truly Multilingual ASR: Generalizing Code-Switching ASR to Unseen Language Pairs
- Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models
- When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents
- Imagine Before You Predict: Interleaved Latent Visual Reasoning for Video Event Prediction
- SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents
- DRIFT: A Residual Flow Adapter for Decoding Continuous Outputs in Vision-Language Models
- Cosine Misleads: Auxiliary Losses Reshape Vision Language Models, Not Their Latents
- Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning
- AdaPlanBench: Evaluating Adaptive Planning in Large Language Model Agents under World and User Constraints
- AsyncWebRL: Efficient Multi-Step RL for Visual Web Agents
- SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations
- AURA: Intent-Directed Probing for Implicit-Need Surfacing in Situated LLM Agents
- ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?
- Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models
- Dream.exe: Can Video Generation Models Dream Executable Robot Manipulation?
- ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment
- BRepCLIP: Contrastive Multimodal Pretraining on BRep Primitives for CAD Understanding
- Agents' Last Exam
- Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference
- What Should Agents Say? Action-state Communication for Efficient Multi-Agent Systems
- Personal AI Agent for Camera Roll VQA
- VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding