AI 论文 · 2026-05
浏览 2026-05 发布的AI 论文内容,第 1 页,共 552 条。
- ClawHub Security Signals: When VirusTotal, Static Analysis, and SkillSpector Disagree
- Decoupled Residual Denoising Diffusion Models for Unified and Data Efficient Image-to-Image Translation
- CARVE: Certified Affordable Repair of Vetoed Maneuvers via Envelopes for Interactive Driving
- OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
- An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
- Agent Skills Should Go Beyond Text: The Case for Visual Skills
- A Formally Verified Library of Mathematical Finance in Lean 4
- ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- Trust Region On-Policy Distillation
- Measuring the Symmetry--Data Exchange Rate
- 3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code
- τ_0-WM: A Unified Video-Action World Model for Robotic Manipulation
- FVSpec: Real-World Property-Based Tests as Lean Challenges
- Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
- Honest Lying: Understanding Memory Confabulation in Reflexive Agents
- SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory
- Confidence-Adaptive SwiGLU for Mixture-of-Experts
- OCC-RAG: Optimal Cognitive Core for Faithful Question Answering
- FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search
- Critic-R: Improving Agentic Search using Instruction-tuned Retrievers with Natural Language Introspective Feedback
- Semi-Supervised Noise Adaptation: Transferring Knowledge from Noise Domain
- Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs
- On the Limits of LLM Adaptability: Impact of Model-Internalized Priors on Annotation Task Performance
- SDR: Set-Distance Rewards for Radiology Report Generation
- AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
- The Shape of Addition: Geometric Structures of Arithmetic in Large Language Models
- αDepth: Learning Single-Pass Soft Boundary Decomposition for Stereo Conversion
- Score-Control for Hallucination Reduction in Diffusion Models
- Model-Based Quality Assessment for Massively Multilingual Parallel Data
- StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement
- MindZero: Learning Online Mental Reasoning With Zero Annotations
- PaintBench: Deterministic Evaluation of Precise Visual Editing
- SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
- SurGe: Improved Surface Geometry in Point Maps
- Functional Attention: From Pairwise Affinities to Functional Correspondences
- RayDer: Scalable Self-Supervised Novel View Synthesis from Real-World Video
- SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
- How can embedding models bind concepts?
- DRIFT: Decoupled Rollouts and Importance-Weighted Fine-Tuning for Efficient Multi-Turn Optimization
- SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
- DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
- Mellum2 Technical Report
- COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
- Emergent Languages in Populations of Language Model Agents: From Token Efficiency to Oversight Evasion
- Trust-Region Behavior Blending for On-Policy Distillation
- SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes
- iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
- Task-Focused Memorization for Multimodal Agents
- Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
- LVSA: Training-Free Sparse Attention for Long Video Diffusion
- GGT-100K: Generative Ground Truth for Generalizable Real-World Image Restoration
- PEEK: Picking Essential frames via Efficient Knowledge distillation
- SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue
- Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer
- MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft
- The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
- dMoE: dLLMs with Learnable Block Experts
- Distilling LLM Feedback for Lean Theorem Proving
- Count Anything
- Send a SCOUT First: Pre-hoc Reasoning for Adaptive Detector Allocation in Prompt-Injection Defense
- MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
- OpenSTBench: Beyond Semantic Evaluation for Speech Translation
- MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
- Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning
- Guidance Contrastive Token Credit Assignment for Discrete Policy Optimization
- The Distillation Game: Adaptive Attacks & Efficient Defenses
- A Multi-AI-agent Framework Enabling End-to-end Finite Element Analysis for Solid Mechanics Problems
- Multimodal Music Recommendation System using LLMs
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Prior Availability in Industrial Visual Sim-to-Real: A Review of CAD-Guided and CAD-Unavailable Regimes
- Memory-Bound but Not Bandwidth-Limited: The Physical AI Inference Gap in Batch-1 LLM Decode
- Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)?
- MAAT: Multi-phase Adapter-Aware Targeted Unlearning
- Linear Ensembles Wash Away Watermarks: On the Fragility of Distributional Perturbations in LLMs
- SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer
- DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation
- AdaState: Self-Evolving Anchors for Streaming Video Generation
- NeuROK: Generative 4D Neural Object Kinematics
- YoCausal: How Far is Video Generation from World Model? A Causality Perspective
- Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
- Colored Noise Diffusion Sampling
- SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?
- PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
- LoMo: Local Modality Substitution for Deeper Vision-Language Fusion
- minWM: A Full-Stack Open-Source Framework for Real-Time Interactive Video World Models
- Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning
- GenClaw: Code-Driven Agentic Image Generation
- Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
- Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection
- Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
- Meta-Cognitive Memory Policy Optimization for Long-Horizon LLM Agents
- PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
- When Cloud Agents Meet Device Agents: Lessons from Hybrid Multi-Agent Systems
- Geometry Matters: 3D Foundation Priors for Learning Semantic Correspondence
- UniSteer: Text-Guided Flow Matching in Activation Space for Versatile LLM Steering
- Towards Consistent Video Geometry Estimation
- REPOT: Recoverable Program-of-Thought via Checkpoint Repair
- VisualThink-VLA: Visual Intermediate Reasoning for Effective and Low-Latency Vision-Language-Action Policies
- EarlyTom: Early Token Compression Completes Fast Video Understanding