AI 论文 · 2026-05
浏览 2026-05 发布的AI 论文内容,第 2 页,共 552 条。
- Discovering Cooperative Pipelines: Autoresearch for Sequential Social Dilemmas
- Adapting Multilingual Embedding Models to Turkish via Cross-Lingual Tokenizer Surgery and Offline Distillation
- Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation
- AgentDoG 1.5: A Lightweight and Scalable Alignment Framework for AI Agent Safety and Security
- SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search
- Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding
- Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering
- Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- UI-KOBE: Knowledge-Oriented Behavior Exploration for Lightweight Graph-Guided GUI Agents
- Xetrieval: Mechanistically Explaining Dense Retrieval
- Access Sets Matter: Budgeting Expert Reads for Scalable Weight-Space Model Merging
- AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
- PhoneWorld: Scaling Phone-Use Agent Environments
- Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents
- Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation
- One Click per Cell Type Suffices: Training-free Group Interaction for Cell Instance Segmentation
- The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
- GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models
- WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
- FreeForm: Reduced-Order Deformable Simulation from Particle-Based Skinning Eigenmodes
- GrepSeek: Training Search Agents for Direct Corpus Interaction
- Diagnosing Harmful Continuation in Answer-Correct Long-CoT Training Traces
- CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval
- ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood
- OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
- CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists
- OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents
- Reducing Political Manipulation with Consistency Training
- Thinking Before Constraining: A Unified Decoding Framework for Large Language Models
- Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
- Parallax: Parameterized Local Linear Attention for Language Modeling
- RUBRIC-ARROW: Alternating Pointwise Rubric Reward Modeling for LLM Post-training in Non-verifiable Domains
- The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure
- Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG
- FRAPPE: Full Input, Residual Output Autoencoding with Projection Pursuit Encoder
- The Hamilton-Jacobi Theory of Deep Learning
- Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization
- Review Arcade: On the Human Alignment and Gameability of LLM Reviews
- PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
- Self-Improving Language Models with Bidirectional Evolutionary Search
- OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration
- Learn from Weaknesses: Automated Domain Specialization for Small Computer-Use Agents
- Agent Explorative Policy Optimization for Multimodal Agentic Reasoning
- Rethinking Memory as Continuously Evolving Connectivity
- CORE: Contrastive Reflection Enables Rapid Improvements in Reasoning
- MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
- AlphaTransit: Learning to Design City-scale Transit Routes
- LiveBrowseComp: Are Search Agents Searching, or Just Verifying What They Already Know?
- DEMON: Diffusion Engine for Musical Orchestrated Noise
- Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
- Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios
- LACUNA: Safe Agents as Recursive Program Holes
- Models That Know How Evaluations Are Designed Score Safer
- A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
- GEM: Generative Supervision Helps Embodied Intelligence
- GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection
- Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
- Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning
- DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes
- HRBench: Benchmarking and Understanding Thinking-Mode Switch Strategies in Hybrid-Reasoning LLMs
- ProRL: Effective Reinforcement Learning for Proactive Recommendation via Rectified Policy Gradient Estimation
- Category-Level 3D Correspondence in Camera Space via Morphable Object Priors
- AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
- Pruning and Distilling Mixture-of-Experts into Dense Language Models
- Joint Training of Multi-Token Prediction in Reinforcement Learning via Optimal Coefficient Calibration
- When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
- Long Live The Balance: Information Bottleneck Driven Tree-based Policy Optimization
- Clark Hash: Stateless Sparse Johnson-Lindenstrauss Quantization for Neural Embeddings
- Show, Don't TELL: Explainable AI-Generated Text Detection
- Frequency-Guided Action Diffusion via Sub-Frequency Manifold Traversal
- ESC-Skills: Discovering and Self-Evolving Skills for Emotional Support Conversations
- AI Research Agents Narrow Scientific Exploration
- The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
- VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
- Revealing Algorithmic Deductive Circuits for Logical Reasoning
- PRISM: A Multi-Dimensional Benchmark for Evaluating LLM Peer Reviewers
- Got a Secret? LLM Agents Can't Keep It: Evaluating Privacy in Multi-Agent Systems
- PEAM: Parametric Embodied Agent Memory through Contrastive Internalization of Experience in Minecraft
- SkillGrad: Optimizing Agent Skills Like Gradient Descent
- Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems
- GE-Sim 2.0: A Roadmap Towards Comprehensive Closed-loop Video World Simulators for Robotic Manipulation
- AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems
- SpatialBench: Is Your Spatial Foundation Model an All-Round Player?
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- MobileMoE: Scaling On-Device Mixture of Experts
- Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
- Guiding LLM Post-training Data Engineering with Model Internals from Sparse Autoencoders
- MERIT: Learning Disentangled Music Representations for Audio Similarity
- How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
- Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini
- SIA: Self Improving AI with Harness & Weight Updates
- Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
- MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale
- Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments
- VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions
- DEI: Diversity in Evolutionary Inference for Quality-Diversity Search
- JLT: Clean-Latent Prediction in Latent Diffusion Transformers
- Trust Region Q Adjoint Matching
- QUACK: Questioning, Understanding, and Auditing Communicated Knowledge in Multimodal Social Deduction Agents