AI 论文 · 2026-05
浏览 2026-05 发布的AI 论文内容,第 6 页,共 552 条。
- ViMU: Benchmarking Video Metaphorical Understanding
- Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards
- LiSA: Lifelong Safety Adaptation via Conservative Policy Induction
- FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale
- BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE
- InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation
- Dynamic Latent Routing
- MMSkills: Towards Multimodal Skills for General Visual Agents
- Overcoming Dynamics-Blindness: Training-Free Pace-and-Path Correction for VLA Models
- Decoupling the Benefits of Subword Tokenization for Language Model Training via Byte-level Simulation
- AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting
- When Vision Speaks for Sound
- CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves
- HodgeCover: Higher-Order Topological Coverage Drives Compression of Sparse Mixture-of-Experts
- Learning POMDP World Models from Observations with Language-Model Priors
- KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
- PanoWorld: Towards Spatial Supersensing in 360^circ Panorama World
- LoREnc: Low-Rank Encryption for Securing Foundation Models and LoRA Adapters
- PRISM: Prior Rectification and Uncertainty-Aware Structure Modeling for Diffusion-Based Text Image Super-Resolution
- CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
- Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
- Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
- Code-Guided Reasoning for Small Language Models: Evaluating Executable MCQA Scaffolds
- Training Large Language Models to Predict Clinical Events
- PreScam: A Benchmark for Predicting Scam Progression from Early Conversations
- Hölder Policy Optimisation
- OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
- Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning
- Follow the Mean: Reference-Guided Flow Matching
- Learning to Communicate Locally for Large-Scale Multi-Agent Pathfinding
- Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection
- Unsupervised Process Reward Models
- Zero-Shot Sim-to-Real Robot Learning: A Dexterous Manipulation Study on Reactive Catching
- Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
- Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning
- ORACLE: Anticipating Scams from Partial Trajectories in Streaming App Usage
- Seeing the Needle in the Haystack: Towards Weakly-Supervised Log Instance Anomaly Localization via Counterfactual Perturbation
- RewardHarness: Self-Evolving Agentic Post-Training
- DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules
- Computer Science Conferences Should Require Nonrepudiable Experimental Results
- Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
- SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild
- Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
- Token Time Continuous Diffusion for Language Modeling
- Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs
- Long Context Pre-Training with Lighthouse Attention
- MobileEgo Anywhere: Open Infrastructure for long horizon egocentric data on commodity hardware
- Steered LLM Activations are Non-Surjective
- Liberating LLM Capabilities in Full-Duplex Speech Models
- Distilling Long-CoT Reasoning through Collaborative Step-wise Multi-Teacher Decoding
- Towards Customized Multimodal Role-Play
- WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild