AI 论文 · 2026-05
浏览 2026-05 发布的AI 论文内容,第 4 页,共 552 条。
- Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
- HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction
- GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction
- PhotoFlow: Agentic 3D Virtual Photography Missions
- CRONOS: Benchmarking Counterfactual Physical Consistency in Video Models
- CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
- StepAudio 2.5 Technical Report
- One-Forcing: Towards Stable One-Step Autoregressive Video Generation
- Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusion
- SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models
- EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
- Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution
- Convex Low-resource Accent-Robust Language Detection in Speech Recognition
- Foundation Protocol: A Coordination Layer for Agentic Society
- FastKernels: Benchmarking GPU Kernel Generation in Production
- AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery
- Decoding the Critique Mechanism in Large Reasoning Models
- EMMA: Extracting Multiple physical parameters from Multimodal Data
- ThriftAttention: Selective Mixed Precision for Long-Context FP4 Attention
- MotiMotion: Motion-Controlled Video Generation with Visual Reasoning
- Sensor2Sensor: Cross-Embodiment Sensor Conversion for Autonomous Driving
- Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention
- DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders
- Understanding Data Temporality Impact on Large Language Models Pre-training
- Diversed Model Discovery via Structured Table Discovery
- Uniform Diffusion Models Revisited: Leave-One-Out Denoiser and Absorbing State Reformulation
- WorldKV: Efficient World Memory with World Retrieval and Compression
- Live Music Diffusion Models: Efficient Fine-Tuning and Post-Training of Interactive Diffusion Music Generators
- AnyMo: Geometry-Aware Setup-Agnostic Modeling of Human Motion in the Wild
- Forecasting Scientific Progress with Artificial Intelligence
- Swift Sampling: Selecting Temporal Surprises via Taylor Series
- SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
- Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
- More Context, Larger Models, or Moral Knowledge? A Systematic Study of Schwartz Value Detection in Political Texts
- Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
- SceneAligner: 3D-Grounded Floorplan Localization in the Wild
- VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis
- FashionLens: Toward Versatile Fashion Image Retrieval via Task-Adaptive Learning
- Segment Anything with Motion, Geometry, and Semantic Adaptation for Complex Nonlinear Visual Object Tracking
- SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation
- TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
- TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
- Bernini: Latent Semantic Planning for Video Diffusion
- Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments
- Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles
- One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems
- Efficient Agentic Reasoning Through Self-Regulated Simulative Planning
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
- From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
- LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
- RiT: Vanilla Diffusion Transformers Suffice in Representation Space
- The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
- ACC: Compiling Agent Trajectories for Long-Context Training
- Next-Acceleration-Scale Prediction for Autoregressive MRI Reconstruction
- The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm
- Representation over Routing: Overcoming Surrogate Hacking in Multi-Timescale PPO
- How Far Will They Go? Red-Teaming Online Influence with Large Language Models
- SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research
- Same Architecture, Different Capacity: Optimizer-Induced Spectral Scaling Laws
- Reflective Prompt Tuning through Language Model Function-Calling
- RankJudge: A Multi-Turn LLM-as-a-Judge Synthetic Benchmark Generator
- Broadening Access to Transportation Safety Data with Generative AI: A Schema-Grounded Framework for Spatial Natural Language Queries
- Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
- GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
- Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models
- PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
- Equilibrium Reasoners: Learning Attractors Enables Scalable Reasoning
- Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories
- DelTA: Discriminative Token Credit Assignment for Reinforcement Learning from Verifiable Rewards
- Mem-π: Adaptive Memory through Learning When and What to Generate
- iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
- SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents
- "I didn't Make the Micro Decisions": Measuring, Inducing, and Exposing Goal-Level AI Contributions in Collaboration
- OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation
- OCTOPUS: Optimized KV Cache for Transformers via Octahedral Parametrization Under optimal Squared error quantization
- RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
- UniT: Unified Geometry Learning with Group Autoregressive Transformer
- ACL-Verbatim: hallucination-free question answering for research
- Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints
- Q-ARVD: Quantizing Autoregressive Video Diffusion Models
- DrawMotion: Generating 3D Human Motions by Freehand Drawing
- FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching
- PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models
- Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment
- Rethinking Cross-Layer Information Routing in Diffusion Transformers
- IndusAgent: Reinforcing Open-Vocabulary Industrial Anomaly Detection with Agentic Tools
- On the limits and opportunities of AI reviewers: Reviewing the reviews of Nature-family papers with 45 expert scientists
- Evaluating Temporal Semantic Caching and Workflow Optimization in Agentic Plan-Execute Pipelines
- HRM-Text: Efficient Pretraining Beyond Scaling
- Generative Recursive Reasoning
- ZeroUnlearn: Few-Shot Knowledge Unlearning in Large Language Models
- AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment
- Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
- Disentangling Sampling from Training Budget in Class-Imbalanced CT Body Composition Segmentation
- ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
- Mix-Quant: Quantized Prefilling, Precise Decoding for Agentic LLMs
- MSAVBench: Towards Comprehensive and Reliable Evaluation of Multi-Shot Audio-Video Generation
- TIDE: Efficient and Lossless MoE Diffusion LLM Inference with I/O-aware Expert Offload
- From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models