AI 论文 · 2026-06
浏览 2026-06 发布的AI 论文内容,第 2 页,共 818 条。
- NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning
- Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
- ZooClaw-FashionSigLIP2: Distilled Fine-tuning for Robust Fashion Retrieval
- When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search
- RocketSmith: Agentic Additive Manufacturing of High-Powered Rockets
- Simplified Sparse Attention via Gist Tokens
- ReFreeKV: Towards Threshold-Free KV Cache Compression
- PolyFlow: Continuous Topology Embedding Flow Matching for Artist-style Mesh Generation
- Qwen-Image-2.0-RL Technical Report
- Ko-WideSearch: A Korean Breadth-Search Benchmark for Exhaustive Set Enumeration by Web Agents
- MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
- Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving
- Delayed Verification Destabilizes Multi-Agent LLM Belief: Instability Thresholds and Optimal Corrector Placement
- DanceOPD: On-Policy Generative Field Distillation
- PhysiFormer: Learning to Simulate Mechanics in World Space
- SAM2Matting: Generalized Image and Video Matting
- Hallucination in World Models is Predictable and Preventable
- ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
- When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
- How Good Can Linear Models Be for Time-Series Forecasting?
- EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
- LISA: Likelihood Score Alignment for Visual-condition Controllable Generation
- Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
- How Much Static Structure Do Code Agents Need? A Study of Deterministic Anchoring
- To Run or Not to Run: Analyzing the Cost-Effectiveness of Code Execution in LLM-Based Program Repair
- RedVox: Safety and Fairness Gaps in Speech Models Across Languages
- Focusing on What Matters: Saliency-Harnessing Accurate Routing for Diffusion MoE
- Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
- Confidence-Aware Tool Orchestration for Robust Video Understanding
- Information-Aware KV Cache Compression for Long Reasoning
- OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning
- LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing
- Boundary-Aware Context Grounding for A Low-Channel EEG Agent
- In-Context World Modeling for Robotic Control
- Large-Scale Tunnel Air-Ground Collaboration With FLISP: Fast LiDAR-IMU Synchronized Path Planner
- JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting
- Running the Gauntlet: Re-evaluating the Capabilities of Agents Beyond Familiar Environments
- Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
- The Verification Horizon: No Silver Bullet for Coding Agent Rewards
- COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami
- Fast LeWorldModel
- TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy
- MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation
- Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents
- DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation
- AI translation of literary texts is "fine", but readers still prefer human translations
- Why Multi-Step Tool-Use Reinforcement Learning Collapses and How Supervisory Signals Fix It
- MIMFlow: Integrating Masked Image Modeling with Normalizing Flows for End-to-End Image Generation
- The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar
- Autodata: An agentic data scientist to create high quality synthetic data
- ShutterMuse: Capture-Time Photography Guidance with MLLMs
- One Model, Many Latencies: Universal Speech Enhancement for Diverse Real-Time Applications
- The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models
- Constraint Tax in Open-Weight LLMs: An Empirical Study of Tool Calling Suppression Under Structured Output Constraints
- Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models
- TheoremGraph: Bridging Formal and Informal Mathematics
- Improved Large Language Diffusion Models
- V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning
- Physics Question Scene Graph: Fine-grained Evaluation of Physical Plausibility in Text-to-Video Generation
- Transition-Aware best-of-N sampling for Longitudinal Chest X-ray Reports
- What Intermediate Layers Know: Detecting Jailbreaks from Entropy Dynamics
- Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
- Do Thinking Tokens Help with Safety?
- DiffusionBench: On Holistic Evaluation of Diffusion Transformers
- InSight: Self-Guided Skill Acquisition via Steerable VLAs
- FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation
- FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation
- OpenThoughts-Agent: Data Recipes for Agentic Models
- IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
- Are We Ready For An Agent-Native Memory System?
- World Value Models for Robotic Manipulation
- DREAM: Dense Retrieval Embeddings via Autoregressive Modeling
- Qwen-AgentWorld: Language World Models for General Agents
- MEMPROBE: Probing Long-Term Agent Memory via Hidden User-State Recovery
- Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning
- NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?
- AGORA: An Archive-Grounded Benchmark for Agentic Workplace Document Reasoning
- Advancing WordArt-Oriented Scene Text Recognition: Datasets and Methods
- Lite Any Stereo V2: Faster and Stronger Efficient Zero-Shot Stereo Matching
- Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
- Trimming the Long-Tail of Visual World Modeling Evaluation
- FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning
- AsyncOPD: How Stale Can On-Policy Distillation Be?
- Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning
- ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection
- CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression
- RoPE-Aware Bit Allocation for KV-Cache Quantization
- SkillHone: A Harness for Continual Agent Skill Evolution Through Persistent Decision History
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- LLM Program Optimization via Retrieval Augmented Search
- The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
- GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
- ChartWalker: Benchmarking the Cross-Chart RAG Task
- Critique of Agent Model
- Mind the Heads: Topological Representation Alignment for Multimodal LLMs
- ABACUS: Adapting Unified Foundation Model for Bridging Image Count Understanding and Generation
- Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
- Semantic Browsing: Controllable Diversity for Image Generation
- Tapered Language Models
- EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions