AI 论文 · 2026-07
浏览 2026-07 发布的AI 论文内容,第 1 页,共 385 条。
- Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers
- GraphVid: Interactive Graph-Controllable Video Generation
- Self-Supervised Learning of Structured Dynamics from Videos
- OpenForgeRL: Train Harness-native Agents in Any Environment
- Visual Contrastive Self-Distillation
- SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation
- Recurrent Sinusoidal INRs for Efficient High-Fidelity Representation
- AREX: Towards a Recursively Self-Improving Agent for Deep Research
- ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
- Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
- Sample-Efficient Learning from Agent Experience
- TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation
- Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
- K12-KGraph: A Curriculum-Aligned Knowledge Graph for Benchmarking and Training Educational LLMs
- Robostral Navigate
- LLMs Get Lost in Evolving User Intent
- NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
- ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion
- Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
- Self Gradient Forcing: Native Long Video Extrapolation
- SeededGrasp: Language-Guided Grasping in Complex Scenes with Multiple Embodiments
- SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
- ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
- ReferTrack: Referring Then Tracking for Embodied Visual Tracking
- Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
- G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
- DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
- Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning
- Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
- SLPO: Scaling Latent Reasoning via a Surrogate Policy
- ISO: An RLVR-Native Optimization Stack
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
- Generative World Renderer at the Speed of Play
- Scaling Laws for Hypernetwork-Based Knowledge Injection in Large Language Models
- Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
- Masked Visual Actions for Unified World Modeling
- Two-Level Meta-Rubrics for Evaluating Open-Ended Generation: GAMUT, a Benchmark for Factual Completeness
- FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
- Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers
- Delineate Anything v2: A Global Foundation Model for Field Delineation
- Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
- Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
- Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
- H^2SD: Hybrid Hindsight Self-Distillation
- Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
- HPD-Parsing: Hierarchical Parallel Document Parsing
- Moving Alphabet: A Controlled Study of Training Data for Text-to-Video Generation
- AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
- Stale but Stable: Staleness-Adaptive Trust Regions for Stabilizing Asynchronous Reinforcement Learning
- AutoIndex: Learning Representation Programs for Retrieval
- NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
- AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
- EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
- FlowMimic: Mask-free Visual Editing and Generation with Pixel-pair Warped Flow Field for Online Video Editing Data Generation and Modality Mimicry
- HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement
- SWE-Pruner Pro: The Coder LLM Already Knows What to Prune
- FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications
- Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices
- Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
- LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks
- SciForma: Structure-Faithful Generation of Scientific Diagrams
- WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
- Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
- DiFA: Inference-Time Forward-Process Alignment for Diffusion Models
- ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
- SLAM in Low-Light Environments: Project Report
- ShotPlan: Cinematic Video Generation with Learnable Planning Token
- ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning
- Token-Level Off-Policy Learning for Faithful Generation Under Distribution Shift
- Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
- Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
- TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs
- EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World
- Distilled Reinforcement Learning for LLM Post-training
- The Geometry of Semantic Space: A Continuous Geometric Framework for the Transformer Architecture
- HarmoHOI: Harmonizing Appearance and 3D Motion for Multi-view Hand-Object Interaction Synthesis
- Environment-free Synthetic Data Generation for API-Calling Agents
- Dataset Distillation by Influence Matching
- Group Entropy-Controlled Policy Optimization
- DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
- Can Multimodal Large Language Models Understand OCT?
- Nonuniformity Principle in Human-AI Coworking
- Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence
- FVAttn: Adaptive Sparse Attention with Runtime Load Balancing for Video Generation
- When Does Muon Help Agentic Reinforcement Learning?
- An Exam for Active Observers
- Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
- Understanding Reasoning from Pretraining to Post-Training
- JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
- Loop the Loopies!
- DSWorld: A Data Science World Model for Efficient Autonomous Agents
- S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
- RecGPT-V3 Technical Report
- SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
- Recursive Harness Self-Improvement
- Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
- Trajectory-aware Cross-view Geo-localization with Sequential Observations
- Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
- Hierarchical Denoising For Multi-Step Visual Reasoning