AI 论文 · 2026-08
浏览 2026-08 发布的AI 论文内容,第 4 页,共 664 条。
- GenRouter: Unified Workflow Routing for Agentic Image Generation
- Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models
- Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
- Towards Real-Time and Adaptable LiDAR Scene Completion
- HiFi-BRep: High-Fidelity Latent Representation for Robust B-Rep Generation
- Drive, Pack, Fly: The Travelling Thief Problem with Drone
- Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs
- GRNEdit: Efficient General Video Editing from a New Binary-Evidence Perspective in Generative Refinement Networks
- StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding
- Advancing Open and Reproducible Relational Learning: RelArena-α, TabPFN-Rel and RPI
- FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution
- AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
- Unifying Graph Neural Networks Through a Common Layer Equation
- Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
- R^3-Bench: LLMs Struggle with Resource-Rational Reasoning under Shared Budgets
- Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
- From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
- A Plug-and-Play 2D Motion Interface for Real-World Motion Language Models
- Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs
- LLMs Get Smarter from Targeted Synthetic Multilingual Data
- UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
- Bounded Agents: Delegation Security for Multi-Agent AI Systems
- GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
- Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
- TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
- ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval
- Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search
- WorldRover: A Scalable Synthetic Video Data Engine for World Exploration with Rich Annotations
- Dynamic Multi-Byte Prediction With Hierarchical Language Models
- Understanding Cognition-Induced Risks in Agentic AI Systems
- VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
- LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- MOSS-VL Technical Report
- Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
- Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
- Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents
- MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided Refinement
- A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
- ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World Models
- Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification
- How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks
- Personalized Auto-Research: Towards a True AI Co-Scientist
- MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling
- CPI-Bench: A Comprehensive,Practical and Intelligent Benchmark for Real-World Image Editing
- Marionette: Predicting World States, Rendering Geometry, Painting Appearance
- PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments
- Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
- Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning
- PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment
- SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
- The More Popular, The Harder to Forget: Adaptive Popularity for LLM Unlearning
- Self-Supervised Visual On-Policy Distillation
- SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
- Forecast Collapse in Time-Series Foundation Models
- Scaling Domain Data Repetition in LLM Pretraining
- Demystifying Agent Skills: Why They Work-Until They Don't
- Nanbeige4.2-3B on Apple Silicon: Fixing Deployment Bugs and Decreasing Looped Transformer Memory Overhead
- Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion
- Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning
- Agentic Transaction: Towards ACID-Compliant Agent Systems
- Second Thought: Reasoning in Parallel as LLM Agents Act and Observe
- Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models
- ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
- AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
- OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
- V-RAE: Rethinking Video Latent Spaces for Generation
- HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
- PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives
- QuoteBench: How Matched Scores Can Hide Command-Path Failures
- Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
- LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
- DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
- Intern-S2-Preview: Scientific Agentic Foundation Model
- DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
- Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
- Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings
- Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation
- NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
- SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
- H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models
- Latent On-Policy Self-Distillation
- PixSDS: Why Latent SDS Makes Noisy Pixels
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
- CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation
- NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- The Embedder's Dilemma: LLMs Are Better, but at What Cost?
- From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options
- CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers
- Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
- UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
- LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time
- DREAM Technical Report
- A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware
- VideoGAIA: A Benchmark for General AI Assistants on Agentic Video Understanding
- Is this Citation on Point?
- Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
- StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
- AVA-Encoder: Towards Agent-Native Video Representation Learning