AI 论文 · 2026-08
浏览 2026-08 发布的AI 论文内容,第 3 页,共 664 条。
- GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?
- GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
- Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention
- Training Agents to Evolve with Their Harness: TaoLive Digital Avatar Agent Technical Report
- Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It?
- SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation
- EXPL-FR: Explaining Face Recognition Models via Vision-Language Alignment
- OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs
- CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
- Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence
- A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans
- Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs
- PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration
- Training, learning and inference: unified dynamics of neural systems
- TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming
- Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
- InfinityEdit: Infinite Video Editing with a Lightweight Edit-Ignition Adapter
- EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking
- Towards Faithful Simulation of Human Shopping Behavior
- AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale
- Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving
- Tomatoes, Potatoes, and Onions: Questioning the Need for Faces in Face Presentation Attack Detection
- FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth
- Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
- Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources
- RISE: Adaptive Imagination for World Action Models
- WithEveryone: Unified Planning and Identity Grounding for Group Image Generation
- 4DAnyone: Create Anyone in 4D from a Casual Monocular Video
- Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
- Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
- MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use
- Towards Quantifying Benchmark Optimization in ASR Models
- EXIMO: VLM Guided Exploration of VLA Policies
- EnvHarness: Awakening Static Worlds for Agent Learning
- Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents
- Repo0: Design-Driven Zero-to-All Code Generation
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
- CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
- GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation
- One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows
- Block3D: Efficient Text-to-3D Generation via Block-Wise Diffusion
- Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
- FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
- Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress
- SPADE: Self-Play in Adaptive Synthetic Executable Environments
- Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
- SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object Detection
- Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
- Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
- SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
- SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
- Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC Planning
- SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
- VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
- FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
- SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code Generation
- DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents
- Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
- FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
- Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
- Hydra-0: Action Flow for Generalist World Modeling and Control
- Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot See
- WorldMind: Decoupled Game World Model for State-Aware NPC Behavior
- Human-Centric Intelligence in the Era of Foundation Models: A Survey
- From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
- EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing
- Chain-of-Experience for Continual LLM Improvement
- GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation
- aDSL: Agentic 3D Creation via Joint Agent-Program Design
- StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
- HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
- CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video Editing
- Agent Lightning v1.0: Towards Harnessed Agentic RL
- Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation
- SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
- MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video Understanding
- LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
- PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
- TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
- Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
- Abra: Scaling Diffusion Image Training
- ASI-Bench: At the Dawn of Artificial Superintelligence
- Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt Injection
- Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems
- ParaTempo: Efficient Parallel Reasoning via Temporal Confidence
- Looped Language Models Improve Compositional Tool Calling
- DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
- Cross-Model Memory Transfer via Target-Side Reader Adaptation
- The Problem Is the Problem: Towards Scalable Mathematical Discovery
- An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models
- τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation
- Improving the matrix multiplication exponent with modern optimization and AlphaEvolve
- HarnessEval-W: Agentifying the Evaluation of Visual Worlds
- Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision
- ClawGym II: Exploring Black-Box RL on Agent Harness
- PixRestore: Unified Image Restoration via Pixel Diffusion Transformer
- TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation