AI 论文 · 2026-08
浏览 2026-08 发布的AI 论文内容,第 1 页,共 664 条。
- AgenticGen: Reward-Guided Agentic Video Generation for Advertising
- NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference
- Group Adaptive Clipping Policy Optimization
- FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos
- Adapting Without Gradients: Affine Statistics Transport and What Its Certificate Can Tell You
- Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation
- Dr. Claw: An AI Scientist Workspace for Vibe Research
- WHALE: A Simple Recipe for Joint Harness-Weight Optimization
- ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
- Recursive Criticality of AI Self-Improvement
- Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
- Safin-1: Safety from Within through Memory-Native State Evolution
- PaperGym: Rubric-Centered Evolution for Research-Plan Generation
- BLARM: Animating 3D Objects from Video via Blending Latent Rigid Motion Primitives
- Aspire: Can Models Self-Evolve from Vague Goals?
- DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
- S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
- Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data
- Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
- Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
- Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
- Normalized Low-Rank Adaptation
- MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
- CogEvol: Towards Efficient and Reliable Learning Environment Generation
- MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval
- LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
- Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
- Uncertainty-Aware End-to-End AI Weather Forecasting: Disentangling Observation and Model Contributions
- E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
- PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
- WebWorld: The Browser as a World Model for Self-Improving Web Code
- Agents in the Large: Perception-Centered Architecture for Persistent Agents
- Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering
- Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
- Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
- Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation
- Using Grounded Theory for Agent Behavior Analysis at Scale
- Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents
- On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
- PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
- DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection
- CAST: Critique-Aware Supervision for Training Reliable Long-Horizon Tool-Calling Agents
- Verification-Aware Training for Speculative Decoding
- RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
- The Mechanics of Democratic Dominance: A System Dynamics Paradigm for Dynamic Consent Engineering
- Small Language Models as Judges for Rubric-Based Reinforcement Learning
- SpanCalib-VLM: Calibrated Hallucination Span Detection in Vision-Language Models
- Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
- ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models
- Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation
- Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations
- MMMMM: A Unified Taxonomy for Investigating the Mechanisms of Multilingual MultiModal Misinformation
- Wasserstein-Barycentric Interaction Fields for Spatial Factor Models: Evidence from Language-Model Representations
- Cross-lingual Functional Vectors for Emotion Detection in Large Language Models
- SnapBench: Benchmarking Snap-and-Ask Multimodal Retrieval for Mobile Interactions
- AgentKernel: The Trust-Native Agentic Operating System
- Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
- EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
- GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
- Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase
- QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
- Dynamic Important Example Mining for Reinforcement Finetuning
- Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
- Chat-Edit-3D++: Interactive 3D and 4D Scene Editing via Large Language Models
- Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions
- SafeAtlas-VL: Beyond Binary Multimodal Safety with Large-Scale Data and Guard Models
- CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
- Scaling Automatic Research Agents via World Models
- To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation
- Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
- Evaluating the Hidden Costs of Personalization in Large Language Models
- Video Generative Models as Geometry Learner
- Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
- ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
- LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
- Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
- Sliding-window beats linear attention
- EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
- LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
- Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
- Agentic Artifact Creation: Systems, Evaluation, Principles, and Opportunities
- Rubric-to-Code Credit Assignment for Reinforcement Learning
- HyQuant: Hybrid-Precision Quantization for LLM Attention
- An Empirical Study on Zero-Data Bootstrapping for Conversational Recommender Systems
- UI-Venus-2 Technical Report
- Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching
- Fast Weight Attention for Continual Learning
- Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
- Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
- Revisiting Local Context for Long-Horizon Streaming 3D Reconstruction
- UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
- CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
- WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
- TTPO: Test-Time Policy Optimization
- Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
- Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO
- PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
- What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
- Magpie: Real-Time World Renderer for Interactive Games
- EditaLive! Unified Character Video Editing for Live Streaming