AI 论文 · 2026-08
浏览 2026-08 发布的AI 论文内容,第 7 页,共 664 条。
- Self-Evolving Coding Agents
- Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging
- DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack
- SKILL-KD: Contrastive Skill Distillation for LLM Agents
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
- TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex
- Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking
- Towards Interpretable Foundation Models for Retinal Fundus Images
- Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning
- What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
- CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
- Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents
- Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation
- Quo Vadis, World Modeling?
- Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
- ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads
- WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity
- AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling
- CAPEval: A Decoupled Caption Evaluation across Understanding and Generation
- GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning
- UEmbed: Unified Sparse and Dense Multimodal Embeddings
- Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data
- SWE-Touch: Benchmarking Coding Agents When Users Touch the Code
- InfiniSplat: Implicit Gaussian Decoding for Large-Baseline Monocular View Synthesis
- GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
- ScrambleToolBench: Agents Search Exhaustively Even When Their Own Map Points to the Next Step
- GEOID-Flood: A Large-Scale Multi-Modal Benchmark Dataset for Flood Segmentation
- SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation
- PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs
- Lossless Tensor Compression as Program Synthesis
- Douyin Multimodal Embedding Model Technical Report
- SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
- Roomer: Reflective Object-Grounded Model Editing and Repair for 3D Indoor Layout Synthesis
- LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
- StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field
- Wnuan: Staged Post-Training for Question Answering over Proprietary Enterprise Knowledge
- Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills
- PCSD: Persistent Consistency for Self-Distillation in Agentic Reinforcement Learning
- DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
- Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
- DAPD: Dual-Anchored Policy Distillation
- Progressive Agent Skill Generation via Reinforcement Learning
- Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations
- Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
- GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization
- Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
- Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
- FATE: Frame-Level Audio-Visual Temporal Embedding
- RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache Eviction
- 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
- MiniWorld: Democratizing the Training of Video World Models from Scratch
- FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds
- Round-Trip Consistency: Bidirectional Diffusion Models Can Predict Their Own Rollout Errors
- CADENA: Stepwise CAD Reverse Engineering
- Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
- Push-Wiper: Toward General-Purpose Robotic Cleaning across Varied Stains and Surfaces with Segmented Pushing Trajectories
- OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution
- Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
- DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion Latents
- Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
- Decoding Children's Gait Behavior
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models