AI 论文 · 2026-10
浏览 2026-10 发布的AI 论文内容,第 1 页,共 207 条。
- OuroWorld: Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
- WorldGuide: Goal-Directed Video World Model for Procedural Task Execution
- OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning
- BrickBench: Evaluating Agentic Brick Design
- One Block, Multiple Depths: Recurrent Vision Transformers with Depth-Programmed Experts
- Pumpire: Unified Benchmark for Metric Distance Estimation
- Beyond Spatio-Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
- OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video
- SpaceFlow: Locally Controllable 3D Generation
- Embodied Turing Machines: Stateful Code for Robot Recursive Self-Improvement
- Accurate but Not Humble: Evaluating Epistemic Humility in LLM Agents under Knowledge Conflict
- Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models
- Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction
- TestPrism: Rethinking Test Evaluation Beyond a Single Reference
- A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
- Predicting Cable Dynamics with Physical Attention Bias
- MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement
- Memento 3: Model-Based Recursive Self-Improvement through Reflective Rulebooks
- Incremental Open-Ended Deep Research with Structured Harness
- REMORY: Learning Residual Memory for Context Compaction
- The Lattice of Transition Laws
- Skill Constellations: Tracing the Supply Chain of Agent Skills on GitHub
- Opera: A Verbal Critic Framework for Long-horizon Coding Agents
- MARGIN: Runtime Confidence Calibration for Multi-Agent Foundation Model Coordination
- Foundations of Large Language Models
- Tetris3D: 3D Scene Generation With Objects That Fit Together
- EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
- GRACE: Generation-aware latent compression for efficient video generation
- RoboJEPA: Scaling Robotic Latent World Models
- QuadTok: Quadtree Visual Tokenizer for Autoregressive Image Generation
- Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
- RunningTab: Direct Workspace Interaction with Environment-Side Tabs
- Q-Learning with Scalar Adjoint Matching
- RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
- RoboQuest: Generalist Physical Agents that Search, Inspect and Test
- MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent
- UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy
- SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
- UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
- From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery
- System Switch: When Should a Fast Decision Model Stop and Think?
- On-Policy Distillation Teaches New Skills but Not New Knowledge
- MIMESIS: Learning User Simulators as Training Environments for Interactive Agents
- DSReg: Provably Recovering Individual World Latents without Reconstruction
- Iris-3B: Going Beyond the Latent with Pixel-Space Diffusion Training, Conversion and Fine-Tuning
- Evaluating the Transfer of Co-Evolved Communication from 2D to 3D Simulation
- SWE-Game: Can Coding Agents Build the Games We Want?
- AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
- Co-Evolving Robot Orchestrators and Policies through Deployment
- Frozen Models, Evolving Expertise: Model-Agnostic Learning from Deployment Experience for Multimodal Medical AI
- CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use
- U-Space: Uncovering When and Why Uncertainty Arises in Language Models
- EDiS: Edge Disjoint Subgraph Sparsification Framework for Graph Neural Networks
- PhysEvo: Astra Can Act, Let It
- SPW-Nav: A Streaming Panoramic World Model for Language-Guided Navigation
- Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station
- CARE: Certifying Acceleration for Vision-Language-Action Inference
- Agent Plasticity: Measuring Self-Improvement Through Experience
- Task-Sufficient Contraction: Source Selection for Machine Information Interfaces
- Sherpa: Teaching LLMs to Teach Adaptively
- CtrlCache: Accelerating Interactive Video World Models with Control-Aware Caching
- Towards In-Parameter Memory Augmentation for Large Language Models
- Incidental information contaminates patient notes and disrupts clinical reasoning in large language models
- UNREAL: Unifying Retrieval and Long-Context with a Single Model
- Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability
- NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale
- Sensor-Language-Action Models
- DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks
- On-Policy Distillation with Negative-Policy Rollouts
- A self-learning scientific agent for X-ray diffraction
- TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models
- From Evidence to Action: How Tool-Using Agents Fail
- DLoop: Looped Speculative Decoding
- CheckerBench: Can Long-Horizon Agents Synthesize Static-Analysis Checkers?
- Improving Proactive AI Assistance with Hierarchical Procedural Understanding
- A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies
- SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
- A GPU-Parallel Framework for Heterogeneous Multi-Task Reinforcement Learning
- MEND: RL For Flow Models via Proximal Velocity Matching
- Technical Report on the Turba Fertilizer Machine Learning Stack in Morocco
- Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
- WildMatch: Weakly Supervised Image Matcher Adaptation for Wildlife Re-Identification
- Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
- Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding
- Structuring MoE Expert Selection for Agentic Reinforcement Learning
- Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents
- Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation
- TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
- S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation
- Learning to Read the Contextual Tokens in Diffusion Transformers
- Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
- MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
- Adapting prior-data fitted networks for tabular anomaly detection
- Closing the Context Gap: Activation Alignment for Tabular In-Context Learning
- What Matters for Latent Reasoning with Flow Matching
- Representation-Space MMD for Diffusion Language Models
- Optimizing the Optimizer: Language Models Discover Faster Molecular Relaxation
- Empirical Variational Autoencoder
- You Changed Your Mind, The Model Didn't: Demystifying Intent in Multi-Turn Dialogue
- SoK: Semantic Decision Engines in Network Control Loops