AI 论文 · 2026-06
浏览 2026-06 发布的AI 论文内容,第 6 页,共 818 条。
- Unstable Features, Reproducible Subspaces: Understanding Seed Dependence in Sparse Autoencoders
- FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents
- World Model Self-Distillation: Training World Models to Solve General Tasks
- On the Limits of LLM-as-Judge for Scientific Novelty Assessment
- Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation
- Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
- Notes2Skills: From Lab Notebooks to Certainty-Aware Scientific Agent Skills
- LLM-Enabled NWDAF: A Step Toward AI-Native 6G Network Intelligence
- Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code
- Lius: Translation Model Based Instructional Lingustic Using Continual Instruction Tuning In Kupang Malay
- ICA Lens: Interpreting Language Models Without Training Another Dictionary
- Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning
- TreeSeeker: Tree-Structured Trial, Error, and Return in Deep Search
- When is Your LLM Steerable?
- RedAct: Redacting Agent Capability Traces for Procedural Skill Protection
- Building Social World Models with Large Language Models
- EEVEE: Towards Test-time Prompt Learning in the Real World for Self-Improving Agents
- Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization
- TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning
- Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
- When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
- How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs
- Towards Diverse Scientific Hypothesis Search with Large Language Models
- One Token per Multimodal Evidence: Latent Memory for Resource-Constrained QA
- Forecasting Future Behavior as a Learning Task
- Risk Under Pressure: Compute-Aware Evaluation of Adversarial Robustness in Language Models
- Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
- i1: A Simple and Fully Open Recipe for Strong Text-to-Image Models
- ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations
- Next Forcing: Causal World Modeling with Multi-Chunk Prediction
- Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories
- The Role of Feedback Alignment in Self-Distillation
- P3D-Bench: Benchmarking MLLMs for Parametric 3D Generation and Structural Reasoning
- UniPET: a universal network for high-quality PET image denoising across varied dose reduction factors
- WorldOlympiad: Can Your World Model Survive a Triathlon?
- IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder
- Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning
- Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
- U-TTT: Towards Generalizable PET Image Denoising via Test-Time Training
- Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models
- Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning
- Flash-GMM: A Memory-Efficient Kernel for Scalable Soft Clustering
- SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning
- N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization
- The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment
- DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch
- FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion
- Decentralized Multi-Agent Systems with Shared Context
- Kwai Keye-VL-2.0 Technical Report
- Dynamic Linear Attention
- ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics
- WebChallenger: A Reliable and Efficient Generalist Web Agent
- Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders
- BrainSurgery: Reproducible and Reliable Declarative Weight Manipulations for Model Editing and Upcycling
- PsychoSafe: Eliciting Psychologically-Informed Refusals in Large Language Models
- Optical Reasoning: Rethinking Images as an Expressive Reasoning Medium Beyond Text
- Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short
- Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory
- Hardening Agent Benchmarks with Adversarial Hacker-Fixer Loops
- τ-Rec: A Verifiable Benchmark for Agentic Recommender Systems
- BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts
- ABot-Earth 0.5: Generative 3D Earth Model
- OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics
- Rethinking the Divergence Regularization in LLM RL
- iMaC: Translating Actions into Motion and Contact Images for Embodied World Models
- AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing
- Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
- iOSWorld: A Benchmark for Personally Intelligent Phone Agents
- SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
- End-to-End Context Compression at Scale
- SwiftVR: Real-Time One-Step Generative Video Restoration
- Leveraging Morphology for Historical Script Metrological Analysis
- WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
- Precision Is Not Faithfulness: Coverage-Aware Evaluation of Grounded Generation with a Complete Oracle
- PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
- TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders
- SG-OPD: Sign-Gated On-Policy Distillation via Sign-Consistency Gating and Phased Teacher Sampling
- Visual Para-Thinker++: A Single-Policy Multi-Agent Framework for Visual Reasoning
- Late-Layer Fusion is Enough: Dual-Path Vision Token Routing for Multimodal Large Language Models under Visual Saturation
- FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
- Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions
- Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating
- MilliVid: Hierarchical Latents for Long-Range Consistency in Video Generation
- Bridging the Agent-World Gap: Text World Models for LLM-based Agents
- TRIAGE: Dialectical Reasoning for Explainable Risk Prediction on Irregularly Sampled Medical Time Series with LLMs
- AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models
- MBench: A Comprehensive Benchmark on Memory Capability for Video World Models
- FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching
- Attacks on Machine-Text Detectors Retain Stylistic Fingerprints
- PaperMentor: A Human-Centered Multi-Agent Writing Tutor for AI Research Papers on Overleaf
- MaskAlign: Token-Subset Representation Alignment for Efficient Diffusion Training
- WaveDiT: Distribution-Aware Wavelet Flow Matching for Efficient 3D Brain MRI Synthesis
- OmniCap-IF: Benchmarking and Improving Instruction Following Abilities for Omni-Video Captioning
- OASIS: From Simulation Data Collection to Real-World Humanoid Loco-Manipulation
- PIPE-Cypher: Automatic Enterprise Benchmark Generation for Text-to-Cypher Systems
- Trajectory-Refined Distillation
- Phase Marginalization for Patch-Grid Instability in Vision Transformers
- EmpiriGraph-Psy: A Dataset and LLM Pipeline for Extracting Empirical Relation Graphs from Psychology Abstracts
- Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses