AI 论文 · 2026-08
浏览 2026-08 发布的AI 论文内容,第 5 页,共 664 条。
- AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
- Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction
- Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
- Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
- Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
- Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
- TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
- Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
- Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill
- ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents
- Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
- MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
- Dion3: Full-Stack Orthogonal Updates
- Hand Visibility Detector: Per-Keypoint Visibility Estimation for Hands
- From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
- Persistent Recursive Worlds Enable Autonomous Software Evolution
- UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
- MobileMem: Learning from a Year of Mobile Experiences
- Gaze Target Estimation Anywhere with Concepts
- Self-Evolving Embodied Agents via Skill-Harness Evolution
- Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
- Agent Safety Should Be a Runtime Contract
- AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss
- SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
- ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization
- ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
- VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
- Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
- Beyond Pixels: From Video Priors to 4D Worlds
- Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
- Self-Geometry: GT-Free and Plug-and-Play Test-Time Adaptation for Geometrically Consistent 3D Vision Foundation Models
- SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
- ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
- DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation
- InSight-doc: Agentic Visual Perception for Long-Document Understanding
- Simplex Relaxation for Discrete Diffusion
- SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
- DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
- Thought-Level Beam Search for Reasoning
- Parameter Exploration for RLVR via Variational Learning
- Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design
- Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference
- Multimodal Model Diffing for Feature Discovery and Control
- Learning How the World Evolves: Extrapolative Video World Models via Latent Dynamics Reasoning
- Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness
- BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
- Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
- Stealing Reasoning Traces from Proprietary LLM APIs
- RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance
- CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
- Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
- Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
- Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation
- UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers
- From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
- Motif 3: Technical Report
- Evo-Bench: Can Language Models Improve Agent Harness?
- Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
- How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
- A^2E : An End-to-End Agent Auditing Engine
- RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections
- RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
- Full-bandwidth transformer
- 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
- SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
- Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
- UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
- UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
- Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
- Mitigating Gender Bias in English to Romanian Machine Translation
- VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use
- Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
- Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
- On-Policy Self-Distillation without Any Supervision
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
- Ego-OSCAR: Egocentric Open source Stereo CAptuRe System
- Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
- A Hybrid Nested Harness for Decoupling Structure and Parameters in LLM-Driven Optimization
- TSDS-Toolbox: A Toolbox for Measuring Time-Series Dataset Similarity
- NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
- OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
- Evidence-RL: Towards Evidence-intensive Visual Reasoning
- Vision-Language Grounding as Bidirectional Concept Correspondence
- Training-Free Speech-Centric Omni Understanding with Frozen VLMs
- Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
- SimWAM: A Simple World Action Model for End-to-End Autonomous Driving
- MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation
- CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
- Addressable Memory for Video World Models
- Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
- Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
- An AI4AI Framework for Visual Token Pruning
- Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
- Modular TTT: Rethinking Test-Time Training as Composable Modules
- YOLO-PEFT: Parameter-Efficient Fine-Tuning on YOLO Family
- When Context Bites: Detecting RAG Poisoning via Document-Level Attention Collapse
- LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers
- Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
- Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence