ADR
面向企业 AI Agent 的安全检测与响应系统,覆盖可观测性、威胁检测和攻防基准测试,已在 Uber 生产环境部署。亮点是自带 300+ 基准任务和 133 个 MCP 服务器,覆盖 17 种 agent 攻击手法,并采用双 agent 检测器提升可疑会话识别精度。目前开源 Sensor、Benchmark 和 Detector,Prevention 组件未包含;基准数据为合成数据,仅限防御性安全研究使用。
README
ADR:Agentic AI(智能体式 AI)检测与响应
ADR(Agentic AI Detection and Response)是一套面向 AI agent(智能体)的企业级安全系统。它帮助组织保护面向员工的 agent(如 Cursor、Claude Code、Codex)以及面向客户的 agent(如 AI 客服 agent)。
ADR 已在 Uber 生产环境部署,配套论文已被 MLSys 2026 接收:论文 PDF · 幻灯片 PDF
ADR 如何保护企业 AI agent
ADR 通过四项互补能力保护企业 AI agent:观察 agent 活动、评估防御、检测威胁和预防不安全行为。
- ADR Observability(可观测性):了解 AI agent 正在做什么以及为什么。 在生产环境中,ADR 会跨 macOS、Linux 和 Windows 从 7+ 个 AI 编码工具以及内部自动化和面向客户的客服 agent 捕获 agent 意图、工具使用和执行轨迹。
- ADR Benchmark(基准测试):在真实企业条件下测试 agent 安全性。 ADR-Bench 包含 300+ 个任务、133 个 MCP 服务器,并覆盖全部 17 种 agent 攻击技术。
- ADR Detection(检测):高效检测有风险的 agent 行为。 其两层架构将高召回率分流(triage)与针对可疑会话的更深层 agentic 推理(agentic reasoning)相结合。
- ADR Prevention(预防):在不安全行为造成危害之前阻止它们。 此组件不包含在当前开源版本中。敬请期待。
仓库结构
本仓库包含论文中描述的开源 ADR Sensor、ADR-Bench 和 ADR Detector。离线 ADR Explorer 引擎不包含在此处,它通过部署前红队演练(red teaming)来强化 ADR Detection。
| 路径 | ADR 组件 | 描述 |
|---|---|---|
| Sensor/ | ADR Observability | 从 Claude Code、Cursor、Codex 等工具收集并标准化 agent 遥测数据 |
| Detection/ | ADR Benchmark + Detection | 双 agent 检测器、133 个 MCP 服务器、303 个基准测试任务、基线、图表脚本 |
| docs/REPRODUCIBILITY.md | Evaluation(评估) | 复现基准检测和论文图表的分步工作流 |
快速开始:ADR Detection
git clone https://github.com/uber/ADR
cd ADR/Detection
uv sync
export ANTHROPIC_API_KEY="..." OPENAI_API_KEY="..."
默认检测器为 adr(ADR dual-agent)。如需无 API key 的冒烟测试,请使用 --detector llamafirewall(参见 Detection/README.md)。
完整的评估工作流(解压打包的基准测试 → 运行检测器 → 绘制图表)请参阅 docs/REPRODUCIBILITY.md。
组件文档:
- Sensor/README.md:遥测收集与统一 Schema
- Detection/README.md:ADR-Bench、检测器基线、MCP 基础设施
引用
@inproceedings{li2026adr,
title={ADR: An Agentic Detection System for Enterprise Agentic AI Security},
author={Li, Chenning and Hu, Pan and Xu, Justin and Ozbas, Baris and Liu, Olivia and Van, Caroline and Li, Manxue and Zhou, Wei and Alizadeh, Mohammad and Zhang, Pengyu and Sriramadhesikan, KK and Zhang, Ming},
booktitle={Proceedings of the Ninth Conference on Machine Learning and Systems},
year={2026}
}
或者使用 CITATION.cff。
许可证
Apache License 2.0。参见 LICENSE。Detection/benchmark/agentdojo/ 为按其自身 LICENSE(MIT)以 vendor 方式引入的第三方代码。
数据声明
Detection/ 中仅包含用于防御性安全研究的合成基准测试夹具(伪造凭据、模拟环境、提示注入(prompt injection)场景)。详情请参阅 docs/OPEN_SOURCE_REVIEW.md。