litellm
统一接入 100+ LLM 的 AI Gateway 与 Python SDK,用 OpenAI 格式调用 OpenAI、Anthropic、Bedrock、Azure、vLLM 等模型,内置虚拟 key、花费追踪、限流、guardrails 和负载均衡。Rust 内核重写后标称 1k RPS 下 P95 约 8ms,支持自托管,Stripe、Netflix 等已在生产使用;还顺带提供 A2A agent 与 MCP 网关,方便把 agent 和工具挂到同一入口。注意企业级功能走 LiteLLM Commercial License。
README
🚅 LiteLLM
LiteLLM AI Gateway
面向 100+ LLM 的开源 AI Gateway。可自托管。企业级就绪。以 OpenAI 格式调用任意 LLM。
LiteLLM Proxy Server(AI Gateway) | 托管 Proxy | 企业版 | 官网
什么是 LiteLLM
LiteLLM 是一个开源 AI Gateway,提供单一、统一的接口,让你以 OpenAI 格式调用 100+ LLM 提供商 —— OpenAI、Anthropic、Gemini、Bedrock、Azure 等。
你可以把它当作 Python SDK 直接集成到代码中,也可以部署 AI Gateway(Proxy Server) 作为团队或组织的集中式服务。
跳转至 LiteLLM Proxy(LLM Gateway)文档
跳转至受支持的 LLM 提供商
为什么选择 LiteLLM
跨提供商管理 LLM 调用很快就会变得复杂 —— 每个模型都有不同的 SDK、认证模式、请求格式和错误类型。LiteLLM 消除了这些摩擦:
- 统一 API —— 一套接口覆盖 100+ LLM,无需在各类提供商 SDK 之间周旋
- 开箱即用的 OpenAI 兼容 —— 无需重写代码即可切换提供商
- 生产级 gateway —— 内置 virtual key、成本追踪、guardrail、负载均衡和管理面板
- 在 1k RPS 下 P95 延迟为 8ms(benchmark)
OSS 采用者
![]() |
![]() |
![]() |
![]() |
![]() |
Netflix |
![]() |
功能特性
LLM - 调用 100+ LLM(Python SDK + AI Gateway)所有支持的 endpoint - /chat/completions、/responses、/embeddings、/images、/audio、/batches、/rerank、/a2a、/messages 等。
Python SDK
uv add litellm
独立的 litellm-core 发行版提供与 Python SDK 相同的
import litellm API 和运行时依赖。它不包含可选扩展、CLI 入口点
或内置 dashboard。每个环境只能安装一种 SDK 发行版,因为
litellm 和 litellm-core 存在重叠的 Python 文件。如果你需要
proxy、CLI 和可选扩展,请使用 litellm
在其发布集成尚在进行期间,可从此检出目录构建并安装 core:
python scripts/build_core_distribution.py --out-dir dist/core
python -m pip install dist/core/litellm_core-*.whl
该构建脚本需要 Git、uv 和 Rust 构建工具链。它从根目录的 pyproject.toml
读取发布版本号,不修改源文件,并在 dist/core 中生成一个 wheel
和一份自包含的 sdist。请在一个未安装 litellm 的干净环境中执行安装
from litellm import completion
import os
os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"
# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])
# Anthropic
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])
AI Gateway(Proxy Server)
快速上手 - 端到端教程 - 配置 virtual key,发出第一个请求
uv tool install 'litellm[proxy]'
litellm --model gpt-4o
import openai
client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "Hello!"}]
)
Agent - 调用 A2A Agent(Python SDK + AI Gateway)受支持的提供商 - LangGraph、Vertex AI Agent Engine、Azure AI Foundry、Bedrock AgentCore、Pydantic AI
Python SDK - A2A 协议
from litellm.a2a_protocol import A2AClient
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4
client = A2AClient(base_url="http://localhost:10001")
request = SendMessageRequest(
id=str(uuid4()),
params=MessageSendParams(
message={
"role": "user",
"parts": [{"kind": "text", "text": "Hello!"}],
"messageId": uuid4().hex,
}
)
)
response = await client.send_message(request)
AI Gateway(Proxy Server)
第 1 步。 将你的 Agent 添加到 AI Gateway —— 为每个 agent 将 protocolVersion 设为 1.0 或 0.3
第 2 步。 通过 A2A SDK 调用 Agent(需要 a2a-sdk>=1.1.0)
import httpx
from a2a.client import A2ACardResolver, ClientConfig, ClientFactory
from a2a.types import Message, Part, Role, SendMessageRequest
from a2a.utils.constants import TransportProtocol
from uuid import uuid4
base_url = "http://localhost:4000/a2a/my-agent" # LiteLLM proxy + agent name
headers = {"Authorization": "Bearer <your-master-key>"} # LiteLLM master key or a virtual key
async with httpx.AsyncClient(headers=headers, timeout=60.0) as http_client:
resolver = A2ACardResolver(httpx_client=http_client, base_url=base_url)
agent_card = await resolver.get_agent_card()
config = ClientConfig(
httpx_client=http_client,
streaming=False,
supported_protocol_bindings=[TransportProtocol.JSONRPC, TransportProtocol.HTTP_JSON],
)
client = ClientFactory(config).create(agent_card)
request = SendMessageRequest(
message=Message(
message_id=uuid4().hex,
role=Role.ROLE_USER,
parts=[Part(text="Hello!")],
)
)
async for event in client.send_message(request):
populated = event.ListFields()
if populated and populated[0][0].name in ("message", "msg"):
print("".join(getattr(p, "text", "") or "" for p in populated[0][1].parts))
MCP 工具 - 将 MCP server 连接到任意 LLM(Python SDK + AI Gateway)Python SDK - MCP Bridge
from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from litellm import experimental_mcp_client
import litellm
server_params = StdioServerParameters(command="python", args=["mcp_server.py"])
async with stdio_client(server_params) as (read, write):
async with ClientSession(read, write) as session:
await session.initialize()
# Load MCP tools in OpenAI format
tools = await experimental_mcp_client.load_mcp_tools(session=session, format="openai")
# Use with any LiteLLM model
response = await litellm.acompletion(
model="gpt-4o",
messages=[{"role": "user", "content": "What's 3 + 5?"}],
tools=tools
)
AI Gateway - MCP Gateway
第 1 步。 将你的 MCP Server 添加到 AI Gateway
第 2 步。 通过 /chat/completions 调用 MCP 工具
curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
-H 'Authorization: Bearer <your-master-key>' \
-H 'Content-Type: application/json' \
-d '{
"model": "gpt-4o",
"messages": [{"role": "user", "content": "Summarize the latest open PR"}],
"tools": [{
"type": "mcp",
"server_url": "litellm_proxy/mcp/github",
"server_label": "github_mcp",
"require_approval": "never"
}]
}'
配合 Cursor IDE 使用
{
"mcpServers": {
"LiteLLM": {
"url": "http://localhost:4000/mcp/",
"headers": {
"x-litellm-api-key": "Bearer <your-master-key>"
}
}
}
}
对于 MCP OAuth,上游可能会声明支持动态客户端注册,但仍以 HTTP 401 或 403 拒绝请求。如果该提供商要求使用预先注册的 OAuth 应用,请在 MCP server 上配置其 credentials.client_id,以及在需要时配置 credentials.client_secret。这会跳过 gateway 登录流程中的动态注册。提供商必须批准该应用以访问 MCP;仅能打开其授权页面并不代表登录或工具调用一定会成功
Python SDK - Agent
import litellm
from litellm import Harness, sandbox
result = litellm.agent(
Harness.CLAUDE_CODE, # or Harness.CODEX, Harness.OPENCODE, Harness.DEEPAGENTS
"Find why tests/test_router.py is flaky and fix it.",
sandbox=sandbox.local("./repo"),
model="litellm_proxy/claude-sonnet-4-5", # a model group on your AI Gateway
)
print(result.text, result.cost, [f.path for f in result.files])
设置 LITELLM_PROXY_API_BASE 和 LITELLM_PROXY_API_KEY 后,agent 发出的每一次模型调用都会经由你的 AI Gateway,并带上 harness,claude_code 标签。去掉 litellm_proxy/ 前缀即可直接调用提供商。安装 starlette uvicorn 以及对应 agent 的 CLI(claude、codex 或 opencode),或者为 Deep Agents 安装 deepagents langchain-litellm。





