开源项目

litellm

litellm

统一接入 100+ LLM 的 AI Gateway 与 Python SDK,用 OpenAI 格式调用 OpenAI、Anthropic、Bedrock、Azure、vLLM 等模型,内置虚拟 key、花费追踪、限流、guardrails 和负载均衡。Rust 内核重写后标称 1k RPS 下 P95 约 8ms,支持自托管,Stripe、Netflix 等已在生产使用;还顺带提供 A2A agent 与 MCP 网关,方便把 agent 和工具挂到同一入口。注意企业级功能走 LiteLLM Commercial License。

README

🚅 LiteLLM

LiteLLM AI Gateway

面向 100+ LLM 的开源 AI Gateway。可自托管。企业级就绪。以 OpenAI 格式调用任意 LLM。

Deploy to Render Deploy on Railway Deploy on AWS Deploy on GCP

LiteLLM Proxy Server(AI Gateway) | 托管 Proxy | 企业版 | 官网
PyPI Version GitHub Stars Y Combinator W23 Whatsapp Discord Slack CodSpeed
LiteLLM AI Gateway

什么是 LiteLLM

LiteLLM 是一个开源 AI Gateway,提供单一、统一的接口,让你以 OpenAI 格式调用 100+ LLM 提供商 —— OpenAI、Anthropic、Gemini、Bedrock、Azure 等。

你可以把它当作 Python SDK 直接集成到代码中,也可以部署 AI Gateway(Proxy Server) 作为团队或组织的集中式服务。

跳转至 LiteLLM Proxy(LLM Gateway)文档
跳转至受支持的 LLM 提供商


为什么选择 LiteLLM

跨提供商管理 LLM 调用很快就会变得复杂 —— 每个模型都有不同的 SDK、认证模式、请求格式和错误类型。LiteLLM 消除了这些摩擦:

  • 统一 API —— 一套接口覆盖 100+ LLM,无需在各类提供商 SDK 之间周旋
  • 开箱即用的 OpenAI 兼容 —— 无需重写代码即可切换提供商
  • 生产级 gateway —— 内置 virtual key、成本追踪、guardrail、负载均衡和管理面板
  • 在 1k RPS 下 P95 延迟为 8ms(benchmark)

OSS 采用者

Stripe image Google ADK Greptile OpenHands

Netflix

OpenAI Agents SDK

功能特性

LLM - 调用 100+ LLM(Python SDK + AI Gateway)

所有支持的 endpoint - /chat/completions、/responses、/embeddings、/images、/audio、/batches、/rerank、/a2a、/messages 等。

Python SDK

uv add litellm

独立的 litellm-core 发行版提供与 Python SDK 相同的 import litellm API 和运行时依赖。它不包含可选扩展、CLI 入口点 或内置 dashboard。每个环境只能安装一种 SDK 发行版,因为 litellm 和 litellm-core 存在重叠的 Python 文件。如果你需要 proxy、CLI 和可选扩展,请使用 litellm

在其发布集成尚在进行期间,可从此检出目录构建并安装 core:

python scripts/build_core_distribution.py --out-dir dist/core
python -m pip install dist/core/litellm_core-*.whl

该构建脚本需要 Git、uv 和 Rust 构建工具链。它从根目录的 pyproject.toml 读取发布版本号,不修改源文件,并在 dist/core 中生成一个 wheel 和一份自包含的 sdist。请在一个未安装 litellm 的干净环境中执行安装

from litellm import completion
import os

os.environ["OPENAI_API_KEY"] = "your-openai-key"
os.environ["ANTHROPIC_API_KEY"] = "your-anthropic-key"

# OpenAI
response = completion(model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}])

# Anthropic  
response = completion(model="anthropic/claude-sonnet-4-20250514", messages=[{"role": "user", "content": "Hello!"}])

AI Gateway(Proxy Server)

快速上手 - 端到端教程 - 配置 virtual key,发出第一个请求

uv tool install 'litellm[proxy]'
litellm --model gpt-4o
import openai

client = openai.OpenAI(api_key="anything", base_url="http://0.0.0.0:4000")
response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Hello!"}]
)

文档:LLM 提供商

Agent - 调用 A2A Agent(Python SDK + AI Gateway)

受支持的提供商 - LangGraph、Vertex AI Agent Engine、Azure AI Foundry、Bedrock AgentCore、Pydantic AI

Python SDK - A2A 协议

from litellm.a2a_protocol import A2AClient
from a2a.types import SendMessageRequest, MessageSendParams
from uuid import uuid4

client = A2AClient(base_url="http://localhost:10001")

request = SendMessageRequest(
    id=str(uuid4()),
    params=MessageSendParams(
        message={
            "role": "user",
            "parts": [{"kind": "text", "text": "Hello!"}],
            "messageId": uuid4().hex,
        }
    )
)
response = await client.send_message(request)

AI Gateway(Proxy Server)

第 1 步。 将你的 Agent 添加到 AI Gateway —— 为每个 agent 将 protocolVersion 设为 1.0 或 0.3

第 2 步。 通过 A2A SDK 调用 Agent(需要 a2a-sdk>=1.1.0)

import httpx
from a2a.client import A2ACardResolver, ClientConfig, ClientFactory
from a2a.types import Message, Part, Role, SendMessageRequest
from a2a.utils.constants import TransportProtocol
from uuid import uuid4

base_url = "http://localhost:4000/a2a/my-agent"  # LiteLLM proxy + agent name
headers = {"Authorization": "Bearer <your-master-key>"}    # LiteLLM master key or a virtual key

async with httpx.AsyncClient(headers=headers, timeout=60.0) as http_client:
    resolver = A2ACardResolver(httpx_client=http_client, base_url=base_url)
    agent_card = await resolver.get_agent_card()
    config = ClientConfig(
        httpx_client=http_client,
        streaming=False,
        supported_protocol_bindings=[TransportProtocol.JSONRPC, TransportProtocol.HTTP_JSON],
    )
    client = ClientFactory(config).create(agent_card)

    request = SendMessageRequest(
        message=Message(
            message_id=uuid4().hex,
            role=Role.ROLE_USER,
            parts=[Part(text="Hello!")],
        )
    )
    async for event in client.send_message(request):
        populated = event.ListFields()
        if populated and populated[0][0].name in ("message", "msg"):
            print("".join(getattr(p, "text", "") or "" for p in populated[0][1].parts))

文档:A2A Agent Gateway

MCP 工具 - 将 MCP server 连接到任意 LLM(Python SDK + AI Gateway)

Python SDK - MCP Bridge

from mcp import ClientSession, StdioServerParameters
from mcp.client.stdio import stdio_client
from litellm import experimental_mcp_client
import litellm

server_params = StdioServerParameters(command="python", args=["mcp_server.py"])

async with stdio_client(server_params) as (read, write):
    async with ClientSession(read, write) as session:
        await session.initialize()

        # Load MCP tools in OpenAI format
        tools = await experimental_mcp_client.load_mcp_tools(session=session, format="openai")

        # Use with any LiteLLM model
        response = await litellm.acompletion(
            model="gpt-4o",
            messages=[{"role": "user", "content": "What's 3 + 5?"}],
            tools=tools
        )

AI Gateway - MCP Gateway

第 1 步。 将你的 MCP Server 添加到 AI Gateway

第 2 步。 通过 /chat/completions 调用 MCP 工具

curl -X POST 'http://0.0.0.0:4000/v1/chat/completions' \
  -H 'Authorization: Bearer <your-master-key>' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "gpt-4o",
    "messages": [{"role": "user", "content": "Summarize the latest open PR"}],
    "tools": [{
      "type": "mcp",
      "server_url": "litellm_proxy/mcp/github",
      "server_label": "github_mcp",
      "require_approval": "never"
    }]
  }'

配合 Cursor IDE 使用

{
  "mcpServers": {
    "LiteLLM": {
      "url": "http://localhost:4000/mcp/",
      "headers": {
        "x-litellm-api-key": "Bearer <your-master-key>"
      }
    }
  }
}

对于 MCP OAuth,上游可能会声明支持动态客户端注册,但仍以 HTTP 401 或 403 拒绝请求。如果该提供商要求使用预先注册的 OAuth 应用,请在 MCP server 上配置其 credentials.client_id,以及在需要时配置 credentials.client_secret。这会跳过 gateway 登录流程中的动态注册。提供商必须批准该应用以访问 MCP;仅能打开其授权页面并不代表登录或工具调用一定会成功

文档:MCP Gateway

Agent - 在任意模型上运行 Claude Code、Codex、OpenCode 或 Deep Agents(Python SDK)

Python SDK - Agent

import litellm
from litellm import Harness, sandbox

result = litellm.agent(
    Harness.CLAUDE_CODE,  # or Harness.CODEX, Harness.OPENCODE, Harness.DEEPAGENTS
    "Find why tests/test_router.py is flaky and fix it.",
    sandbox=sandbox.local("./repo"),
    model="litellm_proxy/claude-sonnet-4-5",  # a model group on your AI Gateway
)

print(result.text, result.cost, [f.path for f in result.files])

设置 LITELLM_PROXY_API_BASE 和 LITELLM_PROXY_API_KEY 后,agent 发出的每一次模型调用都会经由你的 AI Gateway,并带上 harness,claude_code 标签。去掉 litellm_proxy/ 前缀即可直接调用提供商。安装 starlette uvicorn 以及对应 agent 的 CLI(claude、codex 或 opencode),或者为 Deep Agents 安装 deepagents langchain-litellm。

文档:Agent Harness

受支持的提供商(官网支持模型 | 文档)

提供商 /chat/completions /messages /responses /embeddings /image/generations /audio/transcriptions /audio/speech /moderations /batches /rerank
Abliteration (abliteration) ✅
AI/ML API (aiml) ✅ ✅ ✅ ✅ ✅
AI21 (ai21) ✅ ✅ ✅
AI21 Chat (ai21_chat) ✅ ✅ ✅
Aleph Alpha ✅ ✅ ✅
Amazon Nova ✅ ✅ ✅
Anthropic (anthropic) ✅ ✅ ✅ ✅
Anthropic Text (anthropic_text) ✅ ✅ ✅ ✅
Anyscale ✅ ✅ ✅
AssemblyAI (assemblyai) ✅ ✅ ✅ ✅
Auto Router (auto_router) ✅ ✅ ✅
AWS - Bedrock (bedrock) ✅ ✅ ✅ ✅ ✅
AWS - Sagemaker (sagemaker) ✅ ✅ ✅ ✅
Azure (azure) ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
Azure AI (azure_ai) ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅
Azure Text (azure_text) ✅ ✅ ✅ ✅ ✅ ✅ ✅
Baseten (baseten) ✅ ✅ ✅
Bytez (bytez) ✅ ✅ ✅
Cerebras (cerebras) ✅ ✅ ✅
Clarifai (clarifai) ✅ ✅ ✅
Cloudflare AI Workers (cloudflare) ✅ ✅ ✅
Codestral (codestral) ✅ ✅ ✅
Cognition (cognition) ✅ ✅ ✅
Cohere (cohere) ✅ ✅ ✅ ✅ ✅
Cohere Chat (cohere_chat) ✅ ✅ ✅
CometAPI (cometapi) ✅ ✅ ✅ ✅
CompactifAI (compactifai) ✅ ✅ ✅
Custom (custom) ✅ ✅ ✅
Custom OpenAI (custom_openai) ✅ ✅ ✅ ✅ ✅ ✅ ✅
Dashscope (dashscope) ✅ ✅ ✅ ✅ ✅
Databricks (databricks) ✅ ✅ ✅
DataRobot (datarobot) ✅ ✅ ✅
Deepgram (deepgram) ✅ ✅ ✅ ✅
DeepInfra (deepinfra) ✅ ✅ ✅
Deepseek (deepseek) ✅ ✅ ✅
Eden AI (edenai) ✅ ✅ ✅ ✅ ✅ ✅ ✅
ElevenLabs (elevenlabs) ✅ ✅ ✅ ✅ ✅
Empower (empower) ✅ ✅ ✅
Fal AI (fal_ai) ✅ ✅ ✅ ✅
Featherless AI (featherless_ai) ✅ ✅ ✅
Fireworks AI (fireworks_ai) ✅ ✅ ✅
FriendliAI (friendliai) ✅ ✅ ✅
Galadriel (galadriel) ✅ ✅ ✅
GitHub Copilot (github_copilot) ✅ ✅ ✅ ✅
GitHub Models (github) ✅ ✅ ✅
Google - PaLM ✅ ✅ ✅
Google - Vertex AI (vertex_ai) ✅ ✅ ✅ ✅ ✅
Google AI Studio - Gemini (gemini) ✅ ✅ ✅
GradientAI (gradient_ai) ✅ ✅ ✅
Groq AI (groq) ✅ ✅ ✅
Heroku (heroku) ✅ ✅ ✅
Hosted VLLM (hosted_vllm) ✅ ✅ ✅
Huggingface (huggingface) ✅ ✅ ✅ ✅ ✅
Hyperbolic (hyperbolic) ✅ ✅ ✅
IBM - Watsonx.ai (watsonx) ✅ ✅ ✅ ✅
Infinity (infinity) ✅
Jina AI (jina_ai) ✅
Lambda AI (lambda_ai) ✅ ✅ ✅
Lemonade (lemonade) ✅ ✅ ✅
LiteLLM Proxy (litellm_proxy) ✅ ✅ ✅ ✅ ✅
Llamafile (llamafile) ✅ ✅ ✅
LM Studio (lm_studio) ✅ ✅ ✅
Maritalk (maritalk) ✅ ✅ ✅
Meta - Llama API (meta_llama) ✅ ✅ ✅
Mistral AI API (mistral) ✅ ✅ ✅ ✅
ModelScope (modelscope) ✅ ✅ ✅ ✅
Moonshot (moonshot) ✅ ✅ ✅
Morph (morph) ✅ ✅ ✅
Nebius AI Studio (nebius) ✅ ✅ ✅ ✅
NLP Cloud (nlp_cloud) ✅ ✅ ✅
Novita AI (novita) ✅ ✅ ✅
开源项目BerriAI2026-10-09原文

相关内容