行业新闻

Databricks 推出 Proteus:自动生成专用 GPU 内核,推理性能提升最高 5.2 倍

Databricks 用智能体自动生成专用 GPU 内核,配合防作弊校验,将 Qwen 3.5 推理速度提升至 vLLM 的 1.8-5.2 倍。

传统推理系统使用通用 GPU 内核,但不同模型和动态请求导致计算形状各异,效率不高。Databricks 构建 Proteus,让智能体自动生成专用内核,并通过严格校验防止优化作弊。实测中,为 Qwen 3.5 122B 生成的内核比 vLLM 快 1.8 至 5.2 倍。

正文摘录

Reliable, validated GPU kernel generation Traditionally, production inference systems rely on generic kernels to handle diverse models and workloads. This is suboptimal because GPU operation shapes are determined by a combination of static model parameters and dynamic request-time factors; for instance, while a model defines one of the dimensions for a matrix multiplication, the other dimension fluctuates based on the specific token count of each request. There is growing interest in agentic GPU kernel generation, and recent efforts have shown promise. We explored a core question: if kernel generation can be automated, why should models of vastly different sizes (from 1 billion to 1 trillion parameters) rely on the same kernel? By specializing kernels to the specific shapes encountered at …

阅读原文(databricks.com)→

行业新闻2026-09-04原文

相关内容