Databricks 把向量检索做成 Runtime 原生 SQL join
Databricks 将向量检索改造成 Runtime 内的原生 top-k join,配 Photon 内核与 IVF 索引,面向百万级批处理而非在线低延迟。
向量检索一般按在线服务优化:一次查询、几十毫秒返回 top-k。但 Databricks 上很多负载是批处理——用上亿条记录按天匹配上亿个向量,只看作业能否在 SLA 内跑完。他们因此把向量检索做进 Runtime 引擎,用 NEAREST BY 这个一等 SQL join 暴露出来。
正文摘录
How we built vector search into Databricks as a first-class SQL join, with deep kernel optimizations in Photon and a vector index in an open storage format. by Zero Qu (Zero) , Alexis Schlomer , Akash Nayar , Yingyi Bu and Sergei Tsarev Vector search originated as a serving problem. The classical use case is a chatbot or a search bar: one query embedding arrives, and the system is optimized to return the top-k nearest documents within tens of milliseconds. However, a good share of the vector search workloads on our platform are inherently batch-oriented — precomputing exact or approximate nearest neighbors offline rather than looking them up at request time. A payments company matches 100M+ daily transactions against 140M merchant embeddings for entity resolution; a data firm enriches tens…