行业新闻

初创公司竞逐 LLM 的下一代技术

Transformer 架构遭遇瓶颈,初创公司押注下一代技术,这决定了未来大模型的能力上限。

Transformer 是当前所有大语言模型的基础架构,靠密集注意力机制处理文本,但计算成本随文本长度急剧上升,上下文窗口也受限。初创公司正尝试用新架构取代它,以支撑更复杂的推理和更大的数据量。

正文摘录

EXECUTIVE SUMMARY MIT Technology Review ’s What’s Next series looks across industries, trends, and technologies to give you a first look at the future. You can read the rest of them [here](https://www.technologyreview.com/tag/whats-next-in-tech/?) . Way back in the summer of 2017, AI researchers at Google put out a paper called [“Attention Is All You Need,”](https://arxiv.org/abs/1706.03762) in which they described a new type of neural network called a transformer. It proved to be very good at processing long sequences of data, especially text. Nine years on, transformers are the engines inside every major large language model on the market. “The entire AI industry is built on transformers,” says Justin Dangel, cofounder and CEO of the AI startup Subquadratic. “They are one of the most imp…

阅读原文(technologyreview.com)→

行业新闻Will Douglas Heaven2026-08-10原文

相关内容