单卡 H200 训练 3500 万参数 BarunLM 自称最优小模型
3500 万参数小模型用单张 H200 训练,在多项基准上超过数倍于己的模型,其架构和效率值得关注。
独立开发者 Harshal Singh 发布 3500 万参数基础语言模型 BarunLM,据称用一张 H200 GPU 完成预训练,在九项零样本评测中平均准确率 41.01%,超过参数量为 2.3 亿的 Liquid AI LFM2.5 以及 GPT-2 等更大模型。模型采用局部与全局注意力交替设计,降低计算开销,权重已开源。
正文摘录
- title: Using a Single Card to Create the Best Small Model Claimed to Be Under 100M Parameters? - sourcecompany: 机器之心 - bodymarkdown: .jpg) There is a rather retro sense of joy in the AI community: small parameter models beating opponents far larger than themselves. Last week, a developer named slvDev packed a language model with 28.9 million parameters into the ESP32-S3 microcontroller, which sells for about $8. The whole board has only 512KB SRAM, 8MB PSRAM, and 16MB flash memory. Without needing internet access, the model can generate text directly on the chip at approximately 9.5 tokens/second. Of course, this model was only trained on TinyStories. It can't write code, can't call too…