Coinbase 通过使用便宜推理模型和智能路由将 token 支出减半
This is very interesting. Coinbase seems to have lowered their token spend ($$) to about half, by
1) routing to cheap inference like GLM 5.2 and Kimi 2.7 that are still pretty performant
2) Smart routing + caching
They still use the same tokens as before. Start of a trend?
1) routing to cheap inference like GLM 5.2 and Kimi 2.7 that are still pretty performant
2) Smart routing + caching
They still use the same tokens as before. Start of a trend?
How to keep AI spend flat while token usage grows exponentially: Not with friction and spend alerts. With better defaults, routing, and caching.
Better Defaults (not Usage Caps) – Engineers can choose any model they want, but defaults matter. We’re experimenting with defaulting https://t.co/PUV9uQHGO0