热门产品

oMLX

oMLX

oMLX是一款Mac上的LLM推理服务器,通过连续批处理和KV缓存加速,让Claude Code和Cursor等AI编码工具的响应时间从90秒降至5秒。

热门评论

PH 用户
Kept hitting the same wall with local models: the agent loops back, the whole conversation recomputes, 90 seconds of nothing.

oMLX writes the KV cache to SSD. Old context comes back in milliseconds, even after a restart. Claude Code on a local model stops feeling like dial-up.

Jun has been shipping this almost daily since February. 21k stars, and his Show HN still barely got seen. Felt wrong.

If you run models on a Mac: what does your stack look like? Curious what people pair this with.
PH 用户
Recently got my Mac that has enough RAM to run local LLMs and honestly – oMLX was the best solution so far to optimize for context and speed at the same time on Apple silicon. Also huge shoutout to the devs, who are constantly shipping updates at a crazy pace. Love to see it on PH, thanks for hunting!
PH 用户
Cache surviving a restart is the detail I'd have skipped and then regretted. How much SSD does the tiered cache use in practice?
PH 用户
The number that would sell me isn't 90 to 5, it's what happens when the cache is wrong. A KV cache that survives a restart also survives me swapping the model or editing the system prompt, and a stale prefix doesn't crash, it just answers a slightly different question than the one on screen. If the cache key includes the model hash and the full prefix, put that on the page, because that's the thing a dev has to trust before leaving it running for a week. Tiered RAM plus SSD is the right shape though, most local setups throw the whole thing away and pretend prefill is free.
PH 用户
my stack right now is just LM Studio for casual local chat, which never bothered me because that's a one-off question and answer. the thing that actually annoyed me was agent coding loops - watching Claude Code re-read the whole conversation on every follow-up turn. that's a different use case than most "run a model locally" tools are built for, so the KV cache surviving a restart is the part that's actually interesting here, more than raw tokens/sec.
热门产品Rabnoor Singh2026-08-30原文

相关内容