热门产品

Inferock Bench

Inferock Bench

为开发者提供一个本地代理,捕获每次 LLM API 调用的 token 用量和费用,生成独立收据,帮助识别过度支出并优化成本。

热门评论

PH 用户
This solves a problem I didn't know I could solve, I always assumed API billing discrepancies were just the cost of doing business. Turns out I was wrong.
PH 用户
Overpaying for failed calls has quietly cost me more than I'd like to admit. I've had timeouts that still consumed tokens on the provider's side and no clean way to catch that pattern until my bill arrived. If this surfaces failure-related overspend specifically, not just total usage, I'd trust the numbers more than any provider dashboard.
PH 用户
Hey PH

We built inferock-bench because we kept paying for AI answers that died mid sentence, and nobody could tell us where the money went.

Providers give you totals. They don't give you the per call receipt you'd need to prove which answer broke, which retry ran, or which token count changed. The company that charges you also decides what counts as a failure and keeps the only detailed records.

inferock-bench runs locally as a proxy in front of OpenAI, Anthropic, Gemini, OpenRouter shaped calls. Point your existing SDK at it (change two settings: apiKey and baseURL), and it captures every call as an independent, per call record. Your provider key never touches our servers, it's used locally only, attached to provider requests.

What it catches:
- Answers cut off mid stream that still got billed
- Empty replies with billed tokens attached
- Token counts that don't match visible output
- Retries that may have silently doubled a charge
- Cache discounts you may be missing on your invoice

Every run reports a receipt: spend observed, bill-bounded money loss, time loss, and a separate "invoice-check exposure" line that never gets summed into money loss, because we don't want a louder headline at the cost of a weaker claim.

Run it in about a minute: npx inferock-bench

Its open source (FSL-1.1-Apache-2.0, converts to Apache-2.0 in 2 years).

Question for this community: has anyone here actually disputed an AI provider bill and gotten a credit? What worked?
PH 用户
The failed-calls and retries breakdown is the part I'd use first. One case I keep hitting might not show up there: a tool call that returns 200 with a silently corrupted value. I measured this on Anthropic models, 40 calls, none flagged the value was wrong, so it bills as a clean success and the retry logic never fires. Can the receipt catch a call that looked fine but wasn't? Or is that out of scope by design?
PH 用户
Congrats on the launch @himashwetha_gowda. Good find @fmerian.

Question regarding accuracy, is it possible it might not be able to distinguish a genuinely billable provider failure from valid hidden token usage, such as reasoning, refusal, cache, or tool-call tokens?
PH 用户
我之前就被静默重试坑过,OpenAI 账单悄悄变高。如果当时有独立的消费凭证,就能省掉一场痛苦的账单沟通了。
热门产品Hamza Afzal Butt2026-08-15原文

相关内容