热门产品

space ocr

space ocr

这是一款能自动校验识别结果的 OCR 工具,可将发票、收据和表单照片转为可查询表格,适合需要批量录入票据数据的个人和开发者使用。

热门评论

PH 用户
Hi Product Hunt,

I made space ocr because I kept not trusting OCR output.

Reading a document is the easy part now. Knowing whether the number you got back is the number actually printed on the paper is not. If you still have to open the image and check by hand, you haven't really automated anything.

There are two ways in, and both run the same pipeline.

If you don't want to write code, you upload photos into a folder and they become a sheet. Hover any cell and the photo beside it lights up on the exact spot that value was read from, zoomed in, so checking a page takes a second instead of a squint. Cells that failed the check are marked, so you know which ones to look at rather than rereading all of them. Fix a value by hand and your correction sticks. Folders, memos and search across everything you have scanned are in there too.

If you do write code, three endpoints give you structured fields, markdown, or plain text, and all of them come back with the same verification data: where each value sits on the page, whether it passed the check, and what still needs a look.

Either way the results stay somewhere you can use, so there is no database to stand up. A folder and a sheet are the storage. Photos land in the sheet as rows of the columns you asked for, and later you can ask that sheet for the rows over an amount, or from one vendor, newest first, a page at a time. That runs on the server, it does not read the images again, and it is not charged. The rows keep the coordinates and the flags they were stored with, so a filtered answer is as checkable as a single scan.

If you would rather have an agent do the filing, there is a hosted MCP server on the same account. You point an MCP client at one URL with your key and it can make the folders and sheets, upload photos into them, and ask for rows later. Deleting is the one thing it cannot do in one step. The first call removes nothing and reports what would go, so it has to come back to you before anything disappears.

The checking itself is the part I care about. The model never produces coordinates. Every value it returns is matched character by character against what the OCR engine actually saw on the page. Values that fail get flagged instead of quietly passing, and the ones it still isn't sure about are cropped out of the image and read a second time.

I measured this on my own regression corpus, 333 hand graded cells from phone photos rather than flat scans. Turning the checking stages off drops accuracy from 93.7% to 91.3%. They fixed 22 cells and broke none. A value marked unverified turns out to be wrong 6.4 times more often than average, so the flag is worth acting on.

100 pages a month are free and failed scans are never billed. Same price whichever way you use it.

What I would really like to hear: what would make you trust OCR output enough to skip the manual check? That is the part I keep getting wrong.

Yongha
PH 用户
Asad's recall question got the most useful answer on this page, and I don't think the 6-of-39 is a tuning problem.

If the re-read runs the same model over the same pixels, it inherits the same failure mode. A 7 that got read as a 1 because the glyph is genuinely ambiguous will get read as a 1 again — the second pass isn't independent, so it can only catch noise, not systematic misreads. That would explain a recall floor that prompt work won't move.

The cheapest independent signal on invoices isn't another model, it's arithmetic. Line items × qty should reconcile to the subtotal, and subtotal + tax to the total. When the sum doesn't close, you know at least one cell is wrong without trusting any model to tell you. I spent 19 years building banking apps and that's what we called a control total — it isn't clever, it just doesn't share the OCR's blind spot.

Do you reconcile totals already, or is the check purely model-vs-model today?
PH 用户
@yonghahwang The page as a row model is clean for receipts, but a lot of invoices run three or four pages, with line items continuing past the break and the total only on the last page. Does a multi page PDF land as separate rows I'd have to stitch back together, or can the document be the unit with pages underneath it?
PH 用户
The 333 cell corpus and the 3 in 4 flag precision are unusually honest for a launch page. The number I'd want next is the other direction, of the cells that were actually wrong, how many did the check miss. Precision tells me the flags are worth reading, recall tells me whether I can skip the unflagged rows, and skipping is the whole product. If recall is weak then a self check is worse than no check, because it's the thing that stops people looking.
PH 用户
the provenance-per-value thing is the part I'd actually pay for. most OCR tools give you a confident-looking number and no way to tell if it read the receipt correctly or just guessed something plausible from a smudge. what happens when the self-check disagrees with the first pass - does it flag the cell as low-confidence for a human to glance at, or silently pick whichever answer scored higher internally? for invoices specifically the failure mode that costs money is a confident wrong number, not a missing one.
热门产品Yongha Hwang2026-08-04原文

相关内容