Engineering7 min read

A PDF’s Text Layer Has the Right Numbers in the Wrong Order

There are two ways to get the transactions out of a bank statement PDF. You can read the text layer, the characters the PDF actually stores, or you can draw the page and read the picture. Most converters pick one. We use both, and the reason is that each one is right about exactly the thing the other gets wrong.

The text layer is exact about figures. A PDF that says 1,842.17 stores the characters 1, 8, 4, 2, 1 and 7, and no amount of blur or a cheap scanner changes them. What it is not reliable about is order. The characters sit where the PDF put them, and the PDF was written to be printed, not read. A picture is the other way round: the order on the page is the order you see, and every figure in it is a guess by whatever reads it.

How a text layer gets the order wrong

Three things go wrong often enough that we see them every week. Columns run together: a statement that prints the amount and the balance close to each other comes out of the text layer as one long token, so 5,45,047.48 and 500.00 arrive as 5,45,047.48500.00. A line breaks in two: a long description wraps, and the wrapped half lands on its own line with no date, which a parser reads as a new row or drops. And a row’s amount lands on the next row’s line, because the PDF drew the amounts column after the descriptions column and the text extraction stitched them back by height, not by row.

None of these are rare banks or broken PDFs. They are ordinary statements from large banks, produced by the same software that has printed them for twenty years. The text is all there. It is just not in the order a person reads it.

The balance says which page is wrong

Most statements print a running balance, and each balance is the one above it plus the credit or minus the debit. We wrote about using that as an oracle: a reading that adds up from the opening balance to the closing one found every row and put each in the right column. The same arithmetic, done row by row, points at the rows that do not add up, and every row carries the page it came from.

So the first thing a conversion does is count the pages and read the text. If the balances add up, the reading is proved and nothing else runs. That is most statements, and it costs nothing. If they do not, we know which pages hold the rows that broke the chain, and only those pages go any further.

Read the failing page from its picture

A failing page is drawn on the visitor’s own device, cropped from the transaction table’s heading to its last row, and read by an OCR model, which writes the lines in the order they are printed. Cropping is not cosmetic. It means the name, the address and the account number above the table never leave the device; the picture that does is a strip of dates, descriptions and amounts. The PDF itself is still never uploaded.

The OCR text is then read for rows the same way any statement is. What comes back is in the right order, and that is all we trust it for.

The text layer gets a veto over every digit

An OCR model reading a picture will occasionally turn an 8 into a 3. On a statement that is the worst kind of error, because it is a plausible number in a plausible row. So every figure in the rows read from the picture has to appear, character for character, in the same page’s text layer. A digit the picture got wrong is a figure the page does not print, and the whole page is refused.

Then the page’s new rows replace its old ones only if the statement adds up better with them. Not differently: better, measured as fewer rows the balances contradict. A picture reading that is right about order and right about every figure, and still does not help the arithmetic, is thrown away.

Put together, the two sources check each other. The picture fixes the order the text layer scrambled, the text layer catches the digits the picture misread, and the running balance decides whether the result is an improvement. Nothing here asks a model whether it is confident.

Bug one: an honest figure looked invented

The veto compares figures as text, so it has to know every way a figure can be printed. It knew 545047.48, 545,047.48 and the European 545.047,48. It did not know the Indian form, 5,45,047.48, which groups the first three digits and then every two.

On an Indian statement that meant every correct figure the picture read looked like a figure the page did not print, and every page read from its picture was refused. Worse, the same check marks an ordinary model reading as unverified when a figure is missing from the page, so every reading of an Indian statement was flagged and escalated to the most expensive model in the chain, which then failed the same check for the same reason.

The fix was one more printed form. The lesson was that a veto is only as good as its idea of what the page can say. A check that rejects honest input does not look broken; it looks strict.

Bug two: a proof that proved the wrong thing

The second bug went the other way. A 21-page Revolut statement came back proved, every balance accounted for, with its first row, a 1,000 top-up, in the money out column.

A running balance cannot check the first row on its own. The balance printed beside it already includes it, so the chain needs the opening balance from somewhere else. Revolut prints it in a summary table at the top, and the search for it looked three lines below the heading. Revolut puts it five lines down. With no opening found, the opening was worked out backwards from the first row, and a first row worked out backwards fits whichever column you put it in.

Widening the search fixed Revolut. The general rule we took from it is the one from the oracle article, made sharper: a proof covers exactly the rows its arithmetic reaches, and the first row only counts as reached when the opening balance was read off the page rather than derived.

Long statements, a few pages at a time

The same page-level thinking changed how long statements are read. A 200-page statement is not sent to a model in one piece. It is read in chunks of four pages, in parallel, by the cheapest model in the chain. The holder’s details are removed once from the whole text before it is split, and the table heading is copied into every chunk so each one knows its columns.

The chunks are stitched together and checked against the balances. Only a chunk that breaks the chain, or comes back with far fewer rows than the page holds, is read again by the bigger models, and its new reading is kept only if the statement adds up better. On the statements we tested, a 25-page Indian current account was proved for about a cent.

If you are building one of these

Three things we would tell ourselves at the start:

  • Do not choose between the text layer and the picture. The text is the authority on figures, the picture on order, and each can veto the other.
  • Escalate by page, not by document. The balances tell you which page is wrong; reading the other 199 again costs money and can only make them worse.
  • Test every check against honest input from every number format you serve. The failure mode of a strict check is silence, and it will be the most expensive bug you have.

Try it on a statement

Every conversion runs the same pipeline, free rows included. Drop a statement into the converter or export straight to OFX for Xero and Sage, and the preview marks which rows the balances vouch for. The API returns the same per-row verdict alongside the transactions.

Convert a bank statement now

PDF to Excel or CSV in seconds, without ever uploading the file. Free to start, no sign-up required.

Start Converting