How do you know your document actually converted correctly?

You upload a document, an AI answers confidently, and you never find out that the chart it should have read was never there. A conversion that looks clean is not the same as a conversion that kept everything. This page is about how we try to make that difference visible — and where we still cannot.

The failure nobody sees

Converted text usually looks fine. That is exactly the problem: the parts that go missing leave no gap behind.

  • A chart carries the conclusion of the whole document. The surrounding paragraphs convert perfectly; the chart does not, and nothing marks its absence.
  • A page in the middle of a report is a scan. The text around it converts; that page comes back empty, or as recognition guesses nobody flagged.
  • A spreadsheet's numbers are formula results. If the file never stored those results, the cells are simply blank.
  • Once the text reaches an AI, none of this is recoverable — the model answers from what it received, not from what your document said.

What we tell you, on every conversion

These appear on your result page automatically. They are measured from your file, not estimated.

  • Images, charts, SmartArt and embedded files that are in your document but not in the Markdown. For Word, PowerPoint and Excel this count is exact — it comes from the file's own internal parts.
  • For PDFs, the same disclosure, stated as "at least N images". Small marks like logos and rules are deliberately not counted, so the real number can be higher — we would rather understate than overstate.
  • Which recognition engine actually read a scanned page, and how confident it was. If our main engine was unavailable and the weaker backup was used, we say so — it is noticeably worse on Chinese, and you deserve to know that before trusting the text.
  • On a scanned page, when the digit-bearing OCR tokens are at least 10 confidence points lower than the surrounding text, we show both figures and ask you to check the original. That is a warning to review, not proof that a number is wrong or correct.
  • Pages that were not converted at all, when a scanned document exceeds our page limit.
  • Spreadsheet formulas whose result was never stored in the file. We show those cells as empty rather than inventing a number.
  • When large tables were shortened by our own limits.
  • Spreadsheet cells that display a rounded number. Excel can show 1,234.57 over a stored 1,234.56789; we follow what is displayed, and we say which cells that applied to so you can go back to the source.
  • When a number from the extracted source is not found exactly in the final Markdown. We give a count, not your values, and ask you to check the original.
  • When we refuse outright — no readable text, password protected, too many pages, workbook too complex — with the reason named. You never get a blank result presented as a success.

What we cannot tell you yet

This section exists because a verification tool that only lists its strengths is not a verification tool. These are real gaps, and we would rather you hear them from us.

  • Whether a number is correct in the original. We now compare extracted-source number tokens against final Markdown and flag ones not found exactly, but an OCR or text extractor can preserve its own mistake; matching tokens also do not prove the figures still belong to the right labels. Numeric OCR confidence remains a review signal, not ground truth.
  • Garbled text. We now flag explicit unknown-character markers (�) in the final Markdown, but that catches only that marker. A broken PDF text layer can still be made entirely of apparently valid characters that we cannot reliably identify; our confidence score also covers only image-recognition pages, not text already in the file.
  • Whether the converted text is correct. We report what was carried across and what was not. We do not judge whether the meaning survived.
  • Layout and reading order in complex documents. We do our best; we do not currently measure how well we did.

Drop a document — or paste a screenshot

PDF, Word, PowerPoint, Excel, HTML, CSV/TSV, TXT, or a screenshot (PNG/JPG — paste with Ctrl/⌘+V) · up to 100 MB

Why we can claim this at all

  • Office files are archives with a fixed internal structure, so we count the pictures and charts from the file's own contents rather than guessing.
  • PDFs have no list of pictures — only drawing instructions — so we follow those instructions and measure how much of the page each image actually covers. Anything too small to be content is treated as decoration and left out of the count.
  • Every one of these behaviours is covered by automated tests that run before anything ships, including tests written specifically to stop the counts from inflating.
  • When a measurement cannot be completed, we say nothing rather than report a zero. "We did not look" and "there was nothing" are different statements.

This page will change

The list of gaps above is the current, honest state — not a permanent limitation. As detection improves, items move from the second list to the first. If you hit something we failed to warn you about, telling us is the fastest way to make that happen.

Questions

Does this mean my charts and images are lost?

The picture itself is not carried into Markdown, which is text. What we do is make sure you know it was there, so you can supply that information another way instead of assuming the AI already has it.

Why not just say the exact number of images for PDFs too?

Because we cannot honestly get one. A PDF describes drawings, not a list of pictures, and a repeated header logo would otherwise show up as dozens of missing images. We exclude small marks, which means our number can be lower than reality but never invented.

Do you check that the converted text is accurate?

No, and we will not claim otherwise. We report what was carried across and what was not. Judging whether the meaning survived is a different problem, and pretending we had solved it would defeat the point of this page.

File retention

Uploaded originals and conversion results are temporary. They are not kept as a permanent document library.

  • Uploaded originals are used only to produce your Markdown result and are deleted as soon as processing ends, or within 60 minutes at the latest.
  • Conversion results and job records are available for up to 60 minutes, then expire and are automatically removed from the active database.
  • We do not keep a permanent document library. Limited provider-managed infrastructure snapshots may retain residual copies for up to 5 days before automatic expiry.
  • Uploaded originals are held in temporary local storage only while the job is processed. Recurring cleanup removes any expired temporary upload left by an interruption.