Before uploading a PDF anywhere, it's worth two minutes to find out whether there's anything inside it to recover. Three kinds of PDF look identical on screen and behave completely differently when you try to rebuild them. Here is how to tell them apart with nothing but a PDF viewer.
Fine. Exported from a design or office application with its objects intact: text is text, shapes are vectors, photos are embedded images. Everything can be read back out. This is most PDFs that came straight from InDesign, Illustrator, PowerPoint, Word or a browser's "save as PDF".
Flattened. Started life as objects, but somewhere along the way — usually at "print-ready" export with transparency flattening, or by being passed through a preflight tool — parts of the page were rasterised into image slices. Text is often still real; effects, gradients, overlapping transparent objects and sometimes whole regions are pictures. Recoverable, with gaps.
Scanned. A photograph of a page, or a screenshot saved as PDF, or a fax. One big image per page, sometimes with an invisible OCR text layer laid on top. There are no shapes and no real layout to recover. No converter can rebuild what was never there.
Open the PDF in any viewer. Press the text-selection tool and drag across a line of body text.
This single test predicts the fidelity score better than anything else. Selectable text is structure; structure is what gets rebuilt.
Zoom to 800% on a curve — the edge of a logo, a rounded corner, a letter. Vectors stay razor sharp at any zoom. Images go soft or blocky. If a logo blurs while the body text beside it stays crisp, the logo is a placed image and will come back as an image, not as editable paths. That's fine to know in advance; it's the difference between "the logo is a picture in the rebuilt file" and "the logo is missing".
Most viewers show document properties (File › Properties, or ⌘I). The Producer line often says how the PDF was made. "Adobe InDesign", "Adobe PDF Library", "Microsoft: Print To PDF", "Skia/PDF" (Chrome) all suggest an object export. A scanner model name, "ABBYY", "Tesseract" or a phone-app name means a scan with OCR. "Ghostscript" or a print-driver name can mean anything, so rely on the cursor test.
File size helps too. A 40-page text document at 60 MB is almost certainly scanned images. The same document exported from Word is under 2 MB.
| Text | Vectors and shapes | Images | Fidelity score | |
|---|---|---|---|---|
| Fine | live, editable | editable paths | placed, editable frames | usually 80+ |
| Flattened | mostly live | partly pictures | placed | 45–80, varies by page |
| Scanned | none | none | one per page | declined |
If a file is scanned and you need it editable, the honest route is OCR for the words plus a redraw of the layout. The Sourceback preview will decline the pages rather than charge for an empty result — but it's quicker to know before you upload.
The print-ready file is usually the most flattened version of a document that exists. Ask for the proof PDF, the "screen" or "web" PDF, or the file that was sent to the client for approval. They are the same design exported with fewer flattening steps and often recover almost completely.