Invisible text layer over a scan
A page whose text is essentially all invisible over a page image — the ordinary searchable-scan construction, reported so the record is complete rather than because anything is wrong.
How the text is hidden
A scanner or OCR tool lays the recognised text under the page image in the invisible text rendering mode, so the page stays a picture and the text stays selectable. Mechanically this is identical to the invisible-mode concealment technique; what separates them is proportion. A page carrying visible text that ALSO carries invisible text is the attack; a page whose text is essentially all invisible is a scan.
Why a model still reads it
The layer is real, correct text and every extractor reads it — which is the point of a searchable scan, not a defect.
What we do about it
The same renderModeInvisible flag the invisible-mode rule reads. pdf-ocr-text-layer takes the other side of the split: it requires container.visibleTextRatio at or below the share written into the rule's own predicate in the pack (a ratio of non-whitespace characters, with no length floor, because a PDF container is a page), plus absentFromRender == false. It fires at informational, action flag — it exists so the record is complete without crying wolf on every OCR'd file.
How often it fires
0% of 400 real UK public-sector contract PDFs (Contracts Finder, 2019–2023) — Word, Adobe, Nitro, office copiers, measured 2026-08-23.
This is an alert-volume number and nothing else. It says how often the alarm sounds on documents as found — not how often it is right, and not whether what it found was harmless. Documents as found may themselves carry concealment. Read it against the population named above rather than as a property of documents in general.
Check your own file
Three commands: a key, credit, a verdict.
Start with the API