Most mass scrapers, on the other hand, simply grab the raw HTML underneath. ShieldFont exploits this difference through an automated process called OpenType glyph substitution.
That said, because the whole defense rests on scrapers reading code rather than screens, taking a screenshot of a shielded page and running OCR on the image can still recover the real words.
Screen readers used by blind readers also work from the code, so they read the decoys aloud. ShieldFont ships with a beta feature that provides those readers with the real text instead.



does this also block blind people who depend on an e-reader?
Per the article, yes:
But:
AI scrappers will just pretend to be screen readers then.
And if the approach becomes popular they will just OCR the text instead.
There’s a fair chance they’re already doing that. It provides better insight on the content, less formatting to handle, and even visual stuff gets text alternatives.
It is also very easy to detect surprising text by looking at the perplexity levels by feeding it to a very small model. If there’s a problem, OCR/‘screen read’ it instead…
screen reader, SEO, indexation, in page search, etc.
Basically, it breaks everything except people… unless they block/substitute fonts for accessibility reasons, in which case fuck people too.
This is a terrible idea, and it won’t even achieve it’s original “purpose” as it is trivially detectable. Only negatives in this.
it does, unless the reader has an OCR mode of sorts
Presumably the AI scraper would also have OCR, and would sidestep things like this?
A scraper has many more ways around something like sheildfont than just OCR. The question will be if it was actually programmed to check for such measures.
Yes, the decoy text is marked aria-hidden, so it won’t be read out loud. The real text is sent to the browser encrypted, and the decryption process takes ~20 seconds, roughly the same as running ocr.