Most mass scrapers, on the other hand, simply grab the raw HTML underneath. ShieldFont exploits this difference through an automated process called OpenType glyph substitution.

That said, because the whole defense rests on scrapers reading code rather than screens, taking a screenshot of a shielded page and running OCR on the image can still recover the real words.

Screen readers used by blind readers also work from the code, so they read the decoys aloud. ShieldFont ships with a beta feature that provides those readers with the real text instead.

  • FTonsilStones@lemmy.caOP
    link
    fedilink
    English
    arrow-up
    76
    arrow-down
    1
    ·
    6 days ago

    Per the article, yes:

    Screen readers used by blind readers also work from the code, so they read the decoys aloud.

    But:

    ShieldFont ships with a beta feature that provides those readers with the real text instead.

    • ViatorOmnium@piefed.social
      link
      fedilink
      English
      arrow-up
      79
      ·
      6 days ago

      ShieldFont ships with a beta feature that provides those readers with the real text instead.

      AI scrappers will just pretend to be screen readers then.

      And if the approach becomes popular they will just OCR the text instead.

      • cley_faye@lemmy.world
        link
        fedilink
        English
        arrow-up
        9
        ·
        5 days ago

        AI scrappers will just pretend to be screen readers then

        There’s a fair chance they’re already doing that. It provides better insight on the content, less formatting to handle, and even visual stuff gets text alternatives.

      • Sims@lemmy.ml
        link
        fedilink
        English
        arrow-up
        5
        ·
        6 days ago

        It is also very easy to detect surprising text by looking at the perplexity levels by feeding it to a very small model. If there’s a problem, OCR/‘screen read’ it instead…