With the frustrations from using the usual search engines growing I’ve been thinking about this a lot lately.

It seems we’ve poisoned the well by allowing the proliferation of advertising interests to dominate the web.

Like how hard would it be to make your own non-commercial index?

Only human made sites that aren’t related to buying, selling, marketing, etc.

Could that be a federated open-source project?

  • moonshine69@lemmy.nz
    link
    fedilink
    arrow-up
    5
    ·
    7 days ago

    Oops, realized I didn’t answer your question about actually crawling, dig into common crawl documentation, they provide a bunch of technical data and stats that show you the scale…2-4billion pages per month

    And note CC just does a sample of the pages it finds. So the more monthly dumps don’t contain all of the data afaik

    And the number above are for one of the monthly dumps

    https://commoncrawl.github.io/cc-crawl-statistics/plots/crawlermetrics