

3·
7 days agospin up YaCY, see for yourself. It turns into a full time project, because there’s always ranking tweaks and and indexing startpoint changes to be made.
It’s relatively easy to put together a crawler for specific purposes, but a general web search is hard, and requires a lot of storage.
I was in machine learning for a bit, up until venture capitalists eviscerated research teams chasing ways to make money from it. While there’s definitely a degree of overhyping risks for investment, there are absolutely some real, tangible risks. But the dangers aren’t like skynet/terminator, its more of a slow decline in knowledge combined with an increase of predictive surveillance (https://arxiv.org/abs/1902.10739 as an example of why predictive surveillance should worry everyone)
Every mistake and error in token generation is getting confounded with human text, and being fed back as part of the training loop. Combined with the sheer amount of botting of the modern internet, you get a feedback loop of error being used to train, followed by output of that error dominating the next training cycle as bots drown out voices of the knowledgable. For some basic things it’s easy to see the error, but in any expert domain the ratio of bots to experts is large enough that expert knowledge won’t stay within the training distribution.