The core llama.cpp maintainers also work at HF and will now work for Nvidia I guess. Llama.cpp is a pretty significant part of the local LLM stack, especially since other tools like Ollama and LMStudio are just GUIs built on top.
I guess that explains why features like quantized KV caches are lagging behind. Maintainers are purposely dragging their feet.
the community may need to step up our game and get our eggs out of the big tech basket.
The community has chosen to not fight at all, which is worse. Anti-AI sentiment is at an all-time high.
Publicly. Privately, these hypocrites still whisper in ChatGPT’s ear when they get lazy enough. Or use some feature in Photoshop or some other software that they didn’t even understand was AI-driven.
guess that explains why features like quantized KV caches are lagging behind
Yeah, I was using the TheTom fork for a while and not sure why TQ KV cache hasn’t merged yet.
Anti-AI sentiment is at an all-time high
I suppose the tech bros have understandably soured a lot of people on LLMs with all of the negative societal costs due to their greedy and irresponsible bejavior. On the other hand, the concepts of the perceptron and artificial neural networks from the 40s and 50s are finally coming to fruition and we have these things now that are legitimately artificially intelligent that we can run on our gaming PCs. They are overhyped and used in ways they make no sense, but they’re also useful as long as their limitations are kept in mind. From a technology perspective, they’re cool as hell.
I guess that explains why features like quantized KV caches are lagging behind. Maintainers are purposely dragging their feet.
The community has chosen to not fight at all, which is worse. Anti-AI sentiment is at an all-time high.
Publicly. Privately, these hypocrites still whisper in ChatGPT’s ear when they get lazy enough. Or use some feature in Photoshop or some other software that they didn’t even understand was AI-driven.
Yeah, I was using the TheTom fork for a while and not sure why TQ KV cache hasn’t merged yet.
I suppose the tech bros have understandably soured a lot of people on LLMs with all of the negative societal costs due to their greedy and irresponsible bejavior. On the other hand, the concepts of the perceptron and artificial neural networks from the 40s and 50s are finally coming to fruition and we have these things now that are legitimately artificially intelligent that we can run on our gaming PCs. They are overhyped and used in ways they make no sense, but they’re also useful as long as their limitations are kept in mind. From a technology perspective, they’re cool as hell.