Some_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 12 days agoGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comexternal-linkmessage-square187fedilinkarrow-up1778arrow-down120
arrow-up1758arrow-down1external-linkGenerative AI Is an Engineering Disaster. A shockingly inefficient trillion-dollar project.www.theatlantic.comSome_Emo_Chick@lemmy.world to Technology@lemmy.worldEnglish · 12 days agomessage-square187fedilink
minus-squareBrett@programming.devlinkfedilinkEnglisharrow-up4·12 days agoIs that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
minus-squareAsafum@lemmy.worldlinkfedilinkEnglisharrow-up3·12 days agoIt’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
minus-squareDamage@feddit.itlinkfedilinkEnglisharrow-up3·11 days agoMy framework 13 with shared RAM runs qwen quite well
Is that quantized? 4 bit Qwen 3.6 can get 22tps on a 1060.
It’s the q4 quantization, but it requires 20+GB vram and my 5080 only has 16
My framework 13 with shared RAM runs qwen quite well