Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Keep in mind also that LLMs are currenctly heavily subsidized. Once VCs money are gonna run out, you will see the real price, and your 200K per year dev is probably nothing.


A $1200 GPU is all a reasonably capable engineer needs for each work stream. The cloud prices are a scam.


hmm that's a bit low right now, since you need at least 48Gb of VRAM to be confortable. 32Gb might do it (on quantitized models + optimized) but it would be very slow.


That may have been true a few months ago, but shit changes fast in this space!

I run Qwen 3.8 27b on each of my six AMD AI PRO r9700 GPUs at 80tps decode each and they cranks for days with 256k context, doing complex kernel, compiler, debugging, enclave, bootstrapping, pentesting, hardening, and systems work full time.

~$1200-1500/ea on ebay.

Also ~40tps on my strix halo now but with room for several sessions at once.


What is your config for getting 80tps decode?


Qwen 3.8 27b Q4_K_M w/ matching dflash drafter, 256k context, llama.cpp+dflash2, draft-n-max=5

Also using with Charmbracelet Crush with Froggerinc fixed chat templates.


Thank you! Have to look into it.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: