Hacker Newsnew | past | comments | ask | show | jobs | submit | p0u4a's commentslogin

> if software knows in advance which data chunk (expert) it'll need for the next token, it can load that in parallel with computing current token

You could actually use the model's MTP head to make a ~decent prediction on what experts would be activate in future tokens and preload them


more of stuff like this on the web please


I have no idea what's going on but it looks really cool


How much does training such a model cost?


Sadly they don't teach Show Don't Tell in the lectures


Aaaand we’re back


Edit: I was hosting this on Vercel free tier, and since this post made it to the front page, I've hit Vercel's quota limit. I will work on getting around this.


how about static hosting?


I was on Vercel free tier and hit the quota for blob storage. Will probably make it open source so people can self host


Hmm, yeah I suppose you're right. I'm currently using the algolia api https://hn.algolia.com/

Will look into tightening this up


yeah that would be awesome, will land it tomorrow


It doesn’t seem to work anymore, it’s stuck in the same set of text


Looking forward to that, I have added it to my home screen :thumbsup:


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: