Hacker Newsnew | past | comments | ask | show | jobs | submit | theredsix's commentslogin

Congrats on the launch! What's different between Jev and Microsoft's Guidance package? https://github.com/guidance-ai/guidance Is it a diffusion generator under the hood?

Thought I was going crazy, glad to know I wasn't the only one.


If you're getting fired in 2 weeks, it's likely not because of a single comment.


If you’re ever gonna get fired for a single comment, it would be at 2 weeks. I have seen people get fired that early on for a single poor-taste comment to the wrong person. It’s wild, but at 2 weeks you’re completely expendable.


this was literally the reason pointed out in chat. devs liked how I worked, had zero questions on review, and that's despite me seeing their tech stack second time in my entire life


10% cheaper than official! Let the inference pricing wars begin!


I'm really curious how far and how fast prices will drop (if at all)!


Genius I love it!


Are you guys going to follow up with a paper showing EDEN results match or beat turboquant for needle in a haystack benchmarks?


The note includes extensive experiments and reproduces many of the figures from the TurboQuant paper in our Section 5. Honestly, I think our case is pretty clear-cut as is. I am not sure what the overhead for those specific benchmarks would be, but we will look into it.

(In any case, I want to emphasize that TurboQuant quantizer is a private case of EDEN)


with the amount of traction this has gotten... coming with a clear set of experiments even on arxiv paper would be of great help to showcase your improvements. And if they're easily reproducible, they could get integrated in the mainstream inference engines as well, as the main point here is compression with little degradation.


When you use TurboQuant, you are essentially using the EDEN quantizer under a different name applied to KV-cache.

Both EDEN and its 1-bit variant have been implemented in PyTorch, JAX, and TensorFlow across numerous open-source libraries and are used in various applications. I am currently writing a blog post that will document these in detail.

EDEN defines a scale parameter, S, for which we suggest specific optimal values for both biased and unbiased versions. As shown in the note I shared, these values lead to clear empirical improvements. Consequently, users who rely on the less optimal S value and the unbiasing method popularized by TurboQuant will generally see inferior results compared to those using EDEN with the optimal scale values suggested in our original papers.


With AI you actually don't need to choose anymore. Well laid out abstractions actually make AI generate code faster and more accurately. Spending the time in camp 2 to design well and then using AI like camp 1 gives you the best of both worlds.


Extrapolating the benchmarks, this would imply the best RYS 27B is capable of out performing the 397B MoE?


super clever and awesome!


The freeze sometimes does capture in between states. What I've seen the agent does in those cases is that it recognizes it's in between states and calls browser_wait(). Where the agent goes off the rails isn't a snapshot in the middle of a state transition, (it's smart enough to know to retry in that case), it's when the DOM changes after the agent believes the page has settled.

For async, lots of people smarter than me working on the smarter agent problem. Though there's a latency floor with inference due to prompt processing, and output generation. Without tools like ABP, the LLM is always aiming at a moving target.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: