Hacker Newsnew | past | comments | ask | show | jobs | submit | ramon156's commentslogin

i can see the comment now, so this seems undone?

The more I see about how Jev works, the less interested I get. Jev would be cool as a local model for e.g. Home Assistant. Projects like this are at best misleading, this one also happens to be broken.

Strange - I had the opposite experience. The more I learned about Jev the more interesting I thought it was. This Syntax video got me excited about developing with Jev: https://youtu.be/QbYBRjOaGOo?si=Up0QbW8tWCZQzZR4

Surprised by these takes.

Are people not getting that this (Jev) can do classification, programmatic branching, real time decision making (e.g. applicable to robotics) an order of magnitude faster and cheaper?


It seems almost like a smart switch statement.

if you can be replaced by an algorithm, how useful were you really?

Very useful. That’s a weird question.

I aspire to be at least as useful as bogosort.

But the evidence is not there...

Indeed, they talk as skeptics but don’t offer a ton of evidence, other than a couple videos of demos. A live demo would be far more convincing.

They gesture at not using benchmarks for some reason...

https://typesafe.ai/blog/antibenchmaxxing

But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent "global" benchmark might be hard to do


I've implemented tree-sitter in pi before, and while it works, I have no real proof it saves me tokens, or is more accurate. I think a better implementation is a model that's trained for AST's, not just "use tool, see what happens".

I'd love to do research on this when I have the time.


Cool project!

That's what I was insinuating through "better encoder"; the model creating more efficient representations of ASTs using something like JEPA


This sounds good but so far all claims just sound like marketing terms. I'd love to see real proof. e.g. "RLCD" and "parallel sampling" have nothing to back it up.

also "70-500ms vs 3-329 seconds" are apples-to-oranges unless the LLM baseline is doing comparable work (e.g., long chain-of-thought). If Jev is skipping generation entirely for a narrow structured task, of course it's faster.

Nonetheless i want this to be true, so I'm looking forward to Jev

Edit: I really have to say that I like their manifesto https://typesafe.ai/manifesto


They have various benchmarks, e.g. how much time it takes them to do wikipedia page -> page games. Jev seems to take the same or fewer hops but in ~10x less time and for ~10x less money.

It's totally reasonable to compare against LLMs doing chain of thought if it gets comparable performance.


> If Jev is skipping generation entirely for a narrow structured task, of course it's faster

I think this is reasonable if people are actually using LLMs to solve this type of narrow structured task, which they are. The evidence is that every LLM provider has some method of forcing the output to conform to a json schema in their documentation.


love that you love the manifesto! letting the first batches off the waitlist now, but we do have some early users describing their experience (https://x.com/danshipper/status/2099947471518474522)

Congrats ! Really excited for the team.

Did you see the video where it plays Doom, it made it click for me

BTW it was not multi model playing doom, it was passing structured input and getting structured output. Its not what I thought: frames of video passed and real time game play.

so what? put an LLM on Cerebras and get its responses faster, and put Jev on Cerebras and gets its responses even faster

i've seen it play minecraft as well, what im not sure here is what is the thing that produces the JSON and keeps track of the objects

can jev play battlefield six for example


> I really have to say that I like their manifesto

Their manifesto: "you only build on top of it if it's trustworthy." - the irony of this while putting out the most misleading, dishonest marketing campaign I've seen in months for their first public appearance doesn't exactly scream "trustworthy" to me.


What do you find dishonest?

Comments under the x post.

It's not an LLM though it's a frontier model on structured data

Thank you Claude, very cool!

Yes this is true, you even got the model correct. I relied too much on Claude this time. I'm starting to work on the next AI agent lesson this week. I'll do better on the upcoming lesson. The next lesson is A2A focused

name some examples. easy to use is opinioned,

Sure. I tried Open WebUI, SillyTavern, AnythingLLM, etc. I wasn't a fan of installing a 500MB desktop app, or having to run a webhost locally, or having to configure a bunch of settings just to chat with the newest LLM.

i quite liked using strix. last time i tried it, deepseek was a mess and bloated the context with nonsense. that was ~5 months ago, i wonder how it performs now

We've made a lot of awesome changes recently, would love any feedback on the latest version :)

not in this context, though?

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: