Hacker Newsnew | past | comments | ask | show | jobs | submit | redman25's commentslogin

Maybe they’re gunning for speedy non-interactive pricing? Or its a limit of the technology or a business decision?


More like acting without a sense of civic duty is stupid. Or acting without a basis in evidence is stupid.


Deepseek is one of the worst in terms of hallucination rate according to artificial analysis' benchmark: https://artificialanalysis.ai/?omniscience=omniscience-hallu...


That's interesting. The "best" current models, Fable and especially GPT 5.6, are also lying liars that lie all the time. Seems like we're going the wrong way on hallucination.


And I remember Sam Altman saying 2 years ago or more, that hallucination was already "fixed" internally. Well, clearly not.


I've been seeing more "grounding" messages in the thinking log of Claude models. I guess that's "do a web search, read the docs", probably designed to mitigate hallucinations.

I am surprised by how much the really giant models hallucinate, though. My vague feeling was that little models hallucinate a lot because they just don't know anything (the world's knowledge simply does not fit in a few GB) and don't know how to say, "I don't know". But, the big models kinda do know everything, and yet, here we are, they're still making shit up all the time.


> I am surprised by how much the really giant models hallucinate, though. My vague feeling was that little models hallucinate a lot because they just don't know anything (the world's knowledge simply does not fit in a few GB) and don't know how to say, "I don't know". But, the big models kinda do know everything, and yet, here we are, they're still making shit up all the time.

In my mostly unscientific RAG experiments I found that larger models were more likely to hallucinate in a RAG setting. I think that it's because they have more world knowledge. As an example, when asking about safety legislation in Ireland, Claude got hooked on the notion that OSHA was involved, which clearly is not true, while the smaller models just said that they couldn't find any information (which was the correct answer).


One of my tests to see if "we are there", it's to ask a list of mayors from my city. Sounds silly, but most models just create random mayors that never existed.

Without search, of course. In theory this should be easy, because clearly the Wikipedia data it's in the model.


A large portion of the tests are closed source unfortunately which would make it tough to create a port.


What the hell? This is the first I'm learning of this.

Why did they do that? Is it owned by a private company?



That's one way they make money. That's the reason sqlite exists.


In a 2021 podcast interview[0], Dr. Hipp noted they've sold zero copies of the TH3 (extensive, proprietary) test suite in SQLite's history.

By the time TH3 was added in 2008, SQLite had gained a fair bit of traction across multiple industries. Though I totally agree that the comprehensive coverage is a leading reason why they've been so stable over the last ~20 years.

So indirectly, TH3 is why they (continue to) exist and (are able to) make money, but it isn't a direct line as one might assume.

[0] https://corecursive.com/066-sqlite-with-richard-hipp/

[1] https://youtu.be/5zQdYx-fqJg?t=300


They sell several niche things, for hopefully hundreds of thousands of dollars, to just a few people each. TH3 is just one instance of the general pattern.


I've been preferring Mimo recently. Same price as deekseek, more reliable tool calling (subjectively), and has some nice qualities in terms of prose, etc.

I've heard others say that Deepseek tends to be smarter on specific problems but that Mimo tends to more well-rounded.


Exactly, intelligence is limited by cost and physical constraints just as much as anything. That's the thing that seems to always be missing from the run-away singularity discussions, it's treated like a perpetual motion machine.


Typical "runaway" scenarios I see described involve something like the AI designing a worm that it uses to propagate itself across the Internet, hijacking whatever CPU/GPU power it can find, and making itself more powerful in the process. Of course this depends on bandwidth, humans not finding a way to shut it down, etc. There indeed are physical constraints even on the transmission of data.

Some people seem to think that simply uttering these ideas on the Internet is harmful (in the "don't give it ideas!" way); but the MIRI types were expressing them pre-ChatGPT in an attempt to warn people, so there was really never any chance of keeping it out of the training data.

But it's also worth considering here just how awful AI security postures have been. The MIRI types used to speculate about how difficult it would be for AIs to social-engineer users into granting them irresponsible levels of agency. It turns out that they don't even have to try.


IDK this model release is a bit disappointing considering the community has been chomping at the bit for the 124ba4b model. There was some leaked info about it but people suspect it was not released because it was too close to gemini flash in performance.


I too prefer my misinformation delivered with maximum confidence and no accountability.


What prompt had you given it?


I'm not OP but I work outside and use light mode. Macs are generally fairly bright as long as you aren't in direct sunlight. Solarized light mode for the win though.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: