The big difference is that Altman et al aren't just, or even mainly, pundits or prognosticators. Zitron's whole thing is commentary and predictions about AI and his predictions are almost all wrong.
Altman and friends are mainly actually making and delivering AI. If they also hype their timelines/valuations, they also seem to make directional progress on the goals.
Surely this is realistic now (or soon) in the form of LLM curation. A few auto-librarians reading everything, looking at different versions side by side, making choices, etc.
LOL, not realistic at all. The differences between book versions lie in more than the raw text. Moreover, an archival project would be loathe to favour or disfavour a version unless an actual human made the call.
How else is this administration going to make money?!? How dare you...if they do not accept bribes...what is there left for them? This is a premium buy...First one to beat competition gets the worm. So, you pay Trump, trump gives you access...then you pay subscription to SAMA lol.
Hmm. I also recently got a subscription to the NYT. I noticed they sent me a lot of email so I just filtered it with the click of a couple buttons. It's obnoxious but in the minor way that having to click a couple buttons one time is.
The paper says the professors have a median of 200 comparisons each. It also says they only used 2 models because using more models would require more comparisons and they selected Google models because Google was branded/advertised as being education focused. When you see other models show up elsewhere, that's because they extended the main idea to other models but using LLMs to judge instead of human professors.
Sure, but the biggest problem is they have no statistical significance. Variance is too high. How do you distinguish the signal from the noise? Confidence intervals aren't enough.
But is it a surprise law professors aren't great statisticians?
I disagree. 16 isn't necessarily the relevant N here but the number of responses is.
If you have 100 responses from 1 professor, and the AI wins 75% of the time that is very likely a true signal that the AI is better than this prof. It would be incorrect to generalize this to all profs though.
Further, if you sample 16 profs and the AI beats 10 of them you can be fairly certain that the real percentage of profs it beats isn't 10%. Further, when estimating the probability that the AI beats a random prof, it's the relative estimation error that scales with 1/sqrt N. If you have a coin and it lands heads up 16 times, that tells you something quite robust about the coin.
Reasonably estimating confidence intervals at small N and high p is not trivial. But it can be done.
A good heuristic is "add 2 successes and 2 failures" which is due to Agresti & Couli.
I think it is more likely that they selected Gemini because the lead author is a fellow at an institute which receives a lot of their funding from Google.
I don't understand the "deathbed" perspective. Are you going to wish you made more hackernews comments on your deathbed? Probably not. Does this mean you should stop using hackernews?
If you optimized for minimizing deathbed regret perhaps you'd regret that on your deathbed!
If I have cogent thoughts on my deathbed I expect they'll be along the lines of "I wish I wasn't dying" and not regretting the many ways I enjoyed my time on Earth (which includes vibe coding apps nobody uses).
This just says 60% of systems, but not the frequency for those systems. They were evaluating 20 systems, so for 12 systems there were mistakes in the prescriptions, but there isn't information about how common those mistakes were and it's hard to judge relative to a human system.
3 years ago the best model was DaVinci. It cost 3 cents per 1k tokens (in and out the same price). Today, GPT-5.4 Nano is much better than DaVinci was and it costs 0.02 cents in and .125 cents out per 1k tokens.
In other words, a significantly better model is also 1-2 orders of magnitude cheaper. You can cut it in half by doing batch. You could cut it another order of magnitude by running something like Gemma 4 on cloud hardware, or even more on local hardware.
If this trend continues another 3 years, what costs 20k today might cost $100.
Think of it as paying for tokens. The tokens you could buy 3 years ago are better and two orders of magnitude cheaper today. If that happens again over the next 3 years then the tokens you can buy today to do a job for 20k will cost 200.
This isn't optimistic in my opinion. It's not even fully realistic because Gemma 4, which you can run on local hardware, is even better and another few orders of magnitude cheaper. A 20k job today might a few dollars in a few years.
Altman and friends are mainly actually making and delivering AI. If they also hype their timelines/valuations, they also seem to make directional progress on the goals.