Hacker Newsnew | past | comments | ask | show | jobs | submit | scrlk's commentslogin

The UK has mini roundabouts, which are just a painted white circle on the road: https://upload.wikimedia.org/wikipedia/commons/b/b0/Mini-rou...

Drivers treat them as a normal roundabout by giving way to traffic on the right.


I contest that drivers always treat it the same. They often go through the circle/cut th corner when there is low traffic, and I know at least one mini roundabout where traffic is way less frequent in one direction and the other direction is treated as the "main road" without slowing down.

I mostly like them though.


The share of household spending on clothing has gone from 10%+ to under 5% in Western countries. When you're spending that much, it's not a surprise that clothing was "BIFL" both in quality and upkeep (repair rather than replace).

The overall revealed consumer preference has been "more and cheaper", not "less and higher quality", which still exists as a niche if you are willing to go down the tailoring route, not off the peg.


The thing is mainly there are no quality options available. It's mostly all cheap shit or slightly more expensive, but still equally shit. So how can this reveal preference if there are no options?

I would argue the surge in vintage clothing ecosystems is an indication there is an unmet demand.


You can still go to a tailor and commission jackets, trousers and shirts, made with the same quality of fabrics and manufacturing techniques. With shoes, you can still go and buy hand welted shoes.

It all comes at a price but I don't disagree that we have a barbell problem here (ie cheap fast fashion or bespoke, with nothing in the middle).


There’s certainly options above temu junk. Maybe it doesn’t meet your bar but I’ve found the 100% cotton stuff from Uniqlo to be massively higher quality than the average. I’ve got tshirts from them I’d guess I’d worn and washed 80+ times.

There absolutely are still rich-people brands that sell absurd quality at absurd price.

See, for example, Marine Layer in San Francisco, or boutiques like Red Ants Pants.


Not just rich people, I pay maybe 3 times more for my T-shirts than the average person. I'm certain they last me at least twice as long, probably 3 times and they're way nicer to wear.

That's not actually "revealed consumer preference" for quantity over quality. It's just that it's impossible for the average consumer to assess quality.

Might be related to this announcement from Tibo on Sunday:

> We've made some improvements that improve usage on the long tail for power users of Astra when logged in with your ChatGPT account.

> No change in quality and a pure win that on the long tail can result in up to 3-4X less usage being drawn from the subscription.

https://x.com/thsottiaux/status/2096717905614524491 (https://xcancel.com/thsottiaux/status/2096717905614524491)


It seems to me the people working at OAI may believe all other humans must be a little bit behind intellectually.

It always reminds of the story of the creator of counter strike. Every new release he would get a ton of complaints from players about things they didn't even change. Notably that each version had more lag. And he got so fed that he start to negatively subtract peoples pings. And suddenly a ton of players reported back that the change was incredibly good.

Point is, I really don't buy all the stories about a model suddenly being downgraded without at least a modicum of substance. People are grasping at straws in the noise.


I mean, in general they aren't wrong.

You can't fool everybody all of the time, but you can fool almost everybody most of the time.

But most of all, it's easy to fool yourself.


Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)


Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.

ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard

A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!

(I coauthored the linked blog post)


Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?


Those people haven't verified their results against the private set: https://arcprize.org/leaderboard


Astra also not verified using private set, but on "semi-private" set


if that is true then why is astra on the official ARC leaderboard now ?


ARC leaderboard has results from semi-private data for frontier models, they have another competition for private data.

It is described in their methodology: https://arcprize.org/policy

It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.


Where are results for private data?

Which LLMs participate on private set? Open weight LLMs only?


Yes, they run competitions once a year amongst open weight models


Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.


It is about memory retention. No heavy lifting done on the reasoning side so I hardly see anything misleading here.

Edit: update from fchollet https://x.com/fchollet/status/2095598451115614371


yes it is.


Not just speed, also reliability. IME, Gemini's speed and quality doesn't degrade badly during weekday working hours compared to OAI, and especially Anthropic.


I've had Gemini model API use degrade the most out of OAI/Anthropic/Google (often "over capacity" vs true failures)

Not sure on consumer/product use though


That's interesting to hear. I should have added that I use Gemini through Google AI Studio as my general chat model, which probably explains our wildly different experiences.



Thanks, missed that article.



Plus the Samsung Exynos modems that they were using from Pixel 6-10 (11 switched to Mediatek) had worse power efficiency and performance vs Qualcomm.



It was Lenovo ditching the classic ThinkPad 7 row keyboard with the Ivy Bridge models c. 2012 that really pissed people off


Indeed, multiple civil wars have been fought since the Lenovo acquisition.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: