Hacker Newsnew | past | comments | ask | show | jobs | submit | WASDx's commentslogin

I tried calculating historical "intelligence per cost" recently but stopped when I realized intelligence is not linear. For any meaningful "x per y" you can just double "y" if you have a half as efficient system to get the same result but so-called intelligence doesn't work like that.

It works if you neglect code quality and understanding. I've done a few vibe coded projects on my free time but I would feel shame for submitting that kind of code at work where I use LLMs more responsibly.

Likely because it uses fewer thinking tokens (that you don't see anyways).


Give a task you have to 3 different models and see what actually works for you. There are no good benchmarks.


This is like saying mass production is "cheating" against handcraft.


It is, if you are trying to sell all of them in the same market, so 99 stalls with mass produced stuff, one with handcrafted stuff, but you don't know which is which.


Once it figures out a puzzle it could probably be instructed to design a specialized harness for Luna to be able to solve other instances of the same puzzle. Minimum wage workers are not solving novel problems.


They explain it here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...

TLDR: The official ARC harness throws away old context and reasoning. No real-world harness is this bad, the model has to re-learn the game repeatedly. OpenAI basically just added standard compaction. Their harness is still "general".


With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.


and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.

This technology is strictly an extractive parasite on the world. Use it, but don't be excited.


My labor makes other people's lives better, so I would expect something that replaces my labor to do the same.


global development and relief of poverty has relied on there being an economic surplus for all from organized labor. everyone gets a benefit although it is unfairly distributed.

i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.

depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.

the work does nothing or causes net harm.


You’d expect that, wouldn’t you? But, alas…



[flagged]


Would you care to discuss the topic, or just throw grenades? Surely you can come up with something more substantive than this


Ok. This requires the notion of "intrinsic value", which I believe does not exist (all value is subjective), yet is a foundation of all Marxist theory.


it's easy. people are intrinsically valuable.

do you believe that people are not intrinsically valuable? that their value is what they do for others, that it is not they themselves the person.


Buddy, admitting your thought processes forcibly terminate on pre-programmed keywords isn't a flex.


Why "terminate", Marxist philosophy is a legitimate topic, deserving to be studied. Like a rich sci-fi lore or a history of Tarot magic. Deep, fascinating, and wrong.


And yet you terminated. Which would be the correct thing to do if Marxist philosophy would be wrong on all counts, as you explicitly state. Which of course, it isn't.


How dumb are you?


I'm using AI to build things I wouldn't (and/or couldn't) have built before.

That's the opposite of parasitic.


Talking as if you are not disposable. If you are let go from your company, you can be easily replaceable.

People already started using contributor API, and your input is irrelevant.


Don’t you have some looms to break?


I’m retired so it won’t be replacing my labor :)


The sibling reply to this is just such lazy thinking, such a trite cliche. Yes, all members of a generation are bad, end of story. Can we get back to the war between the sexes now?


I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.


3.7 high and 3.8 medium are essentially the same on AA intelligence and cost. Output tokens on DeepSWE gives the same picture. So there might be something to it but they have done other things as well. At least the tokens are really fast.


i find deepswe not very reliable for instance it puts grok 4.6 xhigh over sol medium


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: