Hacker Newsnew | past | comments | ask | show | jobs | submit | eqmvii's commentslogin

In a year or two, articles like this will either be artifacts from peak hype or evidence of the beginning of the singularity. Right?

The singularity, as defined by Hinton (and others) as RSI (Recursive Self Improvement) may actually be beginning already, as OpenAI has announced an AI acting as a "research intern" (!).

How is this different from arguing that Microsoft Clippy was RSI? An AI tool being involved in the process of work can't be the bar for RSI.

I don't think there can be a coherent definition of RSI unless people lay out their theory for how intelligence scales. LLM-assisted coding is great but respectfully optimizing pytorch features or whatever is not gonna lead to exponential improvements. That approach to scaling diminished years ago, leading all the labs to switch to reasoning.

Now it seems reasoning is also yielding diminishing returns, so all the labs are pivoting to specializing in particular fields like math / infosec / biology. They're improving due to accessing new proprietary training data and doing RL with human experts. Again I don't really see any amount of "AI research interns" leading to an exponential improvement to this strategy, they're not the bottleneck in the first place.


>Now it seems reasoning is also yielding diminishing returns

Is the diminishing returns in the room with us?

>so all the labs are pivoting to specializing in particular fields like math / infosec / biology.

They're not pivoting to anything. The goal has always been creating a machine that could automate all or nearly all human work. They're just coming along on that mission.

As for RSI...I think the term is a bit odd in the modern context. It was created at a time when conventional wisdom was that generally intelligent machines would be these logic automatons that could "alter their own code". Instead we have massive neural networks that take months to train.

In this paradigm, the ways a LLM could "improve itself" would be altering its own weights directly or creating and training better, vastly more efficient architectures for the next generation of models.

The former is probably not happening but the latter is possible.


Yes, diminishing returns. Not overall, they've still been able to create more intelligent models even up to today. But the strategy for scaling that intelligence has shifted. From the initial ChatGPT release to GPT-4.1, they were basically scaling up compute training compute / model size. Then 4.5 flopped, while o1 demonstrated that gains could continue by reasoning (scaling up compute at inference time). o1 is now the ancestor of all their flagship models from GPT-5 on.

This is why I'm trying so hard to drill down on the theory of scaling, and not just talk about improvement in general, hand-wavy terms. If the bottleneck of current scaling strategies is training data, or something fundamental about the model architecture, then just throwing more harnessed chatbots at it won't lead to an exponential increase in performance.

Now you could argue that the AI we have now will help us find that change in architecture, and I would agree. But that means we're firmly outside the singularity for the time being, and what people are in fact talking about is a hypothetical.


>Then 4.5 flopped, while o1 demonstrated that gains could continue by reasoning (scaling up compute at inference time). o1 is now the ancestor of all their flagship models from GPT-5 on.

That's not quite right. They are still scaling model size and have had several new base pre-trains, just nothing so big as 4.5 (as far as we're aware). o1/4o has not been the base for some time now.

Data is obviously a bottleneck for some regimes and LLMs will have to get their hands dirty experimenting but it doesn't look like an insurmountable wall either.


> "get their hands dirty" > "insurmountable wall"

This is gibberish, you may as well tell me you've found a load-bearing seam.


Okay?

There's no reason the reinforcement learning that is getting them better at computer use can't be applied to other domains, like biology, chemistry etc. It's just expensive, because the environment often becomes the physical world, it requires creating labs like anthropic are doing here, and gathering a lot of data, it requires llms attempting their own experiments(that's what 'getting their hands dirty' means).

Getting the data and setup will be expensive, but not impossible, and labs are clearly gearing up to do just that. If you can't understand that then that seems like a you problem.


> Now it seems reasoning is also yielding diminishing returns

Not true. On the contrary, LLMs are developing faster than predicted. They were expected to solve a Millennium Prize by 2030... and here we are in 2026. Release cycles are getting faster. Just compare the most recent GPT or Claude with what they were an year ago.

> How is this different from arguing that Microsoft Clippy was RSI?

We can argue about semantics, but that's not really the point. The point is that what started now - which no doubt is in its infancy - will result in full autonomy quite soon (they project an year or so), with the risk of RSI causing agent development to slip (long term) outside human cognitive control/capacity.


Again, can you lay out your theory for how intelligence scales? You're using a lot of terms like "full autonomy" without definitions. Why do you think that just throwing more harnessed LLMs at (something?) will lead to an increase rate of improvement?

I feel like I laid out several cases where other things were the limiting factor on improvement and more agents wouldn't have helped, and I didn't get a response to those cases.

What "they project" (the labs) is of minor interest to me. Aside from their incentives and track record of lying, in recent months they are laying out a story that is pretty much just the plot of Terminator, and directly referencing rationalist beliefs that were published long before LLMs even existed.


How is it improving, that would require rearranging its weights and biases which it cannot do easily or quickly.

Self improvement during training, and AI self training are already happening. Easily/quickly are seemingly a factor of how much power/hardware you want to use at once.

With the level of compute they have they aren't stuck with frozen models like you are.


The infrastructure provisioning alone to train is heavily dependent on humans, as is dealing with failures (training runs fail a ton). Its not as simple as adding another ec2 on your dashboard. < 1k people in the world know how to do this, there will not be "recursive" or looped continual training for a long long long time. There are so many delicate inputs and controls. Not to mention the chains of businesses and the people required to operate them just to obtain the data needed, clean it and hand it to the llms.

The llms are supervising rlhf and creating synthetic data (to an extent) but they're nowhere close to being able to operate the full training stack end to end. This is a fantasy being sold to investors to create fomo.

Remember they're also limited by an effective memory of like 500k words a turn. Memory systems are lossy, so are swarm/sub agent mechanism. Im not worried about llms becoming self powered super entities anytime soon.


Is easily and quickly a requirement? Isn't it enough that over time it improves itself even if the process is complex and slow?

Do we know it's actually improving itself? Perhaps it's just opaquely sorting all ones and zeros for better lookup efficiency.

That's black and white thinking; it will be a midgularity - so neither.

Mehgularity

Whompageddon

Yes. I would bet on the latter.

I think this is key. Lawyers make a lot of money for being experts on the sidelines of disputes with values greatly in excess of their fees. They were never paid for their busy work.

Not to mention CONNECTIONS, of which LLMs will start with 0 and never acquire more.

So. Just like there are a gazillion book keeping tools/saas out that allow you to focus on your primary concern of the business instead of side quests and chores, agents can help with more.

Generating reports, generating ideas, taking things through (even if it is just rubber ducking), updating a picture quickly at 90% of quality that you could have done it but in 1% of time and effort. It allows you to try out new approaches and focus on your main product.

That is the value proposition of AI.


They're doing pretty well on OnlyFans though. You know what they say? Start where you are you. If the only way you can make connections initially is through AI thirst traps, and you're really good at it, then... you got your ins.

Not sure why the downvotes… my company solely survives on the connections and trust made over time. Real human connections, face to face. This in turn established a reputation.

But for a lot of companies, all the connections in the world won't mean shit if someone else can offer a lower cost with a better lead time.

Debatably the people who were getting 20 years of education will still be just as capable, and the 20 years of education was always a side effect of their capabilities, not the cause of them.


7+20=27


Yeah. I buy that there are some use cases where AI is a trap, or not as good as people.

I'm also chatting with my logs and triaging things like 10x faster in multiple open terminal sessions while I post on HN right now. So I'm also certain there are some use cases where not using AI will soon be professional negligence!


I see it in a slightly opposite way: even the good models are relatively cheap, and so I worry what we might miss by spending too much time playing with the Sonnets of the world when the Opuses are still objectively a bargain for the power they bring.


> when the Opuses are still objectively a bargain for the power they bring.

The cost isn't just what you're billed. There are security, privacy etc. concerns.


I know companies that are using github, even using public repo, and request their teams to not use SOTA models, but are ok with local models. Just stupid policy.


If you're writing open-source code then there's obviously nothing wrong with publishing it in a public repo. "Using Github" doesn't require you to use their CI, but even then, a human managing secrets for GitHub CI is worlds apart from trying to make sure an internet-connected agent doesn't leak secrets. And if you have sensitive data that you can't send to a remote model but you would benefit from the technology, then processing it with a local model without network access is the obvious way to address that.


If Orang mane bans GitHub they've got their local clones and can whip out a local server and a CI solution.

If Orang mane bans Claude, they've got their local models.

The latter has already happened too so I'd say their risk modeling is spot on.


hell yeah it does, feels nice at the end of a long day


I like computers more than many members of my family, so...


This is heavily contingent on one's individual experiences. An uncurated modern web/mobile sounds like a hellscape - ads in one's start menu, autoplaying videos, a flood of promotional notifications from shopping apps, people arguing with bots arguing with people about populist political garbage... blech. The real question is how this compares to your family member having a few cocktails then blaming you for their divorce or getting calls from collection agencies on behalf of somebody you haven't seen in five years.

Both of these problem spaces run the gamut from solvable to intractable. Either way you have to put in some work.


something something machine elves...



>They believed their thoughts and bodily sensations were controlled by a machine that defied their technical comprehension and secretly influenced them from a distance, often claiming that it was operated by a group of people who were persecuting them.

That also describes data centers ;)


It’s the Torment Nexus.


Why do people say "something something...", and why does it irk me?


think of it like a placeholder for a great argument/explanation that doesn't necessarily exist, but would be funny if it did and would aledge whatever follows (in this case mechanical elfs) as being highly relevant to the topic at hand.


No tv and no beer make Homer something something [0]

0 https://youtube.com/watch?v=ufDqfQxPnJo


There's a huge market right now for "AI product, without much different than Gemini/Claude/GPT do out of the box, but from somebody you actually trust"


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: