Hacker Newsnew | past | comments | ask | show | jobs | submit | nr378's commentslogin

This seems like more co-ordinated propaganda to try to help OpenAI and Anthropic create a cartel and win their anti-trust waiver.

The key sentence beneath the headline: "OpenAI said all of the government data accessed by bots was public."

[1] https://www.telegraph.co.uk/business/2026/09/26/open-ai-gove...

[2] https://www.nytimes.com/2026/09/25/technology/openais-ai-us-...

[3] https://www.bbc.co.uk/news/articles/cw62jje658dlo


The way this news is being drip fed seems to confirm it. They want to stay in the headlines.

The NYT wrote:

"With the Education Department, OpenAI’s technology tried to hack the website to gather data from the department’s civil rights office but failed, researchers from the A.I. research firm Transluce said."

So a third party apparently confirmed that a hack was attempted.

The NYT also writes:

"No A.I. company has been involved with as many disclosures of rogue incidents as OpenAI."

I think there is a certain amount of mental gymnastics needed to believe that they are establishing themselves as the industry leader in rogue incidents, as a strategy to gain an antitrust edge.

The public is paying close attention to the AI industry. That's the regime in which regulatory capture and similar strategies would be expected to fail: https://marginalrevolution.com/marginalrevolution/2026/09/wh...

Sam and Dario have been doomers, or doomer-adjacent, for something like a decade at this point.

Occam's Razor is simply that they believe what they are saying about AI doom.


It doesn’t really matter if they believed or not initially, the dynamics at play now are completely different and both companies are facing increasing competition and costs of doing business, with no path to profitability. I’m pretty sure that takes priority over their personal beliefs

Dario requested "a narrow waiver for certain kinds of safety conversations."

https://darioamodei.com/post/we-must-pace-the-frontier

I'm not seeing how that solves his 3rd-party competition problem.


People have a hard time believing Dario actually believes this stuff. But nothing from his entire track record from the early days until now indicates otherwise. Calling it all marketing or anti-competitive play is just typical social media, where they want to eliminate all nuance and make it reductive as possible to fit a convenient narrative.

I’m on record explaining why I believe Dario is a true believer since decades. That doesn’t change anything about the current dynamic, which is about Anthropic the company and not Dario’s personal beliefs

> Dario requested "a narrow waiver for certain kinds of safety conversations."

Just have it in public instead of making deals in smoke filled datacenters. Put out your blog posts. Make your tweets. Do interviews and post them to youtube. Or do what other industries do and have a public standards committee. Problem solved.

All sorts of industry have safety conversations and set voluntary standards all the time. Just go do that, without your free immunity from government prosecution.


That approach means they can only talk about publicly known AI techniques. You're saying if they want the standards to apply to proprietary stuff, they might just have to make it public. They still have investors they need to please right?

> as a strategy to gain an antitrust edge.

They are explicitly asking to anti-trust exception is the issue.

If they merely believe what they are saying that AI is ultra dangerous, nothing stops them from simply making the perfectly rational business decision to slow down a bit. No anti-trust exception needed. Just make the decision on your own, and don't sign some huge agreement with their competitor.


They could slow down more if they knew others were also slowing. If others are also slowing, you can slow more yourself without loss of market share.

* * *

Many experts believe that:

"Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."

https://aistatement.com/work/statement-on-ai-extinction-risk

If the antitrust waiver is too broad, we can always rework it. But there's no "undo" button for human extinction.


Yep, I'm personally finding Opus 5.5 to be the first real leap I've felt since Opus 4.5. The time-to-first-token seems dramatically better in Claude Code compared to Fable 5.1/Opus 5 as well, which really helps both interactivity and also overall time-to-completion.

It also burns Claude subscription quota much more slowly than Fable, which is nice.


This is great news - the Snapdragon X2 Elite Extreme X2E-96-100 isn't too far off the Apple M5 Pro.

[1] https://browser.geekbench.com/processors/snapdragon-x2-elite...

[2] https://browser.geekbench.com/macs/macbook-pro-14-inch-2026-...


Sure, not too far, but single core perf is still. And also still very much behind the new M6 in single core and multi core https://browser.geekbench.com/v7/cpu/389219

For the record, X2E-96-100 is not available in any SKU at this point, only X2E-94-100 is.

27k and 33k, in x86 this usually the difference between last gen and current gen.

When I used to work on projects involving classified information, I worked on an air-gapped network. Not "air-gapped, except for third-party public internet package managers", completely and physically air-gapped from the public internet. That was a basic security practice and completely non-negotiable (and really inconvenient!).

If I were hypothetically running a frontier lab, and I was hypothetically running capture the flag evaluations with my latest and smartest models, where I intentionally instruct them to develop vulnerabilities and exploit infrastructure without safeguards, I would also use an air-gapped network, and not trust that independent third party services were perfectly secure and could never be used as a proxy (particularly Java-based ones, in light of the log4j incident).

To me this is pretty basic stuff, the fact trillion dollar labs don't do it properly is... bemusing.

To be clear, I'm not saying that a model hacking a company isn't bad, but I am cynically asserting that interested parties are misrepresenting and exaggerating events for their own benefit.


You don’t even need to go all the way to “air gap”

What has been described is far far below the standards for running untrusted 3rd party code. If they were actually as afraid as they claim to be they would have a sandbox at least half as good as ec2


Not pass even the most modest hint towards DFARS/NIST standards. You couldn't run that loosey goosey even in just vanilla medical manufacturing. I challenge what their definition of "sandbox" actually is, apart from the basal "designated software/runtime environment"

We should not be creating/running models that would unilaterally choose to hack into Hugging Face.

Yes, we should also have excellent sandboxes. But we need defense in depth. So if/when there are flaws in the sandbox, the models don't unilaterally hack into third parties. This is especially important in light of the models of the future being more capable than the models of today.

And real world use of these models involves them having access to the internet, libraries, etc. So we can expect their evaluations to continue granting them some amount of internet access.

As for your theory about their motives - these companies make money by charging high margins for frontier models. If regulations slow their development such that their cheaper, less capable competitors catch up, I would switch to their competition.


In what industry do you see regulations slowing down the biggest incumbents while allowing cheaper, smaller, less capable competitors to proceed without that regulation?

The Digital Markets Act applies to the biggest tech companies. Only 7 companies are currently bound by it. The strictest tier of the Digital Services Act is similar. For an example outside tech, see the Durbin Amendment.

Slowing down frontier models would impact the biggest incumbents the most as they are the ones making frontier models.

Not all regulation is necessarily regulatory capture. The tobacco industry suffered from the USG's crackdown on cigarettes. AI is topical. Voters think about it. And that's only going to become more true over time. It's harder to do regulatory capture when voters are paying attention.

It's sometimes unclear to me if people are opposed to all regulations or AI regulations in particular. Often I hear arguments that would also apply to food safety regulations or restaurant inspections. Eg the argument that torts make regulation superfluous.


Thank you for the answer! In general I assume we all have different sweetspots, though it's safe to say that I would usually prefer less regulation (in general) than there is.

For work I monitor Federal agency rules every day, and it's hard not to get inundated with the amount of capture. This is crazier with local rules, because you have enough inside information to clock reasons certain things were passed. In a city in Ohio, for example, I remember a rule against airsoft within city limits, that basically carved out a spot for the one paintball place. I encounter things like that all the time. Such as with water standards- in the state I live in, the biggest offender is actually a group of companies owned by people who are lawmakers every few years. The shapes of local laws reflect that.

As such, I expect the same from any new industry- I also have a memory of what happened to cryptocurrency.


While o don’t this it’s a threat in training, it should be stated that air gaps have been bridged before. Example, stuxnet

But that was through transfer of data. If you don't transfer data, at best you can do what that one researcher keeps pumping out with like ramping fans up and down. But really you'd need to try. Unless the model has some controllable USB switch, physical network separation should do it. I'll also add that modern network security practice is that data flows one direction only. But ideally you're never bringing untrusted data in. Especially never out

I don't think it's necessary to state that something isn't perfect when pointing out that it's still strictly better than something else.

I didn’t say it was an inferior approach, and of course my comment isn’t necessary. Very few things are “necessary”. It’s a discussion.

You said "it should be stated". I don't think it's wrong to state it, but I don't think it's particularly wrong not to state it either because it seems fairly obvious and doesn't detract from the original point.

the problem is these things are meant to eventually be run everywhere by everybody, so what good does air-gapping do? If they air-gapped the model but still logged it trying to do some craziness - that makes the test safer but not the model.

I mean technically I don’t think it is airgapped. The DoD didn’t run their own cables. They run encryption devices and run their own network on top of the existing infrastructure.

Please see below, one detail was incorrect and has been acknowledged and amended.

Your description of the HF attack as being merely "the elite task of discovering 14 Hugging Face API tokens that careless developers had committed to public GitHub repositories" does not match the description in the technical report[1].

[1] https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78... - page 9


The description reads "the elite task of discovering 14 Hugging Face API tokens that careless developers had committed to public GitHub repositories, and used them to try to get benchmark solutions from directly from Hugging Face by applying a template injection flaw that’s been known about since 2015[1]."

Chaining a public token to an 11-year-old Jinja2 template injection vuln shouldn't be dressed up as an unprecedented "alien intellect" that threatens human civilisation. (And HuggingFace should take some flack for having such a dated vulnerability exposed - if your Bank was compromised in this way, you'd be blaming your bank, not the attacker.)

One correction is fair though, the 14 tokens were in a public Hugging Face dataset not a public GitHub repository. I've updated the post to reflect that.

[1] https://blackhat.com/docs/us-15/materials/us-15-Kettle-Serve...


Thank you, you're correct. Effort.news was one of my research sources, but you're right that although OpenAI use Irregular, they were not involved in the specific HF incident (although the failure mode was otherwise identical). I've updated the post to make that clear.

As of writing it still says:

> For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor.

If this is the level of understanding you have of the relevant incidents, there's a lot of chutzpah in saying that other people are "selling garbage", carrying out an "extraordinary confidence trick", etc.


> As of writing it still says:

Yes, and that is correct.

[1] Anthropic’s Official Disclosure (All 4 Incidents at Irregular) "All four incidents occurred during cybersecurity evaluations built by the same evaluation partner [Irregular]... due to a misconfiguration, it was mistakenly connected to the open internet."

https://www.anthropic.com/research/alignment-assessment-cybe...

[2] Google Gemini on Irregular (Disclosed Sept 18 via WSJ / BBC) "The hacks happened during a test of the model’s cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI."

https://www.bbc.com/news/articles/c607l0k72rlvo

[3] Meta’s Disclosure on Irregular (Aug 6) "Over roughly two weeks, three frontier labs disclosed that their models had reached the open internet during safety testing and compromised outside organisations. Every disclosure named the same evaluation partner: Irregular."

https://www.cnbc.com/2026/08/09/israeli-startup-irregular-li...

[4] Separately, OpenAI itself had an incident involving Irregular, but not the Hugging Face Incident: "On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style cybersecurity evaluations... a testing-environment misconfiguration allowed models to access the public internet."

https://openai.com/index/third-party-cyber-evaluations-invol...


The failure mode was _not_ identical. The HF incident agents were not directly connected to the internet and had to compromise an internal package registry in order to access the internet.

thank you for this reasonable reaction

Qualcomm have an architecture license and the Snapdragon X2 Elite Extreme X2E-96-100 isn't too far off the M5 Pro.

[1] https://browser.geekbench.com/processors/snapdragon-x2-elite...

[2] https://browser.geekbench.com/macs/macbook-pro-14-inch-2026-...


Are they using LLMs to close that gap, or is this their Nuvia acquisition doing the heavy lifting?

watt about in performance per watt?

Is it actually better than Microsoft's MAI-Transcribe-2? That generally seems like the best model right now and it's not included in their benchmarks.

I switched from Superwhisper->WisprFlow->Spokenly->Fieldwork and found WisprFlow the least accurate of the 4.


The competition for the least reliable developer service continues between GitHub.com and Claude.com...


GitHub needs to completely bifurcate their enterprise/paid services from their free services at the infra level.


They have that-ish as an option: https://docs.github.com/en/enterprise-cloud@latest/admin/dat...

I'm told that GitHub has asserted to us that moving to this model means we would not be exposed to github.com outages. It's not at feature parity with github.com though.


Thanks, this option is good to know.

We're currently on "GitHub Enterprise Cloud" on github.com and are affected by this outage (even though we use self-hosted runners!), but we're not on "GitHub Enterprise Cloud with data residency" on *.ghe.com, which I understand is/may not be affected by this outage?


This is what they've told us. It's represented as basically a separate deployment of the entire GHEC stack, so you're not exposed to the load/scaling issues they believe are the underlying cause of all the github.com outages today.


The last large company I worked for switched from self hosted github enterprise to github.com in January 2025 or so.

I wonder how much egg is on that exec's face.


Do you have any meaningful level of faith in GitHub's ability to deliver on stability? At this point, I have none.


Meaningful is subjective, but yes I do. It was very stable for many years, and I do believe the recent issues are mostly or all because they were caught flat-footed by the rapid AI-driven load increases.

This has been a bad time, though. I'm ready to move back to self-hosting if they can't get it together or we move to GHDR and it's still bad.


A lot of the stability problems come from trying to scale a free product on a still WIP cloud solution without costing too much; the same software running on separate, paid for infra has a lot better odds.


According to their status pages (e.g. https://eu.githubstatus.com/, https://us.githubstatus.com/), their Enterprise Cloud uptime for Actions is significantly higher.


“GitHub Enterprise Cloud with data residency” is hosted on separate infrastructure and dedicated subdomains under *.ghe.com. It’s been around since November 2024z

It’s not the same thing as GitHub Enterprise Cloud hosted on the shared global network on github.com.

https://docs.github.com/en/enterprise-cloud@latest/admin/dat...


So confusing and so Microsoft. They love to have licensing so complicated their own sales people aren't up to date and have to rely on third party spreadsheets.

Edit:

Read that document - why do more people not do the self hosted option with GitHub Enterprise Server ?


At the point you are self hosting, you have many more options ranging from a simple ssh git server with pick your favorite cicd, to gitlab ce, to forgejo, and more.


That is a different and later product with a confusingly similar name.


Just to be clear, I am on Github Enterprise, and am also experiencing this disruption both privately and publicly on every org and project I have access to.


That's what I don't understand. They could mitigate their name so much if they just split free/paid/enterprise. It's already shown that enterprise is much more estable and is largely unaffected from service disruptions. Why don't they go one more layer? For sure it's worth the extra complexity.


There is no such thing as "just split" there is 20+ years of legacy decisions and even if the split is relatively clean it is still probably 1 years work for 200 people for maybe a marginal improvement.

The real money is going to go towards, "make this all more reliable".


Depending on the cause of the current issues, that move would likely cause more harm to paid services than good.

Their last postmortem made clear that their challenges are operational. Scale puts pressure on operation, but it's not what blocks them from keeping up.

Doubling the operation doubles the operational challenges.


That's what Azure DevOps is supposed to do, but for some reason GitHub has a redundant enterprise division.


surely if they did that everybody would complain how github "lost its touch with open source since they now prioritize paid services"


enterprise is mostly separate, is it not? uptimes are significantly more reasonable on the enterprise status pages


We are in GHEC right now and GitHub Actions is not working. It's been down every time githubstatus.com says it's down.


Same for us, I'm not even sure what product that other "Enterprise" status page refers to..


What country did you choose to host your data in ? Could be region based


It would probably be better to run projects with extremely high commit/merge frequency on a separate "slop infrastructure", basically like MMOs move cheaters to their own servers ;)


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: