I'm one of the whistleblowers (from GDM [1]). I gave up over a million dollars (compared to quietly switching labs and continuing to work at one) to speak frankly about these issues. I hold no equity and tried to zero out my position before ever joining GDM.[2]
It's wild to me that people think such whistleblowers are fronts for labs to take over or pump valuations. We are trying to call out how these labs will, by uninterrupted AI-race default, concentrate enormous power over the rest of humanity.
Joe Benton left Anthropic a day before Dario's post, to work for METR evaluations. He was with Anthropic for over a year. He did the same thing that Jacob did (big song and dance about AI apocalypse, media interviews all over the place). He managed the Scalable Oversight team at Anthropic and was the research lead for the Anthropic Fellows Program. So he has equity, and likely lots of it.
Then you have Josh Engels quitting DeepMind to work for METR the day before as well, doing the exact same thing. Again, doomer drama all over socials, interviews, and so on.
Did I mention METR is founded by an ex-OpenAI researcher?
Now you have Demis Hassabis, Sam Altman and Dario, all circlejerking eachother on X saying "we all agree with Dario" - while they ask to be "regulated" by the company that has all of their combined equity-holding ex-employees in it.
METR's salaries are listing around 500k/yr. Gee, I wonder where this non-profit with ~35 people is getting all of its money?
So the fact that Dario tries to frame it as an "independent third party" is all the evidence you need to know that Dario is a pathological liar and always will be.
---
Some more info:
Dario's sister, president of Anthropic, is married to the co-founder of Open Philanthropy. The two largest AI doomer NGOs, Center for AI Safety (CAIS) and the Future of Life Institute (FLI), have both received many millions of dollars from them.
Ajeya Cotra worked at Open Philanthropy/Coefficient Giving for roughly nine years, including leading its technical AI-safety program in 2024 and contributing to AI-giving strategy in 2025. She subsequently left Coefficient and joined METR, where she is now technical staff.
Ajeya is married to Paul Christiano, who founded Alignment Research Center (ARC). Alignment Research Center donated ~$4.5mil to METR.
Good Ventures is a funding partner of Open Philanthropy, who funded Jacob Coxon (the first of the Anthropic employees going viral in the media) via a scholarship.
> Good Ventures is a funding partner of Open Philanthropy, who funded Jacob Coxon (the first of the Anthropic employees going viral in the media) via a scholarship.
This conspiracy theory is truly crazy. A $20K scholarship in 2022 is supposed to explain Coxon walking away from unvested equity for a company worth over $950 billion dollars?
Obviously not, and that's not what it demonstrates. It demonstrates relationships, collusion and favoritism.
I don't think Coxon was ever planning on or entitled to taking equity, I think this was the plan from the beginning and why he was hired for 6 weeks to begin with.
OAI could check whether those accounts enabled training data. If "yes", OAI could trace whether that data was used in any related training process. If either of those answers comes out to be "no", then that's sufficient to conclude training data independence.
We wouldn't need a full ablated re-training and solution attempt, contra tedsanders in a sibling comment.
If the model includes unique data from a person then that person can identify the data - the allegedly plagiarised material - and so re-identify it. There doesn't need to be a privacy breach to close that loop as it requires the person to identify the information is associated with them first.
Those topics aren't on-topic for the essay. I've taken a pledge to donate at least 10% of my money to charity / impactful giving. A good sum of my donations have targeted high-impact opportunities to improve life for people in third-world countries.
That's commendable, but I do think it emphasizes my point. If you have such ethical concerns, why develop powerful technology on behalf of shareholders whose identities and intentions are unknown to you? Activation engineering has obvious malicious uses and I can distinctly remember the sense of dread I felt when 'relaxed and expansive' went viral. I believe the 'good AI' future faces political obstacles, but perhaps you take an accelerationist view that technical solutions will pave the way. Given your politics, I wonder if an interview on a show like a Ralph Nader Radio Hour could help you connect with likeminded people from different backgrounds. Best of luck
> I'd guess TurnTrout doesn't agree on that framing, otherwise he probably would not have been at Deep Mind. But clearly he and I agree on other ethical positions; I am nothing but glad to see him stick to his principles here.
FWIW I agree that creators should be compensated (but evidently it wasn't a deal-breaker for me joining). I think it's bad how little that has happened.
I joined GDM to work on AGI safety to reduce existential risk from AI. I consciously avoided work that would improve the raw capabilities of Gemini. When I joined, I made a trade-off and it's valid to disagree with my decision.
As your writing seems to allude to, life is an impossibly complicated series of trade-offs made with imperfect information, so I definitely wouldn't consider joining an AI frontier lab to be a priori bad. Even if I have some qualms with how the models are built.
If anything I was just trying to point out that, even with people we might disagree with, we deserve to show them recognition and respect when they make moves we do agree with.
But I doubt we disagree on much. Your writing is great, and it makes me sad to see this thread somehow disappear from the HN front page so fast. Either it triggered some internal flamewar detector--but without a flamewar--or someone didn't find it convenient.
I wish you the best of luck with your next endeavors, and look forward to seeing the value you add to this world.
>The stereotypical activist action is to make a petition. But Google had already ignored a large petition on this issue. Plus, Google’s executives likely hardened their company against stereotypical organizing tactics. Sit-ins, strikes, even a mass of Google engineers quitting: I deemed all of them ineffective (if I could even pull them off).
I agree that they called many things remarkably well! That doesn't change the fact that AI 2027 is not a thing which happened, so it isn't valid to point out "this killed us in AI 2027." There are many reasons to want to preserve CoT monitorability. Instead of AI 2027, I'd point to https://arxiv.org/html/2507.11473.
No one has empirically validated the so-called "most forbidden" descriptor. It's a theoretical worry which may or may not be correct. We should run experiments to find out.
As someone who did their PhD in RL and alignment, it was not obvious to me a priori if, or when, or how badly obfuscation would be a problem. Yes, it's been predicted (and was predicted significantly before that Zvi post). But many other alignment fears have been _predicted_, and those didn't actually happen.
I don't think the existence of specification gaming in unrelated settings was strong evidence that obfuscation would occur in modern CoT supervision. Speculatively, I think CoT obfuscation happens due to the internal structure of LLMs and it being inductively "easier" to reweight model circuits to not admit wrongthink, rather than to rewire circuits to solve problems in entirely different ways.
reply