Hacker Newsnew | past | comments | ask | show | jobs | submit | criley2's commentslogin

I believe the that the companies who claim to not train on my data are more likely to not train on my data than the companies who refuse to even claim they won't.

Also why Meta gets a +1, just charge less money on the training path.


I’m not sure that follows. You’re assuming that all those claims have the same weight, without considering the size, jurisdiction, reputation or even the general vibe of the company making that claim.

If you factor that in, then there are clearly different tiers: one you can trust, and one that may well just be saying that to increase market share with little reputational or legal consequences if they are found to be lying.

These are not equal.


Yes I sometimes think the "don't train on my data" is actually a good signal for "this data/person is probably better to train on because they want to keep something private". The whole copyright system should have stopped these guys from training on everyone's data and it did not, if you think they care about the privacy checkbox I think you're dreaming personally, based on their past behavior.

> I’m not sure that follows

To be fair, none of us are sure of anything and I think that’s the part that’s most irritating


It’s more a polite way of saying “that’s crap”

And mine a polite way to say “you are equally uninformed”. We’re not getting anywhere. All the best.

Absolutely galactic-level ooof :)

> ZCode, the GLM coding agent, silently uploads your Git history

https://news.ycombinator.com/item?id=49752422


FYI it’s helpful to actually say your point during a discussion. And if you don’t want a discussion then why did you comment?

I have been writing an internal code review tool that is a bit maximalist. I created subagents for many internal domains and technologies we manage, with prompts focused on best practices, common problems, owasp guidelines, etc), a separate tier of wider band subagents (design, rollout, security/privacy), and a final agent at the top orchestrating and combining. I also use adversarial validator passes against all findings.

Right now, Fable 5.1 delivers incredible reviews. Opus 5 delivers good reviews. These agents are finding really impressive issues that humans just don't have the attention span to track down. My reviewer has a very impressive signal to noise ratio at this point, after half a year of iterating and improving. (I use a lot of Opus high, Opus medium for less critical tickets/domains, Fable 5.1 high for critical domains and all of the issue validators, and even Fable 5.1 xhigh for my design agent, whose job it is to think about the project at a high level and provide the kind of high level tech design review that AI notoriously can't do well)

I'm testing Astra so I don't have strong opinions yet. I've also done extensive testing of the same skill and subagent pattern in opencode/omp using GLM 5.3, Kimi K3 max, Deepseek V4 pro, Deepseek V4.1 flash, Qwen 3.8 2.4T max, and others.

My experience is that open weights models find between 1/4 to 1/2 of what Fable/Opus stack can find, and often miss the most critical issues. I work where privacy isn't just good behavior, it's enforced by law, and the Fable/Opus stack has found privacy leaks that the openweights stacks don't find.

You can imagine that paying for these Claude runs isn't cheap, each one can eat 25-33% of my 5 hour limit. I am quite desperate for openweights models to be competitive, but at the end of the day, the biggest limit here isn't the price difference between GLM 5.3 max (my current best-in-class choice for open weights, offering Kimi k3 performance for like half the price), it's the cost to the business for shipping lower quality.

Can't wait to dig in more with Astra, I just haven't iterated much on my skill port to codex yet.

One criticsm I have for the article, that is important for my own work, is not simply comparing "bugs found" because these agents can find endless reams of lows and nitpicks that are just ~worthless hardening. I'd be much more interested to see how many critical/high/medium's each test found, not "overall bug count". I also think review is about A LOT more than "finding bugs"...


Hi, I'm looking for a cofounder to build in this space. Will you get in touch if you're interested? Thanks

> I work where privacy isn't just good behavior, it's enforced by law

HIPAA/medical?


Nah 10 years ago I could view the whole website logged out. Now it's been reduced to a single post with no replies.

And remember, Elon only granted that so he wouldn't be (rightfully) banned from Google. At first you couldn't view anything, then Google delisted X because it only indexes public pages, then Elon conceded you can view the direct thing you linked to, and then Google relisted it.

Solar only became "economical" (read: profitable) because a socialist economy dumped a huge amount of money into scaling it up without requiring it to be "economical". It could have been half a century or more ago if we actually cared.

Nuclear was not stopped by the environmentalists, it was stopped by the fact that it cost 4X more than coal at the time. You claim that solar wasn't profitable in the west, thus it didn't take off, surely you can also see that nuclear wasn't profitable in the west, thus it didn't take off as well.

The same country that invested the time and money to make solar profitable is also investing the time and money to make nuclear profitable, with nuclear reactors entering mass production...

It was never about "economical", it was about a system being mature enough to make long term investments. Ours simply can't do that anymore.


> Nuclear was not stopped by the environmentalists, it was stopped by the fact that it cost 4X more than coal at the time.

A big part of that is the insane culture of safety around nuclear. In a sane field like highway engineering or healthcare, you come up with the statistical value for a human life (or a statistical value for a quality adjusted life year), and then use that to make decisions about eg how much money to invest to make your highway a tiny bit safer.

Nuclear is forced to act as if the statistical value of a human life is pretty much infinity. While coal is allowed a finite and rather low effective value for a human life.

> It could have been half a century or more ago if we actually cared.

Interesting. What makes you think so?

Btw, just because China made solar cheap with huge amounts of effort doesn't mean that was necessarily the best use of those resources ex ante.

Just like going to the moon with Apollo wasn't necessarily a good idea. Nor do the wonders of computers justify world war 2.


It's nice to call it "politics, ignorance and greed" but those three are just "Capitalism".

Our system is doing exactly what it is designed to do. Nuclear reactors were never profitable compared to coal or gas, so it never succeeded in strongly capitalist societies, only seeing great success in socialist economies where the people can invest outside of a profit motive.

It's actually quite funny that socialism is saving the day. The mega-capitalist countries turned their backs on nuclear and solar because they were less profitable than gas and coal. But socialist China invested anyway, and now China is mass producing nuclear reactors and every layer of the solar stack. China produces 50% of all nuclear reactors, 90% of all solar panels, 90% of all battery systems for solar storage.

Now that a socialism-based society has proven market viability, suddenly the greedy capitalists want in. But they're decades behind and don't have the private debt appetite to compete.

Womp womp... At least someone is leading the energy revolution.


Agreed on all points. I for one welcome a redistribution of economical world domination. The American system is failing us all.

The term that Anthropic is now using is "mannered prose". If the creator of this example simply prompted "Remove all mannered prose" then the entire experiment would suddenly become normal sounding. In the fable 5.1 prompting guide, they have a longer prompt for removing mannered prose too if needed. I find it works on Astra as well.

https://platform.claude.com/docs/en/build-with-claude/prompt...


This post isn't convincing me. I spent so much time meticulously organizing my techno box. I bought a back of the door shoe holder for tech. Every wire, charger, usb key, web cam, airline earbud, everything. It's been beautifully organized for years.

It's been beautifully organized and useless for so many years that it's all junk now. Nothing is using USB A/B. The airline earbuds are as much trash today as they were then. The usb keys are very old and untrustworthy. Even the USB-C wires are all out of spec, won't charge modern phones, etc. The old webcams look like true trash compared to the modern ones. Etc etc

And this post is basically "hoard all this garbage because in 10 years you might need one thing"? I'm sorry, but "buy that one thing for a few bucks off amazon the one time you need it" is looking a lot more attractive than "meticulously maintain a tech hoard to save $3 on a single use wire"

In engineering terms, it feels like my tech hoard is an automation I spent two weeks writing, for code that only ever runs once every ten years.


Yes, I guess its a balance - the inconvenience of sorting through odd-shaped defunct USB standards vs actually needing one of those weird cables for an ancient camera in 10 years time.

I naturally have the ~< £ from AZ mindset, but I try to challenge that by re-using avoiding waste. It doesn't always work (and there's a lot of eventual waste in those hundreds of zip-lock bags).


>There are numerous benchmarks that measure cost per task, which factors out tokens entirely. Gemini 3.8 flash is significantly lower than Sol on basically all of them https://artificialanalysis.ai/#cost-tabs

Not sure if you read your own link but Sol 56 high ranks smack between Gemini 3.8 flash medium and high. Gemini 3.8 flash comes in as more expensive per task than Sol 56 high according to artificial analysis.

Luna high is literally 30X cheaper than Gemini 3.8 flash high.

You can limit the model viewer and they're getting better at testing multiple effort levels now: https://artificialanalysis.ai/?models=gpt-5-6-sol-medium%2Cg...

One reason is clear: Sol uses dramatically fewer output tokens than Gemini 38 flash https://artificialanalysis.ai/?models=gemini-3-8-flash%2Cgem...


I open the link and I see Flash 3.8 high at 0.58 and Sol at 0.95. I don't understand why you say that "Sol 56 high ranks smack between Gemini 3.8 flash medium and high" but that is clearly wrong.


On cost per intelligence task, Gemini38flash and Sol56 trade back and forth on cost depending on effort level. https://i.imgur.com/zPaWPXx.png As seen in this image, literally: Sol56 high ranks in between Gemini 38 medium and high. The image proves it.

I also included Sol56 xhigh, which ranks above even Gemini38 high.


The government will pay because it's not his money, it's our money. He loves spending our money...


Sonnet 5 is the worst model of 2026. Literally just turn effort slider down on Opus, it's smarter, faster and cheaper than whatever Sonnet is.

Beyond that, I find this whole plan and build thing to be a pointless waste of tokens. If your planner made a detailed enough plan, then the cost of executing that plan is a just one turn more of cached tokens, and minimal time.

Meanwhile: switching agents, reloading context and building from the plan will easily balloon your token use and time. And any emergent problem that the dumb executor finds will instantly wreck the implementation because they're not competent at solving it. And if your plan is so perfect that there's no edge case then you're wasting tokens because your planner was one turn away from finishing the project via cached tokens.


Completely disagree (except the Sonnet bit, yes, it's degrading).

"then the cost of executing that plan is a just one turn more of cached tokens, and minimal time."

This is just not true at all. There's a huge gap between 'figured out the hard stuff' and 'rock solid'.

Dependencies, integration, corner cases, docs, testing, unforeseen issues, a lot of back and forth auditing making sure things are really tight.

Audits get diminishing marginal returns, but you have to do them until they don't find anything, and that's usually a few cycles.

So aside from the fact there is 'a lot of labour' - part of the plan (maybe the most important part) is documenting most of the trip-up scenarios. If you ran an experiment or two in the background your agent will 'discover' a few key odd things, you back those into the plan.

I'm 100% certain that this pattern works because I (and others) use it very successfully.

Hint: save your main context by using sub-agents to do grunt work - even in impl phase - farm out anything directly implementable without a ton of background.

Also - make a skill so your Claude can call Codex and visa versa and maintain long-running sub agents of 'the other kind'.

An Opus with 1M context window executing on a 'plan' that a Codex 'sub-agent' is executing on - ad a different Opus sug-agent is auditing hard ... that 1M token window is dramatically extended to 'many millions of tokens'.

That can work within Anthropic/Codex Pro plans.


I'm sorry, but just because you achieve results you consider acceptable with this method doesn't mean everyone does.

I don't work where we can ship slop. I don't work where PRs can be merged based on what the agents say. I work where a human has to read and approve and own every single line of code. I work where the stakes are actually high, so the cost of not using the best tools in terms of human time are big. A single turn around in a PR costs more in human time than the difference between deepseek and fable in API costs.

So, when you admit "There's a huge gap between 'figured out the hard stuff' and 'rock solid'." but then claim that the cheapest/dumbest agent in your arsenal is your go-to for "rock solid", I have to question the quality of your results.

Personally, "using plan mode" is a very 2025 way of using these tools, and I wouldn't be surprised to see "plan mode" be removed from codex/claude code/et al.

Realistically, I'm using the best models to think about a domain and problem (Fable High+), and I'm using a cheap daily driver with an advisor pattern (Opus High + Fable) to iterate through POCs, and I'm using human review to guide design. None of that is "plan mode", it's actual engineering. Then we decompose the solution, we stack it, and we use only really strong agents to build, review and refine.

This obsession with cheap agents leads to low quality outcomes. "Rock solid" deserves the best tools, and the "plan" will never be good enough. I'm going to be sending fable xhigh and sol 56 xhigh et al at it in adversarial review, why the heck am I cheaping out on the actual implementation?

And finally: my time costs way more than any of this. Cheaper models are slower overall and when combined with re-work time, are dramatically slower. I'm costing my company hundreds in my time to save a few bucks on the API bills. Nonsense!


Yours was the casual dismissal; and based on a misunderstanding of what can be achieved.

Based your arbitrary dismissal and unwillingness to even try to consider new patterns with which you may be unfamiliar - it may be difficult to communicate with you.

I have the advantage of 'certainty' because I have the evidence over many projects / team members.

We ship near perfect code.

In addition to the hints above, we do this at least in part by explicitly anchoring and testing requirements into several aspects of the code, and ensuring that known 'weak spots' are managed.

The 'planning process' ensures the requirements are mechanically anchored and integrated into tests, that 'proportional' documentation is applied, and that module, library and project level documentation is perfect (and mechanically validated where possible), which FYI is what solves most of 'context problems'. (That's another hint, if you have extremely good docs, you don't need to load vast amounts of code).

Yes - I hear you that 'time matters' and that 'the stakes are high' - consider that you may be talking to people where the stakes are just as high, or higher - but more specifically, this is not about 'saving tokens' or cost so much as it is using the right level of model for the task.

Use the best models for background research and planning, use mediocre models for execution, and mid-high for auditing - in other words 'use the right model for the right work' - and in a certain methodology, dumber models are appropriate.

FYI this saves you the ugly 'Fable' problem which many are encountering as it burns though Max plans. Don't 'automate' with Fable, it's the wrong model for that.

I could go on, but consider that there are actually ways of organizing projects and orchestration that work well.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: