If true, there was no verification of the targets selected by AI, meaning AI is the final arbiter of the deadly use of force. Another Rubicon passed.
Trump pretends like it didn’t happen: “We might never know what really happened there” is a frustrating lack of responsibility.
Still missing, as it always is from all of these kinds of articles, is: How common is this error rate compared to pure human evaluation and decision-making? I’m not saying I support AI-assisted killing, but why is nobody comparing it to human failure rates? If the article is true, then there were 1,000 targets hit in the first 24 hours. If the majority of the legwork and selection was done by AI, how many schools, or equivalents, would one expect to be wrongly hit if all the targeting work were done by humans only?
The correct title: Researcher brakes one specific stubborn historic enigma message with good help from Astra.
Stubborn for a long time because the message used a completely different key from the rest of that day's traffic. Everyone assumed it shared the daily key. The original transcription had errors. The left rotor turned over at letter 72, which is rare and breaks standard crib attacks.
What is cool, if true, is that it was a 2 day collab between the Leffer and Astra. To me this shows the importance of human in the loop, was still all also showing how immensely power of llm tools. But I think it’s getting a bit silly how much anrticles ignores the driving force (the person) in breakthroughs like this.
"Carter Leffer only directed GPT–6 Astra to see if it could break any of the unbroken Enigma messages published on the Crypto Cellar Research web page."
Now we just need this as a service. Another LLM that would encourage your agent like a cheerleader and provide emotional support and reassurance if necessary
I agree with this. I think the researchers who's harnessing the llm's power should be credited more than the model itself. We also need to understand the thought process and the prompts that are given to the model so we can learn and collab to ensure humanity's progress as much as the llm itself.
"After analysing the unbroken messages on the website, it decided that the most promising message was Nr. 172, MVUEH and it also quickly suspected that the plaintext of Nr. 173, SIPVX ..."
Even if this is true, we must avoid falling into the trap of Kasparov of betting on Centaur Chess.
Just like with Kasparov's Centaur Chess, the idea of a 'human in the loop' is just a necessity due to current limitations.
There will hopefully (?) come a time one day when human beings provide only ultimate value judgments, and everything else is done by machines. Or it may not.
But I don't think betting your ego on the idea that you will be useful in the loop for very long is very wise.
While this is a reasonable analogy, engines became better than humans in the late 1990s, and engines became better than centaurs in the early 2020s. Could AI-powered mathematics improve faster than the ~25 years it took for chess? The AI labs are certainly hoping it does, but that's far from a guarantee.
Might be a bad analogy, because early chess computers were purely algorithmic and worked despite their poor board state evaluation heuristics that humans used to be able to augment.
Compare with go (boardgame): Centaur go was basically not ever a thing.
Other applications behaved kinda similarly (AI Starcraft/Dota/...), where we had decent "human-like" heuristics from the get go and the Centaur concept could never really shine, much less for a decade or more.
I'd also like to stress that past progress in this mainly happened for the love of the game, while the (economical) incentives to replace human office workers are... high.
This is a bit too future-oriented. Let's not mix up current capabilities and speculation about future capabilities. For the time being, collaboration works well. What the future brings is uncertain.
I'm still in my 20's so I feel some necessity to be future-oriented.
I do expect centaurs to outperform other systems for many types of tasks for years to come (and am kind of betting on this to keep getting paid).
But what I'm talking about is ego. It your ego is tied up with (a) your intelligence or (b) your ability to perform task X; you will probably be humbled this century.
Why do you want machines to replace human intelligence and ability? I find that dystopian. AI could have meant Augmented Intelligence (which was proposed a long time ago), not let's see when we can replace all human activity.
One is humanist, the other is anti-human, (in the end goal at least).
> There will hopefully (?) come a time one day when human beings provide only ultimate value judgments, and everything else is done by machines.
In their previous grandparent post. They do reserve a place for human value judgement, but I doubt even that remains if everything else has been automated. The machines will decide what we value. We already see that to some degree with algorithmic engagement and targeted advertisements.
The (?) is meant to express doubt that it is a good thing on net. But regardless of whether it is, it is what is likely to happen, and there will certainly be some good things that come out of it.
Seems like machines being able to replace human intelligence and ability will be a dystopia in the short term, but it’s the only route to a long-term utopia.
> If there is one thing I have realized after reading quite a few science fiction books, it is that I don't want to live in any one them - not one.
I find this suprising kinda. If I got a choice right now, there's a bunch of dystopian Scifi that I would instantly go for (out of sheer curiosity), e.g. the Murderbot universe.
Are there fantasy worlds that you would want to live in? I feel this is a bit of a suspect benchmark in the first place because books typically want some kind of tension/conflict which you won't get if everyone is just gratefully living their best life.
It's good to think about where you're going, but you also have to keep track of where you are.
(Also, getting people to think about the future rather than the present is a classic con. Looking at an empty field: "can't you just see the potential here?")
It seems I was wrong in this instance with regard to the "colab" part.
I found that Leffen even said the explanatory website took about 99 times more effort than the codebreaking itself. And he said that he set the direction and pushed, and the model did the execution. How much steering "pushed forward" involved is not disclosed anywhere, but in this instance, it seems to be more a case of "Human pointed at hard task and AI did an awesome job mostly by itself." Tho how much he was a simple meat-ralph-loop is not entirely clear.
> The correct title: Researcher brakes one specific stubborn historic enigma message with good help from Astra.
That's not correct for the content.
"However, the most astonishing thing about this break is that the GPT–6 Astra did it entirely on its own. Carter Leffer only directed GPT–6 Astra to see if it could break any of the unbroken Enigma messages published on the Crypto Cellar Research web page."
*breaks, and also, your conclusion ignores the words typed by the article's author in the piece you presumably read, where it is reported that Astra did it mostly on its own.
Therefore Astra could also have done this comment better
So the LLM would have done all of this on its own? Why is it ok to acknowledge the human was needed but it’s not a collaboration? Is there a defined percentage of ownership required to make the word collaboration valid?
If you get someone to build you a house and they do it on their own, does that not count because they wouldn't have done it if you didn't pay them to do it?
Technically you built it yourself and the builder was just a minor collaborator?
"However, the most astonishing thing about this break is that the GPT–6 Astra did it entirely on its own. Carter Leffer only directed GPT–6 Astra to see if it could break any of the unbroken Enigma messages published on the Crypto Cellar Research web page."
I mean... I'm all for collaboration but I think this case is pretty clear, no?
I mean, the LLM could do it even without all the HUMAN knowledge that was stealed during training about the Enigma machine?
We are fooling to me, there is no intelligence in these models, they just apply methods that were invented by humans without any consciousness on what they are doing.
You could say exactly the same about humans. Ex nihilo nihil fit. Every human depends on a vast corpus of prior human knowledge to be able to accomplish anything. This doesn't mean they have no intelligence.
It is much closer now. Just like how C/C++ devs used to say "JavaScript isn't real code" or "Python script kiddies". Code was literally designed to be just easy enough for people to understand, now it's just even easier to understand.
It is not yet the same as C++ or Python, for two reasons:
1) Ambiguity is still the default. Formal languages force you to resolve it up front.
English lets you paper over it until the model or the compiler (the human) notices.
2) The "compiler" (the LLM) is statistical and non-deterministic. Same prompt, different day, different bugs. A real language has a spec.
The practical move is to treat English as a high-level specification language, keep the generated artifacts inspectable, and "still know enough of the lower layers to notice when the translation went wrong."^1
[1] This is the key that is where humans can still be necessary, or at least another pass through the LLMs to decide on the best path, in the compiled code. Compilers for other languages do the same C -> Binary, etc.
A conventional compiler is bound by an as-if rule. It can take many internal routes, but the observable behavior has to match the language spec. Same source, same defined semantics. If two gcc runs emit different binaries, the program is still supposed to compute the same answers on the same inputs. That is why people treat the source as the artifact and the binary as disposable.
An LLM compiling English has no as-if rule unless you add one. "Sort the users by last active" can become a stable sort, an unstable sort, a SQL order by, an in-memory timsort, or a query that drops people with null timestamps. All of those can look like success. They are different programs. The model is not optimizing under a spec. It is filling in the parts you did not write.
A human who can read the destination language still notices when the chosen path is the wrong program.
A second model pass can compare paths, but only if you give it a way to score them: tests, types, invariant
So the historical analogy still holds, with one correction. JavaScript and Python were dismissed for being too easy, but they already had grammars and evaluators. English is easier still, and the evaluator is a statistical translator that will invent a dialect if you let it. The practical move stays the same: treat English as the spec language, pin the generated artifacts behind tests, and keep enough fluency in the lower layer to see when the translation chose a different program than the one you meant. The human is not required because the computer is weak. The human is required because the source language still leaves room for more than one destination.
A cyclist pedaling up a mountain isn't a "collaboration" between a bicycle and a human. This is the same. You don't see feral bicycles roaming the land. All models are ultimately built and run by humans, with human-provided instructions. And as with any program, it's garbage in, garbage out.
More apt analogy here: a cyclist pushing a bicycle down the mountain and seeing it somehow get down the whole track without falling down, is not a collaboration between a bicycle and a human. The human was not involved beyond giving the initial push.
Cyclist still chooses the time, mountain and direction the bicycle gets pushed in. Bicycles have no agency and only go down because gravity. Bicycle will not "discover" tree or wall, any outcome solely the result of cyclist's decisions even if thrown bicycles don't have generally deterministic paths. Don't anthropomorphize the bicycle.
If I told you to pick an encrypted message from the web site and decrypt it, and you went and did it with no further input from me, would you okay with me calling it a collaboration to decrypt the message?
The bicycle -- a simple method of transport powered entirely by humans -- used analogically to prove a point about [clears throat] automation.
I think this sounds somewhat less silly in English because "automotive" and "automative" don't have the same hyper-visible affinity, but all the same; you may want to consider the car as a more viable analogand.
What do you even mean it wasn't a collaboration. At any meaningful level LLMs just plain out suck when left unguided.
The shortcomings should really be obvious by now to anyone honest. And the marketing distortion being oushed out is just tiresome and detrimental for all of us.
What's tiring is the constant snide dismissiveness of the advances in technology by people who don't even RTFA, but yet show up on every AI thread and spout the same nonsense.
My thinking is that Passkey is an extremely good and, in theory, extremely easy solution to security, authorization, and login issues. The problem is that Google, Apple, and Microsoft have done a just horrendous job, and made it suck as much as possible.
They didn't want to cooperate, and they wanted to make passkeys transferable within you cloud account, while not cooperating with anyone or anything else. The result was that you have no predictable and stable pattern/protocol/interface, or even general description, for how, for instance, a website connects to the passkey or even a hardware key, if you wanted it.
We basically have all the browsers, the operating systems, and the password managers, all fighting over who gets to store and present the passkey. And everybody assumes that they are the only one that exists and actively tries to fight the others is they can.
The basic technology is really good and could work well, but the large asshole tech firms focused on self-interest and walled gardens and made it insufferable.
And nothing of value was lost. God, I hate everything about Salesforce. Sometimes I have to integrate against their services, and it is always a pain, not to mention what the core project actually is: optimization of marketing and spam.
Sandboxes drift from production in both metadata and data. Refreshes change record ids.
API access is gated by edition. Professional edition has none by default. Salesforce sells MuleSoft as the fix to their own shit architecture.
Custom objects and fields mean that connecting successfully doesn’t establish what represents a customer, subscription or completed sale. Even standard objects can be used differently between organizations.
Metadata deployment is SOAP-based and slow. Dependencies between components break deploys and is massive pain to debug.
And notice that I said integration, not API integration. There’s a shit storm of terrible way to integrate beyond just simple API. Have fun with Bulk 1 or 2, Composite, Streaming, Platform Events, Change Data Capture, Tooling, GraphQL. An my most hated item, anything web based.
I seriously want to talk to someone that have a good time integrating against Salesforce. Maybe I will have my mind blown as to how easy it could be, but I suspect I will have the same experience I have every time I critical scrutinize and such situation: it’s just as terrible as it seems and the person claiming it’s easy is hands down producing close to nothing of value.
> Sandboxes drift from production in both metadata and data. Refreshes change record ids.
I find it odd you would expect sandboxes to not drift and record ids to not change on refreshes. People are working on sandboxes so they will naturally diverge from production, same for record ids.
Refreshes mean that data will be wiped out so you have to preserve the data you want first, then restore it after the refresh. Isn't that the same with any other database where you dump a copy of production on an instance?
I don't see a difference between custom objects in SF and any plain SQL database table that needs to be changed to accommodate a new requirement.
Dependencies between components break deployments - isn't that universally true?
What are other platforms where such things don't happen?
No. The whole discussion betray an insane lack of basic understanding of of LLMs and what reasoning, layers and the processing architecture does as opposed to predicted token collapse.
If so, the very action of feeding forward through the layers are hidden reasoning. There is nothing about looping the processing though the same layers a set amount of times, that is any different from copy/pasting the layers and processing it though the same weight. Except it would be stupid waste.
I really don’t understand how this is misunderstood by people that should know better.
Another way to point out the silliness.
Raschka's own argument: his Luna vs Sol point shows that ordinary added depth already shifts computation into latents, and nobody called that hiding.
I’ve never been more baffled by an Apple event in my life.... We in Norway are part of the first batch of the rollout for the new AI/language feature?!??!?! Something’s fishy. Somethings up!
If the fold make up 4% of sales, but with 2x or more margin vs mini, it might be worth it for them. Seems like a product that MIGHT survive due to that.
+ the pricing anchor drags the others models up and makes them look "better value" ... vs a lower price model selling 4% but anchoring the price downwards.
For anybody wondering what this is about... It’s a newspaper article about a famous old artwork coming back to London from France for the first time in 950 years, with some musing about how history gets distorted over time (history is written by the winners and all that).
If true, there was no verification of the targets selected by AI, meaning AI is the final arbiter of the deadly use of force. Another Rubicon passed.
Trump pretends like it didn’t happen: “We might never know what really happened there” is a frustrating lack of responsibility.
Still missing, as it always is from all of these kinds of articles, is: How common is this error rate compared to pure human evaluation and decision-making? I’m not saying I support AI-assisted killing, but why is nobody comparing it to human failure rates? If the article is true, then there were 1,000 targets hit in the first 24 hours. If the majority of the legwork and selection was done by AI, how many schools, or equivalents, would one expect to be wrongly hit if all the targeting work were done by humans only?
reply