Hacker Newsnew | past | comments | ask | show | jobs | submit | z4y5f3's commentslogin

I did SFT / RL post-training on Qwen3 models a bit. This is an issue that dates back long ago.

My favorite theory is that they had many model variants internally, each using a slightly different chat template, so when it comes to the release they are not even sure what to use any more.


Yes, and this is not the first time they messed up. They had tokenizer bugs where the trained weights do not match the template back to Qwen3 series.


Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/

Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high.

I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?


> ... Anthropic's Project Glasswing is supposed to find them quite a while ago?

That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.


> and that Apple would not attribute to GLM

That was a wtf to me, so I checked Apple’s latest iOS release security content and GLM & z.ai is mentioned once (under WebKit), Anthropic is mentioned twice, Codex is mentioned once. Not clear if there are other instances where the model did most of the work but wasn’t credited. I didn’t bother to check other releases.

https://support.apple.com/en-us/128066


> That was my thought too. For all of Anthropic's talk about their "adversaries"

It’s very likely they found all of them, but that the same happened that happened to Microsoft a couple of decades ago: NSA orders not to disclose / fix them so that they can put it in their collection of unfixed zero days.


This is a coherent explanation for why federal model censorship has started with cyber capabilities. But this GLM model release is an in-your-face challenge to that policy. They now have to either set models free or impose a censorship regime that will put anyone not under it at an advantage. Or muddle along in the middle as usual.


> Or muddle along in the middle as usual.

I'm not a gambling person, but if I was this would be my bet.


Then "security through secrecy" is really bad mantra especially in the age of AI: others will find the same zero days very soon. If they attack you, then this loses the whole plot. If they propose a fix, then your arsenal becomes smaller.


Who says they missed them? Could also be sitting pretty in CIA’s long list of ready to go Vault7-like exploits.


Probably Anthropic found them too and promptly got a call from Isreal to stop looking.


> I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?

You have to consider that having an LLM scan for vulnerabilities is hardly infallible. It is a search guided by heuristics and given a large enough codebase, it is unlikely to identify all vulnerabilities.

Personally, I've had Fable 5, GPT 5.6 Sol, and GLM 5.2 all looking for correctness issues in an old abandoned WIP codebase of mine and all of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.


> [A]ll of them found some that the others hadn't discovered. Now, correctness issues aren't the same as vulnerabilities, but the same principle about using heuristics to find defects applies.

This makes perfect sense, but that conflicts with the impression put forward by Anthropic and OpenAI (in particular) that they alone occupy 'frontier model' spots. Frontier models should large dominate their competitors on a capability basis, but if GLM 5.2 (now 5.3) is routinely finding bugs / vulnerabilities missed by Fable and Sol then GLM might be genuinely a frontier-grade model by itself.


“Company hypes own product, downplays competitors” is still a thing with AI


> This makes perfect sense, but that conflicts with the impression put forward by Anthropic and OpenAI (in particular) that they alone occupy 'frontier model' spots.

Not necessarily. Even near the frontier, we don't really have a total ordering of capabilities, but a partial order. And even frontier models make plenty of mistakes. Combined with the randomness inherent in searching large codebases for vulnerabilities or correctness issues, it is entirely plausible that even much weaker models (and GLM-5.2 isn't even weak) can stumble upon issues that stronger models missed.

My current hypothesis – for which I have only limited evidence, unfortunately – is that it is better to have multiple reasonably powerful (but not necessarily frontier) models looking for issues than just one very powerful one. And even then you're likely to miss out on some issues.


Fable 5 is just over two months old.

For normal software it would be as you say, but LLM progress is so ridiculously fast that things go from "bleeding edge" to "eh, you'll do" in about that timeframe, and "eh, you'll do" to "why even bother with this old rubbish?" in the same again.

Or, from a different perspective, we can expect some new frontier model from Anthropic in a week or two, and from OpenAI in a month or so.


There’s no guaranteed progress, tho. Opus 5 is a regression.


I find that LLMs also generate a lot of false positives, or extremely minor issues that don't warrant a fix (that are always overstated by the LLM as very important!). Signal to noise is still not great and requires somebody to wade through and pick out the actual good findings.


> and Anthropic's Project Glasswing is supposed to find them quite a while ago?

We cannot trust a single company to report security issues, it’s good to see competition in that domain


Open source competition no less.


> Anthropic's Project Glasswing is supposed to find them quite a while ago?

Someone still has to run it. The analysis and fix could be someone's machine but not committed / published.


this is impressive and actually matches my expectations in terms of near term AI progress. we are going to continue to seem impressive progress in coding & related, anything where verifiability is scalable in an automated way: https://transitions.substack.com/p/a-quantum-of-ai-progress?...


Interesting... So Chinese models are not so bad?


There's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.


Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.


It's pretty easy for the US to functionally ban chinese models. They only have to target US firms like inference providers or the biggest users, and pretty much the whole domestic market will fall into line. They don't actually care about the final few %.

Regardless of whether or not adversaries are using them, the US has by far the most compute available, and we've now hit the line where major providers are no longer releasing their best models. The public gets the "current" level of intelligence, while the US government gets to control access to the actual frontier of non-public AI. From their perspective, their enemies using GLM5.3 while they have GPT6 and Mythos6 or whatever is a fine trade.

I don't support a ban at all, nor the US's behavior, I'm just pointing out some facts that change the argument.


But the real bad guys will be this final few %, which defeats the purpose. The 99% will be average user which will swing to cheapest AI or easiest to access.


Unfortunately, those in power, pretty much all over the world, lie/deceive themselves and believe they can.


In this case depriving US companies would be the point though, so that's not necessarily a disadvantage.


> Anyone who knows anything realises banning things is a) impossible and

Maybe "It's really hard" is more accurate? We (humanity) for most part basically agreed to ban the usage of various chemical weapons in wartime, which seems to have drastically reduced the usage of it, even though it's still used by shit actors today from time to time. But it's hard to deny that usage didn't decrease after banning it, which makes "banning" maybe not completely useless for certain things.

"Banning" things that can be easily copied over cyberweb transportation pipes feels like an fool's errand though, regardless of what it is. It's just too easy to get around, compared to actual physical items I suppose.


Chemical weapons are not used not because some agreements - they just messy and only good for killing civilians. Also contaminate area and might also kill your own personnel.

If you look at all other banned weapons they are all used against Ukraine by Russia and nobody gives a damn.


> Chemical weapons are not used not because some agreements

We've literally destroyed countless of supplies of chemical weapons because of agreements about not using them, because we all agree they're absolutely horrible:

> The OPCW administers the terms of the CWC to 192 signatories, which represents 98% of the global population. As of June 2016, 66,368 of 72,525 metric tonnes, (92% of chemical weapon stockpiles), have been verified as destroyed.[40][41] The OPCW has conducted 6,327 inspections at 235 chemical weapon-related sites and 2,255 industrial sites. These inspections have affected the sovereign territory of 86 States Parties since April 1997. Worldwide, 4,732 industrial facilities are subject to inspection under provisions of the CWC.

CWC = Chemical Weapons Convention

> If you look at all other banned weapons they are all used against Ukraine by Russia and nobody gives a damn.

If only I had mentioned something about "shit actors using chemical weapons", basically predicting this exact response. But alas, here we are with zero defense about such a sharp point you made.

Of course even agreements are only agreements. What's valuable is what happens afterwards, but it'll take a while to get there in the context of war crimes by Russia and/or Ukraine, but it will be investigated once the conflict is over.


Yes chemical weapons has been destroyed and it's a good thing, but all other efficient and practical weapons weren't successfully banned.

Even EU countries that signed agreements against anti-infantry mines exiting them because how efficient they are at slowing down invasion.


This is different now. US labs and companies are not releasing frontier-level models openly (specially those capable of assisting cyber intelligence work), but commercializing them instead. Thus, any ban would not be symmetrical to begin with, and that is precisely what maintains the balance.


Anyone who knows anything about politics realises that until that actually happens, it works. And you should do things while they work, then stop doing them once they don't work.


I don't think this really works because the Chinese government is going to be incentivised to tip off the US companies to deny the US government those exploits. I guess maybe that's what the open source patch program here is about, making sure banning the models doesn't work because they can just report the exploits without the company running the model themselves.


How does banning the models in the US prevent this?


US companies will have to use models that are restricted from fixing bugs.


Do you actually believe this?


The CIA ran one of the world's largest cryptography companies, for DECADES[1]. Are you truly so naive that you believe intelligence agencies that have more to gain from stifling the discovery of vulnerabilities they know of and use wouldn't do so?

[1] https://www.washingtonpost.com/graphics/2020/world/national-...


You should probably realise that the world has radically changed since then. This kind of thing works when you have a significant lead in the field that makes keeping vulnerabilities open sufficiently low risk for your own side. But if your adversaries have similar capabilities, then the calculation changes.


Has anything changed? Governments are hoarding undisclosed vulnerabilities, using them as they see fit instead of fixing. Every espionage, surveillance, or war campaign (see Russia v Ukraine, US/Israel v Iran etc) is followed by a ton of burned 0-days.

>This kind of thing works when you have a significant lead in the field

No? It works even if the adversary has the same capabilities. It only stops working when everything is fixed.


I believe it is unlikely. (Not because I do not believe NSA is hoarding 0-days, but for many other reasons.)

I'm curious: to any professional vulnerability researchers reading this, what do you think?


I used to call everything a conspiracy theory, but then Glenn Greenwald published "No Place to Hide: Edward Snowden, the NSA and the Surveillance State".

Now i know that reality is worse than the worst conspiracy theorist.


[flagged]


Sadly.


Why would you even believe the opposite? US spooks have been amassing vulnerabilities and relying on them for decades, they literally pioneered it in the 90's if not earlier. Everyone does it now but the US is the biggest of them all. Surely this devalues a lot of what they did. Moreover, the way the US government handled new capabilities, and OpenAI's training policy (they are in bed with the government) just scream "we want to create weapons for cyber-offence and deny them to everyone else"

It might not be the reason, but of course it's a contributing factor.


[flagged]


Good thing I said nothing of that (especially nothing about China). Reread it again to understand you built an incredible strawman and ignored my last sentence.


Well we know that the US government is pushing to restrict access to such models while the Chinese are publishing them for free, so it's mostly a matter of motivations, not the actual facts of the matter. And the USG has a documented history of unsavory behavior (including toward its own citizenry) in that area.

So we might ask if one of the reasons the US is being the bad guy is it's usual spying antics, and we're left asking why China is being the good guy.


Intelligence agencies have been known for exploiting and planting software and hardware Buga for decades, going as far as weakening cryptographic standards or intercepting hardware in transit to implant a backdoor device.

Why do you _not_ believe it's a possibility?


Critical thinking says this is not only possible but likely too.


when in doubt, it's better to assume capitalism than anything else.


They've always been good enough for double digit less money. Always. Anyone thinking "Chinese models fake models built using dirty distillation scam" don't know what they're talking about.

Distillation is just forcing the model to use an exam prep workbook for training instead of generic publicly available textbooks. The models themselves has to be smart enough for that to work. It's the exact same thing as Asian tiger mom double schoolwork strategy, to paint a picture.


To carry on this analogy - do test prep workbooks make you meaningfully more competent in general, or is it benchmaxing? (Versus studying textbooks for a similar time, of course.)


All machine learning is benchmaxxing. It's a big problem.


It all comes down to how aligned the benchmarks are with what you want the model to do. If it's close, good performance. If it's not, unpredictable performance.


Their best models are getting more and more expensive, and still aren't SOTA.

It's almost like there's an actual cost to developing these models, and the Chinese don't have magic dirt that allows them to do it at a fraction of the cost.


Looks like they're going for good PR now, to avoid smearing by the "Western" models. Smart!


I'd love to live in a society where people and corporations do good things for PR.


> I'd love to live in a society where people and corporations do good things for PR.

Maybe so, but I'm not sure I'd like to live in China of all places. (Don't get me wrong. Lotta places I'd like to visit if I ever got the chance, and China's on that list, but to live there? I don't think so.) Maybe one of the Nordic countries?


PR for good things doesn’t make money.


In a similar vein, does anyone know how to classify the kinds of problems that are being found?

Is it possible to build heavier traditional linting to catch whatever is being caught in a more deterministic way? It seems to me that would be far more efficient in the long run (even if the efficiency is only for the AI to know that aspect was already checked).


If you look at the distribution of their findings in the linked post, most of theirs are issues introduced a long time ago, almost all before 2006.

Complete speculation, but I wonder if they and Anthropic are scanning very different codebases and Anthropic's skew would be in the other direction.


amazing! huge clusters in code from the 1980s haha


> but isn't the cost for such a scan getting lower by the week

Not with Anthropic's models!


Wordpress having a high number of vulnerabilities not surprising lol


They will release the weights by 7/27 along with support in vLLM. Stop second guessing. Source: their blog post https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ


Thanks for the link. No need to be so aggressive. The blog with that detail was not live before; and they removed that language from the original link in this post.


They will release the full weights by 7/27 along with support in vLLM.

Source: their release blog on WeChat. https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ


>We are currently working closely with our inference partners and open-source maintainers to align the technical details and ensure the model can be reliably deployed across the ecosystem. The full model weights will be released by July 27, 2026. Further details regarding the architecture, training, and evaluation will be released with the Kimi K3 technical report.

(translated by chrome)

11 days is a long time. It does not take that long to implement inference at providers. In my opinion, seems like they're being pre-emptively cautious about government intervention/review


Actually it does for a massive model, serving it correctly is not easy.

I believe Kimi also does some sort of Q&A and eval for day 0 partners, since early on a long of inference providers just weren’t running their models properly.


Eh, Minimax M2.7 also took a similar amount of time (actually longer) between availability and weights release.


I'm so glad to be wrong!


NVIDIA is likely citing 1 PFlops at FP 4 sparse (they did this for GB200), so that is 128 TFlops BF16 dense, or 2/3 of what RTX 4090 is capable of. I would put the memory bandwidth at 546 GBps, using the same 512 bit LPDDR5X 8533 Mbps as Apple M4 max.


I can't see how this will work in terms of a TDP. 2/3 of the 4090 power would be several times more power than can be effectively cooled in the physical form factor of an Apple Mini. Either there is severe downclocking happening under full throttle, or NVIDIA have come up with more low power design mojo than Apple has been able to muster for the M4 Max.


Based on your evaluation, it sounds like it will run inference at speed similar to an M4 Max and also allow "startups" to experiment with fine tuning larger models or larger context windows.

It's the best "dev board" setup I've seen so far. It might be part of their larger commercial plan but it definitely hits the sweet spot for the home enthusiast who have been pleading for more VRAM.


What they missed is that current scaling laws (OpenAI, Deepmind Chinchilla) are based on the assumption that the model is trained for one epoch. This essentially means that in order to scale compute, you will have to scale the model size and/or the size of the dataset. So Meta cannot simply spend 3.8e25 FLOPs on a 70B model - to do this they must find 86T pretraining tokens which they do not have.

Of course, ultimately we will figure out scaling laws for LLMs trained on multiple epochs of data, but not today.


There is some good published research about doing multiple passes over the training data, and how quickly learning saturates. The TL:DR is that diminishing returns kicks in after about 4 epochs.

https://arxiv.org/abs/2305.16264


Yep I have seen this paper before, and thank you for linking it here for reference. My personal opinion is that compared to single epoch scaling laws, we still need more evidence and literature on effects of multiple epochs, but this paper is one of the best results we have so far on using multiple epochs.


But inside on epoch there is a lot of duplication already.

By duplication I mean if context length is N there is many sequence of N word that are not unique.


My experience is that < 500M models are pretty useful when fine-tuned on traditional NLP tasks, such as text classification and sentence/token level labeling. A modern LM with a 32K context window size could be a nice replacement for BERT, RoBERTa, BART.


NVLink advertises combined bandwidth in both direction, so the 1800 GBps NVLink on Blackwell is actually 900 GBps for everyone else. PCIe can also do multi-node direct transfer via PCIe switches and has been already widely adopted. NVLink still has the power and chip arena advantage even if the bandwidth is similar.


I was just saying they've announced x2 already, around the same pace as PCIe. And it's G-Byte-ps on nvlink, while G-bit-ps on PCIe, right? I'm probably missing something...

Anyway, good read here https://community.fs.com/article/an-overview-of-nvidia-nvlin...


Both the PCIe and NVLink numbers here are full-duplex bandwidth. So no, you're not missing anything, you are correct. NVLink 5.0 is far ahead (but presumably it's not actually shipping yet either). Next-gen PCIe 7.0 512GB/sec full-duplex bandwidth actually sits between NVLink 2.0 (300GB/sec) and NVLink 3.0 (600GB/sec).

NVLink 5.0's 1800GB/sec comes from 18 NVLinks, with each NVLink comprised of 2 lanes of 200Gbps (single-duplex). So in normal single-duplex numbers, each NVLink is 400Gbps, and 18 links together provide 7200Gbps. When advertised as full-duplex bandwidth that's 14400Gbps (or 1800GB/sec).

In comparison, PCIe 7.0 uses only 128Gbps lanes: 128Gbps x 16 lanes x 2 directions = 4096Gbps (or 512GB/sec).

PCIe has always been way behind cutting-edge serdes speeds. NVidia just made it more obvious to those working outside of specialist HPC and Networking. But we must cut PCIe some major slack, it has to (eventually) work on relatively cheapo hardware from hundreds of different vendors, in a restricted power and thermal environment i.e. consumer devices.


Ah thanks for the detailed breakdown. You're right that PCIe has more problems to deal with.


Depends. NVLink advertises bidirectional bandwidth, whereas PCIe and standard networking calculate bandwidth in a single direction. So a 1800 GBps NVLink is actually 900 GBps in PCIe and standarding networking terms.

Therefore, a 512 GBps PCIe would sit between the current H100 NVLink (450 GBps) and next generation B200 NVLink (900 GBps). With that being said, NVLink still has lower power draw and smaller chip area, so it would still have a competitive advantage even if the bandwidth is similar.


I think you've mixed your Bytes and bits here. Not sure where the power and area numbers come from but I'd expect NVLink to be worse on both counts for similar bandwidths since they shipped much, much earlier.

PCIe 7.0 is 512GB/s and not 512Gbps. 512GB/s is the full-duplex bandwidth. It sits between NVLink2 and NVLink3.

I gave a fuller breakdown in a sibling thread.


Checked your numbers in another thread - excellent breakdown. Thanks for the clarification.

I did not read the original anandtech post so I did not realize the 512GBps already refers to the full-duplex bandwidth. You are right that PCIe 7.0 x16 sits between V100 and A100 NVLink.


Wut, I did not know that regarding NVLink marketing gimmicks! Thanks for the info.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: