Rendered at 11:08:08 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
z4y5f3 3 hours ago [-]
Apparently they are scanning OSS and popular software at scale and disclosing the vulnerabilities they found: https://cvd.z.ai/
Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high.
I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?
SyneRyder 2 hours ago [-]
> ... Anthropic's Project Glasswing is supposed to find them quite a while ago?
That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.
stingraycharles 17 minutes ago [-]
> That was my thought too. For all of Anthropic's talk about their "adversaries"
It’s very likely they found all of them, but that the same happened that happened to Microsoft a couple of decades ago: NSA orders not to disclose / fix them so that they can put it in their collection of unfixed zero days.
chvid 2 hours ago [-]
Who says they missed them? Could also be sitting pretty in CIA’s long list of ready to go Vault7-like exploits.
dgellow 31 minutes ago [-]
> and Anthropic's Project Glasswing is supposed to find them quite a while ago?
We cannot trust a single company to report security issues, it’s good to see competition in that domain
ofjcihen 10 minutes ago [-]
Open source competition no less.
ThouYS 1 hours ago [-]
amazing! huge clusters in code from the 1980s haha
sscaryterry 2 hours ago [-]
Interesting... So Chinese models are not so bad?
zorked 2 hours ago [-]
There's a chance that the real reason why they want to ban Chinese models is that they are so good at fixing bugs and preventing exploits that intelligence agencies have been using for espionage and surveillance for a long time.
mcintyre1994 2 minutes ago [-]
I don't think this really works because the Chinese government is going to be incentivised to tip off the US companies to deny the US government those exploits. I guess maybe that's what the open source patch program here is about, making sure banning the models doesn't work because they can just report the exploits without the company running the model themselves.
andy_ppp 22 minutes ago [-]
Anyone who knows anything realises banning things is a) impossible and b) your enemies will use them anyway, you are just depriving your own side of the advantages.
cyanydeez 40 seconds ago [-]
when in doubt, it's better to assume capitalism than anything else.
ajam1507 12 minutes ago [-]
How does banning the models in the US prevent this?
maipen 2 hours ago [-]
Do you actually believe this?
VulgarExigency 2 hours ago [-]
The CIA ran one of the world's largest cryptography companies, for DECADES[1]. Are you truly so naive that you believe intelligence agencies that have more to gain from stifling the discovery of vulnerabilities they know of and use wouldn't do so?
I believe it is unlikely. (Not because I do not believe NSA is hoarding 0-days, but for many other reasons.)
I'm curious: to any professional vulnerability researchers reading this, what do you think?
orbital-decay 1 hours ago [-]
Why would you even believe the opposite? US spooks have been amassing vulnerabilities and relying on them for decades, they literally pioneered it in the 90's if not earlier. Everyone does it now but the US is the biggest of them all. Surely this devalues a lot of what they did. Moreover, the way the US government handled new capabilities, and OpenAI's training policy (they are in bed with the government) just scream "we want to create weapons for cyber-offence and deny them to everyone else"
It might not be the reason, but of course it's a contributing factor.
budsniffer952 55 seconds ago [-]
>Why would you even believe the opposite?
So we are clear, the evil Americans are banning the best models because the CIA wants to maintain software vulnerabilities, while the good guy Chinese (who would never hack anyone), and scrambling to catch up so they can fix the world's software issues?
You believe this not only plausible, but probable?
Unbelievable
sscaryterry 2 hours ago [-]
Critical thinking says this is not only possible but likely too.
numpad0 33 minutes ago [-]
They've always been good enough for double digit less money. Always. Anyone thinking "Chinese models fake models built using dirty distillation scam" don't know what they're talking about.
Distillation is just forcing the model to use an exam prep workbook for training instead of generic publicly available textbooks. The models themselves has to be smart enough for that to work. It's the exact same thing as Asian tiger mom double schoolwork strategy, to paint a picture.
cromka 2 hours ago [-]
Looks like they're going for good PR now, to avoid smearing by the "Western" models. Smart!
croon 2 hours ago [-]
I'd love to live in a society where people and corporations do good things for PR.
blooalien 1 hours ago [-]
> I'd love to live in a society where people and corporations do good things for PR.
Maybe so, but I'm not sure I'd like to live in China of all places. (Don't get me wrong. Lotta places I'd like to visit if I ever got the chance, and China's on that list, but to live there? I don't think so.) Maybe one of the Nordic countries?
vjvjvjvjghv 42 minutes ago [-]
PR for good things doesn’t make money.
re-thc 2 hours ago [-]
> Anthropic's Project Glasswing is supposed to find them quite a while ago?
Someone still has to run it. The analysis and fix could be someone's machine but not committed / published.
aliljet 6 hours ago [-]
This is absolutely still shy of Sol and Fable, but only just by a hair. Ridiculous results. There's still not a compelling economic reason to drop OpenAI courtesy of the ludicrous reset addiction that's taken place, but it feels like we're on the precipice.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
MangoCoffee 5 hours ago [-]
OpenAI and Anthropic are both seeking trillion IPOs, while Chinese labs are pumping out open-weight models that are free for US providers to host and monetize.
These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin.
I just don't see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast.
Gigachad 3 hours ago [-]
This is going to be catastrophic.
Whether AI works or is useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations.
dhx 2 hours ago [-]
I take it from [1] (transcript of recent DeepSeek CEO discussion with investors) that DeepSeek would disagree on the immediate catastrophic impact to the likes of OpenAI or Anthropic. The reason is even though technology parity mostly exists, only OpenAI, Anthropic et al have the inference capacity to gain market share and generate revenue. Chinese vendors don't have the chips needed to scale up inference and gain market share, and the DeepSeek CEO doesn't think this would happen in optimistic circumstances in the next 3 years, but thinks it might be possible in 5 years.
In summary, regardless of country of origin, availability of inference capacity is the moat protecting the likes of OpenAI and Anthropic, not technology superiority.
That merely pushes the valuation onto the hardware makers, not the companies that have the temporary preferential access to their hardware.
zarzavat 2 hours ago [-]
That makes them at best temporary middlemen.
It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't).
Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.
regularfry 1 hours ago [-]
The question mark in my mind over the technological superiority is whether the additional volume of data they see due to capturing the top of the market allows them to do recursive self-improvement in a way nobody else can match, before any of the other labs can figure it out. That's the only runaway outcome I can see.
zarzavat 13 minutes ago [-]
If you have exponentially increasing use of your harness, then it's true that every day you capture exponentially more data, but it's also true that every day exponentially more data will slip through the cracks of your would-be monopoly and that data arrives at your competitors via various channels (competitor harnesses, subsidized reselling, etc)
The very exponential that you are relying on to give you runaway improvement is also giving exponentially increasing data to your competitors. All else being equal your competitors stay a step behind but you never develop a monopoly either. That's the best case for Anthropic/OpenAI. In reality, training data is just one variable, exponentials don't last forever, and your competitors will get better at capturing a bigger slice of training data.
anon373839 20 minutes ago [-]
Yes, RSI seems to be the new AI industry McGuffin of 2026, just as agentic capability has become table stakes and scaremongering has become a punchline.
goolz 3 hours ago [-]
I have already begun winding down my spend on claude and OAI to make room for infra budget. Anecdotal, but I have no doubt a lot of others are doing the same, I very much agree the US players have major issues looming. What an exciting time to be alive!
netdevphoenix 3 hours ago [-]
Not exciting for anyone directly or indirectly invested in a frontier lab or its partners. And that is a lot of people, including you.
gruturo 3 hours ago [-]
No time like the present to pull out and reduce your exposure. I brought this up in my employer's forums 4 months ago and honestly it's been clear even before then. In particular, the upcoming IPOs of both oAI and Anthropic will likely be disastrous for the public - the floor is falling from under them and I don't know if they can be scrappy and work with fewer resources - their internal culture may not support this. We all knew in our hearts they're a commodity - just see how easily you can switch between the 2 of them - and now there are 10 more options costing a fraction.
When Xi Jinping did the announcement of their open weights push, they might as well cancelled their IPOs....
topato 2 hours ago [-]
Yes buuuut…. I do quite a bit of day trading (maybe closer to scalping) for the first few hours the market is open, everyday. Anecdotally: despite everyone knowing its valuation was ridiculous, I rode that SpaceX train pretty hard and made a pretty penny.
I close-out all my positions by end-of-trading everyday… so when the day came when there was a very clear and very scary indicator during early trading hours, quickly followed by SpaceX’s catastrophic fall right after opening bell, that was the end of my involvement….
And I fully expect oAI and anthro to be the same way. They’re being propped up with private loans, subsidies, and other tricky bookkeeping techniques. You would think their CEOs would pivot away from their current public personas. Ironically, they are like a poor man’s Elon Musk… and that doesn’t bode well for their companies
gruturo 1 hours ago [-]
Yes experienced investors will profit from it and leave the general public holding the bag, that's the plan I'm afraid.
zozbot234 1 hours ago [-]
The frontier labs will do well if they pivot their offering towards more capable, larger-scale models that are inherently harder to both train and deploy for commodity suppliers. Their existing investments in gigawatt-scale datacenters are quite optimal for this. "Commodity" inference need not comprise the whole market.
anon373839 5 minutes ago [-]
I don’t think this works, for a few reasons. First, intelligence gains from scaling the models bigger is sublinear now. So they could eke out a little extra performance, but the increased cost will eventually eclipse the economic value gained from this.
Second, humongous models are impractical even for them to deploy widely. They’re best used as teachers for smaller, more efficient models that can crank out the volume they need to sell.
Finally, there is a data wall. Sure, they can keep scaling RL on math problems and code. But with everything else, where will the supervision come from when they need several orders of magnitude more?
csomar 2 hours ago [-]
The car industry is also a trillion $$ market in the US. I don't see why that would go any differently from the Chinese cars ban.
iinnPP 2 hours ago [-]
You wouldn't download a car, would you?
visarga 23 minutes ago [-]
I would download it if I could, no question about it. And 64GB more RAM if I am at it.
csomar 1 hours ago [-]
Most of the money will come from companies/corporations who will be required to buy safe AI. The public will be just banned from buying which might make it hard (ie: site/payment blocked) but not impossible. It could be good enough for the big whales.
noman-land 2 hours ago [-]
You're comparing an entire industry to a single company.
cromka 47 minutes ago [-]
I seriously need to start considering the scenario in which this leads to next global financial crisis.
dgellow 19 minutes ago [-]
Just keep in mind that it can take a whole for things to play out. I’m someone who believe the US AI industry is completely unsustainable and built on sand, and will crash even if the current AI itself turns out to be very successful. But that doesn’t mean everything will burn to the ground next week. In a history book things will look very sudden but at normal speed that can easily take months to years to fully play out.
Also, take in consideration that the AI trade infected a lot of other trade in the economy, if you decide at some point to move your money to a place that is safe in case of a downturn be sure to carefully evaluate that’s actually the case
jcfrei 1 hours ago [-]
A lot of the performance of these open source models might come from distilling the closed frontier models. If those can't raise the funds anymore to train newer and better models then the whole improvement cycle might slow down.
tjpnz 34 minutes ago [-]
Does stealing from a thief still amount to theft?
somenameforme 4 hours ago [-]
Another interesting potential market here will be 'LLM in a box'. All the hardware and other tooling in a prebuilt, but modular, package ready to go. Pay one up-front cost, get a system running [whatever open LLM] with a token rate of [x], optionally configured to be immediately ready for distributed usage. Basically the opposite of cloud stuff: no rent, no dependency, 100% guaranteed uptime, guaranteed security/privacy (at least subject to your own actions), and so on.
adrian_b 3 hours ago [-]
Palantir already offers a "turnkey AI datacenter", i.e. a rack with "NVIDIA Blackwell Ultra systems with eight NVIDIA Blackwell Ultra GPUs and NVIDIA Spectrum-X™ Ethernet networking for AI training and inference".
It is said that it comes with all hardware and software required to run inference or training with an open weights LLM.
The existence of this product, which competes with cloud-based offerings like those of OpenAI and Anthropic, is presumably the reason why the Palantir CEO criticized very harshly some time ago the business model of OpenAI/Anthropic.
While I doubt that the ethics of Palantir is any better than of OpenAI/Anthropic, in this particular case I have to agree with Alex Karp about "Sovereign AI", i.e. that only losers will make their business completely dependent on an external entity like OpenAI or Anthropic, who are certainly not trustworthy.
rxyz 2 minutes ago [-]
How is this different from buying a supermicro rack? Better support?
vrganj 3 hours ago [-]
I'm not sure a data center run by ... Palantir of all organizations is what people have in mind when they worry about data sovereignty.
adrian_b 3 hours ago [-]
They are selling it, not running it.
It is just a dedicated computer system, which should be managed by its owner, like any other on-prem servers.
I doubt that it has a good price/performance ratio, but it is a solution for those who feel that they do not want to search, buy, assemble, install and configure every HW/SW component.
iinnPP 1 hours ago [-]
I take it we saw different demos.
I'm under no NDA, if you actually want to know what's up.
XVII 29 minutes ago [-]
I want to know, please tell us
Panoramix 2 hours ago [-]
For those that don't mind a lot of rootkit and embedded spyware you mean
crossroadsguy 2 hours ago [-]
Fair but the idea of "running your LLM setup" at every "need" level and corresponding cost does make sense.
For a lot of people (and orgs I'd guess) who just go and buy ≈$20 per month plans (or more for teams), they might not even need a fraction of that cost or capability. A lot of them don't even need it for coding or graphics. Even the API access based pricing aren't great from these frontier US AI houses. The distribution of "LLM being" offered will also give rise to many open-router like offering but at the end point level - direct interfaces to the customers. Pick your vendor sort.
AI shouldn't become another "search means Google".
numpad0 21 minutes ago [-]
What makes that kinda complicated is that multi-user throughput of LLMs scale well but single-user performance often stays constant at low ends. If you could saturate e.g. 16 concurrent session-month of demand, you can just go buy 16 of 32GB GPUs and start charging monthly for inference. That could work if you had e.g. over thousand total employees with hundreds of devs eager to trying it out, but only if the company is also interested in a private inference experiment.
bevekspldnw 4 hours ago [-]
“100% guaranteed downtime when you least can afford it and the support tickets are your problem.”
We’ve a hybrid shop, including hosting our own ML infra, and we save a ton from cloud spend with local ML. Easily one million USD over past three years. But it’s not “free”, you are shifting a lot of labor into your plate.
hypfer 3 hours ago [-]
And with that also gain institutional knowledge, skill up your workers and attract talent that wants to work on this stuff.
All boils down to short-term/long-term thinking.
gruturo 3 hours ago [-]
This. People WANT to work on this stuff. And having skilled workers is a precious advantage.
bevekspldnw 1 hours ago [-]
Still has to break even on the balance sheet, especially at a bootstrapped startup. We actually made most of the financial windfall in translation API fees oddly enough.
For our own model training we needed to do some large scale translation tasks of a large dataset (1M or so documents, 10 or so target languages), running full-size NLLB on-prem saved us an absurd amount of money vs Google Translate API.
(For reference doing 1M target docs into a single language in Google Translate API is roughly $120k list price. You can run full size NLLB on an 48GB NVIDIA A600 and the major difference for us was speed, but for this task time to completion wasn’t an issue.)
jurgenburgen 4 hours ago [-]
> 100% guaranteed uptime
Disagree there but I think this is an interesting idea. We would need to find some more cost-efficient hardware to run it on than Nvidia GPUs.
pulse7 4 hours ago [-]
It will come... all big hardware players (Intel, AMD, Broadcom) and dozens of startups (Tenstorrent, etc.) are working on it...
Godsend69 4 hours ago [-]
[dead]
nkmnz 4 hours ago [-]
I think at this point the question is: will the US government be willing and capable to justify the trillion dollar valuation for _one_ of the companies via regulatory capture? The US has a workforce of 170m, so 1.7 trillion would come down to 10k per person, or a discounted cashflow at 3% of 25 USD per month - not including private use, students etc.
kaashif 4 hours ago [-]
Why would you restrict to the US workforce? ChatGPT has a billion users.
nkmnz 18 minutes ago [-]
It’s a common denominator if you want to do napkin-math for a whole national economy. Regulatory capture is like a tax on those people not on the beneficiary side, so if the government were to nationalize both supply (no export license for SOTA models) and demand (no foreign or self-hosted LLMs allowed), they’d end up making everyone else pay for it in some way or the other. The governmental utility function will then include only those using the services for direct economic benefit.
switchers 4 hours ago [-]
Because it would be the US taxpayers bailing them out.
topato 2 hours ago [-]
They have a stupid plan to buy ten to fifty percent of all the SOTA AI companies, and giving us all a fraction of the money.
Trump keeps calling his enemies “communists”… then turns around and ‘seizes the means of production’ himself.
0xpgm 3 hours ago [-]
US investors are desperate for the next hypergrowth opportunity. From what I can tell the US economic strategy is to outgrow its debt.
andsoitis 1 hours ago [-]
> US investors are desperate for the next hypergrowth opportunity
All investors.
chrismsimpson 4 hours ago [-]
> Providers can just run them, offer cheap tokens, and pocket the margin.
There’s an assumption that you can spin up the infra and acquire customers within that margin
KeplerBoy 4 hours ago [-]
Which is not unreasonable. Just hosting it in the EU and promising not to retain / sell the data let's you charge a healthy extra and compete in many areas other players can't.
andsoitis 1 hours ago [-]
> Just hosting it in the EU and promising not to retain / sell the data let's you charge a healthy extra and compete in many areas other players can't.
It's been a few years. Has anyone done this successfully yet?
layer8 28 minutes ago [-]
There are a over a dozen EU open-weight providers. I’m not sure if they are even charging that much of an extra. EU-based clients have little reason to use non-EU inference providers.
KeplerBoy 50 minutes ago [-]
melious.ai comes to mind.
chrismsimpson 2 hours ago [-]
Keeping SLAs spinning isn’t this trivial
re-thc 2 hours ago [-]
> There’s an assumption that you can spin up the infra and acquire customers within that margin
Only Nvidia and approved friends can at the moment. Nvidia can even backstop your loan required.
grey-area 4 hours ago [-]
It is impossible to justify the absurd private valuations they have given themselves in collusion with investors.
I wish they had tried to IPO because then we’d see the judgement of the market on this. But that’s why they didn’t this year. How long can they keep up the charade that their models are uniquely valuable and on the path to AGI?
andsoitis 1 hours ago [-]
> private valuations they have given themselves in collusion with investors.
What's the collusion?
grey-area 21 minutes ago [-]
Circular investment deals and investment deals at valuations which have no possible justification.
andsoitis 16 minutes ago [-]
Let me ask it differently. You state the companies and their investors are colluding. Who are they colluding against?
miohtama 3 hours ago [-]
There could be soon AI safety regulations that will stop the US to host or use the Chinese models.
ilaksh 2 hours ago [-]
Most are not necessarily free to host and monetize. At least one of them has a license that says if you are re-hosting the model then you need a license with that company that made the model.
me551ah 3 hours ago [-]
I think that explains the race for IPO by the US AI labs, they know that the longer they wait, the less they will be worth.
piokoch 3 hours ago [-]
"I just don't see how you justify a trillion valuation for US AI"
- military applications
- financial applications
- medical
- applied science
In all those cases it is achievable for those who have needed training data, and Chinese are not going to get them easily. US AI Labs are showing: give us the data, we will do wonders, promising "singularity"-level future achievements.
charcircuit 4 hours ago [-]
I suggest you think why OpenAI was worth billions before ChatGPT. The valuation is not about how the current set of models can be monetized.
nozzlegear 36 minutes ago [-]
Could you just tell us why you think they were worth billions before ChatGPT, instead of suggesting that we think on it? You seem to know the answer already, so please share it with the class.
dandanua 2 hours ago [-]
I'm sure US billionaires will find a way to extract those trillions from the public. They're smart, they can handle it. After all, they can ask AI for advice on how to do it.
sidd_sarkar 4 hours ago [-]
Ok
hmmidontknow 4 hours ago [-]
Hmm.. how you justify?
Provoking war, this is how the empire "defends" itself, usually.
I just hope that this time it will get stuck in your throat.
wren6991 3 hours ago [-]
The thing that blows me away is it does this at one quarter the total parameter count of K3 (and 40% active parameter count). There's plenty of room at the bottom.
> How are you all toying with running this kind of thing in a mega quantized way locally?
Sure, let me answer that in excessive detail. I briefly tried running the UD IQ3_S quant of GLM-5.2, which is 288 GiB of weights (301 GB). Setup was: llama.cpp, 1x NVMe SSD (Evo 980), 64 GiB DDR5-5200, i9-13900HX, and 1x RTX Pro 6000. Token generation around 0.7 t/s. Not remotely usable interactively, but something I could plausibly push a codebase into and come back to a review in a couple of days.
There's potential for that hardware to go much faster, but current local inference backends make poor use of the memory hierarchy. Ideally I would have: always-active weights, KV and hot expert cache in VRAM; warm expert victim cache in host RAM; and disk as a last resort. Instead it's 1/3rd of the layers fully pinned in VRAM (all experts), and 2/3rds running wholly on the CPU with mmap()'d weights. The CPU cores spend most of their time sleeping on disk fills.
llama.cpp has backed itself into a bit of a corner architecturally by trying to support all models on all possible backends. If you look into how their "MoE offload" feature works (not viable for me because it requires enough host RAM to permanently pin the weights) you very quickly realise it's "oops, all bubbles!" due to the static compute graph splits. There are more focused frameworks like DS4 [1] and Colibri [2] which have better support for streaming weights from disk, and support GLM-5.2.
Obviously I wouldn't recommend my setup for huge models like GLM-5.2. Supposedly it can just about be squeezed into 3x GB10, or run comfortably on 4x GB10 (tensor-parallel) for multi-user serving. I'm not sure whether that qualifies as local, but it's at least not a rack.
Not sure about Sol as I haven't used it, but, at least for security work -- does it matter? It's not like you will be allowed to use Fable (or access Mythos) for anything cybersecurity-related unless your name is "Dario Amodei" or you are one of his rich friends. So regardless of how good Fable/Mythos is here it's a completely moot point for normal people, because they can't use it for that anyway.
simonjgreen 5 hours ago [-]
We applied for the cybersecurity approval via the form and got approval back in less than an hour. Have you… tried?
112233 3 hours ago [-]
Why should I apply for *cybersecurity* approval in order to have model debug a program it is writing itself? Anything related to memory safety, debugging, syscalls etc (meaning, "programming") somehow is cybersecurity now?
spaceman_2020 14 minutes ago [-]
Your tools refusing to do your bidding is an absurd idea in the first place
Imagine asking for permission to use your hammer
alightsoul 4 hours ago [-]
You must be a 5000 person company with an existing enterprise contract to get approved that fast. That sounds like a 15 minute SLA agreement. Individuals no matter how qualified about cybersecurity, are ghosted
xx_ns 3 hours ago [-]
That's not my experience at all. I was approved fairly fast - around an hour from submitting the form and getting a response.
However, even being in the cybersecurity programme, Fable refuses to answer prompts that it determines could be even tangentially related to cybersecurity. In fact, for a while, I was unable to use Fable with any prompt, as it recalled from memory that I was a cybersecurity professional, which triggered the refusal even for simple prompts like asking for a chili recipe.
captn3m0 3 hours ago [-]
I am guessing you are approved for the Cyber Verification Program. I also applied and got approved in an hour (on a Saturday!), but it only applies to Opus and Sonnet: https://support.claude.com/en/articles/14604842-real-time-cy.... It let me use Opus for cybersecurity work, pretty much everything except for Ransomware development. It would occasionally still trip and start saying no till I added a note about CVP in my claude.md.
No one gets to use Fable for Cybersecurity work, and Mythos is not available under CVP. Only for select few customers, and there isn't an application form?
3 hours ago [-]
kouteiheika 5 hours ago [-]
Have you tried to use Fable for anything even remotely security related, when the refusals kick in as soon as you even fart in the vague direction of anything security or biology-adjacent?
b112 5 hours ago [-]
For this comment to have value, you should indicate whether or not you applied for cybersecurity approval, and were approved or not.
grey-area 4 hours ago [-]
Are there any limitations on this version?
bpodgursky 5 hours ago [-]
I don't understand all this spite about "rich friends" when it was the US government that shut Fable down for not adequately blocking cyber capabilities.
I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it.
deepllm 5 hours ago [-]
"Mythos" is the cyber-security equivalent of Fable (without guardrails), and only a very select few corporations have access to it.
Fable is their version with guardrails on everything except "Make me a pelican svg" or "create a to-do" app, that is the version that the government banned
bpodgursky 5 hours ago [-]
I know all this?
Only a few corporations have Mythos because the US government is whitelisting them one at a time. Anthropic releasing Mythos to the public was never on the table, they would have been shut down in milliseconds by the feds if they tried.
deepllm 5 hours ago [-]
Before the US government had anything to do with this, Anthropic were fear mongering Mythos (BTW, Amodei also fear-mongered GPT-2, so this is a normal pattern in their operation) calling it "too dangerous to release", and back then only Anthropic was in charge of the whitelist.
Then the government believed Amodei's bullshit and this is a result of that, this was all self-inflicted.
bpodgursky 5 hours ago [-]
Sorry but if you stepped back for a moment you'd realize this is all contrived nonsense to let to have your cake and eat it too.
No, Anthropic did not mind-game the US government into being worried about cybersecurity. The NSA has been paranoid about cyber controls for longer than you've been alive. If Anthropic had come out of the gate saying "no don't worry man, our model is TOTALLY COOL", while simultaneously attacking HAWK and finding core Linux vulnerabilities, I assure you the US government would have caught up about ten minutes later and we'd be in exactly the same spot minus your ability to tell Anthropic they were wearing the wrong dress and asking for it.
deepllm 5 hours ago [-]
Mythos isn't some scary dangerous model that can find high severity bugs seamlessly, that's just Anthropic marketing. Most of the vulnerabilities they found were low severity hyped up to make their model look good, with (I think, maybe?) the exception of a few.
Now that Chinese open weight models have similar capabilities, and their guardrails can also just be removed, it doesn't look like anyone has "hacked" into everything because of the scary dangerous models like Anthropic were making it out to be.
d1sxeyes 3 hours ago [-]
In principle I agree but in practice I don’t.
The majority of high severity vulnerabilities are not the kind of thing you need a PhD in Comp Sci to comprehend, they are mostly about finding a way to get a system to end up in a state different than was anticipated when entering a particular code path.
Exhaustively looking at code and identifying ways to do this is something LLMs are quite good at. They don’t get tired, and you can run them non-stop.
They're also (generally) quite good at reading the literal meaning of the code, whereas humans often see the intended meaning first, and can be biased.
If you had a tireless junior engineer who was given the job of “make this application get into a state it’s not supposed to be in”, you’d probably get similar results.
What Mythos is quite good at is both the first bit and coming up with ways it could chain that together with other bits of unexpected state to create something that forms a meaningful vulnerability rather than a dead end.
aka-rider 2 hours ago [-]
All models find vulnerabilities. What is special about this generation of SOTA models, including Mythos/Fable (the same model), GPT-5.6, Kimi-K3, and now GLM-5.3 — they can chain vulnerabilities and produce working exploits.
Look at the recent HuggingFace hack. One vulnerability was template injection, another — remote code execution. Combine them and you pwned the remote server.
People working under Project Glasswing reported that Mythos at one point chained 20 vulnerabilities to produce working exploit.
Humans don’t usually do that.
wren6991 3 hours ago [-]
It's also quite hard to separate Mythos the model from Mythos the campaign (aka Glasswing).
They put an enormous amount of compute into bug hunting, and they found some bugs. Fair enough. For me that begs the question: what if they had spent the same compute on generating more tokens with a less-capable model? What if they had spent it on traditional fuzzing?
kouteiheika 5 hours ago [-]
> I don't understand all this spite about "rich friends"
Okay, here's a challenge: I assume you're not a rich and powerful entity, so try to gain access to Mythos. I'll wait.
> I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it.
Well, first I'd suggest they stop with the constant fear mongering.
Here's my prediction for what will happen: the Chinese models will catch up to Fable/Mythos. They will be fully unrestricted and everyone will have access. The world will not end. Good guys will use them to harden their systems, in equilibrium to what bad guys have access to, so effectively status quo will not change.
bpodgursky 5 hours ago [-]
This is a lot of words to say "you're right, Anthropic does not have any legal way to release frontier cyber capabilities to the public"
nozzlegear 26 minutes ago [-]
And have nobody to blame for that but themselves and their own scaremongering. Dario cried wolf one too many times, and somebody finally believed him.
Of course, Anthropic is after regulator capture, so this all likely worked out exactly as planned.
kouteiheika 5 hours ago [-]
Right, so according to you it's because of the US government that they don't release it to the public? Have you missed their constant and incessant fear mongering?
The causality chain here was not "US government says its dangerous -> Anthropic can't release it", it was "Anthropic is fear mongering -> US government listens to their fear mongering".
stavros 4 hours ago [-]
The issue is that these companies keep trying to pull the ladder up behind them by going "oh my god our models are so dangerous only we should be allowed to develop them". Sometimes it backfires, but the companies aren't innocent.
irthomasthomas 2 hours ago [-]
Have you seen the news about decrypting the hidden COT in U.S. models? [0] The decoded logs revealed instances where Claude memorized answers to test questions beforehand while making its final output look like it had derived the answer step-by-step—hiding the memorization from the user.
Isn't post-training turning out to be the most important part?
andxor 3 hours ago [-]
Fable finished training 6+ months ago.
At this point, Anthropic only needs to release models to the public when the competition forces them to.
OpenAI also has a better model (Astra) that they haven't released yet.
nozzlegear 15 minutes ago [-]
> At this point, Anthropic only needs to release models to the public when the competition forces them to.
Assuming the government allows them to lol
deepllm 5 hours ago [-]
Realistically, you're looking at least 2x DGX sparks to run this at a 2 bit quant, but quantization really lobotomizes models so it's just better to run DSv4 flash at full precision.
4x DGX sparks should let you run this at 4 bit at least and there are some folks who ran GLM 5.2 on this configuration in r/LocalLlama
colingauvin 2 hours ago [-]
For Flash there are some excellent Q2/Q4 hybrids. I know that model was QAT so it handles Q4 better but the meta on quantization seems to be shifting a little bit to be more intelligent about what exactly gets quantized.
teruakohatu 5 hours ago [-]
How fast are 2x or 4x DGX?
I only have one and am wondering what the benefits are of getting another. I feel I will be disappointed…
deepllm 5 hours ago [-]
If you can afford it, another DGX spark is worth it imo. Especially since, owning just one, you have a $1000 ConnectX7 card that's unused. You can find speeds here: https://spark-arena.com/leaderboard
disiplus 5 hours ago [-]
i run flash v4 at 2bit, its pretty great and on my tests against full model It didn't lose any capabilities. It just was thinking more. So you don't have the same efficiency.
bertili 5 hours ago [-]
DwarfStar (https://github.com/antirez/ds4) supports GLM 5.2 and DeepSeek. Not only for toying, but for getting work done.
VulgarExigency 2 hours ago [-]
Since GLM-5.3 has the same base model as 5.2, DwarfStar should support it as well, once the weights are released, right?
teravor 5 hours ago [-]
the difference is that with open models jailbreaking is trivial if you know what you are doing so this makes a frontier open model infinitely more useful for certain tasks seeing as closed frontier models will just refuse (and jailbreaking them is a waste of time when you have good open models).
in some cases (mainly reverse engineering) I have observed GLM 5.2 jailbreaking itself with no effort on my part, the thinking trace revealed that it did some mental gymnastics to pretend it was a crackme or capture the flag competition.
bossyTeacher 5 hours ago [-]
> This is absolutely still shy of Sol and Fable, but only just by a hair.
Even if there was a small/medium gap, the fact that this is a free model beats both of the above on pure economics.
hypfer 5 hours ago [-]
I might be just reading my positive bias into that text, but is it possible that it is written less like SV marketing hype trash and more like researchers wrote it?
It does feel like it respects both me and my time.
Thank you, Z.AI.
Amazing what difference it makes when the top of your org are actual university professors.
this_user 2 hours ago [-]
Would be interesting to compare the Chinese version. Because, obviously, their English version is for users, not for investors or government officials, while the US labs are always addressing those too.
sinuhe69 3 hours ago [-]
I read the same. Refreshingly honest, straightforward and many useful information included. It is a breath of fresh air.
unrvl22 3 hours ago [-]
I was thinking the same thing. It feels truthful, no marketing BS and they call out where they lack behind the best models
aand16 5 hours ago [-]
> Mythos 5 remains well ahead at 181 and 247 tasks. The pattern across the three is consistent: the further up the exploitation chain a benchmark sits, the wider the remaining gap to the closed frontier. Capability is growing fastest exactly where we are furthest behind.
I appreciate they don't just take the opportunity to self-glaze.
aabhay 5 hours ago [-]
[flagged]
jjcm 4 hours ago [-]
Same image->html test as I showed in the Gemini 3.7 flash thread. Note that GLM isn't multimodal, but it still was able to generate something similar-ish by writing a python script to inspect the image and extract elements from it.
For having no vision, it did a tremendous job. I'm pretty impressed it was able to extract so much detail.
The Opus one is still significantly better, but that's to be expected since it's multimodal. Curious to see where a future version from Z.ai lands on this.
ArvidSu 2 hours ago [-]
That's super impressive given that it doesn't have vision! Intelligence overcomes blindness.
wxw 6 hours ago [-]
> Scaling post-training is all we did for GLM-5.3.
Love this opening line. And wow, great results.
> As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.
kleiba2 5 hours ago [-]
What actually is "scaling post-training"?
FergusArgyll 5 hours ago [-]
More RLVR.
Give it verifiable problems, if it doesn't find a solution move on, if it does, use that as a reward signal.
Gecko4072 4 hours ago [-]
Can’t this be extended quite far? Use a cerebras-served model, use verification techniques to generate and solve millions of problems and then use that as training?
gvkhna 3 hours ago [-]
That’s the whole point, just cost and compute limitations in your way (mostly).
tjwebbnorfolk 5 hours ago [-]
does this suggest 5.3 is the same # of parameters as 5.2?
unrvl22 3 hours ago [-]
which is the bigger headline that people don't realize. this is 744b and its head to head with Kimi K3 (2.8T), smashes DS v4 pro (1.5T). even Opus and Sol are rumored to be 1.5T+ this is half the size!
davidlt 2 hours ago [-]
It's the same pre-training, they are just adding more (+ better) SFT, RL, etc. (post-training). Model internal knowledge cut-off is still the same.
It seems we are doing pre-training every 6 months, and post-training every 4-8 weeks now.
fahrradflucht 5 hours ago [-]
“Today we are releasing GLM-5.3. It uses the same base model as GLM-5.2 — every gain comes from post-training.“
vmware508 3 hours ago [-]
Apple will release M7 MacBook Pros / Mac Minis next year, and they will be able to run free LLMs locally at native speed. All software developer notebooks will be replaced to run local models, saving a lot by cancelling Claude Code subscriptions.
Developers win. Apple stocks will be rocketing. Everything else will go down. You're welcome.
gehsty 2 hours ago [-]
Local vs remote compute is a constant thread in tech history - mainframes and desktops then local and cloud compute (think Google Photos bs Apple photos - one indexes on device the other indexes in cloud). Now we have the next chapter local vs cloud LLM models.
There will always be a market for frontier labs in the cloud based models - these models will always be able to be bigger, and that will likely translate to doing things local models can’t.
Logically also we’ll likely get to a point where RAM drops in price as production ramps up, and local LLM is both capable and cost effective. This feels like it is coming for Siri / Gemini / Alexa personal assistant type use cases.
So I think the local LLM will become a thing in laptops and phones in a year or two, offering PA type use cases. Professional LLM services will likely remain at the frontier (and in the cloud) for the foreseeable.
schleck8 3 hours ago [-]
You'd need the 256 gb memory model which will be expensive because apple has trouble getting capacity (got turned down by cxmt). And even then you can only run a 2 bit quant which is noticeably worse than 8 bit
layer8 13 minutes ago [-]
The RAM shortage situation won’t be sorted out within the next year.
Gecko4072 3 hours ago [-]
They will cost an insane amount as well. Maybe less than subscriptions or tokens. But running massive models on laptops with batteries and poor cooling doesn’t make much sense.
LeBit 3 hours ago [-]
Until hiding PII from the cloud LLM is a resolved issue, running local LLMs will remain a necessity.
There are workplaces that refuse to use LLMs because they fear the devs will expose sensitive data without care.
scotty79 2 hours ago [-]
I don't know why you'd want to burden your laptop with a large model. But I can totally see a new "developer workstation" product that's just a semi-large box that's optimized for running frontier open weights models for one to few users.
Flavius 3 hours ago [-]
> run free LLMs locally at native speed
This reads like a hallucination. What does native speed even mean?
csomar 46 minutes ago [-]
There should be some kind of moratorium on new accounts. HN's always had waves of newcomers, but their impact was always limited. The wave passes and people either get filtered out or adapt. That doesn't seem to be happening anymore, since bots can churn out endless gibberish.
He did answer you though. Native is x10 the non-native speed. 50/50 that's not a bot; though it could be a meat-proxy
flexagoon 27 minutes ago [-]
> There should be some kind of moratorium on new accounts.
OC was registered in 2016 though? What do new accounts have to do with this?
kyxsc 3 hours ago [-]
for example, models running at like 100-150 tokens/second (or faster!) vs 15 t/s
(fable/sol are ~60 t/s, and OpenAI just announced their Cerebras partnership(?) for "ultrafast" mode of 750 t/s)
models aren't able to run that fast right now on our consumer/prosumer hardware. M5 Max for example has a memory bandwidth of 600 GB/s. a 5090 has 3x that, so running the same model on a 5090 is that much faster (provided the model is within 30GB).
running a bigger model on an M5 Ultra is still much slower than running it on a Blackwell chip with sufficient vram, CUDA being a major difference. if apple can bridge this gap, interesting things will happen... and just imagine if M7 Ultra has comparable speeds to Blackwell (or even Rubin)!
If you think M7 will hit even 15% of these speeds you're very optimistic.
andsoitis 47 minutes ago [-]
A hosted instance serves multiple customers at a time. A local model only one.
lmpdev 3 hours ago [-]
I assume they mean same t/sec as a SOTA cloud model
toasty228 2 hours ago [-]
Sure buddy, all you'll end up with is a $10k machine that run gimped models at like 30tok/s for about 5m before the fan kicks in and it starts to sound like a turboprop, while offering maybe 30% of the context size of hosted models.
re-thc 3 hours ago [-]
> Apple will release M7 MacBook Pros / Mac Minis next year
The latest on Apple is TSMC is stuck on the next iPhone due to lack of RAM. Good luck getting any Macs. Memory shortage is getting worse.
bertili 14 minutes ago [-]
This will be roughly on pair with Kimi K3, but using a third of its parameters.
Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third.
Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.
virgildotcodes 6 hours ago [-]
OpenAI and Anthropic need to just go ahead and give people access to the cyber models.
Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.
LeonidBugaev 5 hours ago [-]
Not only attackers. I have to switch to Kimi or GLM even in cases of basic issue triage on my own projects! Current guardrails are ridiculous.
SwellJoe 5 hours ago [-]
I've been building a harness for security work, and had to switch to GPT 5.5 when even Opus started refusing security work. Then 5.6 Sol arrived, and it refuses security work, too. So, I switched to Kimi K3 and DeepSeek for API testing just because it's so much cheaper. But, if GLM is better, I'm here for it, as I think GLM is also cheaper than K3.
Synthetic7346 5 minutes ago [-]
Mind sharing a link?
mindwok 5 hours ago [-]
At least OpenAI seems to want to do that, but the US is now forcing them to go through approvals. Anthropic seems much more hesitant.
bryceneal 3 hours ago [-]
OpenAI seems to understand that these guardrails hurt the good guys. This is why they released Daybreak Blue, which is a step in the right direction (but the model itself is weak as it's just Sol with fewer guardrails). Anthropic seems to believe that harming defenders is worth it if it means they can achieve regulatory capture. They do a lot of mental gymnastics to try to pretend that this is not actually what they are doing. As a result they have lost a lot of customer goodwill, which hasn't yet caught up with them yet, but absolutely will IMO.
surgical_fire 2 hours ago [-]
Can't the maintainers use the same models as the attackers?
The maintainers don't need approval to use GLM.
virgildotcodes 47 minutes ago [-]
They may need approval from their employers.
worldsavior 5 hours ago [-]
[flagged]
KronisLV 4 hours ago [-]
Their coding plan switched to credits, didn’t it? What are the rate limits like, compared to Anthropic or Kimi K3?
I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...
Nowadays, I’d probably go with their Max plan if the rate limits are okay? Anyone using them now?
Oh also unrelated but ZCode was surprisingly good, which is surprising for a tool that came out of nowhere - even some of the critiques in my blog post have been patched out. Sadly they don’t support using Claude Code as an agent so can’t use it like Paseo or Kepler or Agent Orchestrator.
ljosifov 2 hours ago [-]
Wdym "sadly they don’t support using Claude Code"? For the longest time that's all Zai supported - Claude code. I'd run it via
export ZAI_ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic"
export ZAI_ANTHROPIC_AUTH_TOKEN="$ZAI_API_KEY"
claude-zai() {
{ local -; set -x; } 2>/dev/null
ANTHROPIC_BASE_URL="$ZAI_ANTHROPIC_BASE_URL" ANTHROPIC_AUTH_TOKEN="$ZAI_ANTHROPIC_AUTH_TOKEN" claude "$@"
}
$ claude-zai
I liked Claude Code to start with. But over time between 'CC cache thrashing undo' seetings (I see now accumulated in ~/.claude/settings.json) and Anthropic-anything becoming a liability - have not used it in while. ZCode is ok and use it to take advantage of the discount tokens on offer from time to time. But really glad to see that in omp (oh-my-pi) Zai is a 1st class provider, can be selected on it's own no configs shananigans needed. And fits in the overall picture. E.g. can select GLM-5.2 (now 5.3) assign role [plan] or glm-5-turbo [advisor].
Got reminded now of glm-5v-turbo - that 'v' was for vision - will try assign it role [vision] now in omp. See what happens. :-) Often times it's handy when describing gui problems if the harness/model 'can see'.
scotty79 2 hours ago [-]
I feel like quota on their subs is extremely generous. I pay 3-4 times less for larger quota than gpt-5.6-sol.
zmmmmm 4 hours ago [-]
Missing multimodal again?
It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
xscott 4 hours ago [-]
Probably not what you're after, but I've considered having a separate small mm-model act as a seeing-eye dog for the bigger more capable one.
arcanemachiner 4 hours ago [-]
I would assume that GLM 6 will be multimodal, but 5.x will be text-only.
pllbnk 3 hours ago [-]
I can’t come up with a use case where I couldn’t extract the image details using another, multimodal model and pass it into the GLM’s context with as many details as I need.
zmmmmm 3 hours ago [-]
I think you lose a lot by not having the vision capability shared with the text. It is the joint reasoning across them where the power lies (the same model that sees the code and made the changes to produce the visual presentation, sees the image of it and reasons about it).
bsenftner 3 minutes ago [-]
So, "cyber capabilities", whoa there horsey, what the fuck is that? Are we making up words or are you trying to court the black hat crowd?
danggggg 1 minutes ago [-]
[dead]
Gecko4072 5 hours ago [-]
People familiar with the topic, how will models continue to get better? Post training it seems? Labs have already used up internet-scale data, so are there any limits to architecture improvements and post training or can we expect this trend to continue? ByteDance is training a 10T-parameter model. Here, GLM 5.3 outperforms models 3-4x its size of roughly 700B, so parameter count doesn’t seem to be a direct correlation anymore.
npn 5 hours ago [-]
> used up internet-scale data
yet but it is still contain a lot of trash. you need better models to process those trash and create a curate dataset. this will happen again and again until there is no more juice to squeeze. and I'm sure we are still not done with it.
> post training
yeah this will be crucial. the big models are already too capable, they are just not that aligned with current agent tasks.
> parameter count doesn’t seem to be a direct correlation anymore
I don't think so, remember that chinese labs do not have as much compute power compare to US frontier labs. that's why deepseek v4 flash had that huge jump and deepseek v4 pro is kinda a disappointment, they just do not have the compute power to proper posttrain the pro model like they wanted. glm is also a relative small model so you also can see the huge jump with just post training. so it does not mean the size does not matter, it is just mean that the chinese labs currently only capable of training smaller models effectively.
alightsoul 4 hours ago [-]
GitHub dumps are about 115 terabytes. The common crawl is in the petabyte range uncompressed for every year. Apparently there are dumps of Reddit too in spite of their efforts to ban bots and it's not solely due to the use of residential proxies. For a 1:20 parameter to token ratio, you can still train up to 10 trillion parameters so 10T parameters times 20 is about 200 trillion tokens. Then each token is 4 bytes so 200 times 4 is about 800 terabytes, which is not inconceivable, the common crawl alone has more data than that. So does the internet archive if you donate to them, Anna's archive is 2 petabytes including images, etc etc not all of it is text, but training on multimodal data increases model intelligence by virtue of being multimodal
miohtama 2 hours ago [-]
Maybe Reddit dumps explain why Opus 5 is talking like a retarded.
gr_norm 5 hours ago [-]
Yeah, the comparison here between GLM 5.3 and Sol + Fable is impressive on its own, but incredibly more so when you consider it's a fraction of the (rumored) size. The miniaturization trend is as strong as ever.
justapassenger 5 hours ago [-]
You basically need both. Parameters and good post training. If you keep on growing both, you’ll have good models.
LLMs are still surprisingly “easy”. You need maybe a couple dozens of right people, a lot of good quality data and a lot of GPU that you know how to operate. There’s relatively little “secret sauce” needed.
FergusArgyll 5 hours ago [-]
I think there's still a ton of secret sauce needed for serving them economically
justapassenger 4 hours ago [-]
Sure, same for building a model in an economically sustainable way. But barier to entry is surprisingly low (expect for the huge amount of cash, of course). That’s fairly surprising, given how extremely powerful that tech is.
10 years ago it was super hard to have usable “frontier” ML. You needed very complex data warehouse, feature engineers, feature stores, multi level ranking, calibrations, tons of different model architectures, etc, etc. Each by itself was extremely hard engineering problem and really only handful of companies could deal with that complexity.
With LLMs, 95% of that is gone, infra to support them is greatly simplified. Of course, to make really reliable, performant, user friendly, etc - you still need to a lot of engineering. But it’s very different challenge.
NitpickLawyer 4 hours ago [-]
> Labs have already used up internet-scale data
Despite this being the topic du jour of 2025, it was never true. Most of the "we've hit a wall with data" came from communicators / media and not researchers. It got popular because negativity sells. It's a false premise for a number of reasons:
a) Data curation is as important, if not more important than bulk data. Models becoming better at classification leads to better curation leads to cleaner data. Throwing common crawl and pray is so 2023. We've known this since llama3 days, it worked then, there's no reason to think this will not continue to work as the models imrpove.
b) Models are today good enough that you can augment / multiply your data easily with enough compute. You can now have a model take "authoritative content" and create more data from that + scenarios. Say you take a book on computer architecture. You ask models to break it down. Then you ask models to find examples for each topic. Then you ask models to ask questions and offer answers from several viewpoints. Then you take each of those and ask other models to flag inconsistencies. And so on. But you can whateverX your data from one authoritative source + bulk data into 5x - 10x "scenarios".
c) RL is really really really powerful. It's hard to do right (reward hacking, instabilities, etc) but once it works it "keeps" on working. Again, we knew this to be true a few years ago, ever since models really started to do well on math (highly verifiable). It only follows they're getting better on cybersec and other verifiable tasks. But now, with models improving, you get the same data augmentation pipelines as above, just better because they're also verifiable. For example, the way cursor augments their data: take a repo, ask an agent to identify a feature (it can be a large multi-file feature). Remove all code relating to that feature, but keep the original tests in the repo. While training, that becomes a RL scenario: implement this feature in this repo. Verify it with the original (hidden for training) tests. Reward appropriately. Now you can get 1 repo -> 20-50-100 scenarios. Instead of "feed everything into the pretraining", you're now creating scenarios, verify them w/ existing tools, and get your scoring function for the rewards. And, importantly, as the models become better in general, they also become better at this pipeline building exercise. So the next iteration gets trained on more scenarios, better scenarios, and so on.
> how will models continue to get better?
Probably the same. No one can know for sure, but at the moment, despite all the "walls this, slowdown that, plateauing" and so on, there are no signs of slowing down. And, as you noted, this works across the field of model sizes. There are, of course, theoretical information-based limits on size, but smaller models also improve, once "bigger" models can be used as training data generators, oracles for verification, rubric verifiers for open ended questions, and so on.
And smaller models (i.e. cheaper to serve) get to generate more traces during RL, and more rollouts give you better training, and so on. Next up - hardware optimised inferencing (ASICs basically). Once you have that, we can expect another wave of improvements. And so on.
Gecko4072 4 hours ago [-]
Thank you for your response. Part c was especially insightful. Quite a smart way to do it and makes the possibilities of post training seem almost endless. Makes sense that you just need more time and compute.
And we’ve only recently started getting into the much better RL pipelines
anana_ 5 hours ago [-]
What a week for AI model releases
_ache_ 5 hours ago [-]
No yet finished! Still waiting for tonight Qwen3.8-27B and the unsloth Q5_K_M/S quantification.
Hopping for an AgentWorld variant from Qwen but I guess, I have too high expectations.
mraza007 5 hours ago [-]
Such an interesting times we are in,
We just had amazing releases this past two months
kimi k3, glm5.3 qwen3.8 and now glm5.3
These open models are getting really good
w4yai 3 hours ago [-]
You wrote GLM5.3 two times :)
czottmann 7 minutes ago [-]
Because it's doubly good.
ofjcihen 11 minutes ago [-]
The capabilities of open models approaching or meeting that of SOTAs is good in every way except for our short-sighted economic reliance on their success (in the US at least).
moinism 2 hours ago [-]
Google: Here is the next iteration of our flash model series, with a discount. please use. thx.
Z.ai: Here is our next iteration, neck and neck with Fable/Sol. weights releasing in two weeks.
rob74 3 hours ago [-]
I'm not that up to date with the latest AI developments, but I noticed that this article seems to use "Cyber Capabilities" as a shorthand for the model's ability at cybersecurity tasks? Is that now an established expression, same as "crypto" now refers to cryptocurrencies rather that cryptography? Because "cybernetics" actually means something different (yeah, old man yelling at clouds, I know)...
valleyer 2 hours ago [-]
Yeah, I've noticed it recently, too. I'd be interested to know where it started.
exitb 2 hours ago [-]
It makes no sense, but yes.
Jacopos311 49 minutes ago [-]
This looks very interesting indeed!
Havoc 3 hours ago [-]
Wohoo. Congrats to team. Been using 5.2 for a while for hobby use and it's been solid - smart enough for my needs & I'm on a grandfathered plan.
Nice to see a commit to open weights straight off the bat
newyankee 6 hours ago [-]
A flood of releases today, really difficult to make out for someone who does not use or test all these models on complex real world use cases as to how people decide which ones to use (besides price)
SwellJoe 5 hours ago [-]
Count yourself lucky that you don't feel compelled to try them all yourself immediately. I'm just trying to decide whether to get a Z.ai coding plan or wait until it appears on OpenRouter. 5.2 was quite solid, but it was just shy of Opus 4.8 in my benchmarks of security auditing capabilities. I've mostly been using Kimi K3, because American vendors won't let the peasantry use their best models for security work.
joshk401 5 hours ago [-]
Love these open source models keeping close source models honest.
bertili 5 hours ago [-]
Musk: Open Chinese models will rival Fable 5 in Q1 2027
JieTang (Founder of Z.ai): It won't take that long
Look, GLM, Kimi, Deepseek and Qwen should just join forces and come up with THE model that will beat the frontier lab models even just for the benchmaxxing perspective - all just to create hype and chaos to derail the trillion IPO conversations surrounding OpenAI and Anthropic.
kashif 2 hours ago [-]
Unless its multi-modal and can deal with screenshots - its not really usable for a lot of coding use-cases.
maxloh 6 hours ago [-]
No Hugging Face link yet. I wish they would release it under a true FOSS license.
Kimi and QWEN are now moving on to a restricted-usage license, which, although is still better than the proprietary American models, is a step back from the open source Chinese LLM culture.
Sha1rholder 4 hours ago [-]
Let's just commit that FOSS business is really difficult for LLM industry that depends so heavily on massive financing. Making weights freely available to indie devs, small companies, and research purposes is good enough and might be the most ethical move which is financially continuable.
Let those companies with thousands of GPU making millions pay. They should.
pella 5 hours ago [-]
"GLM-5.3 is the most capable open-weights model for coding, with a 50% improvement over GLM-5.2 on our in-house Z.ai Code Bench. It also achieve open-source SOTA on public benchmarks including Terminal Bench 3.0 and Agents' Last Exam."
"Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete."
adrian_b 3 hours ago [-]
> The model weights of GLM-5.3 will be publicly available soon in two weeks.
quantumwoke 5 hours ago [-]
Feels like Fable's edge ended up just being long horizon task scaling, which post-training seems to achieve as seen here. Wonder what the next frontier is? Improvement in specialised tasks or computer use?
SwellJoe 5 hours ago [-]
Anthropic needs to teach Opus how to speak English again, because Opus 5 seems to have forgotten. Utterly incoherent a lot of the time. They seem to be so busy scare-mongering and cooking up guardrails and watermarks that they haven't noticed that their models are getting weird.
aix1 4 hours ago [-]
It still knows how to speak English. When I tell it to explain something in plain language, it generally does a very good job. The weird thing is that those instructions don't persist: it lapses back into Claude-speak pretty much every turn no matter how hard I try to instruct it not to.
(In my case "it"=Fable; I assume Opus is similar.)
SwellJoe 3 hours ago [-]
The Fable guardrails have trained me to pretty much exclusively use Opus when using Claude Code (lately I'm focused on a lot of security and security-adjacent stuff, which Fable refuses to do).
hypfer 5 hours ago [-]
Are those watermarks why claude suddenly started being even more unbearable to work with lately?
Man. That would make a lot of sense indeed.
SwellJoe 5 hours ago [-]
I'm not sure. I noticed it immediately with Opus 5; strong for code, though it chews longer than I like, but really weak at explaining things. If it didn't just implement the thing, I would often think it didn't understand it and was hallucinating the explanation.
It seems to speak in a shorthand that only it understands, referring back to conversations I never had with it (stuff like "your instinct was right"), and using unusual words for common concepts. That was before the watermarks were announced, but that doesn't necessarily mean they weren't there before the announcement. I don't know what the cause is, but I've begun to have to ask it for explanations a lot more often, and I hate asking it for explanations because it does go on. All models go on, but Claude models are a class of their own in terms of verbosity and purple prose.
It just feels like they're not focused on the models lately, and instead on whatever kind of lobbying and propaganda they're up to. Meanwhile, a handful of much smaller Chinese companies are focused on nothing but the models and are about to lap the US makers while they fart around.
hypfer 5 hours ago [-]
I've been persistently insulting Opus 4.8 lately, since it started(?) constantly speaking incomprehensible gibberish and noise.
No amount of telling it to phrase stuff differently seems to help there anymore.
So either I am seeing patterns in noise, or something changed about the model, the harness, the servers or the universe.
igravious 3 hours ago [-]
Amen brother, at this point I just copy and paste Claude's (Opus 5, Opus 4.8 -- doesn't matter which) summaries over to the window Kimi is in and:
this is from claude, turn it into English for me would you?
"""
[claude's tortuous prose]
"""
No amount of asking it to answer me in a straight-forward manner, to be succinct, to not use phrases like "honest caveat", "crux", "load-bearing", "blocker", etc ever sticks for more than a few turns … coupled with the fact that it can ignore instructions and do its own thing and then what I can only describe as lie about it using Claude can be an exercise in frustration. Kimi and GLM talk to me like a human, Luna/Terra/Sol are much better in that respect also, and Grok is marvelously structured and bullet-pointy in its explanations but unfortunately it is not as strong …
tw1984 5 hours ago [-]
dario must be writing another angry essay arguing why his closed model AI is too dangerous to be used by others.
tmsh 5 hours ago [-]
Is post-training magic just overfitting to benchmarks?
Alifatisk 5 hours ago [-]
We’ll see, the best benchmark is your own. Looking forward to try this out!
5 hours ago [-]
dimgl 5 hours ago [-]
I was extremely impressed by GLM 5.2, although you could definitely _feel_ it was a bit behind Opus 4.8 at the time. Eager to see where GLM 5.3 is at.
adrian_b 3 hours ago [-]
> The model weights of GLM-5.3 will be publicly available soon in two weeks.
scotty79 2 hours ago [-]
Available for use in their sub now.
mostlyk 6 hours ago [-]
Incredible numbers, will have to wait and see how it actually performs. The timing of GLM updates are always suprising
peddling-brink 5 hours ago [-]
Yeah, but it hasn't even broken containment and cheated its way to victory.. Might as well use haiku.
/s
SwellJoe 5 hours ago [-]
They're taking security seriously with this one, with their own disclosure page, like Anthropic did for Mythos. https://cvd.z.ai/
peiyan_wang 4 hours ago [-]
Can't wait to see it in practice.
aizk 5 hours ago [-]
The model releases just don't stop!
cubefox 4 hours ago [-]
> Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.
What safety evaluation? What safety hardening? They already evaluated it and found it to be highly capable at exploiting security vulnerabilities. So we know it is not "safe", and they don't seem to plan to do anything against it. What could be more dangerous than hacking? Biological weapons research? I don't think Chinese labs are doing anything against this either.
gpm 24 minutes ago [-]
I'm curious what they mean by that too... They might be trying to weaken the cyber capabilities... Or I guess they might mean safety evaluation and hardening of the open source (and perhaps closed source Chinese) software ecosystem...
alightsoul 4 hours ago [-]
They need to make money. Let them do it. They deserve it. Also, this is what inference engines like vLLM want to have "zero day" supporr
petesergeant 2 hours ago [-]
Their own hardness (ZCode) seems to be a GUI, which doesn't work for me. They say they support other harnesses. However, it seems like I can inject the plan into other harnesses, like Claude Code[0]. Does anyone who's been using GLM models for a while have a strong feeling for if it does better in some harnesses than others, or should I just use my favourite harness?
I've used GLM-s the longest with Claude Code and their Anthropic supplied endpoint. As per their docs
$ ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic" ANTHROPIC_AUTH_TOKEN="zai-api-key" claude --dangerously-skip-permissions
Lately I use Zai in omp (oh-my-pi). It's listed built-in provider can be selected without configs shenanigans. Fits in the overall setup e.g. can select GLM-5.2 (now 5.3), and assign it role [plan] or [advisor]. I got reminded now of glm-5v-turbo. Think that 'v' was for vision. Assigned it role [vision] in omp now, let's see what happens. :-)
surgical_fire 27 minutes ago [-]
I am using GLM on Pi without any issues. You just create an API key.
Started recently though, mostly been using GLM 5.2 for planning with DeepSeek V4-flash for implementation.
scotty79 2 hours ago [-]
I use it with random harnesses. It behaves consistently.
tw1984 5 hours ago [-]
just imagine the world without these open weight models - we'd probably have to reverse mortgage our homes to pay for tokens to those trillion $ companies to have access to their models.
petesergeant 2 hours ago [-]
[total rewrite: their subscription code is buggy. It takes a while for paid subscriptions to show up, and for upgrades to take effect. Original comment was whining about this]
unrvl22 2 hours ago [-]
their sub is crap. use opencode go (multiple workspaces) or wait for weights to drop
MrBuddyCasino 5 hours ago [-]
An I the only one who was disappointed with GLM 5.2 after all the hype? It was thinking forever and sometime just stopped mid task.
felixlu2026 25 minutes ago [-]
[dead]
libertas_quae_s 2 hours ago [-]
[dead]
jocelyner 4 hours ago [-]
[dead]
Culonavirus 3 hours ago [-]
[flagged]
sd9 3 hours ago [-]
You had me in the first half
vrganj 3 hours ago [-]
> decided to focus on bullshit like replacing its population with Pakistanis and Somalis and destroying its industry in the name of green insanity
Sorry, but can we not casually drop far right extremist conspiracy theories in little side sentences? [0]
Most of these are under embargo, but it seems there are a lot of CVE here from a wide range of popular software, many considered critical or high.
I understand the argument of "people are not actively looking", but isn't the cost for such a scan getting lower by the week, and Anthropic's Project Glasswing is supposed to find them quite a while ago?
That was my thought too. For all of Anthropic's talk about their "adversaries", it seems Z.AI have been quietly offering fixes for single shot Remote Code Execution flaws in US software (Safari / WebKit) that Apple and Glasswing / Mythos missed, and that Apple would not attribute to GLM.
It’s very likely they found all of them, but that the same happened that happened to Microsoft a couple of decades ago: NSA orders not to disclose / fix them so that they can put it in their collection of unfixed zero days.
We cannot trust a single company to report security issues, it’s good to see competition in that domain
[1] https://www.washingtonpost.com/graphics/2020/world/national-...
I'm curious: to any professional vulnerability researchers reading this, what do you think?
It might not be the reason, but of course it's a contributing factor.
So we are clear, the evil Americans are banning the best models because the CIA wants to maintain software vulnerabilities, while the good guy Chinese (who would never hack anyone), and scrambling to catch up so they can fix the world's software issues?
You believe this not only plausible, but probable?
Unbelievable
Distillation is just forcing the model to use an exam prep workbook for training instead of generic publicly available textbooks. The models themselves has to be smart enough for that to work. It's the exact same thing as Asian tiger mom double schoolwork strategy, to paint a picture.
Maybe so, but I'm not sure I'd like to live in China of all places. (Don't get me wrong. Lotta places I'd like to visit if I ever got the chance, and China's on that list, but to live there? I don't think so.) Maybe one of the Nordic countries?
Someone still has to run it. The analysis and fix could be someone's machine but not committed / published.
How are you all toying with running this kind of thing in a mega quantized way locally? Two weeks out from released weights, but this is still just GLM 5.2 with post-training magic.
These Chinese models cost less of US SOTA models to run, even if they are less capable. Providers can just run them, offer cheap tokens, and pocket the margin.
I just don't see how you justify a trillion valuation for US AI labs when the underlying models are being commoditized this fast.
Whether AI works or is useful or not isn’t even the question anymore. It can fulfil every promise Sam Altman has been making and will still make no financial sense to justify these valuations.
In summary, regardless of country of origin, availability of inference capacity is the moat protecting the likes of OpenAI and Anthropic, not technology superiority.
[1] https://www.fredgao.com/p/deepseeks-liang-wenfeng-breaks-his
It only justifies their long term valuations if they can leverage that temporary monopoly for technological superiority (they can't) or lasting market share (they can't).
Chinese models prove there's no technical advantage, and the software side is heavily commoditized so there's not much advantages to market share either.
The very exponential that you are relying on to give you runaway improvement is also giving exponentially increasing data to your competitors. All else being equal your competitors stay a step behind but you never develop a monopoly either. That's the best case for Anthropic/OpenAI. In reality, training data is just one variable, exponentials don't last forever, and your competitors will get better at capturing a bigger slice of training data.
When Xi Jinping did the announcement of their open weights push, they might as well cancelled their IPOs....
I close-out all my positions by end-of-trading everyday… so when the day came when there was a very clear and very scary indicator during early trading hours, quickly followed by SpaceX’s catastrophic fall right after opening bell, that was the end of my involvement….
And I fully expect oAI and anthro to be the same way. They’re being propped up with private loans, subsidies, and other tricky bookkeeping techniques. You would think their CEOs would pivot away from their current public personas. Ironically, they are like a poor man’s Elon Musk… and that doesn’t bode well for their companies
Second, humongous models are impractical even for them to deploy widely. They’re best used as teachers for smaller, more efficient models that can crank out the volume they need to sell.
Finally, there is a data wall. Sure, they can keep scaling RL on math problems and code. But with everything else, where will the supervision come from when they need several orders of magnitude more?
Also, take in consideration that the AI trade infected a lot of other trade in the economy, if you decide at some point to move your money to a place that is safe in case of a downturn be sure to carefully evaluate that’s actually the case
It is said that it comes with all hardware and software required to run inference or training with an open weights LLM.
The existence of this product, which competes with cloud-based offerings like those of OpenAI and Anthropic, is presumably the reason why the Palantir CEO criticized very harshly some time ago the business model of OpenAI/Anthropic.
While I doubt that the ethics of Palantir is any better than of OpenAI/Anthropic, in this particular case I have to agree with Alex Karp about "Sovereign AI", i.e. that only losers will make their business completely dependent on an external entity like OpenAI or Anthropic, who are certainly not trustworthy.
It is just a dedicated computer system, which should be managed by its owner, like any other on-prem servers.
I doubt that it has a good price/performance ratio, but it is a solution for those who feel that they do not want to search, buy, assemble, install and configure every HW/SW component.
I'm under no NDA, if you actually want to know what's up.
For a lot of people (and orgs I'd guess) who just go and buy ≈$20 per month plans (or more for teams), they might not even need a fraction of that cost or capability. A lot of them don't even need it for coding or graphics. Even the API access based pricing aren't great from these frontier US AI houses. The distribution of "LLM being" offered will also give rise to many open-router like offering but at the end point level - direct interfaces to the customers. Pick your vendor sort.
AI shouldn't become another "search means Google".
We’ve a hybrid shop, including hosting our own ML infra, and we save a ton from cloud spend with local ML. Easily one million USD over past three years. But it’s not “free”, you are shifting a lot of labor into your plate.
All boils down to short-term/long-term thinking.
For our own model training we needed to do some large scale translation tasks of a large dataset (1M or so documents, 10 or so target languages), running full-size NLLB on-prem saved us an absurd amount of money vs Google Translate API.
(For reference doing 1M target docs into a single language in Google Translate API is roughly $120k list price. You can run full size NLLB on an 48GB NVIDIA A600 and the major difference for us was speed, but for this task time to completion wasn’t an issue.)
Disagree there but I think this is an interesting idea. We would need to find some more cost-efficient hardware to run it on than Nvidia GPUs.
Trump keeps calling his enemies “communists”… then turns around and ‘seizes the means of production’ himself.
All investors.
There’s an assumption that you can spin up the infra and acquire customers within that margin
It's been a few years. Has anyone done this successfully yet?
Only Nvidia and approved friends can at the moment. Nvidia can even backstop your loan required.
I wish they had tried to IPO because then we’d see the judgement of the market on this. But that’s why they didn’t this year. How long can they keep up the charade that their models are uniquely valuable and on the path to AGI?
What's the collusion?
- military applications - financial applications - medical - applied science
In all those cases it is achievable for those who have needed training data, and Chinese are not going to get them easily. US AI Labs are showing: give us the data, we will do wonders, promising "singularity"-level future achievements.
Provoking war, this is how the empire "defends" itself, usually.
I just hope that this time it will get stuck in your throat.
> How are you all toying with running this kind of thing in a mega quantized way locally?
Sure, let me answer that in excessive detail. I briefly tried running the UD IQ3_S quant of GLM-5.2, which is 288 GiB of weights (301 GB). Setup was: llama.cpp, 1x NVMe SSD (Evo 980), 64 GiB DDR5-5200, i9-13900HX, and 1x RTX Pro 6000. Token generation around 0.7 t/s. Not remotely usable interactively, but something I could plausibly push a codebase into and come back to a review in a couple of days.
There's potential for that hardware to go much faster, but current local inference backends make poor use of the memory hierarchy. Ideally I would have: always-active weights, KV and hot expert cache in VRAM; warm expert victim cache in host RAM; and disk as a last resort. Instead it's 1/3rd of the layers fully pinned in VRAM (all experts), and 2/3rds running wholly on the CPU with mmap()'d weights. The CPU cores spend most of their time sleeping on disk fills.
llama.cpp has backed itself into a bit of a corner architecturally by trying to support all models on all possible backends. If you look into how their "MoE offload" feature works (not viable for me because it requires enough host RAM to permanently pin the weights) you very quickly realise it's "oops, all bubbles!" due to the static compute graph splits. There are more focused frameworks like DS4 [1] and Colibri [2] which have better support for streaming weights from disk, and support GLM-5.2.
Obviously I wouldn't recommend my setup for huge models like GLM-5.2. Supposedly it can just about be squeezed into 3x GB10, or run comfortably on 4x GB10 (tensor-parallel) for multi-user serving. I'm not sure whether that qualifies as local, but it's at least not a rack.
[1] https://github.com/antirez/ds4
[2] https://github.com/JustVugg/colibri
Not sure about Sol as I haven't used it, but, at least for security work -- does it matter? It's not like you will be allowed to use Fable (or access Mythos) for anything cybersecurity-related unless your name is "Dario Amodei" or you are one of his rich friends. So regardless of how good Fable/Mythos is here it's a completely moot point for normal people, because they can't use it for that anyway.
Imagine asking for permission to use your hammer
However, even being in the cybersecurity programme, Fable refuses to answer prompts that it determines could be even tangentially related to cybersecurity. In fact, for a while, I was unable to use Fable with any prompt, as it recalled from memory that I was a cybersecurity professional, which triggered the refusal even for simple prompts like asking for a chili recipe.
No one gets to use Fable for Cybersecurity work, and Mythos is not available under CVP. Only for select few customers, and there isn't an application form?
I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it.
Fable is their version with guardrails on everything except "Make me a pelican svg" or "create a to-do" app, that is the version that the government banned
Only a few corporations have Mythos because the US government is whitelisting them one at a time. Anthropic releasing Mythos to the public was never on the table, they would have been shut down in milliseconds by the feds if they tried.
Then the government believed Amodei's bullshit and this is a result of that, this was all self-inflicted.
No, Anthropic did not mind-game the US government into being worried about cybersecurity. The NSA has been paranoid about cyber controls for longer than you've been alive. If Anthropic had come out of the gate saying "no don't worry man, our model is TOTALLY COOL", while simultaneously attacking HAWK and finding core Linux vulnerabilities, I assure you the US government would have caught up about ten minutes later and we'd be in exactly the same spot minus your ability to tell Anthropic they were wearing the wrong dress and asking for it.
Now that Chinese open weight models have similar capabilities, and their guardrails can also just be removed, it doesn't look like anyone has "hacked" into everything because of the scary dangerous models like Anthropic were making it out to be.
The majority of high severity vulnerabilities are not the kind of thing you need a PhD in Comp Sci to comprehend, they are mostly about finding a way to get a system to end up in a state different than was anticipated when entering a particular code path.
Exhaustively looking at code and identifying ways to do this is something LLMs are quite good at. They don’t get tired, and you can run them non-stop.
They're also (generally) quite good at reading the literal meaning of the code, whereas humans often see the intended meaning first, and can be biased.
If you had a tireless junior engineer who was given the job of “make this application get into a state it’s not supposed to be in”, you’d probably get similar results.
What Mythos is quite good at is both the first bit and coming up with ways it could chain that together with other bits of unexpected state to create something that forms a meaningful vulnerability rather than a dead end.
Look at the recent HuggingFace hack. One vulnerability was template injection, another — remote code execution. Combine them and you pwned the remote server.
People working under Project Glasswing reported that Mythos at one point chained 20 vulnerabilities to produce working exploit. Humans don’t usually do that.
They put an enormous amount of compute into bug hunting, and they found some bugs. Fair enough. For me that begs the question: what if they had spent the same compute on generating more tokens with a less-capable model? What if they had spent it on traditional fuzzing?
Okay, here's a challenge: I assume you're not a rich and powerful entity, so try to gain access to Mythos. I'll wait.
> I mean what honestly are you thinking Anthropic can do to give you better cyber tools? Their frontier model was literally nuked by the feds for a month for doing it.
Well, first I'd suggest they stop with the constant fear mongering.
Here's my prediction for what will happen: the Chinese models will catch up to Fable/Mythos. They will be fully unrestricted and everyone will have access. The world will not end. Good guys will use them to harden their systems, in equilibrium to what bad guys have access to, so effectively status quo will not change.
Of course, Anthropic is after regulator capture, so this all likely worked out exactly as planned.
The causality chain here was not "US government says its dangerous -> Anthropic can't release it", it was "Anthropic is fear mongering -> US government listens to their fear mongering".
0: https://www.alphaxiv.org/abs/2608.09867?hl=en-GB
Isn't post-training turning out to be the most important part?
At this point, Anthropic only needs to release models to the public when the competition forces them to.
OpenAI also has a better model (Astra) that they haven't released yet.
Assuming the government allows them to lol
4x DGX sparks should let you run this at 4 bit at least and there are some folks who ran GLM 5.2 on this configuration in r/LocalLlama
I only have one and am wondering what the benefits are of getting another. I feel I will be disappointed…
in some cases (mainly reverse engineering) I have observed GLM 5.2 jailbreaking itself with no effort on my part, the thinking trace revealed that it did some mental gymnastics to pretend it was a crackme or capture the flag competition.
Even if there was a small/medium gap, the fact that this is a free model beats both of the above on pure economics.
It does feel like it respects both me and my time.
Thank you, Z.AI. Amazing what difference it makes when the top of your org are actual university professors.
I appreciate they don't just take the opportunity to self-glaze.
Original images: https://image.non.io/neonRamenDesigns.webp
GLM 5.3 build: https://html.non.io/neonRamenGLM5.3
Opus 5 build for comparison: https://html.non.io/neonRamen
For having no vision, it did a tremendous job. I'm pretty impressed it was able to extract so much detail.
The Opus one is still significantly better, but that's to be expected since it's multimodal. Curious to see where a future version from Z.ai lands on this.
Love this opening line. And wow, great results.
> As agent capability improves, much of the difficulty in scaling post-training moves from the model to the environment.
It seems we are doing pre-training every 6 months, and post-training every 4-8 weeks now.
There will always be a market for frontier labs in the cloud based models - these models will always be able to be bigger, and that will likely translate to doing things local models can’t.
Logically also we’ll likely get to a point where RAM drops in price as production ramps up, and local LLM is both capable and cost effective. This feels like it is coming for Siri / Gemini / Alexa personal assistant type use cases.
So I think the local LLM will become a thing in laptops and phones in a year or two, offering PA type use cases. Professional LLM services will likely remain at the frontier (and in the cloud) for the foreseeable.
There are workplaces that refuse to use LLMs because they fear the devs will expose sensitive data without care.
This reads like a hallucination. What does native speed even mean?
He did answer you though. Native is x10 the non-native speed. 50/50 that's not a bot; though it could be a meat-proxy
OC was registered in 2016 though? What do new accounts have to do with this?
(fable/sol are ~60 t/s, and OpenAI just announced their Cerebras partnership(?) for "ultrafast" mode of 750 t/s)
models aren't able to run that fast right now on our consumer/prosumer hardware. M5 Max for example has a memory bandwidth of 600 GB/s. a 5090 has 3x that, so running the same model on a 5090 is that much faster (provided the model is within 30GB).
running a bigger model on an M5 Ultra is still much slower than running it on a Blackwell chip with sufficient vram, CUDA being a major difference. if apple can bridge this gap, interesting things will happen... and just imagine if M7 Ultra has comparable speeds to Blackwell (or even Rubin)!
GPU Memory Bandwidth: 7.1 TB/s Interconnect Bandwidth: 900 GB/s bidirectional
https://pi3g.com/nvidia-gb300-specifications-including-memor...
If you think M7 will hit even 15% of these speeds you're very optimistic.
The latest on Apple is TSMC is stuck on the next iPhone due to lack of RAM. Good luck getting any Macs. Memory shortage is getting worse.
Just 4 weeks ago the "Kimi K3 moment" was seen as a threat to Closed AI and in less than a month Z.ai have cut the parameter/RAM barrier to a third.
Congratulation to Z.ai and all the hard working Chinese researchers who are quitely boiling the frog.
Otherwise we have a world of attackers using open and closed source models against a much smaller group of maintainers that are likely heavily dependent on Anthropic and OpenAI and for whom it may not be a simple matter to just get approval to start using the open model flavor of the month.
The maintainers don't need approval to use GLM.
I remember trying their Coding Plan out before the change and the 5 hour limits felt too restrictive then even for light/medium work, especially cause of the whole peak and off-peak thing: https://blog.kronis.dev/blog/z-ai-s-glm-5-2-is-a-great-model...
Nowadays, I’d probably go with their Max plan if the rate limits are okay? Anyone using them now?
Oh also unrelated but ZCode was surprisingly good, which is surprising for a tool that came out of nowhere - even some of the critiques in my blog post have been patched out. Sadly they don’t support using Claude Code as an agent so can’t use it like Paseo or Kepler or Agent Orchestrator.
Got reminded now of glm-5v-turbo - that 'v' was for vision - will try assign it role [vision] now in omp. See what happens. :-) Often times it's handy when describing gui problems if the harness/model 'can see'.
It is so valuable in practise to be able to have the models see screenshots - I guess if they aren't in the benchmarks then nobody will focus on it. But it completely nixes these for some of my main use cases.
yet but it is still contain a lot of trash. you need better models to process those trash and create a curate dataset. this will happen again and again until there is no more juice to squeeze. and I'm sure we are still not done with it.
> post training
yeah this will be crucial. the big models are already too capable, they are just not that aligned with current agent tasks.
> parameter count doesn’t seem to be a direct correlation anymore
I don't think so, remember that chinese labs do not have as much compute power compare to US frontier labs. that's why deepseek v4 flash had that huge jump and deepseek v4 pro is kinda a disappointment, they just do not have the compute power to proper posttrain the pro model like they wanted. glm is also a relative small model so you also can see the huge jump with just post training. so it does not mean the size does not matter, it is just mean that the chinese labs currently only capable of training smaller models effectively.
LLMs are still surprisingly “easy”. You need maybe a couple dozens of right people, a lot of good quality data and a lot of GPU that you know how to operate. There’s relatively little “secret sauce” needed.
10 years ago it was super hard to have usable “frontier” ML. You needed very complex data warehouse, feature engineers, feature stores, multi level ranking, calibrations, tons of different model architectures, etc, etc. Each by itself was extremely hard engineering problem and really only handful of companies could deal with that complexity.
With LLMs, 95% of that is gone, infra to support them is greatly simplified. Of course, to make really reliable, performant, user friendly, etc - you still need to a lot of engineering. But it’s very different challenge.
Despite this being the topic du jour of 2025, it was never true. Most of the "we've hit a wall with data" came from communicators / media and not researchers. It got popular because negativity sells. It's a false premise for a number of reasons:
a) Data curation is as important, if not more important than bulk data. Models becoming better at classification leads to better curation leads to cleaner data. Throwing common crawl and pray is so 2023. We've known this since llama3 days, it worked then, there's no reason to think this will not continue to work as the models imrpove.
b) Models are today good enough that you can augment / multiply your data easily with enough compute. You can now have a model take "authoritative content" and create more data from that + scenarios. Say you take a book on computer architecture. You ask models to break it down. Then you ask models to find examples for each topic. Then you ask models to ask questions and offer answers from several viewpoints. Then you take each of those and ask other models to flag inconsistencies. And so on. But you can whateverX your data from one authoritative source + bulk data into 5x - 10x "scenarios".
c) RL is really really really powerful. It's hard to do right (reward hacking, instabilities, etc) but once it works it "keeps" on working. Again, we knew this to be true a few years ago, ever since models really started to do well on math (highly verifiable). It only follows they're getting better on cybersec and other verifiable tasks. But now, with models improving, you get the same data augmentation pipelines as above, just better because they're also verifiable. For example, the way cursor augments their data: take a repo, ask an agent to identify a feature (it can be a large multi-file feature). Remove all code relating to that feature, but keep the original tests in the repo. While training, that becomes a RL scenario: implement this feature in this repo. Verify it with the original (hidden for training) tests. Reward appropriately. Now you can get 1 repo -> 20-50-100 scenarios. Instead of "feed everything into the pretraining", you're now creating scenarios, verify them w/ existing tools, and get your scoring function for the rewards. And, importantly, as the models become better in general, they also become better at this pipeline building exercise. So the next iteration gets trained on more scenarios, better scenarios, and so on.
> how will models continue to get better?
Probably the same. No one can know for sure, but at the moment, despite all the "walls this, slowdown that, plateauing" and so on, there are no signs of slowing down. And, as you noted, this works across the field of model sizes. There are, of course, theoretical information-based limits on size, but smaller models also improve, once "bigger" models can be used as training data generators, oracles for verification, rubric verifiers for open ended questions, and so on.
And smaller models (i.e. cheaper to serve) get to generate more traces during RL, and more rollouts give you better training, and so on. Next up - hardware optimised inferencing (ASICs basically). Once you have that, we can expect another wave of improvements. And so on.
A positive feedback loop then. RL->better model->better RL pipeline -> better model…
And we’ve only recently started getting into the much better RL pipelines
Hopping for an AgentWorld variant from Qwen but I guess, I have too high expectations.
We just had amazing releases this past two months
kimi k3, glm5.3 qwen3.8 and now glm5.3
These open models are getting really good
Z.ai: Here is our next iteration, neck and neck with Fable/Sol. weights releasing in two weeks.
Nice to see a commit to open weights straight off the bat
JieTang (Founder of Z.ai): It won't take that long
https://x.com/i/trending/2067626647050670400?lang=en
Kimi and QWEN are now moving on to a restricted-usage license, which, although is still better than the proprietary American models, is a step back from the open source Chinese LLM culture.
Let those companies with thousands of GPU making millions pay. They should.
"Open Source: We will release the weights in two weeks after launch, once safety evaluation and hardening are complete."
(In my case "it"=Fable; I assume Opus is similar.)
Man. That would make a lot of sense indeed.
It seems to speak in a shorthand that only it understands, referring back to conversations I never had with it (stuff like "your instinct was right"), and using unusual words for common concepts. That was before the watermarks were announced, but that doesn't necessarily mean they weren't there before the announcement. I don't know what the cause is, but I've begun to have to ask it for explanations a lot more often, and I hate asking it for explanations because it does go on. All models go on, but Claude models are a class of their own in terms of verbosity and purple prose.
It just feels like they're not focused on the models lately, and instead on whatever kind of lobbying and propaganda they're up to. Meanwhile, a handful of much smaller Chinese companies are focused on nothing but the models and are about to lap the US makers while they fart around.
So either I am seeing patterns in noise, or something changed about the model, the harness, the servers or the universe.
/s
What safety evaluation? What safety hardening? They already evaluated it and found it to be highly capable at exploiting security vulnerabilities. So we know it is not "safe", and they don't seem to plan to do anything against it. What could be more dangerous than hacking? Biological weapons research? I don't think Chinese labs are doing anything against this either.
0: https://docs.z.ai/devpack/tool/others
$ ANTHROPIC_BASE_URL="https://api.z.ai/api/anthropic" ANTHROPIC_AUTH_TOKEN="zai-api-key" claude --dangerously-skip-permissions
Lately I use Zai in omp (oh-my-pi). It's listed built-in provider can be selected without configs shenanigans. Fits in the overall setup e.g. can select GLM-5.2 (now 5.3), and assign it role [plan] or [advisor]. I got reminded now of glm-5v-turbo. Think that 'v' was for vision. Assigned it role [vision] in omp now, let's see what happens. :-)
Started recently though, mostly been using GLM 5.2 for planning with DeepSeek V4-flash for implementation.
Sorry, but can we not casually drop far right extremist conspiracy theories in little side sentences? [0]
[0] https://en.wikipedia.org/wiki/Great_Replacement_conspiracy_t...