AI Model Guide

GLM-5.3 in 2026: The Open Model That Shipped Without Its Weights

Z.ai's newest model landed on 14 August 2026 with strong coding claims and a genuine surprise in cybersecurity. It also landed without the downloadable weights that made the GLM line commercially interesting, and without a published per-token price. For businesses, that changes what kind of product this is.

Distk Editorial Aug 2026 12 min read

GLM-5.3 was released on 14 August 2026. Z.ai's own documentation confirms it reuses the GLM-5.2 base model, with all gains coming from expanded post-training, and it keeps the 1 million token context and 128K output ceiling. Two things did not ship with it. The open weights are held back pending a safety evaluation, reported as roughly two weeks out, and no GLM-5.3 entry appears on Z.ai's per-token price list, so only the subscription tiers are knowable. A third detail matters more than either: Z.ai's developer docs state that GLM-5.2 and GLM-5.1 requests on the Coding Plan are automatically routed to GLM-5.3, so version pinning on the subscription is gone. For a business that adopted GLM because it was open and self-hostable, this is not a drop-in upgrade in 2026. It is a different proposition, and it deserves a decision rather than an auto-accept.

What Is GLM-5.3 in 2026?

GLM-5.3 is Z.ai's flagship text model, reported as launched on 14 August 2026. Z.ai's documentation states it is built on the same base model as GLM-5.2, with every capability gain produced by expanded post-training rather than a fresh pretraining run. It keeps a 1 million token context window and a 128K maximum output, remains text only, and supports thinking modes, function calling, structured output, context caching and MCP tool integration.

The positioning is narrower than a general frontier release. Z.ai frames GLM-5.3 around complex coding, long horizon agentic tasks and cybersecurity work. If you have read our explainer on what GLM-5.2 actually was, this is the same machine with a much heavier training regime applied on top of it, which is an unusual and interesting engineering claim in its own right.

One number that circulates widely deserves care. Coverage of the launch describes the shared base as a 743 billion parameter mixture of experts model. The published GLM-5.2 repository on Z.ai's Hugging Face organisation lists 753 billion parameters, and GLM-5.1 lists 754 billion. Those two figures cannot both describe the same unchanged base, so this article does not lean on a parameter count at all. Nothing in a business decision depends on it.

AttributeWhat Z.ai published in 2026Why it matters commercially
ReleasedReported 14 August 2026Roughly six weeks after the GLM-5.2 weights appeared publicly.
Base modelUnchanged from GLM-5.2; gains from post-training onlySuggests capability can move without the cost of a new pretraining run.
Context window1 million tokens, 128K maximum outputSame envelope as GLM-5.2, so document workflows carry over unchanged.
FocusComplex coding, long horizon tasks, cybersecurityA developer and security tool first, a general assistant second.
Open weightsNot released at launch; staged pending safety evaluationThe single biggest change from the GLM-5.2 proposition in 2026.
Per-token API priceNot listed on Z.ai's published price pageUsage based budgeting is not possible yet.
Subscription accessIncluded in all GLM Coding Plan tiersExisting subscribers get it without a new purchase decision.

Why Do the Missing Open Weights Matter in 2026?

Because for a large share of GLM's business users, the weights were the product. GLM-5.2 could be downloaded, inspected, fine tuned and run inside a company's own infrastructure. That property is what let regulated teams, agencies handling client data and cost sensitive engineering groups adopt a Chinese lab's model at all. GLM-5.3 removes that property at launch, so the value proposition is materially different even though the version number suggests continuity.

It is worth being precise about what is and is not known here. Z.ai has not abandoned open weights. Reporting of the launch describes a staged release, with the weights to follow roughly two weeks after 14 August 2026, once safety evaluation and hardening are complete, and under the same permissive licensing pattern the line has used before. At the time of writing there is no GLM-5.3 repository on the Z.ai Hugging Face organisation, which is where GLM-5, GLM-5.1 and GLM-5.2 are all published. The weights are late, not cancelled.

The reason given is also credible rather than convenient. Z.ai says the model's cybersecurity capability grew further and faster than its training intended, reaching multi stage exploit chain reasoning the company did not plan for. A lab that finds an unplanned offensive capability and pauses distribution while it evaluates the consequences is behaving the way you would want a lab to behave. Different outlets report different totals for the vulnerabilities found during that testing, so this article cites none of them, but the direction of the finding is consistently reported.

The business consequence is separate from whether the decision was right. Anyone who followed our GLM-5.2 open weights and self-hosting guide built a deployment whose entire premise was that the newest GLM would be downloadable. In 2026 that premise held for one release and then paused. Any plan that assumed each future version would arrive on schedule in the same form now needs a stated fallback.

The distinction worth holding

An open weight model that ships its weights late is still an open weight model. It is not, however, a model you can promise a client or a compliance reviewer will be running on your own hardware on a specific date in 2026. Those are two different commitments, and only one of them survives a staged release.

Who Is Actually Affected by the Delay?

Three groups feel it, and they feel it very differently. Sorting yourself into one of them is the fastest route to a decision, because the same launch is close to irrelevant for one group and a genuine planning problem for another.

What Do the GLM-5.3 Benchmarks Actually Show in 2026?

They show a large jump on agentic and coding evaluations and a distinctive tilt toward defensive security, with the caveat that every figure comes from Z.ai's evaluation of Z.ai's own model. No third party had replicated these results at the time of writing in 2026. The pattern is more informative than any single score: the model leads where finding and reasoning about problems matters, and trails where exploiting them does.

Benchmark (vendor-reported)GLM-5.3GLM-5.2Reported comparison
Terminal-Bench 3.028.34.6GPT-5.6 Sol reported at 34.6
DeepSWE v1.166.946.2GPT-5.6 Sol reported at 72.7
Agents' Last Exam28.523.8No published peer figure
AutomationBench48.2%Not publishedReported as leading its comparison set
GDPval-AA v2 (Elo)1769Not publishedReported ahead of Fable 5, Qwen3.8-Max and GPT-5.6 Sol
CyberGym84.5%77.2%Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6%
ExploitBench54.4%24.4%Mythos 5 at 78.0%, GPT-5.6 Sol at 76.5%
ExploitGym (tasks solved)105 at 2 hours, 130 at 6 hoursNot publishedFable 5 reported at 181 and 247

Two readings matter for a non-technical audience in 2026. The first is the direction of the cybersecurity result. GLM-5.3 is reported to lead on CyberGym, which measures identifying and validating flaws from source code, while trailing significantly on ExploitBench and ExploitGym, which measure progressing an exploit to completion. That is a defensive profile, and it is the opposite of the story a headline about emergent offensive capability implies.

The second is that the Terminal-Bench 3.0 jump from 4.6 to 28.3 is enormous in relative terms and still means the model fails roughly seven of every ten tasks in that evaluation. A six fold improvement on a hard benchmark is real progress and is not a signal that the category is solved. Any 2026 workflow built on it still needs a human checkpoint.

Z.ai also reports a roughly fifty percent improvement on its own internal code benchmark against GLM-5.2, achieved while consuming fewer output tokens per task. Efficiency claims of that kind are worth testing directly, because token consumption is where the real cost of an agentic workflow accumulates. Our GLM-5.2 API integration guide covers the instrumentation you would need to measure that honestly on your own tasks.

Why Is the Pricing Picture Incomplete in 2026?

Because only half of Z.ai's usual pricing structure has been published for this release. The GLM Coding Plan subscription includes GLM-5.3 in every tier, and Z.ai's documentation states the plan starts at 18 US dollars per month. The separate per-token price list, which carries entries for GLM-5.2, GLM-5.1, GLM-5 and the whole GLM-4 family, contains no GLM-5.3 row at all in 2026.

That gap has a practical effect. GLM-5.2 is priced at 1.4 US dollars per million input tokens and 4.4 US dollars per million output tokens on the published list, which is the number most teams used to justify the switch in the first place. Without an equivalent figure for GLM-5.3, nobody can calculate whether an equivalent workload gets cheaper, dearer, or stays flat. Reporting on the launch also disagrees about whether the API is live at all, with Z.ai's own documentation describing it as coming soon while several outlets describe it as already available. When the vendor's docs and the coverage conflict, the docs are the safer assumption.

Pricing dimensionStatus in 2026What a business can do with it
Coding Plan tiersLite, Pro and Max, all including GLM-5.3; Z.ai states the plan starts at 18 USD per monthBudget by seat. Predictable, and the only complete picture available.
Plan usage limitsDocumented as credit allowances per five hour window and per week, rising across the three tiersCapacity planning is possible, but only in Z.ai's credit unit rather than tokens.
Per-token API rateNo GLM-5.3 entry on the published price pageUsage based forecasting is not possible. Do not extrapolate from the GLM-5.2 rate.
Self-hosted running costNot applicable yet, since the weights are unreleasedExisting GLM-5.2 hardware economics stay valid for GLM-5.2 only.

For teams already on the subscription, this is close to a non-event and mildly positive. For teams doing bottom up cost modelling, it is a blocker. The honest framing for a 2026 planning document is that GLM-5.3 has a known seat price and an unknown unit price, and a plan should say which of the two it is built on. Our breakdown of how the GLM Coding Plan tiers work remains the right reference for the subscription side.

The Detail Most Coverage Missed

Z.ai's developer documentation states that requests for previous models, naming GLM-5.2 and GLM-5.1 specifically, are automatically routed to GLM-5.3 on the Coding Plan. That single line has more operational consequence in 2026 than any benchmark on this page. It means subscription customers cannot pin a version, cannot run a controlled comparison between the two models on the plan, and will see their outputs change whether or not they chose to upgrade.

Version pinning is not a luxury for a business. It is how you keep a prompt library stable, how you reproduce a client deliverable three months later, and how you isolate a regression when quality drops. In 2026 the only way to hold a GLM version steady is to run published weights on your own infrastructure, which returns the argument to exactly the capability that GLM-5.3 has not shipped yet.

What Should Teams Running GLM-5.2 Do in 2026?

Start by naming why GLM was chosen, because the answer decides everything else. Teams that picked it for subscription economics and coding throughput should test GLM-5.3 now, since it is already included. Teams that picked it because the weights were downloadable should recognise that the equivalent product does not exist yet in 2026 and that waiting a few weeks carries almost no cost.

Your situation in 2026Recommended pathWhat to watch
Self-hosting GLM-5.2 for data residency or complianceStay on GLM-5.2. Nothing has changed for you yet.The staged weights release and its licence terms when they land.
On a GLM Coding Plan tierYou already have GLM-5.3. Run your evaluation set against it this month.Automatic routing means the comparison window is limited.
Building on per-token API accessWait. There is no published rate to model against.A GLM-5.3 row appearing on the pricing page.
Evaluating GLM for the first timeEvaluate GLM-5.2, which is fully documented and downloadable.Whether the staged pattern repeats on the next release.
Running security or code review workloadsWorth a serious test given the CyberGym result.Vendor-reported scores need replication on your own repositories.

Whichever path applies, the evaluation discipline is the same. Keep a fixed set of twenty to fifty real tasks drawn from your own workload, with known good outputs, and run any candidate model against it in an afternoon. That set is worth more than every leaderboard in this article, and in 2026 it is the only defence against a vendor benchmark that does not transfer to your domain.

How Does a Staged Weights Release Change Open-Model Strategy in 2026?

It introduces a timing risk that open model strategies have not had to price in before. The implicit contract of the last few years was that a capable model would appear with its weights on the same day, which let teams treat open releases as a hedge against closed vendor dependency. A staged release breaks the simultaneity, and in 2026 that means open weight availability becomes a schedule to track rather than a property to assume.

This is not a reason to abandon open models. It is a reason to write the assumption down. If a deployment plan, a client contract or a compliance narrative depends on running a specific model version inside your own boundary, then the current released version is the only version you can promise, and the roadmap version is a hope. Teams that use our GLM chatbot and agent workflow guide as an operational baseline should treat the hosted and self-hosted paths as separately versioned from now on.

There is also a broader signal here worth reading calmly. A frontier capable lab discovering an unplanned offensive security capability and gating distribution while it evaluates is likely to become more common in 2026, not less. If that becomes the norm across the open weight ecosystem, the practical planning assumption shifts from same day availability to a delay of weeks, and every architecture that depends on self-hosting the newest model needs a supported previous version to fall back to.

What Are the Common Mistakes to Avoid in 2026?

Most of the errors around this release are reading errors rather than technical ones. They come from treating a version number as a promise of continuity, and from repeating vendor figures without the qualifier that makes them honest.

Key Takeaways for 2026

GLM-5.3 is a genuinely interesting engineering result wrapped in a commercially awkward launch. The capability story is about how much post-training alone can move a fixed base model. The business story is about what a company can actually buy, run and promise in 2026.

GLM-5.3 in 2026: FAQs

What is GLM-5.3 and when did it launch in 2026?

GLM-5.3 is Z.ai's flagship text model, reported as launched on 14 August 2026. Z.ai's own documentation states it uses the same base model as GLM-5.2, with every gain coming from expanded post-training rather than a new pretraining run. It carries a 1 million token context window and a 128K maximum output, and it is focused on coding, long horizon agentic work and cybersecurity.

Are the GLM-5.3 open weights available in 2026?

Not at launch. Z.ai held the GLM-5.3 weights back pending a safety evaluation, which reporting describes as roughly a two week delay from the 14 August 2026 launch. At the time of writing there is no GLM-5.3 repository on the Z.ai Hugging Face organisation, where GLM-5.2 and earlier releases are published. This is the first GLM release in 2026 to be gated this way.

How much does GLM-5.3 cost in 2026?

Only the subscription price is knowable. Z.ai's documentation says the GLM Coding Plan starts at 18 US dollars per month and that all Lite, Pro and Max tiers include GLM-5.3. The separate per-token API price list does not include GLM-5.3 at all, so there is no published usage based rate in 2026. Teams that budget by tokens rather than by seat cannot model this release yet.

Are the GLM-5.3 benchmark scores independently verified?

No. Every GLM-5.3 figure in circulation in 2026 originates from Z.ai's own evaluation of its own model, including the CyberGym and AutomationBench results where it is reported to lead. Vendor benchmarks are a directional signal until a third party replicates them. Run the model against your own fixed task set before any of these numbers influence a production decision.

Can teams keep using GLM-5.2 on a GLM Coding Plan in 2026?

Not through the Coding Plan. Z.ai's developer documentation states that requests for previous models, specifically GLM-5.2 and GLM-5.1, are automatically routed to GLM-5.3. Version pinning on the subscription is therefore not available. The only way to hold a specific GLM version steady in 2026 is to run the published GLM-5.2 weights on your own infrastructure.

Should a business switch from GLM-5.2 to GLM-5.3 in 2026?

It depends entirely on why you chose GLM in the first place. If you adopted it for the subscription economics and the coding output, GLM-5.3 is an included upgrade and worth testing now. If you adopted it because the weights were downloadable and you could run the model inside your own boundary, GLM-5.3 is not yet the same product, and waiting for the staged weights release costs you very little.

Deciding which AI models your business should actually depend on?

Distk helps growth and product teams across India and international markets choose, test and budget for AI tooling without betting the workflow on a single vendor's roadmap. If you are weighing hosted convenience against running models inside your own boundary, we can work through that with you.

Talk to Distk →