≫ Home / LLM API pricing / GLM 5.3

GLM 5.3 API pricing

Of the 8 providers checked, DeepInfra lists the lowest blended rate for GLM 5.3: $0.5625 input and $2.50 output per 1M tokens, on its pricing page as of 25 September 2026.

Z.ai's model for coding and long-running agent work. It shares GLM 5.2's base model; Z.ai credits the improvements to post-training alone. The weights use a custom GLM-5.3 license, not MIT like GLM 5.3 Flash.

Creator
Z.ai
Released
18 August 2026
Weights
Open
License
GLM-5.3 License
Context
1M tokens
Input
text
Model card
Hugging Face

Facts from the model card, checked 25 Sep 2026.

Across 8 providers, the most expensive charges 2.1× the cheapest on a blended basis. 1 provider is on a limited-time promotion.

Prices by provider

Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 DeepInfra $0.5625 $0.125 $2.50 25 Sep 2026 Promotional price, 38% off list $0.90 / $4.00
2 Alibaba Cloud Model Studio $1.40 – $4.40 25 Sep 2026 International region
3 Baseten $1.40 $0.14 $4.40 25 Sep 2026
4 Cloudflare Workers AI $1.40 $0.26 $4.40 25 Sep 2026
5 Fireworks AI $1.40 $0.26 $4.40 25 Sep 2026 Standard tier
6 Novita AI $1.40 $0.26 $4.40 25 Sep 2026
7 SiliconFlow $1.40 $0.26 $4.40 25 Sep 2026
8 Together AI $1.40 $0.26 $4.40 25 Sep 2026

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. The same data is open under CC-BY at /api/query/prices.

Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.

Is your product missing?

Add it here →