≫ Home / LLM API pricing / GLM 5.3 Flash

GLM 5.3 Flash API pricing

Of the 8 providers checked, DeepInfra lists the lowest blended rate for GLM 5.3 Flash: $0.075 input and $0.25 output per 1M tokens, on its pricing page as of 25 September 2026.

The first GLM 5 model to take images as input. Only 18B of its 320B parameters are active per token. Z.ai prices it at about a tenth of GLM 5.3 and the providers here follow: $0.15 / $0.50 per 1M against $1.40 / $4.40.

Creator
Z.ai
Released
26 August 2026
Weights
Open
License
MIT
Size
320B total, 18B active
Input
text, image
Model card
Hugging Face

Facts from the model card, checked 25 Sep 2026.

Across 8 providers, the most expensive charges 3.0× the cheapest on a blended basis. 1 provider is on a limited-time promotion.

Prices by provider

Per 1M tokens, on-demand, sorted cheapest first by blended cost (3:1 input to output). Off-peak and long-context rates are listed on each provider's page.

# Provider Input / 1M Cached / 1M Output / 1M Checked Notes
1 DeepInfra $0.075 $0.015 $0.25 25 Sep 2026 Promotional price, 50% off list $0.15 / $0.50
2 Baseten $0.15 $0.03 $0.50 25 Sep 2026
3 Cloudflare Workers AI $0.15 $0.03 $0.50 25 Sep 2026
4 Fireworks AI $0.15 $0.03 $0.50 25 Sep 2026 Standard tier
5 Novita AI $0.15 $0.03 $0.50 25 Sep 2026
6 SiliconFlow $0.15 $0.03 $0.50 25 Sep 2026
7 Together AI $0.15 $0.03 $0.50 25 Sep 2026
8 Berget AI €0.25 ≈ $0.2842 €0.25 €0.50 ≈ $0.5684 27 Sep 2026 Cached input is billed at the input price

Each price was read off the provider's own pricing page on the date shown; the date links to that page. Providers that publish no per-token price for this model are not listed. EUR prices are ranked at the ECB reference rate of 24 September 2026. The same data is open under CC-BY at /api/query/prices.

Compare every inference provider on free tiers, hosting and EU availability in the inference APIs table.

Is your product missing?

Add it here →