Fireworks AI
The production AI platform built for developers.
Fireworks AI offers a portfolio of LLMs and image models for a range of applications, like natural language processing, coding, and image generation. Model examples include Mixtral MoE 8x7B Instruct for instruction following, FireFunction V1 for function calling, and Llama 2 70B Chat optimized for dialogue applications. They also provide models specifically tailored to language tasks in French, Chinese, and Japanese, as well as StarCoder models for programming languages and Stable Diffusion for image generation.
Pricing: Per token usage
Fireworks AI prices by model
Per 1M tokens, read off Fireworks AI's own pricing page on the date shown.
| Model | Input / 1M | Cached / 1M | Output / 1M | Checked | Notes |
|---|---|---|---|---|---|
| DeepSeek V4 Pro | $1.32 | $0.044 | $3.96 | 25 Sep 2026 | Standard tier; listed as DeepSeek V4 Pro (0813) |
| DeepSeek V4.1 Flash | $0.30 | $0.006 | $1.20 | 25 Sep 2026 | Standard tier |
| GLM 5.3 | $1.40 | $0.26 | $4.40 | 25 Sep 2026 | Standard tier |
| GLM 5.3 Flash | $0.15 | $0.03 | $0.50 | 25 Sep 2026 | Standard tier |
| Kimi K3 | $3.00 | $0.30 | $15.00 | 25 Sep 2026 | Standard tier |
| MiniMax M3 | $0.30 | $0.06 | $1.20 | 27 Sep 2026 | Standard tier |
Fireworks AI Alternatives
Explore 115 products in the Inference APIs category. View all Fireworks AI alternatives.
CheapestInference
Flat-rate unlimited inference on open-weight models, sold in daily 8-hour windows
DeepInfra
Run the top AI models using a simple API, pay per use. Low cost, scalable and production ready infrastructure.
Groq
LPU-powered inference API for LLMs, speech, and vision models with usage-based pricing
BentoML
BentoML is the platform for software engineers to build AI products.
Compare
Work on Fireworks AI? Feature it at the top of Inference APIs.
Is your product missing?