Infer by Flow7

Responses API gateway for coding agents, with a public model catalog and per-request spend ceilings

Infer by Flow7 is an API gateway that routes requests to models from OpenAI, Anthropic, Google, xAI, DeepSeek, Moonshot AI and Qwen through a single key and a prepaid balance.

It exposes a Responses-compatible endpoint at /v1/responses rather than chat completions, and documents verified setups for Codex CLI, OpenCode, Pydantic AI, LangChain and the Vercel AI SDK. Four routing options (Low Cost, Balanced, Stable, and Official API to a model developer's first-party endpoint) let you choose between price and route stability per request.

Spend ceilings are checked before a request is dispatched, and each completed call records a receipt with the resolved model, route tier, token counts and price version. The public catalog lists 21 models with 79 published prices, including labels on the routes that cost more than the model developer's own rate.

Funding starts at a USD 20 minimum, with a USD 2,000 cap on account balance.

Pricing: Per token usage

HQ 🇺🇸 United States
Screenshot of Infer by Flow7 webpage

Work on Infer by Flow7? Feature it at the top of Inference APIs.

Is your product missing?

Add it here →