Model ID
deepseek-v4-flashProvider
DeepSeek — served by DeepInfra, DigitalOcean Inference, BytePlus (Ark), Atlas Cloud
Context
1,048,576-token context window · 65,536 max output tokens
Capabilities
Streaming, JSON mode, Tool calling, Reasoning, Long context, Multilingual
Pricing
Typical request (5K input / 1K output tokens): ~$0.001. Billing is metered per token. Prices are live at
GET /api/v1/models/deepseek-v4-flash; the values above were read from that endpoint when this page was generated.
Use this model
The SDK examples assume you have created a server-side client.Supported parameters
messages, temperature, max_completion_tokens, top_p, stop, frequency_penalty, presence_penalty, seed, stream, user, routing, tools, tool_choice, response_format, reasoning, reasoning_effort
Send only parameters from this list. routing.require_parameters (default true) skips any provider rail that cannot honor every requested parameter, so unsupported parameters reduce the rails eligible to serve the request.
Routing & fallbacks
Pindeepseek-v4-flash with model when you want this model’s behavior, or list it in an ordered models array. With ninja/auto first, the router’s top pick runs first and deepseek-v4-flash is the explicit fallback; ninja/auto acts as the router only when it is the first entry.
Check live support
TypeScript SDK