Model ID
nemotron-3-nanoProvider
NVIDIA — served by DeepInfra, Novita AI
Context
262,144-token context window · 16,384 max output tokens
Capabilities
Streaming, JSON mode, Tool calling, Reasoning, Long context, Multilingual
Pricing
Typical request (5K input / 1K output tokens): ~$0.0005. Billing is metered per token. Prices are live at
GET /api/v1/models/nemotron-3-nano; the values above were read from that endpoint when this page was generated.
Use this model
The SDK examples assume you have created a server-side client.Supported parameters
messages, temperature, max_completion_tokens, top_p, stop, frequency_penalty, presence_penalty, seed, stream, user, routing, tools, tool_choice, response_format, reasoning, reasoning_effort
Send only parameters from this list. routing.require_parameters (default true) skips any provider rail that cannot honor every requested parameter, so unsupported parameters reduce the rails eligible to serve the request.
Routing & fallbacks
Pinnemotron-3-nano with model when you want this model’s behavior, or list it in an ordered models array. With ninja/auto first, the router’s top pick runs first and nemotron-3-nano is the explicit fallback; ninja/auto acts as the router only when it is the first entry.
Check live support
TypeScript SDK