NVIDIA
nemotron-nano-12b-v2-vlContext window128,000 tokens
StreamingJSON modeTool callingVisionReasoningLong context
Maximum output: 16,384 tokens.
Pricing
Typical request (5K input / 1K output tokens): ~$0.0016. Live pricing.
Use this model
Set up your SDK client, then run:Supported parameters
messagestemperaturemax_completion_tokenstop_pstopfrequency_penaltypresence_penaltyseedstreamuserroutingtoolstool_choiceresponse_formatimage_url content partsreasoningreasoning_effortrouting.require_parameters (default true) skips any provider rail that cannot honor every requested parameter, so requests keep the behavior you asked for.
Routing & fallbacks
Pinnemotron-nano-12b-v2-vl with model when you want this model’s behavior, or list it in an ordered models array. With ninja/auto first, the router’s top pick runs first and nemotron-nano-12b-v2-vl is the explicit fallback; ninja/auto acts as the router only when it is the first entry.
Check current availability and limits
Check current availability and limits
TypeScript SDK