Nemotron Nano 9B V2 (Free)
nemotron-nano-9b-v2NVIDIA Nemotron Nano 9B v2 is a 9B-parameter language model trained from scratch by NVIDIA, designed to handle both reasoning and non-reasoning tasks. It can generate an internal reasoning trace before producing a final answer, and this behavior is configurable via system prompts—allowing developers to enable or suppress visible reasoning as needed.
- Context
- 131.1K tokens
- Endpoint
Pricing
Quick Start
Select an endpoint and copy a working example for this model.
from openai import OpenAI client = OpenAI( api_key="YOUR_API_KEY", base_url="https://api.apertis.ai/v1") response = client.chat.completions.create( model="nemotron-nano-9b-v2", messages=[ {"role": "user", "content": "Hello!"} ], max_tokens=1024, temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(# model="nemotron-nano-9b-v2",# messages=[{"role": "user", "content": "Hello!"}],# extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )Supported Parameters
API docsmodelmessagesmax_tokenstemperaturetop_pstreamtoolsreasoning_effortstream_optionsthinkingextra_bodyCursor IDE Model IDs
Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.
Compare with Other Models
See how this model compares to others from the same provider.
Nemotron 3.5 Lightning
NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model with 30B total parameters and 3B active per token, optimized for high-throughput agentic workloads and efficient inference. Its lightweight active compute and open design make it well suited for specialized agents, domain-specific customization, and scalable production deployments where speed, cost efficiency, and adaptability are key.
- Context
- 1M
- Input
- $0/M
- Output
- $0/M
Nemotron 3.5 Lightning (Free)
NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model with 30B total parameters and 3B active per token, optimized for high-throughput agentic workloads and efficient inference. Its lightweight active compute and open design make it well suited for specialized agents, domain-specific customization, and scalable production deployments where speed, cost efficiency, and adaptability are key.
- Context
- 1M
- Input
- $0/M
- Output
- $0/M
Nemotron 3 Nano Omni (Free)
NVIDIA Nemotron 3 Nano Omni is an open 30B-A3B multimodal model designed as a perception and context sub-agent for enterprise agent systems. It supports text, image, video, and audio inputs with text output, enabling unified multimodal reasoning within a single inference loop. Built on a hybrid MoE Transformer–Mamba architecture with Conv3D video layers and Efficient Video Sampling (EVS), it delivers significantly improved efficiency for video reasoning—achieving ~2× higher throughput and 2.5× lower compute compared to separate pipelines. With up to 300K context length and extended thinking support, it is well suited for scalable, multimodal agent workflows.
- Context
- 256K
- Input
- $0/M
- Output
- $0/M
Llama 3.3 Nemotron Super 49B V1.5
Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B reasoning and chat model derived from Llama-3.3-70B-Instruct, tuned for agent workflows like RAG and tool calling with a 128K context window. It combines supervised training with multiple RL stages to improve alignment, step-by-step reasoning, and tool use, while a NAS “Puzzle” architecture reduces memory and boosts throughput so it can run on a single H100/H200. It delivers strong results across math and coding benchmarks, supports toggleable reasoning modes, and is designed for efficient, reliable agent systems and long-context retrieval where accuracy and cost balance matter.
- Context
- 131.1K
- Input
- $0.05/M
- Output
- $0.20/M