Skip to content
NVIDIAVoice

Nemotron 3.5 ASR Streaming Multilingual 0.6B

nemotron-3.5-asr-streaming-multilingual-0.6b

NVIDIA Nemotron 3.5 ASR multilingual speech-to-text (0.6B). Billed per second of audio.

Context
N/A
Endpoint

Get API KeyCompare

Pricing

Input$0 / 1M
Output$0 / 1M
Cache Write (5m)Not applicable
Cache Write (1h)Not applicable
Cache ReadNot applicable
Web Search$0 / 1M

Customer audio pricing

USD

Base rates shown at GroupRatio=1. Your account's group multiplier may change the final charge.

/v1/audio/transcriptionsSpeech to text
RateUSD per unit
InputAudio$0.00000333 USD per audio second

Actual settled usage is deducted from your Apertis credit quota.

Quick Start

Select an endpoint and copy a working example for this model.

Endpoint
python
from openai import OpenAI client = OpenAI(    api_key="YOUR_API_KEY",    base_url="https://api.apertis.ai/v1") response = client.chat.completions.create(    model="nemotron-3.5-asr-streaming-multilingual-0.6b",    messages=[        {"role": "user", "content": "Hello!"}    ],    max_tokens=1024,    temperature=0.7) print(response.choices[0].message.content) # Optional: Enable context compression to reduce token usage# response = client.chat.completions.create(#     model="nemotron-3.5-asr-streaming-multilingual-0.6b",#     messages=[{"role": "user", "content": "Hello!"}],#     extra_body={"compression": {"enabled": True, "model": "gpt-4.1-mini"}}# )

Supported Parameters

API docs
Common5 params
modelfilelanguagepromptresponse_format
Extended2 params
temperaturetimestamp_granularities

Cursor IDE Model IDs

Use these namespaced identifiers in Cursor IDE to avoid conflicts with built-in models.

nemotron-3.5-asr-streaming-multilingual-0.6b

Compare with Other Models

See how this model compares to others from the same provider.

Nemotron 3 Nano Omni (Free)

NVIDIA Nemotron 3 Nano Omni is an open 30B-A3B multimodal model designed as a perception and context sub-agent for enterprise agent systems. It supports text, image, video, and audio inputs with text output, enabling unified multimodal reasoning within a single inference loop. Built on a hybrid MoE Transformer–Mamba architecture with Conv3D video layers and Efficient Video Sampling (EVS), it delivers significantly improved efficiency for video reasoning—achieving ~2× higher throughput and 2.5× lower compute compared to separate pipelines. With up to 300K context length and extended thinking support, it is well suited for scalable, multimodal agent workflows.

Context
256K
Input
$0/M
Output
$0/M

Llama 3.3 Nemotron Super 49B V1.5

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B reasoning and chat model derived from Llama-3.3-70B-Instruct, tuned for agent workflows like RAG and tool calling with a 128K context window. It combines supervised training with multiple RL stages to improve alignment, step-by-step reasoning, and tool use, while a NAS “Puzzle” architecture reduces memory and boosts throughput so it can run on a single H100/H200. It delivers strong results across math and coding benchmarks, supports toggleable reasoning modes, and is designed for efficient, reliable agent systems and long-context retrieval where accuracy and cost balance matter.

Context
131.1K
Input
$0.05/M
Output
$0.20/M

Nemotron 3.5 Content Safety (Free)

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, designed for content moderation, safety classification, and AI policy enforcement. Supporting text and image inputs with text output, it evaluates both user prompts and model responses, providing safe/unsafe classifications, safety category labels, and optional reasoning traces. Fine-tuned from Gemma-3-4B and supporting 12 languages with a 128K-token context window, the model is well suited for prompt moderation, response filtering, content classification, and enterprise safety pipelines. As part of the NVIDIA Nemotron family, it offers a configurable reasoning mode and integrates easily into agentic AI systems requiring robust guardrails and compliance controls.

Context
128K
Input
$0/M
Output
$0/M

Nemotron 3 Ultra (Free)

NVIDIA Nemotron 3 Ultra is an open frontier reasoning and orchestration model featuring a 550B-parameter Mixture-of-Experts (MoE) architecture with 55B active parameters per token. Built on a hybrid Transformer–Mamba design, it supports text input and output with a 1M-token context window, enabling large-scale reasoning and long-horizon task execution. Optimized for agent orchestration, coding agents, deep research, and complex enterprise workflows, the model excels at multi-step reasoning, planning, and sustained execution. With high-throughput inference designed for large-scale agent pipelines, Nemotron 3 Ultra serves as a powerful foundation for advanced agentic AI systems.

Context
1M
Input
$0/M
Output
$0/M