GitHub hosts HunyuanVideo, Tencent's open-source framework for large-scale video generation models, enabling AI-driven video creation.
NVIDIA Nemotron

About NVIDIA Nemotron
NVIDIA Nemotron is an open family of reasoning and multimodal models designed for agentic AI workflows. It provides transparent datasets, flexible deployment options, and policy-based routing to ensure production reliability. The models are built for high throughput and long context, supporting transparent datasets and reproducible reports. Nemotron supports deployment via NVIDIA NIM microservices or open runtimes, with TensorRT-LLM acceleration on NVIDIA GPUs to meet throughput and budget targets. It integrates retrieval, parsing, speech, and safety components, enabling robust agent stacks for enterprise and research use cases. Teams can leverage Nemotron for customer service automation, document intelligence, secure IT workflows, supply chain operations, coding assistants, and voice interfaces. The models prioritize verifiability and control, making them suitable for platform teams, enterprise application owners, AI researchers, and startups needing transparent, tunable models. Careful evaluation, routing policy design, and data curation are required to meet domain, safety, and latency targets.
Nvidia
Santa Clara, United States · Founded 1993
- Founders
- Jensen Huang, Chris Malachowsky, Curtis Priem
- Founded
- 1993
- Headquarters
- Santa Clara, United States
- Legal status
- Public company
Key features
- Open weights, data, and recipes for transparency and reproducibility
- Hybrid Mamba-Transformer MoE architectures with up to 1M-token context windows
- NeMo Switchyard for policy-based routing by accuracy, latency, and cost
- Deployment via NVIDIA NIM microservices or open runtimes (vLLM, SGLang, Ollama, llama.cpp)
- TensorRT-LLM acceleration on NVIDIA GPUs for high throughput
- Built-in safety models and guardrails for production reliability
- Multimodal models (Nano Omni) for unified video, audio, image, and text processing
- Speech models for high-throughput ASR, TTS, and S2S
- Retrieval, parsing, and embedding models for RAG and data pipelines
- Reference stack for edge, cloud, and data center deployments
Use cases
- Customer service automation with voice and text agents
- Document intelligence and secure IT workflows
- Coding assistants and supply chain operations
Pros
- Open weights, training data, and technical reports for transparency and reproducibility
- Supports multimodal inputs including video, audio, image, and text for agentic workflows
- Optimized for high throughput and long context (up to 1M tokens) with efficient architectures like hybrid Mamba-Transformer MoE
- Flexible deployment options via NVIDIA NIM microservices or open runtimes on NVIDIA GPUs
- Includes specialized models for retrieval, parsing, speech, and safety tailored for agentic AI applications
Cons
- Requires careful evaluation and routing policy design to meet domain-specific, safety, and latency targets
- Deployment and optimization may demand expertise in GPU acceleration and model tuning
NVIDIA Nemotron videos
Frequently asked questions about NVIDIA Nemotron
What are NVIDIA Nemotron models?
NVIDIA Nemotron is a family of open models with open weights, training data, and recipes designed for building specialized AI agents. They emphasize transparency, efficiency, and high throughput for agentic workflows.
Who should use NVIDIA Nemotron models?
The models suit platform teams, enterprise application owners, AI researchers, and startups needing transparent, tunable models for agentic AI workflows such as customer service automation or document intelligence.
How can I deploy NVIDIA Nemotron models?
Models can be deployed using open frameworks like vLLM, SGLang, Ollama, or llama.cpp on any NVIDIA GPUs, or as NVIDIA NIM microservices for easy integration on GPU-accelerated systems.
What types of models are included in the Nemotron family?
The family includes reasoning models (Nano, Super, Ultra), multimodal models (Nano Omni), retrieval models (Nemotron Retriever), document parsing models (Nemotron Parse), speech models (Nemotron Speech), and safety models (Nemotron Safety).
Are there any prerequisites for using Nemotron models?
Users should have familiarity with NVIDIA GPUs and AI deployment frameworks. Careful evaluation, routing policy design, and data curation are required to meet domain, safety, and latency targets.
Where can I access Nemotron models and resources?
Models and technical reports are available on Hugging Face, and endpoints can be experienced via NVIDIA NIM APIs or OpenRouter demos.