Nemotron

A series of open source large language models launched by NVIDIA, with open weights and training data used to build dedicated AI agents

In-depth Report

  • NVIDIA Nemotron is the most open series of AI large language models, equipped with open weights, training data and schemes, providing industry-leading efficiency and accuracy for building purpose-built AI agents.

  • The core innovation of the Nemotron series lies in its openness. The training data and model weights are completely open and can be downloaded for free on Hugging Face. The architecture adopts a hybrid Mamba-Transformer MoE architecture and supports context windows of up to 1 million tokens. In terms of deployment, it supports open source frameworks such as vLLM, SGLang, Ollama, and llama.cpp, providing ultra-high throughput reasoning and reducing reasoning costs.

  • The Nemotron product line is divided into multiple models to meet different needs: Nemotron 3 Nano 30B is positioned as cost-effective and suitable for the highest accuracy and efficiency in target tasks; Nemotron 3 Super 120B strikes a balance between efficiency and accuracy and is suitable for multi-agent environments to handle complex tasks; Llama Nemotron Ultra 253B provides the highest accuracy and is suitable for multi-agent enterprise workflows; Nemotron Nano VL 12B focuses on visual language and is suitable for document intelligence and video understanding; Nemotron RAG provides retrieval enhancement functions, including extraction, embedding, and reordering; Nemotron Safety provides security audit functions, including jailbreak detection, content audit, and PII detection; Nemotron Voice provides complete voice AI capabilities, including ASR, TTS, and voice translation.

  • Nemotron's multi-modal capabilities cover visual understanding, information retrieval, speech processing and security functions, supporting scenarios such as RAG and agent applications.

  • No specific pricing information is available on this page. Available deployment and trial methods include free trials of some models through OpenRouter, NVIDIA NIM providing inference endpoint API services, and third-party inference providers including Baseten, DeepInfra, Fireworks AI, FriendliAI, Together AI, etc.

User Reviews

  • 头像
    Jason_221
    Nemotron 3 Ultra's SWE-Bench score of 71.9% is quite strong. The key is that the training recipes and data sets are all open source. NVIDIA is much more generous than Meta in this regard.

  • 头像
    MAjac
    I tried Nano Omni. It can run multi-modal by activating 30B. The video understanding throughput is 9 times that of the benchmark model. It can run locally with 25GB of video memory. This efficiency is really outrageous.

  • 头像
    MarilynCooper00728
    The Nemotron committee model was criticized by SemiAnalysis, saying that groupthink slowed down progress. But the model itself is good, and just because the committee screwed it up doesn't mean the technology isn't good.

  • 头像
    史鹏
    Using Nemotron 3 Super for code review, I reviewed the PRs of 19 files in 12 seconds and grabbed a CORS regression, which was much faster than my manual work. However, the suggestions for cross-file context are shallow.

  • 头像
    Amanda.Cruz_2021
    I have been running it locally on Nano 30B for a while. The coding task is indeed strong, but the chat is too robotic, like talking to a serious engineer, without emotion.

  • 头像
    AbigailSchuster
    Ultra's PinchBench 91% is tied with Kimi K2.6, but the inference speed is more than 4 times faster. In terms of practicality, I think Nvidia's optimization ideas are more practical.

  • 头像
    Pamela.Scott_20237
    The Mamba ecosystem is still not mature enough. The support for Nemotron by llama.cpp and Ollama is half-complete, so non-N card users should be cautious.

  • 头像
    James185
    Nemotron 3 Super's SWE-Bench Multilingual 45.78% beats GPT-OSS's 30.8%. The gap is really big in multi-language code repair scenarios.

  • 头像
    ABell520155
    After running a 1M context, I can really fit it into the entire code base for analysis. The official 1 million tokens are not exaggerated. However, the loading time is a bit long.

  • 头像
    Linda.Roberts_Plus
    Speaking of which, Nemotron 3 Super is free on OpenRouter, with a price of 0 USD/million tokens. It’s quite fun to engage in agent development.

  • 头像
    crazybird811
    This model is too long-winded. It outputs a lot more tokens than other models for the same task. When using it in a production environment, you need to consider the cost of tokens. I have used Super to build a small agent service before. For the same customer service work order classification task, Super outputs 800 tokens, while Qwen3.5 only costs 300 tokens. The magnitude difference is nearly three times. Although the accuracy is similar, the cost is doubled.

  • 头像
    JanetNeal
    NVFP4 量化真的是黑科技,精度几乎不掉,但吞吐翻了好几倍。Blackwell 上跑应该更爽。

  • 头像
    RobertSimmonsJr
    The Ultra running robot can last for 80 rounds without losing its chain, and there is no hallucination of the title of the paper. In the same task, DeepSeek V4 compiled a paper that does not exist, and we can judge the difference. However, I used DeepInfra for a fee, and it cost about a dozen dollars to run the complete research pipeline. If I build it myself, the initial infra cost is not small. It depends on whether the team has the deployment capabilities.

  • 头像
    dOUGLAS892
    Nemotron 3 Ultra paired with LangChain's Harness Engineering solution makes the inference cost 10 times cheaper than Opus 4.8, with a difference of only 0.01. It is enough for small and medium-sized enterprises.

  • 头像
    RebeccaGomez1685
    Using Cascade to run HumanEval 97.6%, ClassEval 88%, 30B total parameter 3B activation can have this coding ability. The guy on Reddit is right, it is indeed underestimated.

  • 头像
    CosmosCris915
    搞深度研究的瓶颈已经不在模型本身了,英伟达把整个工具箱都开源了,什么 Nemoclaw、OpenShell 一套下来,搭 agent 比之前舒服太多。之前搭多智能体系统从零开始搞 agent 框架搭了好几天,现在直接拿 OpenShell 的沙盒加编排框架,半天就搭起来了,门槛确实降了很多。

  • 头像
    JoseBaileyII
    中文理解能力跟国产模型比差距还是挺明显的,毕竟主要优化的是英文 Agent 场景,做中文产品的话优先考虑 Qwen 或者 DeepSeek。

  • 头像
    Patricia.Cox
    After Greptile tested it, he said that Super is a "first-class first-round review tool." I also use it myself. The code review scenario is really powerful.

  • 头像
    MelissaGray
    Internet search is run locally with Nemotron Nano 30B. The review on OpenWebUI is reasonable, and the experience is better than many small cloud models.

  • 头像
    IFlores168
    英伟达搞的这个开源协议比 Llama 那些伪开源清爽多了,没有月活限制,商业部署不用提心吊胆。之前公司想用 Llama 做产品,法务一看那个商用协议里一堆附加条款,折腾了好久才放行。Nemotron 的 OpenMDW-1.1 虽然也有 patent termination clause,但整体限制少很多。

  • 头像
    A_Lriv
    Super 在 RULER @1M 上 91.75%,对比 GPT-OSS 的 22.30% 简直是降维打击,长上下文场景选 Nemotron 没毛病。

  • 头像
    purplegoose369
    I wrote a small FPS game using Ultra to test the waters. It took several days to generate code from scratch before it could run. The debugging ability is good, but creative coding is really not its strong point.

  • 头像
    AmellyDa luz
    The model reply smells like over-optimization of RLHF, which is too conservative and neat. It is suitable for enterprise deployment but not fun. Right out of the box, this is the right choice.

  • 头像
    MStephens_77
    Cadence 用 Nemotron 3 Ultra 做芯片验证,把 RTL 验证从几周缩短到几小时,这才是模型的正确用法。

  • 头像
    JoseRoss_Max3
    Mamba-2 + Transformer 混搭架构真的是被低估的设计,长序列内存效率拉满,纯 Transformer 做 1M context 根本吃不消。我试过用同样参数量级的纯 attention 模型加载 500K context,显存直接爆炸,Nemotron 的 Mamba 层在长序列上只存固定大小的状态,这个设计差异在实际部署时是天壤之别。

  • 头像
    MsChristopherCarpenter_dev
    Nemotron 3 Nano Omni ranks first in six multi-modal lists. The video annotation throughput on MediaPerf is higher than GPT 5.1. Can you bear this?

  • 头像
    DorothyBrooksK
    I tried the role-playing function - I originally thought that this kind of enterprise-level model would not be able to write stories, but I ended up writing a cyberpunk plot about repairing computers. The details were so rich that I was a little surprised. From the jingling store door bell to the broken Seagate hard drives on the shelves, to the boss wearing reading glasses, the entire atmosphere is described online. I never expected that a model for enterprise deployment could have such narrative capabilities.

  • 头像
    redmouse606
    Ultra 的 structured output 可靠性真不错,200 次 JSON 请求只有 1 次格式错误,做 pipeline 很稳。

  • 头像
    Eugene.Evans_Pro678
    卡在 Arena-Hard 73.88% 这个成绩上,跟 GPT-OSS 的 90.26% 差了快 17 个点,聊天质量确实不如其他模型。

  • 头像
    Harold84
    Super's 120B model Q4_K can run 1M context on M1 Ultra after quantization. Although the speed is greatly reduced, it can at least run. This is already ahead of many models of the same size. According to actual measurements, prompt processing dropped from 255 t/s in zero context to 142 t/s in 200K context, and token generation also dropped from 26.7 to 20.6 t/s. Although it is slow, it is usable. This was something that I would not have thought of half a year ago.