All news briefs
Developer platform

Nvidia positions Vera Rubin for continuous post-training of AI agents

The story in five sentences

  1. Nvidia describes continuous post-training as a central computing workload for agentic AI.
  2. The company proposes intelligence per dollar as a metric alongside cost per token.
  3. NeMo Gym and NeMo RL are intended to standardize distributed training environments and reinforcement learning.
  4. Nvidia claims that Vera Rubin can train the largest models with one quarter of the Blackwell GPUs.
  5. Most performance and efficiency figures come from Nvidia or the partner companies mentioned.

Full explanation

Nvidia presents continuous post-training as a recurring production cycle for AI agents. Tools and simulated or production-related environments generate rollouts, those attempts are evaluated, and the results are used for further weight updates. NeMo Gym supplies environments for that process, while NeMo RL coordinates distributed reinforcement learning. The stack is intended to make repeated improvement a standard workload.

Alongside cost per token, Nvidia promotes intelligence per dollar to assess both the initial creation and continued improvement of a model. The company cites Nemotron 3 Ultra, a 550-billion-parameter mixture-of-experts model, and reports 71.7 percent on SWE-bench. This is an Nvidia figure. Nvidia also claims that Vera Rubin can train the largest models with one quarter of the GPUs required by Blackwell.

Other figures come from named partners. Prime Intellect reports a 30 percent increase in CPU throughput, while Perplexity describes synchronization for trillion-parameter models in under two seconds. These remain vendor or partner statements. Plans involving Together and Prime Intellect describe intended future use rather than established production. The evidence supports Nvidia’s product positioning, not a neutral comparison proving its claimed advantage.