All news briefs
Product update

NeMo AutoModel scales Diffusers fine-tuning from one GPU to a cluster

The story in five sentences

  1. Nvidia and Hugging Face are integrating NeMo AutoModel more closely with Diffusers-format models.
  2. Models should train without checkpoint conversion or a custom rewrite.
  3. Configurable parallelism methods control scaling instead of separate training programs.
  4. Ready-made recipes cover several open image and video model families as well as full fine-tuning and LoRA.
  5. A typed Python API is planned but is not available in the described version.

Full explanation

NeMo AutoModel is an open-source, PyTorch DTensor-based training library that now works directly with models in the Hugging Face Diffusers format. Checkpoints can be loaded from and written back to that ecosystem without a separate conversion step. Supporting a model is meant to require a contained adapter rather than a fully rewritten training program. The integration is published under the Apache 2.0 license.

Scaling is controlled through configuration. The available approaches include FSDP2 plus tensor, expert, context, and pipeline parallelism. The current implementation supports only flow-matching models. Existing recipes cover Wan 2.1, Wan 2.2, FLUX.1, FLUX.2, HunyuanVideo 1.5, and Qwen-Image. Depending on the model, users can perform complete fine-tuning or use LoRA.

The walkthrough fine-tunes FLUX on 78 public-domain Rider-Waite images for 200 steps. It demonstrates the workflow but does not establish general output quality. Reported measurements use eight H100 GPUs with 80 gigabytes of memory and come from the participating vendors. Nvidia and Hugging Face announce a typed recipe API as future work, so that interface should not be described as available in the current release.

Privacy setting

With your consent, PostHog EU measures which pages are opened and additionally records your session: mouse movement, clicks, scrolling and the rendered page content are stored as a replayable reconstruction. Typed input is masked before sending. The contents of sign-in, sign-up, password recovery, the founder chat, the newsletter field, your email display, your quiz answers and the checklists are not recorded at all; on the sign-in and sign-up pages recording is paused. Addresses are always stored without query parameters. A random device identifier and technical connection data such as the IP address are processed as well. In addition, strictly defined usage events are measured: how far a post was read, which article element was used, which call to action was clicked, how a newsletter sign-up ended, how far you get in a course or in the feed (recording only whether a task was correct or incorrect, never your answer itself), which video you start, whether you switch the language, and how a sign-in or sign-up attempt ended — without the email address, without the password, and without the error message. Only values from a fixed list and whole numbers from fixed ranges are transmitted — no input and no free text. We also label production, internal, and automated test visits separately and derive whether a visit came from a known interface such as ChatGPT, Claude, or Perplexity from a known referring domain or a strictly allowed campaign value. The full referring address and query parameters are not sent, and this cannot identify a specific AI model. If PostHog's IP-based geo enrichment is enabled in the project, the service can derive an approximate country, continent, and region; we do not request browser or GPS location. Every visit is also assigned to a page area from a fixed list (for example home, blog, tools, account) — the area is sent, not the address. Page loading and stability metrics are measured as well (Core Web Vitals: LCP, CLS, INP, FCP), without network payloads. A click on a link leading away from the site is recorded only as the target domain from a fixed list, without path, query parameters, or link text. For the Idea Wheel, AI Labelling and Limit Reset tools only the kind of action is measured — spin, option changed, copy or export, for instance — never your input and never a result. Learn more about privacy

Without your consent no analytics code is loaded and nothing is recorded. You can withdraw at any time: withdrawal ends collection immediately; data already collected may have been transmitted by then and is deleted after the storage period.