ResearchResearch paperReinforcement Learning · Training & Scaling1 source · Oct 6, 2026

NeMo-DCR: Bit-Exact Delta-Compressed Refit for Scalable Agentic RL at Trillion-Parameter Scale

We present NeMo-DCR (Delta-Compressed Refit), which sends only changes yet is bit-exact: receivers obtain the same parameter and buffer bits as a dense refit.

Key points

  • Agentic reinforcement learning (RL) disaggregates training from rollout, so each policy update must reach the rollout clusters before the next batch.
  • Transferring a full 1T checkpoint for such weight synchronization (refit) takes 87.5 min between two AWS regions.
  • Recent systems exploit this sparsity but fall short on placement, exactness, or efficiency: they reimplement placement rules, assemble full tensors, rebuild values arithmetically, or use a cross-cluster collective, and none fully recovers from mid-refit failures.
  • A 1T relay-tree refit at 3% takes 150 s instead of 87.5 min, making refits practical for cross-cluster agentic RL at trillion-parameter scale.

Sources (1)

Extractive summary: sentences quoted from the sources.

Before this

  1. Oct 5, 2026Introducing GLM 5.3 on Amazon Bedrock
  2. Oct 5, 2026Supercharge regulated workloads with Claude Code and Amazon Bedrock
  3. Oct 5, 2026New agent skill: Amazon SageMaker optimized generative AI inference for your coding agent
  4. Oct 3, 2026We're going to need default hard budget caps on pretty much everything
  5. Oct 3, 2026LWiAI Podcast #258 - Opus 5.5, Sol and Luna, Muse, DeepSeek-V4.1-Flash, Xi
  6. Sep 30, 2026[AINews] OpenAI DevDay 2026: Dots, 6.1 Sol, Ultrafast, Decisions API, Agents API, Spaces, Marketplace, and 1.2 Billion ChatGPT WAU

Related