Data Center Richness

Data Center Richness

Google Explores Distributed AI Training

Google DiLoCo Seeks to Tap Idle Compute Across An Entire Data Center Network

Rich Miller's avatar
Rich Miller
Jul 21, 2026
∙ Paid
AI infrastructure inside a Google data center. (Photo: Google)

Google is researching new ways to tap idle compute capacity spread across its global data center network to train AI models.

The approach, detailed in a paper from Google DeepMind and Google Research, introduces an architecture called Decoupled DiLoCo (Distributed Low-Communication) that allows AI training to run across remote facilities connected by standard networking, while tolerating hardware failures that would shut down conventional training runs.

The research points toward a future where distributed compute can supplement the massive clusters that power today’s frontier AI training, though not replace them. The tightest training loops for the largest models still require enormous, purpose-built facilities with high-bandwidth interconnects.

What Decoupled DiLoCo offers is a way to put additional capacity to work, wherever it sits, using the networking already in place between data center sites.

“As frontier models continue to grow in scale and complexity, we’re exploring diverse approaches to train models across more compute, locations and varied hardware,” the Google team wrote.

How Training Works Today

To understand the shift, it helps to understand today’s approach. Frontier AI models are typically trained using a programming model called Single Program, Multiple Data (SPMD) combined with synchronous distributed training. Thousands of accelerators execute the same training program on different portions of the workload, periodically synchronizing their results before advancing to the next step.

That tight synchronization is one of the primary reasons today’s AI training campuses concentrate massive GPU clusters behind ultra-high-bandwidth networking fabric and enormous power infrastructure. It also means that a single slow accelerator, or in many systems a hardware failure, can delay or interrupt work across tens of thousands of GPUs, making resilience and cluster-wide coordination critical to efficient training.

User's avatar

Continue reading this post for free, courtesy of Rich Miller.

Or purchase a paid subscription.
© 2026 Miller Webworks LLC · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture