Shai Magzimof
EN עב

Distributed data centers are more secure

Fire and black smoke rising from an AWS data center in the UAE after Iranian drone strikes, March 2026.
Dubai, UAE, March 1, 2026. Iranian drones struck two Amazon AWS data centers in the ME-CENTRAL-1 region. Two of three availability zones went offline. Photo: Getty Images / BBC.

What worries me about a gigawatt on one site is how much could disappear in one event.

A data center needs a substation, fiber connections, and industrial cooling as well as GPUs. Cybersecurity protects one part of that system; it does not resolve exposure to physical damage. The sums now being committed make that exposure harder to ignore. OpenAI's Stargate was announced as $500B over four years, later framed as a $500B, 10GW program. Meta announced a $10B, 4 million square foot AI campus in Louisiana. TSMC put its US investment at $165B.

One campus

When we started planning gigawatt-scale AI capacity in Israel, we first considered a single campus. It made the plan easier to explain and the GPUs easier to connect. It also concentrated the risk.

Then we asked one security question. How much of the country's AI capacity are we putting inside one failure domain?

That question killed the plan. In Israel, drones, rockets, grid failures, and war-risk insurance change where we can build and what we can insure.

The UAE made it concrete. On March 1, 2026, Iranian drones struck two Amazon AWS data centers in Dubai directly, two of three availability zones in the ME-CENTRAL-1 region went down, 109 services failed, and banks, payments platforms, and ride-hailing apps went dark. The buildings had every standard protection: fences, cameras, redundant power, backup generators. None of it was rated for a drone strike. AWS told customers to migrate workloads to another region and warned that recovery would take weeks, given the physical damage involved. BBC, Data Center Dynamics.

Russia's war in Ukraine pushed the same move at larger scale: a smaller target per site, spread across real failure domains.

Power in, heat out

Strip the marketing away and a data center is a machine for two jobs: push power into racks, pull heat out of them, safely, for years. Most of the cost and most of the risk lives in that plumbing, not in the walls.

Power comes in high and is stepped down stage by stage: the grid delivers 138 to 345 kV, an on-site substation steps it to medium voltage, transformers near the hall drop it to about 415 V, the rack takes it lower, and the chip runs at under a volt. Voltage stays high as long as possible because power lost as heat in a wire rises with the square of the current; higher voltage means lower current, less heat, and less copper. One equation shapes the whole building.

The power chain, outside in: grid, substation, transformer, rack, chip. Each stage steps voltage down and current up. The whole design exists to deliver a lot of power to a tiny chip without melting the wires.

The big step-down transformers run 50 to 100 MVA each; they are custom-built, and lead times can pass a year because they depend on a special grain-oriented steel with few suppliers. So operators buy redundancy and name it precisely: N is exactly enough gear to run, N+1 is one spare, 2N is a full second set, and the Uptime Institute turns this into Tiers. Inside, the building is already broken into pieces on purpose. A hall is split into pods of roughly 1.6 to 2.5 MW, each with its own transformers, switchgear, and a diesel generator the size of a locomotive engine, with UPS batteries carrying the load for the sixty or so seconds it takes the generator to start. Engineers will tell you an electrical fault has a smaller "blast radius" than a cooling failure. Blast radius is shop language, borrowed here. Even one building is built around dividing risk.

Every watt you put into a GPU comes back out as about a watt of heat; a 100 MW hall is a 100 MW heater, and cooling is half the machine. Efficiency gets one number, PUE, total facility power divided by IT power: the industry average is around 1.6, the best hyperscalers run near 1.1. Water gets its own number, WUE, and a mid-size site can drink hundreds of millions of liters a year if it leans on evaporative cooling, which dry air makes efficient, and which is one consideration when choosing a desert site. As racks get denser, air alone stops working, and direct-to-chip liquid cooling becomes mandatory.

The heat path: chip, cold plate, coolant loop, chiller, tower. Power becomes heat at the chip, and the building's job is to move that heat outside, loop after loop, without ever letting the silicon get too hot.

The short wire

The case for concentration is real. Training a large model is synchronous: tens of thousands of GPUs work in lockstep and wait on each other, so the links between them have to be fast and short, and high-speed copper only reaches a couple of meters. The in-rack NVLink fabric is far faster than the network between racks. The cheaper your tokens need to be, the tighter you pack the GPUs.

That is why rack power keeps climbing. Racks sat under 10 kW for years; Nvidia's GB200 NVL72 packs 72 GPUs into one liquid-cooled rack at 120 kW and up, and the 2027 roadmap, built on a new 800 VDC power architecture, points at 600 kW to a megawatt per rack. Deliver 600 kW at the old 54 V and you need around 11,000 amps; do it at 800 V and you need about 750. Putting GPUs close together helps training. It also puts more of the investment within reach of the same failure.

Power per rack is rising as more GPUs share shorter, faster connections.

Concentration does not stop at the fence. A training run is synchronized, so its power draw is jagged: Meta's Llama 3 paper describes tens of thousands of GPUs swinging power at the same instant, "on the order of tens of megawatts," enough to stress the grid, and engineers have literally run fake workloads just to keep the swings from hurting the power system. In July 2024, a single fault knocked roughly 1.5 GW of Virginia data centers off the grid at once. At gigawatt scale the failure domain extends to the grid itself; put a gigawatt behind one interconnection and a local hiccup becomes a regional one.

A normal big factory draws a steady line. A synchronized training run jumps between full power and near idle, tens of megawatts in a fraction of a second. At gigawatt scale, a swing in one place can shove the whole grid.

Failure domains

A failure domain is a part of a system that can fail on its own without taking the rest down. Google Cloud uses exactly this language: a zone or region is a failure domain, and spreading across them improves availability. Google's own example is the arithmetic of the idea: two things at 99.9% availability, placed in separate failure domains, drop the odds of both failing at once toward one in a million, if they are independent.

Other fields use related measures to plan for failures.

FieldWhat they already call it
CloudFailure domains, zones, regions
InsuranceProbable maximum loss
Power gridsN-1 contingency planning
DefenseMission assurance
Critical infrastructureConsequence-based risk (CISA)
Data centersUptime Tier III and Tier IV

How much value, power, and national capability can be lost in one local event, and how fast can the rest keep running? Take a $40B AI platform. One site holds the whole investment; four equally sized sites each hold a quarter; twelve 100 MW modules each hold about a twelfth. Separate sites can limit what one local event takes down, provided they do not depend on the same vulnerable systems. The figures below assume that independence.

ArchitectureFailure domainsValue per domainConcentration
One giant site1$40B1.00
Four sites4$10B0.25
Twelve modules12$3.3B0.08

Eleven of twelve

With independent sites, a local failure leaves the rest of the capacity running. The long-run average loss can stay the same: if every site has the same independent failure probability, the expected loss looks similar whether you have one site or four. I am asking a different question: how much capacity survives a single event, and how long recovery takes.

This example assumes the event affects one site and the modules have independent supporting systems. Under those conditions, eleven of the twelve modules remain available.

Distribution works only when the sites can fail independently. A shared substation, fiber route, cooling dependency, control plane, or security gap turns several buildings back into one failure domain; the Virginia event above is exactly that, many buildings and one failure domain. Distance by itself does not create independence.

Proposals for tightly clustered data centers in orbit raise the same concern. I wrote about them in They're Made Out of Compute.

The price of distance

Distribution is more expensive. You duplicate substations, cooling plants, security, and staff, lose economies of scale, and pay a networking penalty when sites sit far apart. Whether that price is worth paying depends on the risks at those sites.

I model single-event disruption using the measures insurers, cloud architects, and grid planners already use: failure domains, probable maximum loss, and continuity of operations. The extra cost of independent sites has to be weighed against the capacity we could lose and the time it would take to recover.

A small country

Israel is a small country with real advantages for this: sun, talent, urgency, and a genuine need for its own compute. If AI compute becomes critical to defense, the economy, science, and government, then data centers are national infrastructure, and national infrastructure should not hang on one building. Small countries do not get to treat infrastructure risk as theoretical.

Gigawatt-scale AI is the goal. We build it as a network of smaller factories: modular, hardened, connected, spread across real failure domains. A GPU hall wired to a few hundred megawatts, replicated across sites that can each fail independently.

In April 2024 I stopped my bike near home and filmed interceptors and incoming fire lighting up the sky over Israel. I have thought about it ever since. When I look at a site plan now, I ask what would keep running if that site were hit.



Sources

  • AWS ME-CENTRAL-1 UAE drone strikes, March 1-3, 2026: two facilities directly struck, two of three availability zones offline, 109+ services degraded, recovery expected to take weeks. BBC, Data Center Dynamics.
  • OpenAI Stargate, announced at $500B over four years and later framed as a $500B, 10GW program. OpenAI, OpenAI.
  • Meta's $10B, 4 million square foot AI data center in Richland Parish, Louisiana. Opportunity Louisiana.
  • TSMC's planned US investment of up to $165B. TSMC.
  • Nvidia GB200 NVL72: 72 GPUs in one liquid-cooled rack, 120 kW and up, single NVLink domain. Nvidia.
  • Nvidia 800 VDC architecture for 1 MW racks and beyond, starting 2027. Nvidia.
  • Synchronized training power swings on the order of tens of megawatts. Meta, The Llama 3 Herd of Models.
  • How a data center's power and cooling actually work (power chain, PUE, WUE, pods, generators, redundancy). SemiAnalysis, Datacenter Anatomy Part 1 and Part 2.
  • Grid cascade risk from large synchronized loads, and the July 2024 Virginia event where ~1.5 GW of data centers disconnected at once. SemiAnalysis.
  • April 28, 2025 Iberian blackout: about 2,200 MW lost in seconds, full collapse in under half a minute. ENTSO-E final report.
  • Failure domains, zones, and regions as the unit of reliability. Google Cloud.
  • Uptime Institute Tier classification (Tier III concurrently maintainable, Tier IV fault tolerant). Uptime Institute.
  • Critical infrastructure framed by the consequence of disruption. CISA.