Title: When the Lights Go Out: Internet Resilience During the 2025 Iberian Power Outage
Authors: Oliver Hohlfeld (University of Kassel), Yasin Alhamwy (University of Kassel), Tobias Striffler (DE-CIX), Daniel Wagner (DE-CIX / MPI-INF), Matthias Wichtlhuber (DE-CIX)
Scribe: Mengrui Zhang (Xiamen University)
Introduction
At 10:33 UTC on April 28, 2025, a nationwide power outage affected Spain, Portugal, and part of France. In most regions the electrical outage lasted about ten hours, and Spain’s grid was fully restored after approximately 23 hours. The event disrupted transport, retail, public services, and end-user connectivity. Unlike outages caused by censorship, war, or a single natural disaster, this event was a multi-country cascading grid failure affecting both access networks and core infrastructure.
The paper presents a systematic measurement study using four independent views: traffic exchanged at four major Iberian Internet Exchange Points (IXPs), BGP control-plane observations, active IP-liveness probes, and the PSKReporter amateur-radio platform. The measurements reveal an immediate 33–70% traffic reduction at the four IXPs, while traffic traversing Spain and Portugal remained comparatively stable. A second, partially unreported data-center power incident later caused another traffic drop and had a larger impact on BGP stability than the nationwide outage itself.
Key idea and contribution
IXP traffic – traffic changes. The authors combine aggregated traffic levels from four major Iberian IXPs with sampled flow data from three of them. Traffic dropped almost immediately by 55%, 33%, 70%, and 40% at the four IXPs. At 14:00 UTC, 3.5 hours into the outage, the reductions relative to the reference day were 50%, 35.4%, 49%, and 44%. Traffic did not vanish completely; it fell toward nighttime-typical levels, indicating that core infrastructure and some data-center-hosted services remained operational. A /24-level comparison separates local traffic, whose endpoints are both in Spain or Portugal, from traversing traffic. Local traffic shows a pronounced reduction and a stall in newly observed prefixes, while traversing traffic is less affected. Application classification reinforces this distinction: Enterprise and Gaming traffic drops by more than 80%, whereas CDN, Video on Demand, and International ISP traffic drops by about 40%.
BGP – BGP stability. The study combines IXP route-server sessions with globally distributed RIPE RIS observations. The nationwide outage caused a roughly 10% decline in route-server sessions within 4.5 hours, indicating a measurable but limited control-plane effect. During recovery, a partial outage in an unnamed core data center began around 15:30 UTC; another approximately 24% of route-server sessions disappeared within 2.5 hours. Diesel generators had been running after the nationwide outage, but a subset failed, leaving several floors without power. This second incident produced sustained increases in BGP announcements and withdrawals visible at global collectors. The affected IXP’s route-server monitoring had to pause for 24 hours after its disk filled with unexpectedly high BGP traffic. The delayed data-center event therefore had a larger effect on routing stability than the nationwide outage itself.
ICMP probes – host availability. The authors use ZMap to probe all IP prefixes geolocated to Spain or Portugal every 20 minutes at a low 200 Mbps scan rate. The fraction of responding hosts is approximately 35% in Portugal and 23% in Spain, with Portugal reaching a stable recovery plateau sooner. Router interfaces identified from RIPE Atlas traceroutes provide a separate view of core infrastructure: responsive router interfaces increase from 7.7% at 18:00 UTC during the outage to approximately 18.4% after full grid recovery. These measurements show that host and router availability recover on different timelines and that the outage affects core infrastructure as well as end-user access.
PSKReporter – beyond the Internet. PSKReporter adds an observation channel that can separate electrical power from Internet connectivity. A station can transmit an amateur-radio spot with electricity alone, while uploading a received spot to PSKReporter requires both electricity and Internet access. Comparing sent and received spots shows that active stations fall by 69.6% by 18:00 UTC but never reach zero; some stations recover power quickly, and a few remain active throughout. This auxiliary dataset confirms that power recovery and Internet recovery do not follow exactly the same timeline. Together, the four perspectives show a layered failure: the initial outage primarily removes powered end-user traffic, while the later data-center incident affects interconnection and routing beyond the immediate geographic footprint.
Q&A
Q1: What should the networking community take away from this outage beyond the need for more resilient data centers?
A1: Our main takeaway is that networks should not depend on a single data center. We need tested backup generators and redundancy across multiple data centers or locations, so that a system can continue operating from a backup site when one facility loses power. End users have limited options during a long outage, but network operators can improve resilience through geographic redundancy and verified backup power.
Q2: Which cloud-scale reliability mechanisms should be redesigned to prevent the escalation observed here?
A2: We cannot identify one specific CDN or operational mechanism that would solve the problem. If a data center loses power across the equipment needed by a service, changing the behavior of that service alone cannot restore the affected infrastructure. Our practical recommendation is to test backup generators and maintain redundant locations; the partial data-center failure shows why redundancy must include the facilities that host core interconnection systems.
Q3: What did the operator observations add to the measurements?
A3: We contacted operators in the Iberian Peninsula after observing unusual BGP activity and path changes. Their reports helped us identify the partial data-center power loss, failed generators, interrupted transport circuits, and lost interconnection with European exchange points. We used these reports as qualitative context rather than as a statistically representative survey, and then checked the reported event against the BGP and IXP measurements.
Personal thoughts
The paper’s strongest feature is its layered measurement design. No single dataset can distinguish powered hosts from reachable hosts, local user traffic from transit traffic, or a nationwide event from a later data-center failure. IXP traffic, BGP sessions, active probing, and amateur-radio reports provide complementary evidence that turns an exceptional outage into a structured resilience study.
The central engineering lesson is that backup power is necessary but not sufficient. The first outage mostly removed end-user activity while leaving much of the core operational; the later partial data-center outage created disproportionate BGP instability and cross-border effects.





