Skip to content

Advertisement

DevOps Society
CloudAnalysis

Cloudflare Logged Four Incidents on 30 September. Uptime Monitors Showed 100%.

Four Cloudflare incidents in one day, one lasting 43 hours, and an outside monitor still showed 100% uptime. Headline uptime is the wrong number to report.

Published your local timeupdated

Cloudflare Logged Four Incidents on 30 September. Uptime Monitors Showed 100%.

On 30 September, Cloudflare opened four incidents on its status page. One covered network performance in Madrid and stayed open for a little over 43 hours. Another covered account permission changes that took time to reach some services, open for about four and a half hours. A third covered congestion in Los Angeles that touched R2, Durable Objects, Vectorize, Hyperdrive, Stream and Cloudchamber for about five and a half hours. The fourth covered delayed Workers builds and ran for almost 30 hours.

UptimeRobot, which watches Cloudflare from the outside, wrote up the day. Its checks hit three Cloudflare endpoints every 60 seconds from Europe and North America. Over 24 hours, seven days and 30 days, those checks recorded 100% uptime.

Both things are true at once. That is the problem worth talking about.

Why green and broken can coexist

A probe in Frankfurt or Virginia asking the edge for a response gets a clean answer. A customer in Madrid on a congested path does not. A developer waiting on a Workers build does not care whether a DNS endpoint answers in 20 milliseconds. UptimeRobot's own write-up makes the same point: this is what a partial, component-level outage usually looks like.

Large providers rarely go fully dark. They fail in slices. One city, one product, one control plane function such as permissions. A global uptime figure averages those slices into a big denominator, and the slice that hurt your customers disappears.

The status updates did not help much either. According to UptimeRobot's summary, the nine updates across the four incidents named regions and products but gave no root cause. That is normal for a live status page. It is still not enough for you to judge whether your own users were in the blast radius.

What the arithmetic hides

Here is an illustrative example. Say 5% of your traffic lands on one city's data centre, and that city is degraded for three and a half hours. That is 210 minutes. A 99.9% monthly target allows about 43 minutes of downtime in a 30 day month (43,200 minutes times 0.001). For users in that city, you have used roughly five months of budget in one afternoon.

Now blend it globally. Weight 210 minutes by 5% of traffic and you get about 10.5 minutes of equivalent downtime. The monthly number still reads as comfortably above 99.9%. The board sees green. The sales team in Spain hears from angry customers.

Neither number is wrong. They answer different questions. The global figure tells you about the provider. The regional one tells you about your customers.

What to measure on your side

The Google SRE book draws a line that helps here. A service level indicator (SLI) is the thing you measure, such as the share of requests that succeed. A service level objective (SLO) is the target you hold it to. The book also warns that server-side numbers can miss what users actually feel, and that measuring from the client side often matters more.

For most teams that means three changes.

First, measure success rate from your own side of the dependency. Count the requests your application sent to the CDN, the storage bucket or the queue, and how many came back good and fast enough. Your logs know when R2 calls slowed in Los Angeles even if the vendor's global number does not.

Second, cut that success rate by region and by dependency. One SLO per critical vendor function (object storage reads, edge requests by region, build pipeline time) shows you the slice. A single blended SLO cannot.

Third, keep the outside checks, but treat them as a smoke alarm, not a scorecard. A synthetic probe from two continents tells you whether the whole thing is on fire. It says little about a partial failure in a third.

What to ask vendors for

Read the actual SLA. Cloudflare's published Business plan SLA promises 100% uptime, but credits are weighted by an affected customer ratio, the share of unique visitors hit by the outage. Customers must tell support within five business days and file a claim by the end of the following month with descriptions, duration and affected URLs. Enterprise customers have separate terms. In practice, if you cannot show your own per-region evidence, a partial incident is hard to claim for.

When you negotiate, ask for three things. Incident notices that state which regions, products and functions were affected and when impact started and ended, rather than when the ticket was opened and closed. A written post-incident report within an agreed number of days for anything above a set severity, including root cause. And SLA terms measured per region or per product, so a long failure in one city is not diluted by the rest of the network.

Also subscribe to the vendor's status feed by component, not just the top-level page, and pipe it into the same channel your on-call team watches.

The honest caveat

Incident duration on a status page is not the same as impact duration. The Madrid incident was identified within an hour, and a fix was in place and being monitored by 14:06 UTC, about three and a half hours after it opened. The 43 hours mostly reflects how long Cloudflare kept watching before closing it. Likewise, a 30 hour Workers Builds incident may have meant slow builds for some customers, not a full stop for all.

UptimeRobot also only checks a few endpoints from two regions. It never claimed to measure every product in every city. The lesson is not that outside monitoring is useless or that the vendor hid anything. It is that nobody outside your stack measures your users for you.

If your traffic is concentrated in one region and you use only the core CDN, a global figure may be close enough. Most teams with customers in several countries and a handful of edge products are not in that position.

What to do this quarter

Pick your three most important external dependencies and write down, for each, the success rate your application saw last month by region. If you cannot produce that number, that is the first piece of work.

Change the reliability slide for the board. Show user-facing success rate for your key journeys, the worst region for the month, and the minutes of budget each critical dependency consumed. A single uptime percentage invites false comfort.

Before your next vendor renewal, ask for sample post-incident reports and the regional breakdown of their SLA. How a vendor answers tells you how the next bad afternoon will go.

Finally, run a quick exercise with your on-call lead. Take 30 September, pretend your users sat in Madrid and Los Angeles, and ask how long it would have taken you to notice. The answer is your real monitoring gap.

Sources: UptimeRobot, "Is Cloudflare down? 30 September 2026" incident summary; IsDown, "Network performance issues in Madrid" Cloudflare incident timeline, 30 September 2026; Cloudflare status page incident history; Cloudflare Business SLA; Google, Site Reliability Engineering book, chapter on service level objectives. Checked on 6 October 2026.

Advertisement

Follow DevOps Society on LinkedIn

Practical infrastructure engineering in your feed.

Follow

Written by

DevOpsSociety Editorial Team

Editorial Team

The DevOpsSociety Editorial Team covers DevOps, cloud infrastructure, Kubernetes, AI infrastructure, platform engineering, cybersecurity, FinOps, and modern engineering practices. We publish practical insights, technical guides, architecture analysis, and research for engineers and technology leaders.

More from DevOpsSociety →
The Infrastructure Briefing

Get the infrastructure briefing.

Practical DevOps, cloud, AI infrastructure and engineering insights, delivered weekly. Read by engineers and engineering leaders.

No spam. Unsubscribe anytime.

Cloudflare Logged Four Incidents on 30 September. Uptime Monitors Showed 100%., DevOps Society