
The workload
When Cloudflare's network failed to route traffic on 18 November 2025, the immediate workload was diagnosis under pressure: engineers first suspected, then ruled out, a hyper-scale DDoS attack before finding the real cause. Cloudflare's own postmortem, written by CEO Matthew Prince and published the same day, states the company's practice directly: publish an in-depth recount of what happened and what systems and processes failed, naming times to the minute rather than describing the incident only in general terms. That is verified-tier in the sense this site's research uses the term: Cloudflare reports on itself, but with enough independently checkable specificity, exact timestamps, a named cause, a described fix, to be checked against its own internal consistency and customer-observed downtime.
What the documents show
Verified: the postmortem states the outage began at 11:20 UTC on 18 November 2025, when a change to a database permissions system caused a feature file used by Cloudflare's Bot Management system to double in size, exceeding a hard limit in the software reading it and causing that software to fail network-wide. The post states the file was regenerated every five minutes by a query running across a database cluster being gradually migrated, so a good or bad version could be produced at random, which the post says is why engineers' first hypothesis, an attack, was wrong. Verified: the post states core traffic was largely normal by 14:30, and that as of 17:06 all systems were functioning normally, a span of roughly five hours forty-six minutes from detection to resolution. Cloudflare's own status page, as retrieved 16 September 2026, shows the same real-time incident-and-resolution format the postmortem describes, applied to whatever is currently active, evidence of a standing practice rather than of that specific incident.
The operating cost
The postmortem states no dollar cost for the outage, only the technical and time cost above. Estimated: for any customer whose site was unreachable during that window, the cost was whatever revenue or work it would ordinarily process in five to six hours, a figure this entry cannot calculate without that customer's own traffic pattern, and Cloudflare's post does not attempt to either.
The stop condition
The postmortem states its stop condition for the immediate incident precisely: resolution when the bad file's propagation was stopped and an earlier, known-good version restored, timestamped 14:30 UTC, normal operation confirmed at 17:06 UTC. Verified: a longer-term stop condition is not yet reached; the post calls itself the beginning, not the end, of changes meant to prevent a repeat, without naming a date by which that work concludes.
- Does an incident report name exact detection and resolution times, or only approximate ones?
- Was an initial hypothesis, later shown wrong, disclosed in the final report or quietly dropped?
- Does a company's own status page retain enough history to check a year-old postmortem against it directly?
Cloudflare's account is specific enough to be checked, the standard this site treats as the difference between a postmortem and a general incident notice. It remains the company's own account, and this entry does not extend it into a broader claim about Cloudflare's overall reliability.
Sources & reading trail
Cloudflare's own detailed, timestamped account of detection, root cause and remediation for the 18 November 2025 outage.
Source published: 18 November 2025 · Retrieved: 16 September 2026
Vendor's own live status-page format for real-time incident and resolution reporting, as currently maintained.
Source published: Not established · Retrieved: 16 September 2026
Vendor documentation, regulator records and founder-published documents establish the entry; the workload reading and the stop condition are Solo Product Office editorial analysis. This retrospective draft does not imply the site published on the event date.