Tag: Incident
Huawei Cloud Incident: Was IAM Involved?

At 02:49 GMT+8 on July 26, 2026, Huawei Cloud said that it had detected abnormalities affecting some accounts on its International Site, disrupting access to related services. Huawei Cloud later marked the notice as resolved, saying that the affected …
At 02:49 GMT+8 on July 26, 2026, Huawei Cloud said that it had detected abnormalities affecting some accounts on its International Site, disrupting access to related services. Huawei Cloud later marked the notice as resolved, saying that the affected …
When AI Gets the Power to Gridlock a City

On the night of March 31, a large number of Apollo Go robotaxis in Wuhan failed at the same time. The real concern is not merely that autonomous driving failed, but that a centrally controlled, cloud-based architecture may amplify a single-vehicle …
On the night of March 31, a large number of Apollo Go robotaxis in Wuhan failed at the same time. The real concern is not merely that autonomous driving failed, but that a centrally controlled, cloud-based architecture may amplify a single-vehicle …
Drones Took Out Three AWS AZ: Into the Era of Bombable Data Centers

On March 1, 2026, Iranian drones reportedly hit AWS facilities in the UAE and Bahrain. If the public reporting is accurate, this may be the first time a hyperscale cloud provider has suffered direct military damage to physical data-center …
On March 1, 2026, Iranian drones reportedly hit AWS facilities in the UAE and Bahrain. If the public reporting is accurate, this may be the first time a hyperscale cloud provider has suffered direct military damage to physical data-center …
Claude's Global Outage: Missiles or a Success Tax?

Claude went down globally on March 2, 2026. The dramatic narrative that immediately formed was irresistible: AWS facilities in the Middle East had just been hit by drones, and now Claude was collapsing too. The story spread fast. But it was probably …
Claude went down globally on March 2, 2026. The dramatic narrative that immediately formed was irresistible: AWS facilities in the Middle East had just been hit by drones, and now Claude was collapsing too. The story spread fast. But it was probably …
Alipay, Taobao, Xianyu Went Dark. Smells Like a Message Queue Meltdown.

Taobao, Alipay, and Xianyu simultaneously faceplanted on Dec 4. Cards were debited, yet orders sat frozen at “pending.” The story rhymes perfectly with the 2024 Double-11 outage, so the leading suspect is once again the message queue or whatever …
Taobao, Alipay, and Xianyu simultaneously faceplanted on Dec 4. Cards were debited, yet orders sat frozen at “pending.” The story rhymes perfectly with the 2024 Double-11 outage, so the leading suspect is once again the message queue or whatever …
Don't Run Docker Postgres for Production!

Back in 2019, I wrote about “Is running postgres in docker a good idea?” — Don’t run PostgreSQL in containers for production, because you’ll likely hit a pile of issues that simply don’t exist on bare metal or VMs. Well, users of Docker’s “official” …
Back in 2019, I wrote about “Is running postgres in docker a good idea?” — Don’t run PostgreSQL in containers for production, because you’ll likely hit a pile of issues that simply don’t exist on bare metal or VMs. Well, users of Docker’s “official” …
Cloudflare’s Nov 18 Outage, Translated and Dissected

Yesterday the “cyber Bodhisattva” Cloudflare suffered its worst incident since 2019. For six hours, core network traffic couldn’t be delivered. ChatGPT, X, Spotify, Uber—everyone felt it. Root cause: a permission change on ClickHouse made the …
Yesterday the “cyber Bodhisattva” Cloudflare suffered its worst incident since 2019. For six hours, core network traffic couldn’t be delivered. ChatGPT, X, Spotify, Uber—everyone felt it. Root cause: a permission change on ClickHouse made the …
AWS’s Official DynamoDB Outage Postmortem

AWS just released the official postmortem for the Oct 20 us-east-1 meltdown. It’s one of the rare times we get first-hand detail, so I translated it to Chinese and sprinkled in commentary. Here’s the English recap with my notes. Incident page: …
AWS just released the official postmortem for the Oct 20 us-east-1 meltdown. It’s one of the rare times we get first-hand detail, so I translated it to Chinese and sprinkled in commentary. Here’s the English recap with my notes. Incident page: …
How One AWS DNS Failure Cascaded Across Half the Internet

On Oct 20, 2025, AWS’s crown jewel region us-east-1 spent fifteen hours flailing. More than a thousand companies went dark worldwide. The root cause? An internal DNS entry that stopped resolving. From the moment DNS broke, DynamoDB, EC2, Lambda, and …
On Oct 20, 2025, AWS’s crown jewel region us-east-1 spent fifteen hours flailing. More than a thousand companies went dark worldwide. The root cause? An internal DNS entry that stopped resolving. From the moment DNS broke, DynamoDB, EC2, Lambda, and …
How Many Shops Has etcd Torched?

A few days ago Yingshi Hurricane shared their Pigsty/PostgreSQL HA incident. The root cause? etcd hit its default 2 GB limit because auto-compaction wasn’t enabled. As @ayanamist put it on X: “Let’s see how many companies this stupid 2 GB design can …
A few days ago Yingshi Hurricane shared their Pigsty/PostgreSQL HA incident. The root cause? etcd hit its default 2 GB limit because auto-compaction wasn’t enabled. As @ayanamist put it on X: “Let’s see how many companies this stupid 2 GB design can …
OpenAI Global Outage Postmortem: K8S Circular Dependencies

On December 11th, OpenAI experienced a global service outage affecting ChatGPT, API, Sora, Playground, and Labs. The outage lasted from 3:16 PM to 7:38 PM PT, spanning over four hours with significant impact. According to OpenAI’s incident report …
On December 11th, OpenAI experienced a global service outage affecting ChatGPT, API, Sora, Playground, and Labs. The outage lasted from 3:16 PM to 7:38 PM PT, spanning over four hours with significant impact. According to OpenAI’s incident report …
Alibaba-Cloud: High Availability Disaster Recovery Myth Shattered

On September 10, 2024, Alibaba-Cloud’s Singapore Availability Zone C data center experienced a fire caused by lithium battery explosion. It’s been a week now and services have not been fully restored yet. According to the monthly SLA availability …
On September 10, 2024, Alibaba-Cloud’s Singapore Availability Zone C data center experienced a fire caused by lithium battery explosion. It’s been a week now and services have not been fully restored yet. According to the monthly SLA availability …
What Can We Learn from NetEase Cloud Music's Outage?

This afternoon around 14:44, NetEase Cloud Music experienced an outage, recovering at 17:11. The rumored cause was infrastructure/cloud/ disk storage related issues. Incident Timeline During the outage, NetEase Cloud Music clients could normally play …
This afternoon around 14:44, NetEase Cloud Music experienced an outage, recovering at 17:11. The rumored cause was infrastructure/cloud/ disk storage related issues. Incident Timeline During the outage, NetEase Cloud Music clients could normally play …
Blue Screen Friday: Amateur Hour on Both Sides

Recently, due to a configuration update released by cybersecurity company CrowdStrike, countless Windows computers worldwide fell into blue screen death, causing endless chaos — airlines grounded flights, hospitals canceled surgeries, supermarkets, …
Recently, due to a configuration update released by cybersecurity company CrowdStrike, countless Windows computers worldwide fell into blue screen death, causing endless chaos — airlines grounded flights, hospitals canceled surgeries, supermarkets, …
CVE-2024-6387 SSH Vulnerability Fix

Vulnerability description, CVE-2024-6387: https://nvd.nist.gov/vuln/detail/CVE-2024-6387 This basically affects newer versions of operating systems. Older systems like CentOS 7.9, RockyLinux 8.9, Ubuntu 20.04, Debian 11 escaped this due to older …
Vulnerability description, CVE-2024-6387: https://nvd.nist.gov/vuln/detail/CVE-2024-6387 This basically affects newer versions of operating systems. Older systems like CentOS 7.9, RockyLinux 8.9, Ubuntu 20.04, Debian 11 escaped this due to older …
Database Deletion Supreme - Google Cloud Nuked a Major Fund's Entire Cloud Account

Due to an “unprecedented configuration error”, Google Cloud mistakenly deleted UniSuper’s cloud account. The Australian pension fund executive and Google Cloud’s global CEO issued a joint statement apologizing for this “extremely frustrating and …
Due to an “unprecedented configuration error”, Google Cloud mistakenly deleted UniSuper’s cloud account. The Australian pension fund executive and Google Cloud’s global CEO issued a joint statement apologizing for this “extremely frustrating and …
What Can We Learn from Tencent Cloud's Major Outage?

Eight days after the outage, Tencent Cloud published a postmortem report for the April 8th major outage. I think this is a good thing, because Alibaba-Cloud’s Double 11 major outage official postmortem is still overdue. If public cloud vendors want …
Eight days after the outage, Tencent Cloud published a postmortem report for the April 8th major outage. I think this is a good thing, because Alibaba-Cloud’s Double 11 major outage official postmortem is still overdue. If public cloud vendors want …
From Cost-Reduction Jokes to Real Cost Reduction and Efficiency

Year-end is performance rush time, but internet giants are having major incidents one after another. They’ve turned “cost reduction and efficiency improvement” into literal “cost reduction jokes” — this is no longer just a meme, but official …
Year-end is performance rush time, but internet giants are having major incidents one after another. They’ve turned “cost reduction and efficiency improvement” into literal “cost reduction jokes” — this is no longer just a meme, but official …
What Can We Learn from Alibaba-Cloud's Global Outage?

A year after the last major incident, Alibaba-Cloud suffered another massive outage, creating an unprecedented record in the cloud computing industry — simultaneous failures across all global regions and all services. Since Alibaba-Cloud refuses to …
A year after the last major incident, Alibaba-Cloud suffered another massive outage, creating an unprecedented record in the cloud computing industry — simultaneous failures across all global regions and all services. Since Alibaba-Cloud refuses to …
How to Use pg_filedump for Data Recovery?

Backups are a DBA’s lifeline — but what if your PostgreSQL database has already exploded and you have no backups? Maybe pg_filedump can help you! Recently encountered a rather outrageous case. The situation was this: a user’s PostgreSQL database was …
Backups are a DBA’s lifeline — but what if your PostgreSQL database has already exploded and you have no backups? Maybe pg_filedump can help you! Recently encountered a rather outrageous case. The situation was this: a user’s PostgreSQL database was …
Incident-Report: Patroni Failure Due to Time Travel

Summary: Machine restarted due to failure, NTP service corrected PG time after PG startup, causing Patroni to fail to start. The failure information in Patroni is shown as follows: Process %s is not postmaster, too much difference between PID file …
Summary: Machine restarted due to failure, NTP service corrected PG time after PG startup, causing Patroni to fail to start. The failure information in Patroni is shown as follows: Process %s is not postmaster, too much difference between PID file …
Incident: PostgreSQL Extension Installation Causes Connection Failure

Author: Vonng (@Vonng) Today encountered an interesting case where a customer reported database connection issues. The error was: psql: FATAL: could not load library "/export/servers/pgsql/lib/pg_hint_plan.so": …
Author: Vonng (@Vonng) Today encountered an interesting case where a customer reported database connection issues. The error was: psql: FATAL: could not load library "/export/servers/pgsql/lib/pg_hint_plan.so": …
Incident-Report: Connection-Pool Contamination Caused by pg_dump

PostgreSQL is great, but that doesn’t mean it’s Bug-Free. This time in the production environment, I encountered another very interesting case: a production incident caused by pg_dump. This is a very subtle bug triggered by Pgbouncer, search_path, …
PostgreSQL is great, but that doesn’t mean it’s Bug-Free. This time in the production environment, I encountered another very interesting case: a production incident caused by pg_dump. This is a very subtle bug triggered by Pgbouncer, search_path, …
PostgreSQL Data Page Corruption Repair

PostgreSQL is a very reliable database, but even the most reliable database will struggle when faced with unreliable hardware. This article introduces methods for dealing with data page corruption in PostgreSQL. The Initial Problem A statistics …
PostgreSQL is a very reliable database, but even the most reliable database will struggle when faced with unreliable hardware. This article introduces methods for dealing with data page corruption in PostgreSQL. The Initial Problem A statistics …
Incident-Report: PostgreSQL Transaction ID Wraparound

Encountered a transaction wraparound failure caused by disk bad blocks: Primary database (PostgreSQL 9.3) disk bad blocks caused VACUUM FREEZE execution failure on several tables. Unable to reclaim old transaction IDs, causing database transaction …
Encountered a transaction wraparound failure caused by disk bad blocks: Primary database (PostgreSQL 9.3) disk bad blocks caused VACUUM FREEZE execution failure on several tables. Unable to reclaim old transaction IDs, causing database transaction …
Incident-Report: Integer Overflow from Rapid Sequence Number Consumption

0x01 Overview Incident symptoms: A table using auto-increment columns had sequence numbers reach the integer limit, preventing writes. Discovered large gaps in auto-increment columns, with many sequence numbers consumed without corresponding …
0x01 Overview Incident symptoms: A table using auto-increment columns had sequence numbers reach the integer limit, preventing writes. Discovered large gaps in auto-increment columns, with many sequence numbers consumed without corresponding …
Incident-Report: Uneven Load Avalanche

Author: Vonng (@Vonng) Recently there was a perplexing incident where a database had half its data volume and load migrated away. Everything else remained unchanged, and it was fine before. The pressure decreased, yet it fell into a near-death state …
Author: Vonng (@Vonng) Recently there was a perplexing incident where a database had half its data volume and load migrated away. Everything else remained unchanged, and it was fine before. The pressure decreased, yet it fell into a near-death state …