Skip to content

Tag: Incident

  • Huawei Cloud Incident: Was IAM Involved?

    By Ruohang Feng In Cloud-Exit 938 words 5 min

    Ruohang FengCloudIncident

    Huawei Cloud Incident: Was IAM Involved?

    At 02:49 GMT+8 on July 26, 2026, Huawei Cloud said that it had detected abnormalities affecting some accounts on its International Site, disrupting access to related services. Huawei Cloud later marked the notice as resolved, saying that the affected …

    At 02:49 GMT+8 on July 26, 2026, Huawei Cloud said that it had detected abnormalities affecting some accounts on its International Site, disrupting access to related services. Huawei Cloud later marked the notice as resolved, saying that the affected …

  • When AI Gets the Power to Gridlock a City

    By Ruohang Feng In Cloud-Exit 2437 words 5 min

    Ruohang FengAIIncidentSociety

    When AI Gets the Power to Gridlock a City

    On the night of March 31, a large number of Apollo Go robotaxis in Wuhan failed at the same time. The real concern is not merely that autonomous driving failed, but that a centrally controlled, cloud-based architecture may amplify a single-vehicle …

    On the night of March 31, a large number of Apollo Go robotaxis in Wuhan failed at the same time. The real concern is not merely that autonomous driving failed, but that a centrally controlled, cloud-based architecture may amplify a single-vehicle …

  • Drones Took Out Three AWS AZ: Into the Era of Bombable Data Centers

    By Ruohang Feng In Cloud-Exit 580 words 3 min

    Ruohang FengCloud-ExitAWSIncident

    Drones Took Out Three AWS AZ: Into the Era of Bombable Data Centers

    On March 1, 2026, Iranian drones reportedly hit AWS facilities in the UAE and Bahrain. If the public reporting is accurate, this may be the first time a hyperscale cloud provider has suffered direct military damage to physical data-center …

    On March 1, 2026, Iranian drones reportedly hit AWS facilities in the UAE and Bahrain. If the public reporting is accurate, this may be the first time a hyperscale cloud provider has suffered direct military damage to physical data-center …

  • Claude's Global Outage: Missiles or a Success Tax?

    By Ruohang Feng In Cloud-Exit 666 words 4 min

    Ruohang FengAIClaudeIncident

    Claude's Global Outage: Missiles or a Success Tax?

    Claude went down globally on March 2, 2026. The dramatic narrative that immediately formed was irresistible: AWS facilities in the Middle East had just been hit by drones, and now Claude was collapsing too. The story spread fast. But it was probably …

    Claude went down globally on March 2, 2026. The dramatic narrative that immediately formed was irresistible: AWS facilities in the Middle East had just been hit by drones, and now Claude was collapsing too. The story spread fast. But it was probably …

  • Alipay, Taobao, Xianyu Went Dark. Smells Like a Message Queue Meltdown.

    By Ruohang Feng In Cloud-Exit 680 words 4 min

    Ruohang FengCloud-ExitAlibaba CloudIncident

    Alipay, Taobao, Xianyu Went Dark. Smells Like a Message Queue Meltdown.

    Taobao, Alipay, and Xianyu simultaneously faceplanted on Dec 4. Cards were debited, yet orders sat frozen at “pending.” The story rhymes perfectly with the 2024 Double-11 outage, so the leading suspect is once again the message queue or whatever …

    Taobao, Alipay, and Xianyu simultaneously faceplanted on Dec 4. Cards were debited, yet orders sat frozen at “pending.” The story rhymes perfectly with the 2024 Double-11 outage, so the leading suspect is once again the message queue or whatever …

  • Don't Run Docker Postgres for Production!

    By Ruohang Feng In Database 1868 words 9 min

    Ruohang FengPostgreSQLContainersIncident

    Don't Run Docker Postgres for Production!

    Back in 2019, I wrote about “Is running postgres in docker a good idea?” — Don’t run PostgreSQL in containers for production, because you’ll likely hit a pile of issues that simply don’t exist on bare metal or VMs. Well, users of Docker’s “official” …

    Back in 2019, I wrote about “Is running postgres in docker a good idea?” — Don’t run PostgreSQL in containers for production, because you’ll likely hit a pile of issues that simply don’t exist on bare metal or VMs. Well, users of Docker’s “official” …

  • Cloudflare’s Nov 18 Outage, Translated and Dissected

    By Ruohang Feng In Cloud-Exit 1432 words 7 min

    Ruohang FengCloudflareIncident

    Cloudflare’s Nov 18 Outage, Translated and Dissected

    Yesterday the “cyber Bodhisattva” Cloudflare suffered its worst incident since 2019. For six hours, core network traffic couldn’t be delivered. ChatGPT, X, Spotify, Uber—everyone felt it. Root cause: a permission change on ClickHouse made the …

    Yesterday the “cyber Bodhisattva” Cloudflare suffered its worst incident since 2019. For six hours, core network traffic couldn’t be delivered. ChatGPT, X, Spotify, Uber—everyone felt it. Root cause: a permission change on ClickHouse made the …

  • AWS’s Official DynamoDB Outage Postmortem

    By Ruohang Feng In Cloud-Exit 942 words 5 min

    Ruohang FengCloud-ExitAWSIncident

    AWS’s Official DynamoDB Outage Postmortem

    AWS just released the official postmortem for the Oct 20 us-east-1 meltdown. It’s one of the rare times we get first-hand detail, so I translated it to Chinese and sprinkled in commentary. Here’s the English recap with my notes. Incident page: …

    AWS just released the official postmortem for the Oct 20 us-east-1 meltdown. It’s one of the rare times we get first-hand detail, so I translated it to Chinese and sprinkled in commentary. Here’s the English recap with my notes. Incident page: …

  • How One AWS DNS Failure Cascaded Across Half the Internet

    By Ruohang Feng In Cloud-Exit 2354 words 12 min

    Ruohang FengCloud-ExitAWSIncident

    How One AWS DNS Failure Cascaded Across Half the Internet

    On Oct 20, 2025, AWS’s crown jewel region us-east-1 spent fifteen hours flailing. More than a thousand companies went dark worldwide. The root cause? An internal DNS entry that stopped resolving. From the moment DNS broke, DynamoDB, EC2, Lambda, and …

    On Oct 20, 2025, AWS’s crown jewel region us-east-1 spent fifteen hours flailing. More than a thousand companies went dark worldwide. The root cause? An internal DNS entry that stopped resolving. From the moment DNS broke, DynamoDB, EC2, Lambda, and …

  • How Many Shops Has etcd Torched?

    In Database 373 words 2 min

    BlogDatabaseDistributed SystemsContainersIncident

    How Many Shops Has etcd Torched?

    A few days ago Yingshi Hurricane shared their Pigsty/PostgreSQL HA incident. The root cause? etcd hit its default 2 GB limit because auto-compaction wasn’t enabled. As @ayanamist put it on X: “Let’s see how many companies this stupid 2 GB design can …

    A few days ago Yingshi Hurricane shared their Pigsty/PostgreSQL HA incident. The root cause? etcd hit its default 2 GB limit because auto-compaction wasn’t enabled. As @ayanamist put it on X: “Let’s see how many companies this stupid 2 GB design can …

  • OpenAI Global Outage Postmortem: K8S Circular Dependencies

    By Ruohang Feng In Cloud-Exit 1921 words 10 min

    Ruohang FengCodexIncident

    OpenAI Global Outage Postmortem: K8S Circular Dependencies

    On December 11th, OpenAI experienced a global service outage affecting ChatGPT, API, Sora, Playground, and Labs. The outage lasted from 3:16 PM to 7:38 PM PT, spanning over four hours with significant impact. According to OpenAI’s incident report …

    On December 11th, OpenAI experienced a global service outage affecting ChatGPT, API, Sora, Playground, and Labs. The outage lasted from 3:16 PM to 7:38 PM PT, spanning over four hours with significant impact. According to OpenAI’s incident report …

  • Alibaba-Cloud: High Availability Disaster Recovery Myth Shattered

    By Ruohang Feng In Cloud-Exit 2382 words 12 min

    Ruohang FengCloud-ExitAlibaba CloudIncident

    Alibaba-Cloud: High Availability Disaster Recovery Myth Shattered

    On September 10, 2024, Alibaba-Cloud’s Singapore Availability Zone C data center experienced a fire caused by lithium battery explosion. It’s been a week now and services have not been fully restored yet. According to the monthly SLA availability …

    On September 10, 2024, Alibaba-Cloud’s Singapore Availability Zone C data center experienced a fire caused by lithium battery explosion. It’s been a week now and services have not been fully restored yet. According to the monthly SLA availability …

  • What Can We Learn from NetEase Cloud Music's Outage?

    By Ruohang Feng In Cloud-Exit 780 words 4 min

    Ruohang FengIncident

    What Can We Learn from NetEase Cloud Music's Outage?

    This afternoon around 14:44, NetEase Cloud Music experienced an outage, recovering at 17:11. The rumored cause was infrastructure/cloud/ disk storage related issues. Incident Timeline During the outage, NetEase Cloud Music clients could normally play …

    This afternoon around 14:44, NetEase Cloud Music experienced an outage, recovering at 17:11. The rumored cause was infrastructure/cloud/ disk storage related issues. Incident Timeline During the outage, NetEase Cloud Music clients could normally play …

  • Blue Screen Friday: Amateur Hour on Both Sides

    By Ruohang Feng In Cloud-Exit 1125 words 6 min

    Ruohang FengCloud-ExitIncidentCloudflare

    Blue Screen Friday: Amateur Hour on Both Sides

    Recently, due to a configuration update released by cybersecurity company CrowdStrike, countless Windows computers worldwide fell into blue screen death, causing endless chaos — airlines grounded flights, hospitals canceled surgeries, supermarkets, …

    Recently, due to a configuration update released by cybersecurity company CrowdStrike, countless Windows computers worldwide fell into blue screen death, causing endless chaos — airlines grounded flights, hospitals canceled surgeries, supermarkets, …

  • CVE-2024-6387 SSH Vulnerability Fix

    By Ruohang Feng In Database 188 words 1 min

    Ruohang FengSecurityLinuxIncident

    CVE-2024-6387 SSH Vulnerability Fix

    Vulnerability description, CVE-2024-6387: https://nvd.nist.gov/vuln/detail/CVE-2024-6387 This basically affects newer versions of operating systems. Older systems like CentOS 7.9, RockyLinux 8.9, Ubuntu 20.04, Debian 11 escaped this due to older …

    Vulnerability description, CVE-2024-6387: https://nvd.nist.gov/vuln/detail/CVE-2024-6387 This basically affects newer versions of operating systems. Older systems like CentOS 7.9, RockyLinux 8.9, Ubuntu 20.04, Debian 11 escaped this due to older …

  • Database Deletion Supreme - Google Cloud Nuked a Major Fund's Entire Cloud Account

    By Ruohang Feng In Cloud-Exit 533 words 3 min

    Ruohang FengCloud-ExitCloudIncident

    Database Deletion Supreme - Google Cloud Nuked a Major Fund's Entire Cloud Account

    Due to an “unprecedented configuration error”, Google Cloud mistakenly deleted UniSuper’s cloud account. The Australian pension fund executive and Google Cloud’s global CEO issued a joint statement apologizing for this “extremely frustrating and …

    Due to an “unprecedented configuration error”, Google Cloud mistakenly deleted UniSuper’s cloud account. The Australian pension fund executive and Google Cloud’s global CEO issued a joint statement apologizing for this “extremely frustrating and …

  • What Can We Learn from Tencent Cloud's Major Outage?

    By Ruohang Feng In Cloud-Exit 2048 words 10 min

    Ruohang FengCloud-ExitCloudIncident

    What Can We Learn from Tencent Cloud's Major Outage?

    Eight days after the outage, Tencent Cloud published a postmortem report for the April 8th major outage. I think this is a good thing, because Alibaba-Cloud’s Double 11 major outage official postmortem is still overdue. If public cloud vendors want …

    Eight days after the outage, Tencent Cloud published a postmortem report for the April 8th major outage. I think this is a good thing, because Alibaba-Cloud’s Double 11 major outage official postmortem is still overdue. If public cloud vendors want …

  • From Cost-Reduction Jokes to Real Cost Reduction and Efficiency

    By Ruohang Feng In Cloud-Exit 1871 words 9 min

    Ruohang FengCloud-ExitAlibaba CloudIncident

    From Cost-Reduction Jokes to Real Cost Reduction and Efficiency

    Year-end is performance rush time, but internet giants are having major incidents one after another. They’ve turned “cost reduction and efficiency improvement” into literal “cost reduction jokes” — this is no longer just a meme, but official …

    Year-end is performance rush time, but internet giants are having major incidents one after another. They’ve turned “cost reduction and efficiency improvement” into literal “cost reduction jokes” — this is no longer just a meme, but official …

  • What Can We Learn from Alibaba-Cloud's Global Outage?

    By Ruohang Feng In Cloud-Exit 2783 words 14 min

    Ruohang FengCloud-ExitAlibaba CloudIncident

    What Can We Learn from Alibaba-Cloud's Global Outage?

    A year after the last major incident, Alibaba-Cloud suffered another massive outage, creating an unprecedented record in the cloud computing industry — simultaneous failures across all global regions and all services. Since Alibaba-Cloud refuses to …

    A year after the last major incident, Alibaba-Cloud suffered another massive outage, creating an unprecedented record in the cloud computing industry — simultaneous failures across all global regions and all services. Since Alibaba-Cloud refuses to …

  • How to Use pg_filedump for Data Recovery?

    By Ruohang Feng In PGSQL 2801 words 14 min

    Ruohang FengPostgreSQLPG AdminIncident

    How to Use pg_filedump for Data Recovery?

    Backups are a DBA’s lifeline — but what if your PostgreSQL database has already exploded and you have no backups? Maybe pg_filedump can help you! Recently encountered a rather outrageous case. The situation was this: a user’s PostgreSQL database was …

    Backups are a DBA’s lifeline — but what if your PostgreSQL database has already exploded and you have no backups? Maybe pg_filedump can help you! Recently encountered a rather outrageous case. The situation was this: a user’s PostgreSQL database was …

  • Incident-Report: Patroni Failure Due to Time Travel

    By Ruohang Feng In PGSQL 145 words 1 min

    Ruohang FengPostgreSQLPG AdminIncident

    Incident-Report: Patroni Failure Due to Time Travel

    Summary: Machine restarted due to failure, NTP service corrected PG time after PG startup, causing Patroni to fail to start. The failure information in Patroni is shown as follows: Process %s is not postmaster, too much difference between PID file …

    Summary: Machine restarted due to failure, NTP service corrected PG time after PG startup, causing Patroni to fail to start. The failure information in Patroni is shown as follows: Process %s is not postmaster, too much difference between PID file …

  • Incident: PostgreSQL Extension Installation Causes Connection Failure

    By Ruohang Feng In PGSQL 433 words 3 min

    Ruohang FengPostgreSQLPG AdminExtensionIncident

    Incident: PostgreSQL Extension Installation Causes Connection Failure

    Author: Vonng (@Vonng) Today encountered an interesting case where a customer reported database connection issues. The error was: psql: FATAL: could not load library "/export/servers/pgsql/lib/pg_hint_plan.so": …

    Author: Vonng (@Vonng) Today encountered an interesting case where a customer reported database connection issues. The error was: psql: FATAL: could not load library "/export/servers/pgsql/lib/pg_hint_plan.so": …

  • Incident-Report: Connection-Pool Contamination Caused by pg_dump

    By Ruohang Feng In PGSQL 1232 words 6 min

    Ruohang FengPostgreSQLPG AdminIncident

    Incident-Report: Connection-Pool Contamination Caused by pg_dump

    PostgreSQL is great, but that doesn’t mean it’s Bug-Free. This time in the production environment, I encountered another very interesting case: a production incident caused by pg_dump. This is a very subtle bug triggered by Pgbouncer, search_path, …

    PostgreSQL is great, but that doesn’t mean it’s Bug-Free. This time in the production environment, I encountered another very interesting case: a production incident caused by pg_dump. This is a very subtle bug triggered by Pgbouncer, search_path, …

  • PostgreSQL Data Page Corruption Repair

    By Ruohang Feng In PGSQL 2725 words 13 min

    Ruohang FengPostgreSQLPG AdminIncident

    PostgreSQL Data Page Corruption Repair

    PostgreSQL is a very reliable database, but even the most reliable database will struggle when faced with unreliable hardware. This article introduces methods for dealing with data page corruption in PostgreSQL. The Initial Problem A statistics …

    PostgreSQL is a very reliable database, but even the most reliable database will struggle when faced with unreliable hardware. This article introduces methods for dealing with data page corruption in PostgreSQL. The Initial Problem A statistics …

  • Incident-Report: PostgreSQL Transaction ID Wraparound

    By Ruohang Feng In PGSQL 956 words 5 min

    Ruohang FengPostgreSQLPG AdminIncident

    Incident-Report: PostgreSQL Transaction ID Wraparound

    Encountered a transaction wraparound failure caused by disk bad blocks: Primary database (PostgreSQL 9.3) disk bad blocks caused VACUUM FREEZE execution failure on several tables. Unable to reclaim old transaction IDs, causing database transaction …

    Encountered a transaction wraparound failure caused by disk bad blocks: Primary database (PostgreSQL 9.3) disk bad blocks caused VACUUM FREEZE execution failure on several tables. Unable to reclaim old transaction IDs, causing database transaction …

  • Incident-Report: Integer Overflow from Rapid Sequence Number Consumption

    By Ruohang Feng In PGSQL 645 words 4 min

    Ruohang FengPostgreSQLPG AdminIncident

    Incident-Report: Integer Overflow from Rapid Sequence Number Consumption

    0x01 Overview Incident symptoms: A table using auto-increment columns had sequence numbers reach the integer limit, preventing writes. Discovered large gaps in auto-increment columns, with many sequence numbers consumed without corresponding …

    0x01 Overview Incident symptoms: A table using auto-increment columns had sequence numbers reach the integer limit, preventing writes. Discovered large gaps in auto-increment columns, with many sequence numbers consumed without corresponding …

  • Incident-Report: Uneven Load Avalanche

    By Ruohang Feng In PGSQL 1318 words 7 min

    Ruohang FengPostgreSQLPG AdminIncident

    Incident-Report: Uneven Load Avalanche

    Author: Vonng (@Vonng) Recently there was a perplexing incident where a database had half its data volume and load migrated away. Everything else remained unchanged, and it was fine before. The pressure decreased, yet it fell into a near-death state …

    Author: Vonng (@Vonng) Recently there was a perplexing incident where a database had half its data volume and load migrated away. Everything else remained unchanged, and it was fine before. The pressure decreased, yet it fell into a near-death state …