[주의!] 문서의 이전 버전(에 수정)을 보고 있습니다. 최신 버전으로 이동
분류
사건 사고
① 주의. 사건·사고 관련 내용을 설명합니다.
② 실제로 발생한 사건 사고에 관련된 내용을 다룹니다.
1. 개요2. 이유3. 에러 상황4. 피해 상황
4.1. 위키4.2. 커뮤니티4.3. 게임4.4. 정부 기관4.5. 대기업4.6. 기타
5. 대응

1. 개요[편집]

2025년 11월 18일 Cloudflare Global Network 오류로 인하여 전세계의 Cloudflare을 쓰는 웹 및 앱이 사용 불능이 된 사태.

2. 이유[편집]

원본.

파일:클플 오류1.png

2025년 11월 18일 11:20 UTC에(이 블로그의 모든 시간은 UTC 기준), Cloudflare 네트워크에서 핵심 트래픽 전달에 심각한 장애가 발생하기 시작했습니다. 이로 인해 인터넷 사용자가 고객 사이트에 접속하려고 하면, Cloudflare 네트워크 내부에서 문제가 발생했다는 오류 페이지가 표시되었습니다.

파일:클플 오류2.png

이 문제는 사이버 공격이나 악의적 활동과는 직접적·간접적으로도 관련이 없었습니다. 대신, 데이터베이스 시스템 중 하나의 권한 변경으로 인해, Bot Management 시스템에서 사용하는 “feature file”에 여러 항목이 출력되면서 문제가 발생했습니다. 그 결과, 해당 feature file의 크기가 두 배로 늘어났고, 이 커진 파일이 네트워크를 구성하는 모든 머신에 전파되었습니다.

이 머신들에서 트래픽을 라우팅하는 소프트웨어는 Bot Management 시스템을 최신 위협에 맞게 유지하기 위해 이 feature file을 읽습니다. 그런데 이 소프트웨어에는 파일 크기에 제한이 있었고, 두 배로 커진 파일 크기가 그 한계를 초과하면서 소프트웨어가 실패하게 된 것입니다.

처음에는 우리가 관찰한 증상이 대규모 DDoS 공격 때문이라고 잘못 의심했지만, 이후 핵심 원인을 정확히 파악하여 예상보다 커진 feature file의 전파를 중단하고 이전 버전으로 교체할 수 있었습니다. 14:30경에는 핵심 트래픽이 대부분 정상적으로 흐르기 시작했고, 그 후 몇 시간 동안 네트워크 각 부분에 급격히 몰린 트래픽을 완화하는 작업을 진행했습니다. 17:06에는 Cloudflare의 모든 시스템이 정상적으로 작동했습니다.

고객과 인터넷 전반에 끼친 영향에 대해 사과드립니다. Cloudflare가 인터넷 생태계에서 중요한 역할을 하는 만큼, 시스템 장애는 용납될 수 없습니다. 네트워크가 일정 시간 동안 트래픽을 라우팅할 수 없었던 상황은 팀의 모든 구성원에게 깊은 고통이었습니다. 오늘 여러분을 실망시킨 것을 알고 있습니다.

이 글은 정확히 어떤 일이 일어났고, 어떤 시스템과 프로세스가 실패했는지를 자세히 설명하는 내용입니다. 또한 이번과 같은 장애가 다시 발생하지 않도록 하기 위한 계획의 시작점이기도 하지만, 끝은 아닙니다.

장애 상황
아래 차트는 Cloudflare 네트워크에서 발생한 5xx HTTP 상태 코드의 양을 보여줍니다. 정상적으로는 이 수치가 매우 낮아야 하는데, 장애가 시작되기 전까지는 정상적으로 낮은 상태였습니다.

파일:클플 오류3.png

11:20 이전의 수치는 네트워크 전반에서 관찰된 5xx 오류의 정상 기준선입니다. 이후 급증과 변동은 잘못된 feature file을 로드하면서 시스템이 실패한 것을 보여줍니다. 흥미로운 점은 시스템이 잠시 회복되기도 했다는 것으로, 내부 오류에서는 매우 드문 현상이었습니다.

설명하자면, 이 파일은 ClickHouse 데이터베이스 클러스터에서 5분마다 쿼리를 통해 생성되었고, 권한 관리를 개선하기 위해 점진적으로 업데이트되고 있었습니다. 잘못된 데이터는 업데이트된 클러스터 일부에서 쿼리가 실행될 때만 생성되었습니다. 따라서 매 5분마다 정상 또는 잘못된 구성 파일 세트가 생성되어 빠르게 네트워크 전체에 전파될 가능성이 있었던 것입니다.

이런 변동 때문에 정확히 무슨 일이 일어나고 있는지 파악하기 어려웠습니다. 시스템이 때로는 정상 구성 파일을, 때로는 잘못된 구성 파일을 네트워크에 배포하면서 회복했다가 다시 실패했기 때문입니다. 처음에는 공격 때문일 수 있다고 생각했지만, 결국 모든 ClickHouse 노드가 잘못된 구성 파일을 생성하게 되었고, 변동은 실패 상태로 안정화되었습니다.

오류는 근본 원인이 확인되고 해결되기 시작한 14:30까지 계속되었습니다. 문제는 잘못된 feature file의 생성과 전파를 중단하고, 알려진 정상 파일을 feature file 배포 큐에 수동으로 삽입한 후, 핵심 프록시를 강제로 재시작함으로써 해결했습니다.

위 차트에서 남아 있는 긴 꼬리는, 문제가 발생한 나머지 서비스들을 팀이 재시작하는 과정이며, 17:06에 5xx 오류량이 정상으로 돌아왔습니다.

영향을 받은 서비스는 다음과 같습니다:

틀(나중에 만들 예정).

HTTP 5xx 오류가 발생한 것뿐만 아니라, 장애 기간 동안 CDN 응답 지연(latency)도 크게 증가했습니다. 이는 디버깅 및 모니터링 시스템이 많은 CPU를 사용했기 때문인데, 이 시스템들은 잡히지 않은 오류를 자동으로 추가 디버깅 정보와 함께 처리하도록 설계되어 있습니다.

Cloudflare가 요청을 처리하는 방식과, 오늘 이 과정에서 무엇이 잘못되었는지에 대한 내용.

Cloudflare로 들어오는 모든 요청은 네트워크에서 정해진 경로를 거칩니다. 요청은 브라우저에서 웹페이지를 불러오거나, 모바일 앱이 API를 호출하거나, 다른 서비스에서 자동으로 들어오는 트래픽일 수 있습니다.

이 요청들은 먼저 HTTP와 TLS 계층에서 처리되고, 그다음 핵심 프록시 시스템(“Frontline”, 줄여서 FL)을 통과하며, 필요하면 Pingora를 통해 캐시를 조회하거나 원본 서버에서 데이터를 가져옵니다.

핵심 프록시가 작동하는 방식에 대해서는 여기서 이전에 더 자세히 공유한 바 있습니다.

파일:클플 오류4.png

요청이 핵심 프록시를 통과하는 동안, Cloudflare는 네트워크에서 제공하는 다양한 보안 및 성능 기능을 실행합니다. 프록시는 각 고객의 고유한 구성과 설정을 적용하며, WAF 규칙 적용, DDoS 방어, 트래픽을 Developer Platform이나 R2로 라우팅하는 작업 등을 수행합니다. 이는 도메인별 모듈을 통해 이루어지며, 모듈이 프록시를 통과하는 트래픽에 구성과 정책 규칙을 적용합니다.

그 모듈 중 하나인 Bot Management가 오늘 장애의 원인이었습니다.

3. 에러 상황[편집]

Cloudflare Global Network experiencing issues

Resolved - This incident has been resolved.
Nov 18, 2025 - 19:28 UTC

Update - Cloudflare services are currently operating normally. We are no longer observing elevated errors or latency across the network.

Our engineering teams continue to closely monitor the platform and perform a deeper investigation into the earlier disruption, but no configuration changes are being made at this time.

At this point, it is considered safe to re-enable any Cloudflare services that were temporarily disabled during the incident. We will provide a final update once our investigation is complete.
Nov 18, 2025 - 17:44 UTC

Update - We continue to monitor the system through recovery and we are seeing errors and latency return to normal levels. A full post-incident investigation and details about the incident will be made available asap.
Nov 18, 2025 - 17:14 UTC

Update - We continue to see errors drop as we work through services globally and clearing remaining errors and latency.
Nov 18, 2025 - 16:46 UTC

Update - We continue to see errors and latency improve but still have reports of intermittent errors. The team continues to monitor the situation as it improves, and looking for ways to accelerate full recovery.
Nov 18, 2025 - 16:27 UTC

Update - Bot scores will be impacted intermittently while we undergo global recovery. We will update once we believe bot scores are fully recovered.
Nov 18, 2025 - 16:04 UTC

Update - The team is continuing to focus on restoring service post-fix. We are mitigating several issues that remain post-deployment.
Nov 18, 2025 - 15:40 UTC

Update - We are continuing to monitor for any further issues.
Nov 18, 2025 - 15:23 UTC

Update - Some customers may be still experiencing issues logging into or using the Cloudflare dashboard. We are working on a fix to resolve this, and continuing to monitor for any further issues.
Nov 18, 2025 - 14:57 UTC

Monitoring - A fix has been implemented and we believe the incident is now resolved. We are continuing to monitor for errors to ensure all services are back to normal.
Nov 18, 2025 - 14:42 UTC

Update - We've deployed a change which has restored dashboard services. We are still working to remediate broad application services impact
Nov 18, 2025 - 14:34 UTC

Update - We are continuing to work on a fix for this issue.
Nov 18, 2025 - 14:22 UTC

Update - We are continuing working on restoring service for application services customers.
Nov 18, 2025 - 13:58 UTC

Update - We are continuing working on restoring service for application services customers.
Nov 18, 2025 - 13:35 UTC

Update - We have made changes that have allowed Cloudflare Access and WARP to recover. Error levels for Access and WARP users have returned to pre-incident rates.
We have re-enabled WARP access in London.

We are continuing to work towards restoring other services.
Nov 18, 2025 - 13:13 UTC

Identified - The issue has been identified and a fix is being implemented.
Nov 18, 2025 - 13:09 UTC

Update - During our attempts to remediate, we have disabled WARP access in London. Users in London trying to access the Internet via WARP will see a failure to connect.
Nov 18, 2025 - 13:04 UTC

Update - We are continuing to investigate this issue.
Nov 18, 2025 - 12:53 UTC

Update - We are continuing to investigate this issue.
Nov 18, 2025 - 12:37 UTC

Update - We are seeing services recover, but customers may continue to observe higher-than-normal error rates as we continue remediation efforts.
Nov 18, 2025 - 12:21 UTC

Update - We are continuing to investigate this issue.
Nov 18, 2025 - 12:03 UTC

Investigating - Cloudflare is experiencing an internal service degradation. Some services may be intermittently impacted. We are focused on restoring service. We will update as we are able to remediate. More updates to follow shortly.
Nov 18, 2025 - 11:48 UTC


SYD (Sydney) on 2025-11-18

In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.
Nov 18, 2025 - 15:01 UTC

Scheduled - We will be performing scheduled maintenance in SYD (Sydney) datacenter on 2025-11-18 between 15:00 and 19:00 UTC.

Traffic might be re-routed from this location, hence there is a possibility of a slight increase in latency during this maintenance window for end-users in the affected region. For PNI / CNI customers connecting with us in this location, please make sure you are expecting this traffic to fail over elsewhere during this maintenance window as network interfaces in this datacentre may become temporarily unavailable.

You can now subscribe to these notifications via Cloudflare dashboard and receive these updates directly via email, PagerDuty and webhooks (based on your plan): https://developers.cloudflare.com/notifications/notification-available/#cloudflare-status.
Nov 18, 2025 15:00-19:00 UTC


PPT (Tahiti) on 2025-11-18

In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.
Nov 18, 2025 - 12:00 UTC

Scheduled - We will be performing scheduled maintenance in PPT (Tahiti) datacenter on 2025-11-18 between 12:00 and 16:00 UTC.

Traffic might be re-routed from this location, hence there is a possibility of a slight increase in latency during this maintenance window for end-users in the affected region. For PNI / CNI customers connecting with us in this location, please make sure you are expecting this traffic to fail over elsewhere during this maintenance window as network interfaces in this datacentre may become temporarily unavailable.

You can now subscribe to these notifications via Cloudflare dashboard and receive these updates directly via email, PagerDuty and webhooks (based on your plan): https://developers.cloudflare.com/notifications/notification-available/#cloudflare-status.
Nov 18, 2025 12:00-16:00 UTC


ATL (Atlanta) on 2025-11-18

In progress - Scheduled maintenance is currently in progress. We will provide updates as necessary.
Nov 18, 2025 - 07:02 UTC

Scheduled - We will be performing scheduled maintenance in ATL (Atlanta) datacenter between 2025-11-18 07:00 and 2025-11-19 22:00 UTC.

Traffic might be re-routed from this location, hence there is a possibility of a slight increase in latency during this maintenance window for end-users in the affected region. For PNI / CNI customers connecting with us in this location, please make sure you are expecting this traffic to fail over elsewhere during this maintenance window as network interfaces in this datacentre may become temporarily unavailable.

You can now subscribe to these notifications via Cloudflare dashboard and receive these updates directly via email, PagerDuty and webhooks (based on your plan): https://developers.cloudflare.com/notifications/notification-available/#cloudflare-status.
Nov 18, 2025 07:00 - Nov 19, 2025 22:00 UTC


Support Portal Availability Issues

Update - We are continuing to investigate this issue.
Nov 18, 2025 - 15:16 UTC

Investigating - Our support portal is currently experiencing issues, and as such customers might encounter errors viewing or responding to support cases. Responses on customer inquiries are not affected, and customers can still reach us via live chat (Business and Enterprise) through the Cloudflare Dashboard, or via the emergency telephone line (Enterprise).
We are working to understand the full impact and mitigate this problem.
Nov 18, 2025 - 11:17 UTC

4. 피해 상황[편집]

4.1. 위키[편집]

4.2. 커뮤니티[편집]

  • 아카라이브: 로그인 불가.
  • X: 전반적인 시스템 이용 불가.

4.3. 게임[편집]

  • 리그 오브 레전드: 접속 불가.
  • Geometry Dash: 온라인 서버 이용 불가.
  • 던전앤파이터: 접속 불가.

4.4. 정부 기관[편집]

4.5. 대기업[편집]

  • 맥도날드: 키오스크 이용 불가.
  • Chat GPT: 웹사이트에 한하여 이용 불가.
  • Canva: 웹사이트, 앱 이용 불가.
  • Spotify

4.6. 기타[편집]

  • 불법 OTT 사이트: 접속 불가.

5. 대응[편집]