Identify Your 5xx Error's Fault Tier
Work through which layer of your stack actually produced the failure before you pick a 5xx code, because each one points debugging to a different team. The 5xx class covers every situation where the server received and understood a request but could not process it due to a fault on the server side. Unlike 4xx errors, 5xx errors are not the client's fault: the same request may succeed on retry if the server condition resolves.
Yet the specific 5xx code matters: 500, 502, 503, and 504 each point to a different layer of the stack and a different team's responsibility, and picking the right one on the first attempt routes debugging efforts to the right place immediately. What follows surveys the four most common 5xx codes, explains when to use each, and covers how to monitor and alert on 5xx error rates effectively.1
5xx codes by fault tier
- 500 application code crashed or threw an unhandled exception
- 502 proxy received an invalid or empty response from the upstream
- 503 server intentionally refusing requests — overloaded or under maintenance
- 504 proxy's request to the upstream timed out waiting for a response
502 and 504 are both proxy-layer errors but point to different root causes: 502 to an upstream crash, 504 to a latency problem.
Opens the HTTP Status Code Reference with this section's reference values shown at the top of the tool.
Open in the tool →Identifying the fault tier
Before choosing a 5xx code, identify which layer of the stack produced the failure, because each code corresponds to a distinct tier of the deployment architecture and routes debugging to a different team. Application code that crashes or throws an unhandled exception produces 500 Internal Server Error. A reverse proxy or load balancer that receives an invalid or empty response from an upstream application produces 502 Bad Gateway. A server that intentionally refuses requests because it is overloaded or undergoing maintenance produces 503 Service Unavailable: the server is functioning correctly but is applying backpressure to protect itself from further degradation.
Debugging by layer
Each 5xx code points to a different log source, so picking the right code on the first attempt saves an average of one to two hours of triage during a production incident. 500 points to application logs: search the application stack trace for the unhandled exception. 502 points to proxy logs and upstream health: check whether the upstream returned an empty or malformed response before the proxy forwarded it. 503 points to capacity metrics and maintenance schedules: confirm whether the server is applying backpressure intentionally. Misidentifying the fault tier delays resolution by sending developers to the wrong log source, which is why every operations runbook should include a status-code-to-log-source mapping table.
500 vs 502 vs 503 vs 504 decision tree
Four questions identify the correct 5xx code for a given failure. Walking through them in order prevents you from defaulting to 500 for every unexpected condition, which is a common anti-pattern that hides the real root cause of server-side incidents and makes post-incident reviews harder because the original error signal has been lost.
Did the application receive the request and execute code but threw an unhandled error? Use 500. Did a proxy receive a response from the upstream that was invalid, empty, or not parseable as HTTP? Use 502 Bad Gateway. Is the server intentionally refusing requests because capacity is exhausted or maintenance is in progress? Use 503 Service Unavailable with a Retry-After header. Did a proxy send a request to the upstream and wait for a response that never arrived within the configured timeout? Use 504 Gateway Timeout.
504 and 502 are both proxy-layer errors but with different root causes: 504 points to latency or performance issues at the upstream, while 502 points to an upstream crash or protocol error. Treating them as interchangeable in dashboards and alerts masks the distinct operational responses required: a 504 spike demands a performance investigation of the upstream service, whereas a 502 spike after a deployment strongly suggests a misconfigured health check or a broken upstream process.2
Both codes also share one property: the edge could not complete a round trip to the origin, so the first diagnostic step is always the same, determine whether the failure originates at the origin server or the proxy in front of it.3
Monitoring and alerting on 5xx rates
Effective 5xx monitoring uses rate-based alerts rather than absolute counts, because the absolute count of 5xx responses is meaningless without knowing total request volume. A 500 rate above 0.1 percent of total requests is a meaningful threshold for most production APIs, though the right threshold depends on your baseline error rate and the tolerance defined by your error budget. Setting the threshold too low creates alert fatigue during normal traffic fluctuations, while setting it too high delays detection of genuine incidents.
Alert separately on each 5xx code rather than aggregating them into a single server-error metric: a 502 spike after a deployment points to a proxy configuration change, while a 503 spike points to capacity saturation. Use structured logging with request IDs so that each 5xx response in your dashboard links directly to the corresponding server log entry, which reduces the time an on-call engineer spends correlating client reports with internal logs during an active incident.
Aggregating all 5xx codes into a single "server error" metric obscures which layer of the stack is degrading. Tag each log event with the originating service, the upstream dependency name, and the HTTP method to make filtering and grouping possible in your monitoring tool. Without this granularity, a 502 spike from a misconfigured health check and a 504 spike from a slow database query look identical in a blended metric, yet they require completely different remediation steps: one demands a rollback of the proxy configuration, while the other requires a query optimisation or an index addition.
Circuit breakers and 503 responses in microservice architectures
Circuit breakers prevent cascading failures by stopping requests to a failing dependency before they exhaust your application's thread pool or connection pool. When a downstream service starts returning errors or timing out consistently, the circuit breaker opens and your service returns 503 Service Unavailable immediately, without forwarding the request to the failing dependency. This converts a slow degradation into a fast fail, which is easier to detect and handle than requests that stall for seconds before timing out.
Returning 503 from an open circuit, rather than 500, signals to the caller that the failure is temporary and expected to resolve. To stop a slow dependency from cascading into every endpoint, return 503 instead of 500 when the circuit opens. Include a Retry-After header on the 503 to indicate when the circuit will attempt to close again. A standard circuit breaker half-opens after a configured timeout, allowing one test request through. If the test succeeds, the circuit closes and normal traffic resumes; if it fails, the circuit reopens and the timeout resets.4
Resilience4j and circuit breaker configuration in Spring Boot
Resilience4j is the standard circuit breaker library for Spring Boot applications. Configure a circuit breaker with a sliding window of 10 requests and a failure rate threshold of 50%: when more than 5 of the last 10 requests fail, the circuit opens. The @CircuitBreaker annotation on a service method automatically intercepts calls and redirects to a fallback when the circuit is open. The fallback method should build and return a 503 response, not throw an exception that would produce a 500.
Error budget and SLO monitoring for 5xx rates
Service Level Objectives express reliability as a percentage of successful requests over a time window. A 99.9% availability SLO means no more than 0.1% of requests can return 5xx responses. Every 5xx response burns a portion of your error budget: the total allowable failures before the SLO is breached. When your error budget is nearly exhausted, teams typically freeze non-emergency deployments and focus on reliability until the budget recovers at the start of the next window.
Setting up SLO monitoring for 5xx responses requires tagging each log event with the originating service, the HTTP method, and the status code class. Track the 5xx rate per service separately rather than aggregating across all services: a 5xx spike in one downstream service should not inflate the SLO for your API service. Tools including Datadog's SLO widget, Google Cloud Monitoring, and Grafana's SLO dashboards all support 5xx-rate-based SLOs with configurable burn rate alerts.5
Burn rate alerts for fast SLO degradation
A burn rate alert fires when the rate of SLO consumption is fast enough to exhaust the error budget before the SLO window ends. A 5xx rate of 2% with a 99.9% monthly SLO burns the budget in approximately 2.5 days rather than the full month. Multi-window burn rate alerts (a short window for fast detection and a longer window for sustained degradation) reduce false positives while ensuring the on-call team is paged early enough to prevent a full budget exhaustion.
Configuring the right burn rate thresholds requires understanding your typical traffic patterns. A service that receives 100 requests per second has a different baseline than one that receives 10,000, so the same absolute 5xx rate represents a different proportion of total traffic. Start with a multi-window configuration: a 5-minute window with a 14.4x burn rate threshold for fast detection of acute incidents, and a 1-hour window with a 6x threshold for sustained degradation. Tune these values after observing your service's normal 5xx variance for at least one full traffic cycle, because setting them too aggressively creates alert fatigue that trains the on-call team to ignore pages.
When to use this
Work through this guide the moment you are triaging a production 5xx incident. Use it to identify which layer of the stack is failing, which team should respond, and which logs to inspect first.
Examples
Application throws an unhandled NullPointerException
Return 500 Internal Server Error. Log the full stack trace and correlate it to the request ID. The client request was valid; the application code failed.
nginx receives a connection reset from the Node.js upstream
nginx returns 502 Bad Gateway to the client. Check nginx error logs for "upstream prematurely closed connection" and correlate with the application restart timestamp.
Application is shut down for a rolling deployment
Return 503 Service Unavailable with Retry-After set to the expected restart time. The load balancer sees the 503 health check response and stops routing traffic to the shutting-down instance.
- 1.
R. Fielding, Ed., M. Nottingham, Ed., and J. Reschke, Ed., "HTTP Semantics," RFC 9110, IETF, June 2022. https://www.rfc-editor.org/rfc/rfc9110.txt
- 2.
Mozilla Developer Network, "HTTP response status codes," developer.mozilla.org, accessed June 2026. https://developer.mozilla.org/en-US/docs/Web/HTTP/Status
- 3.
Cloudflare, "Error 502 or 504," developers.cloudflare.com, July 2026. https://developers.cloudflare.com/support/troubleshooting/http-status-codes/cloudflare-5xx-errors/error-502-504/
- 4.
Mozilla Developer Network, "503 Service Unavailable," developer.mozilla.org, accessed October 2026. https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Status/503
- 5.
Wikipedia, "Service-level objective," en.wikipedia.org, accessed June 2026. https://en.wikipedia.org/wiki/Service-level_objective
Debug Your Cloudflare Error With the Ray ID
Find the Ray ID on your Cloudflare error page first: it is your primary debugging artifact for any code above 520. These codes are not part of the IANA HTTP status registry; they are vendor extensions that Cloudflare generates at its edge when the connection between Cloudflare and the origin server fails in a specific way.1 Seeing a 52x error means Cloudflare received the request successfully but could not successfully retrieve a response from your origin.
Consequently, the problem is never on the client side and never inside Cloudflare's own network: it is always between Cloudflare and your origin server. Matching your specific 52x code to its failure mode, using the Ray ID from your error page, points debugging directly at the correct origin-side fix.2
Cloudflare 52x codes
- 520 Unknown Error origin returned something that doesn't conform to HTTP
- 521 Web Server Is Down origin actively refused the TCP connection
- 522 Connection Timed Out TCP connection to the origin timed out
- 524 A Timeout Occurred origin accepted the connection but didn't respond in time
- 526 Invalid SSL Certificate origin's SSL certificate is invalid, expired, or untrusted
Opens the HTTP Status Code Reference with this section's checklist shown at the top of the tool.
Open in the tool →Reading Cloudflare error pages vs origin errors
Cloudflare error pages differ visually from origin error pages in several diagnostic ways: they display a "Cloudflare" branding bar at the top, an error number (520, 521, etc.) in the page title, and a Ray ID in the bottom-right corner that uniquely identifies the failing request. An origin error page, by contrast, is served directly from your application without Cloudflare branding and typically does not include the Ray ID, which makes correlating the error with Cloudflare's edge logs significantly harder.
When you see a 52x code, the error page was generated by Cloudflare, not your application. The root cause always lies in the connection or response between Cloudflare and your origin. Each 52x code represents a different failure mode: 520 means the origin returned an unknown or invalid response; 521 means the origin actively refused the connection; 522 means the TCP connection to the origin timed out; 524 means the origin accepted the connection but did not respond in time.1
The Ray ID on the error page is your primary debugging artifact: log it, search for it in Cloudflare's logs or your origin access logs, and correlate it to the exact request. Without the Ray ID, you are left comparing timestamps and request paths across two separate log systems, which is slow and error-prone during an incident that demands a rapid resolution.2
52x codes vs standard 502 and 504
Cloudflare's 52x codes replace the standard 502 Bad Gateway and 504 Gateway Timeout that a generic reverse proxy would produce. Cloudflare generates 52x codes specifically so that Cloudflare-proxied infrastructure can distinguish between failures at the Cloudflare edge and failures between Cloudflare and your origin. This distinction matters because the remediation path is completely different: an edge failure requires Cloudflare support, while an origin failure requires your own infrastructure team to investigate.
52x to standard code mapping
A 520 Unknown Error is the broadest category: the origin returned something that does not conform to HTTP. A 521 Web Server Is Down maps roughly to a TCP connection refused error. A 522 Connection Timed Out maps roughly to a TCP connection timeout. A 524 A Timeout Occurred maps roughly to an HTTP read timeout after the connection was established: Cloudflare connected successfully but the origin sent no response before the default 125 second Proxy Read Timeout elapsed.3
526 Invalid SSL Certificate and 527 Railgun Listener to Origin Error are specific to Cloudflare's SSL validation and the deprecated Railgun product respectively. Knowing this mapping helps bridge the gap between Cloudflare-specific documentation and standard HTTP debugging knowledge. For most debugging scenarios, focus on 520 through 524 first: these five codes cover the overwhelming majority of origin connectivity issues and each points to a distinct failure mode that maps to a specific fix on your origin server.
Debugging with the Cloudflare Ray ID
Every request Cloudflare processes receives a unique Ray ID, which appears in the error page HTML and in the CF-Ray response header. Using this ID in Cloudflare's dashboard, you can find the specific request in the Cloudflare Logs or Logpush stream and see the exact error message, origin IP address, and connection details. Treat the Ray ID as the single source of truth during an incident: it lets you confirm whether Cloudflare saw the request, whether the origin responded at all, and what specific error the edge encountered when forwarding the connection.
Correlating Ray IDs across log systems
Configure Cloudflare Logpush to stream logs to your storage or SIEM so that Ray IDs are searchable without accessing the Cloudflare dashboard manually. On the origin server, log the CF-Connecting-IP and CF-Ray headers on every incoming request: this links your origin access log entry to the Cloudflare log entry for the same request. Without this correlation, you can see that Cloudflare rejected a request but cannot determine what the origin actually sent, which turns every 52x investigation into a guessing game.
When a 52x error is reported with a Ray ID, you can correlate the Cloudflare edge log with the origin access log to determine exactly what the origin returned and why Cloudflare rejected it. tracing a Cloudflare Ray ID back to the origin response is the fastest way to confirm whether the failure is a connection timeout, an invalid response, or an SSL mismatch at the edge. This cross-system correlation is the fastest path to root cause: it tells you whether the origin sent an invalid response, timed out entirely, or refused the connection, without requiring you to reproduce the issue from outside the Cloudflare network.
Cloudflare Always Online and origin connectivity issues
Cloudflare's Always Online feature serves cached versions of your pages when Cloudflare cannot reach your origin server. When a 52x error occurs because the origin is completely unreachable, Always Online checks Cloudflare's cache and serves the last known good version of the page with a banner indicating the site is temporarily offline. This feature activates for 502, 504, and Cloudflare-generated 52x status codes (520-527) when the origin is unreachable.4
Always Online has important limitations you must understand before relying on it. Pages that require authentication, pages with forms that submit to the origin, and pages generated from real-time data are served from cache in a broken state: the cached version exists, but any user action that reaches the origin will fail. Always Online is designed for informational pages, not for application workflows. Disable it for routes that must show a genuine error when the origin is down, rather than a stale cached version.
Verifying Always Online cache coverage
Cloudflare caches pages for Always Online based on your Cache-Control headers and Cloudflare's caching tier. Pages without any Cache-Control headers, or with Cache-Control: no-store, are not eligible for Always Online. Check which pages Cloudflare has cached by reviewing the cache hit rate for your domain in Cloudflare Analytics. For high-traffic public pages, set Cache-Control: public, max-age=3600 to ensure Cloudflare has a cached copy available for Always Online to serve when the origin becomes unreachable.5
When to use this
Grab the Ray ID from your error page and use this guide when diagnosing errors on Cloudflare-proxied infrastructure. Match your specific 52x code to understand what it means about the origin-to-Cloudflare connection, and determine which logs to inspect next.
Examples
521 Web Server Is Down appears after deployment
The origin is not listening on the expected port (80 or 443). Check that the application process started successfully and is bound to the correct port. Verify Cloudflare is pointing to the correct origin IP address.
522 Connection Timed Out appearing intermittently
Cloudflare could not establish a TCP connection to the origin within 15 seconds. Check origin server firewall rules: Cloudflare IP ranges must be allowed. Also check if the origin is under high load causing connection queue buildup.
526 Invalid SSL Certificate
The SSL certificate on the origin server is invalid, expired, or untrusted by Cloudflare. Either install a valid CA-signed certificate on the origin, or change the Cloudflare SSL mode to "Flexible" (not recommended for production) or "Full" (accepts self-signed certs from origin).
- 1.
Cloudflare, "Cloudflare 5xx errors," developers.cloudflare.com, accessed June 2026. https://developers.cloudflare.com/support/troubleshooting/http-status-codes/cloudflare-5xx-errors/
- 2.
Stack Overflow, "What is a Ray ID (Cloudflare)?," stackoverflow.com, accessed June 2026. https://stackoverflow.com/questions/49968948/what-is-a-ray-id-cloudflare
- 3.
Cloudflare, "Error 524," developers.cloudflare.com, July 2026. https://developers.cloudflare.com/support/troubleshooting/http-status-codes/cloudflare-5xx-errors/error-524/
- 4.
Cloudflare, "Always Online because downtime sucks," blog.cloudflare.com, accessed June 2026. https://blog.cloudflare.com/always-online-because-downtime-sucks/
- 5.
Mozilla Developer Network, "Cache-Control," developer.mozilla.org, accessed June 2026. https://developer.mozilla.org/en-US/docs/Web/HTTP/Reference/Headers/Cache-Control
A 520 Unknown Error is generated by Cloudflare when the origin returns a response that is invalid or does not conform to HTTP (blank response, wrong protocol, connection resets after sending headers). A 502 on a Cloudflare-proxied site typically indicates a Cloudflare edge issue rather than an origin issue. If you see 520, the problem is your origin, not Cloudflare.
Look up the Ray ID in Cloudflare Logs or Logpush. The log entry shows the origin IP, the error type, and the connection details. On your origin server, search your access logs for the corresponding CF-Ray header value to see what your application received and responded with at that moment.
524 A Timeout Occurred means Cloudflare connected to your origin but the origin did not send a complete HTTP response within Cloudflare's timeout window (100 seconds by default). The origin accepted the TCP connection but was too slow to respond. Fix by optimising the long-running operation on the origin or enabling Cloudflare's proxy read timeout extension for enterprise plans.
Cloudflare uses multiple anycast edge locations. A 522 from one edge location may indicate a routing or firewall issue specific to the Cloudflare IP ranges used by that edge location. Check that all Cloudflare IP ranges are allowlisted in your origin firewall, not just the ranges you have seen in practice.
Yes, if it persists. Googlebot treats 52x codes as server errors. A brief 52x spike during a deployment does not trigger immediate deindexing, but sustained 52x responses over multiple crawl cycles signal to Google that the site is unreliable. Fix origin connectivity issues promptly and use 503 with Retry-After for planned maintenance to protect crawl signals. CapyToolkit allows you to verify the exact status codes your origin returns, so you can confirm whether a 52x is originating from your infrastructure or from Cloudflare's edge before opening a support ticket.