Summary

  • Cloudflare’s published Enterprise SLA makes the mechanism unusually visible. Assume a 30-day month with 43,200 scheduled minutes, a 60-minute qualifying outage, an affected-customer ratio of 100%, and no deductions for customer-planned downtime or force majeure. Its published formula gives ((60 × 5) × (1 × 5)) / 43,200 = 0.0347222, or about 3.47% of the relevant monthly recurring fee.

    That is an illustrative calculation, not a real customer claim, and it says nothing about business losses caused by the outage. In a 31-day month, with 44,640 minutes in the denominator, the same assumptions would produce about 3.36%. The headline promise is therefore only the first input into a larger contractual machine.

  • Across Cloudflare, AWS, Google Cloud, Oracle and IBM, the economically important questions sit below the percentage: what constitutes downtime, which deployment qualifies, which project, region, instance or service is deemed affected, what fees form the credit base, how incidents are evidenced, how quickly claims must be filed, which exclusions apply, whether overlapping remedies can be stacked, and where liability stops.

    These public terms should not be read as a reliability ranking. They illustrate different ways of drawing the boundary between provider-controlled service failure and the much larger end-to-end exposure retained by the customer.

Sixty minutes inside the formula

Start with the numerator.

Cloudflare publicly commits to 100% uptime for the covered Enterprise Service. But the monetary consequence of missing that level is not calculated as the customer’s lost sales, foregone transactions, remediation cost or reputational damage. It is calculated through a service-credit formula.

The affected-customer ratio is the number of unique visitors affected by an unscheduled service outage divided by total unique visitors. The service-credit ratio then takes outage minutes multiplied by five, multiplies that by the affected-customer ratio multiplied by five, and divides the result by scheduled-availability minutes.

With the assumptions above, a full-hour outage affecting every visitor produces a ratio of about 3.47% in a 30-day month. That percentage is then applied only to monthly recurring fees associated with the covered Service.

The same incident placed in a longer month produces a slightly lower result because the denominator is larger. At 44,640 minutes, the ratio is about 3.36%. The arithmetic is elementary, but its implication is not: the economic value of an uptime SLA depends partly on a denominator that is invisible in the headline.

Scheduled availability is itself defined rather than assumed. Cloudflare starts with the total minutes in the month and subtracts customer-planned downtime and force-majeure downtime. The outage must also survive the SLA’s definitions, exclusions and validation process before a credit becomes due.

So the 3.47% example is not an estimate of what Cloudflare would pay any particular customer. It is a reconstruction of the published mechanism under specified assumptions. A real outcome would depend on the actual outage period, affected-customer ratio, applicable fees, exclusions, evidence and claim validation.

That distinction is the foundation of sensible SLA analysis. A service level is a nominal promise. A service credit is a bounded remedy. Neither is automatically equivalent to the customer’s workload availability or economic loss.

The percentage sits inside a perimeter

Every SLA needs a perimeter. The difficult question is not simply whether infrastructure was “up”; it is which infrastructure, observed from where, under which configuration and against what failure definition.

Cloudflare’s formulation connects the outage calculation to affected visitors and scheduled availability. AWS draws a different boundary. Its current EC2 SLA offers a 99.99% region-level commitment only when all running instances are concurrently deployed across at least two availability zones in the region. A single EC2 instance instead falls under a 99.5% instance-level commitment.

That is not merely a difference between two percentages. It is an architectural prerequisite embedded in the contract. The more demanding region-level commitment applies to a deployment that has already distributed its running instances across failure domains.

Google Compute Engine similarly distinguishes multi-zone and single-instance configurations, with commitments varying by deployment and network tier. Its monthly uptime calculations are made per project and region, or per single instance. A period of intermittent downtime shorter than one minute does not count as a downtime period.

Oracle’s PaaS and IaaS framework illustrates another form of perimeter. Depending on the service, the SLA may concern availability, manageability or performance. The relevant object is the specific non-compliant service, not an undifferentiated customer estate.

IBM makes the boundary explicit through its distinction between an SLO and an SLA. An SLO is an objective, not a contractual guarantee that itself grants credits. Its resilience guidance also states that a workload must be deployed accordingly to take full advantage of the cited 99.999% VPC SLO: the example requires three virtual servers, one in each of three zones, plus a load balancer. IBM describes resilience of the cloud as its responsibility while leaving resilience and recovery of customer workloads with the customer.

These are contrasting measurement designs, not evidence that one provider is inherently more reliable than another. A 99.99% figure attached to a multi-zone regional construction cannot be compared mechanically with a percentage attached to one instance, one service, one project or one customer-facing traffic ratio.

Before comparing the number, a buyer has to normalize the perimeter.

Credits price fees, not consequences

The most important denominator may not be the uptime denominator at all. It may be the money denominator.

Cloudflare calculates credits only against monthly recurring fees associated with the covered Service. AWS uses the affected-region EC2 bill for a region-level claim or the relevant single instance for an instance-level claim, excluding one-time payments such as upfront Reserved Instance payments. Google’s financial credits are bounded by the affected covered-service amount. Oracle calculates credits from net fees for the quantity of the specific non-compliant service actually used during the measured period.

These structures do something economically coherent: they price the provider’s contractual failure against the commercial value of the service the provider sold.

They do not price the customer’s business interruption.

Suppose a cloud component carries a modest monthly infrastructure bill but supports a payment flow, a trading process, an identity system or another workload whose value is many times the infrastructure fee. Even a 100% service-credit tier can still mean, at most, a credit tied to the defined service charges. The 100% refers to the credit base, not to the customer’s downstream loss.

This is the gap that procurement discussions often blur. A high SLA percentage may reduce the probability of an eligible breach. A generous credit tier may increase the remedy once a threshold is crossed. But neither tells the buyer how much end-to-end economic exposure has been transferred.

The retained exposure includes everything outside the SLA’s fee base and failure perimeter. That can include application dependencies, customer architecture, other cloud services, third-party systems and the economic consequences of the workload being unavailable. The public SLA normally does not convert those consequences into damages.

Cloudflare makes the limit particularly visible by making service credits the sole and exclusive SLA remedy and imposing an annual cap of six months of cumulative monthly service fees. AWS and Google likewise describe SLA credits as the exclusive remedy within their published terms. Oracle’s framework describes service credits as the exclusive remedy and Oracle’s entire liability for the relevant service commitment failure.

The economic instrument is therefore bounded twice: first by what counts as a breach, and again by what money can be credited after the breach is established.

Eligibility is an architecture choice

The contract can reward a resilient architecture without financing it.

AWS’s region-level EC2 commitment is a clear example. To receive the 99.99% region-level treatment, all running instances must be concurrently deployed across at least two availability zones. A customer running a single instance has a different commitment and a different measurement boundary.

The architecture therefore precedes the remedy. The customer must bear the cost and operational complexity of multi-zone deployment before the stronger regional commitment is relevant.

IBM’s guidance reaches the same issue from the resilience side rather than the credit side. To take full advantage of its cited 99.999% VPC SLO, the example workload needs three virtual servers across three zones and a load balancer. The objective describes a platform capability that the workload architecture must actually use.

This matters for investment decisions. The price of a stronger availability posture is not simply the premium, if any, embedded in an enterprise support or service contract. It can include duplicate or triplicate compute, load balancing, data replication, application changes, testing, observability and operational procedures.

That cost is part of the risk transfer equation.

A buyer who reads only the headline percentage can make the wrong comparison. A nominally stronger commitment may require more customer-funded redundancy. A weaker single-resource commitment may sit inside a workload that is highly resilient because the customer has engineered around the component. Neither fact can be inferred from the percentage alone.

The useful procurement question is therefore not, “Which provider promises the most nines?” It is, “What architecture must we operate before this particular promise applies, and what risk remains ours even after we do?”

That reframes availability from a marketing attribute into a joint production problem. Providers control part of the failure surface. Customers control another part. The SLA is most informative when those boundaries are visible rather than blended.

A claim is an operational workflow

A theoretically valuable credit can disappear through procedure.

Cloudflare requires the customer to notify support within five business days following an incident before becoming eligible to submit a claim. The customer must provide reasonable detail, including the incident description and duration, network traceroutes, affected URLs and attempts to resolve the issue. Sufficient evidence supporting the claim must be submitted by the end of the billing month following the incident month.

AWS gives a longer window, but the burden is still explicit. EC2 credit requests must be received by the end of the second billing cycle after the incident. The required material includes dates and times, the affected region or availability zone as relevant, resource IDs and logs or other data corroborating the outage.

Google requires a customer seeking Compute Engine financial credit to notify technical support within 60 days from eligibility and provide logs showing the downtime periods and when they occurred.

Oracle’s March 2026 PaaS and IaaS pillar likewise requires a claim within 60 calendar days and asks for the region, relevant OCIDs, a description of attempts to resolve the issue and supporting documentation or logs.

These are not administrative footnotes. They create an operational option that expires.

A customer can suffer a qualifying event and still fail to realize the contractual remedy because the incident was not escalated quickly enough, logs were not retained, resource identifiers were missing or the relevant evidence could not be reconstructed within the claim window.

That makes observability part of contract management. Incident timestamps, synthetic monitoring, request logs, topology records, resource inventories and support tickets do more than help engineers diagnose failures. They determine whether an economic claim can be proven.

Cloudflare states this division directly: comprehensive monitoring of customer content is the customer’s responsibility. It will consider commercially reasonable independent measurement and use available information to calculate the affected-customer ratio.

The party seeking the credit therefore needs its own evidence system. A monitoring stack designed only to restore service, but not to preserve reconstructable evidence, may leave money on the table.

Measurement has to survive disagreement

An SLA is useful when two parties can reconstruct the same event from sufficiently thin, deterministic evidence.

That is harder than it sounds.

Cloudflare’s affected-customer ratio depends on unique visitors affected during the outage relative to total unique visitors. Google defines downtime at the project-region or instance level and disregards partial minutes or intermittent outages shorter than one minute. AWS distinguishes region-level unavailability from the state of a single instance. Oracle may measure availability, manageability or performance depending on the covered service.

Each design answers a different question.

The danger for the customer is to monitor something broader than the contract and assume the two measures will converge. An application can be unusable even while an underlying compute instance remains within its contractual definition of availability. A regional component can suffer degradation without meeting the SLA’s precise downtime threshold. A customer may observe user failure at the edge while the provider attributes part of the incident to an excluded dependency or customer-controlled system.

The answer is not to measure everything jointly. It is to keep the shared contractual measurement small enough that both sides can reconstruct it, while the customer separately measures the full workload.

That produces two ledgers.

The first ledger is contractual: qualifying minutes, covered resources, affected traffic, logs, claim dates and eligible fees.

The second is economic: failed transactions, degraded user journeys, operational labor, recovery time and wider business exposure.

The two should be connected, but they should not be confused.

The SLA governs the first ledger. Resilience engineering, insurance, continuity planning and management decisions address the second.

Non-stacking is part of the price

Overlapping promises do not necessarily create overlapping compensation.

AWS states that region-level and instance-level claims cannot be combined or stacked for a particular EC2 instance. Google similarly provides that a customer cannot receive both single-instance and multi-zone financial credit for the same downtime.

Oracle’s framework is explicit where more than one SLA provision could cover the same incident: the customer generally receives only the highest applicable credit for the service, rather than recovering under multiple SLA provisions for the same event.

These clauses matter because modern cloud workloads are layered. A single incident may touch compute, networking, storage, control-plane operations and application availability at once. Without a non-stacking rule, one physical event could create multiple contractual claims on the same underlying failure.

Providers largely avoid that result by defining affected resources narrowly and restricting overlapping credits.

For the buyer, this means redundancy in contractual labels does not necessarily equal redundancy in remedies. Two relevant commitments may describe the incident, but only one credit may survive.

Caps reinforce the same principle. Cloudflare’s annual credit ceiling, Google’s affected-service ceiling and Oracle’s service-specific limits all stop the remedy from expanding with the customer’s wider loss.

Again, this is not evidence that the promises are meaningless. It shows what they are designed to do: discipline service delivery and return part of the service fee when defined performance falls outside a boundary.

They are not designed as open-ended business-interruption insurance.

The retained gap needs its own price

Once that distinction is clear, SLA evaluation becomes more practical.

Start with the service label, but do not stop there. Trace the dependency the workload actually needs. Identify the deployment required to qualify for the relevant commitment. Reconstruct the downtime definition. Determine the affected-resource boundary. Calculate the credit against the correct fee base. Record the claim clock. Confirm the evidence requirements. Map exclusions, caps and non-stacking provisions.

Then compare the resulting transfer with the workload’s real failure cost.

The difference is the untransferred gap.

That gap can be reduced through architecture: multi-zone or multi-region deployment, failover, replicated data, independent monitoring and tested recovery. It can be financed through insurance where appropriate. It can be negotiated through procurement. It can also be managed through exit options that reduce dependence on a single service or architecture.

None of those measures changes what the public SLA says. They address what the public SLA does not transfer.

This is also why renewal matters. Cloudflare states that the then-current SLA version applies when a renewed term begins. Contract monitoring therefore cannot end after the initial signature. A change in definitions, claim procedures or remedies can alter the effective allocation even if the headline service appears unchanged.

The relevant unit of analysis is consequently not “99.99% uptime.” It is a chain:

measurement boundary, deployment prerequisite, qualifying event, evidence, claim procedure, eligible resource, credit denominator, exclusions, non-stacking rule and cap.

Only after that chain is reconstructed does the percentage become economically meaningful.

Sources