# Should You Promise an SLA Before You Can Measure Reliability?

> Decide whether a customer-facing service promise is ready by measuring the workflow, defining a realistic target, and attaching consequences you can actually honor.

By buildpurdue Team · October 1, 2026 · 4 min read

Source: https://www.buildpurdue.org/blog/promise-sla-before-measuring-reliability

---

An enterprise buyer asks for 99.9% uptime, and the founder says yes before checking whether the team can measure an outage.

Maya runs a small workflow product for logistics teams. A prospect is ready to sign, but its procurement form asks for an uptime commitment, response times, and service credits. Maya has application logs and a few support emails. She does not have a defined service boundary, an uptime calculation, or a reliable incident clock.

She should not sign the SLA yet. She should define and measure the service level first, then make the smallest promise her operation can keep.

## What are you actually promising?

An SLA is a documented customer commitment. It normally defines the service, performance targets, measurement method, and what happens when the targets are missed. [Atlassian's SLA guide](https://www.atlassian.com/itsm/service-request-management/slas) describes examples such as uptime, first-response time, resolution time, and remedies including service credits.

That list exposes the first problem with a vague promise. “The product will be reliable” does not tell Maya what counts as an outage. Does a failed login count? What about a delayed background export, one broken customer workflow, or a third-party dependency? Which customers and endpoints are included? When does the clock start, and who can verify the result?

Before she negotiates a percentage, Maya writes one sentence: “A customer can open a shipment, save the required fields, and export the handoff file during business hours.” That is a customer outcome she can test. A promise about the whole system would be harder to measure and easier to misunderstand.

## Build the internal target before the external promise

Google's SRE guidance separates a service level objective, or SLO, from an SLA. An SLO is an internal reliability target used to guide engineering decisions; an SLA is a business agreement that can require compensation when the target is missed. [The SRE Workbook recommends defining SLOs before general availability](https://sre.google/workbook/engagement-model/) so a team has an objective way to evaluate reliability. Its [SLO guidance](https://sre.google/workbook/implementing-slos/) also warns that 100% reliability is not a reasonable target.

Maya starts with a smaller internal target: 99.5% successful shipment exports over a calendar month, measured from the customer-visible endpoint. She adds a response-time target for urgent support and records which failures are excluded, such as a customer's own invalid file. The exact number is a working assumption, not a universal benchmark. She can revise it after seeing real behavior.

The important part is that the target has an owner and a consequence. If Maya misses it, she pauses a risky release, investigates the failure, and tells affected customers what changed. Google SRE describes an error budget as the room between perfect reliability and the chosen target, with reliability work taking priority when that budget is spent. [Its error-budget guidance](https://sre.google/workbook/error-budget-policy/) shows how a target can change release decisions instead of sitting in a dashboard nobody uses.

## Can you measure the promise without arguing later?

An SLA becomes dangerous when the customer and the provider calculate it differently. Maya needs a simple record for every measurement period: total eligible requests, successful requests, downtime or failed-workflow minutes, incident start and end times, planned maintenance, and excluded events.

She also runs the calculation against the last four weeks of data. If the logs cannot answer the question, she does not fill the gap with an estimate. She instruments the missing event and waits for enough observations to understand normal behavior. A customer-facing target that cannot be reconstructed after an incident is a negotiation trap for both sides.

The same rule applies to support. “Critical tickets answered quickly” needs a severity definition, an intake channel, a response clock, and an escalation owner. If Maya is the only person watching the inbox, she should not promise a 24-hour response seven days a week until she has a coverage plan.

## What should you offer the buyer today?

Maya can still move the deal forward. She can share the service boundary, the current measured target, the incident-notification process, and a review date for a formal SLA. If the buyer requires an SLA immediately, she can negotiate a narrower commitment around one workflow instead of promising the whole product.

She should also price the promise. A tighter target creates monitoring, on-call, redundancy, support, and credit exposure. If the buyer wants custom coverage, that cost belongs in the deal rather than hidden inside an optimistic contract.

Before signing, write down the workflow, SLI, target, measurement window, exclusions, notification duty, remedy, and owner. Then test the calculation on a past incident. If two people cannot reach the same answer, you are still defining the service.

Bring that draft promise to the [buildpurdue cohort](/cohort) and ask someone to challenge the measurement before you put it in a customer contract.
