Defining a service level agreement in customer contracts
A four-hour response promise means nothing if the contract never says when the clock starts. How to write SLA clauses you can measure and defend.
Monday morning brings an email from the largest customer: the contract says four hours, and two days have passed with no answer. The support lead opens the record. The ticket arrived on Friday, forty minutes after the office closed; an automated acknowledgment went out within a minute, and a human reply was written on Monday morning. Nowhere does the contract define what counts as a response, what the working hours are, or whether the weekend runs the clock. Nobody is lying. The person who wrote that sentence never had to measure it. The result is a credit claim and a renewal conversation that opens from a defensive position.
A service level agreement is the clause that turns a promise to a customer into something measurable and arguable. What follows covers what separates an SLA from a statement of good intent, what measurability actually requires, which metrics belong in a contract, who assigns priority, when the clock pauses, how to write exceptions honestly, what makes a remedy real, the operation behind the promise, tiered commitments, the backfire of targets that are too tight, and the reporting rhythm that keeps the clause alive.
An SLA is not a statement of good intent
For a clause to qualify as an SLA it needs three parts: a measured indicator, a numeric threshold and a consequence that triggers when the threshold is missed. Without the third you have goodwill, not an agreement. That is precisely what most contracts lack. A sentence promising a prompt response binds nobody, cannot be measured, and gives no signal at the moment it is breached, so both parties read the same text and understand different things.
The less discussed function of an SLA is internal. The response time you promise a customer is really a capacity plan written for your own team: it dictates how many people must be available, in which hours, on which channel. It is more useful to treat the clause as the beginning of a staffing model than as a legal artifact. Signing a window you cannot hold does not damage the sales conversation; it damages the next twelve months of how your team works.
A clause you cannot measure is not an SLA
Measurability does not mean a number appears in the contract. It means the document states when the clock starts, when it stops and what counts. Which event starts it — the customer sending an email, the ticket landing in the system, or the ticket reaching the right queue? What counts as a response — an automated acknowledgment, or the first sentence a human writes? Every clause written before those two questions are answered will be read two different ways at the first dispute.
The same problem appears on the measurement side. If first response time on the support dashboard is calculated differently from first response time in the contract, the report you send will confuse your own team as much as the customer. Keeping one place where each metric name is defined is the cheapest fix available; we describe how to build that record in a metric dictionary and definition standard.
Working hours are the quietest clause in the contract
Business hours, public holidays and time zones go unwritten in most agreements, because everyone assumes their own calendar is universal. For two parties in different cities or countries that assumption is expensive. A four-hour response commitment inside an eight-hour day is one thing; the same four hours across a twenty-four hour window is an entirely different cost structure. Knowing which one you signed matters more than debating the number itself.
Which metrics belong in the contract?
There are dozens of indicators you could measure and very few that belong in a contract. To qualify, a metric has to measure something the customer genuinely feels and something you control on your own. Committing to an indicator that moves when the customer is slow to answer is buying an argument in advance.
| Metric | What it measures | Must be defined in the contract |
|---|---|---|
| First response time | Time until the first human touch on a ticket | Channel, working hours, exclusion of auto-replies |
| Resolution time | Time until the ticket is closed | Closing criteria and whether a workaround counts |
| Availability | Share of time the service is operational | Measurement point and exclusion of planned maintenance |
| Escalation time | Time to move from one tier to the next | Trigger threshold and who initiates the escalation |
| Maintenance window | The period in which downtime is permitted | Advance notice period and the hours it covers |
Try to keep the contract to three metrics at most. Five separate commitments produce a list neither side can follow, and an unwatched commitment is remembered only when it carries bad news. First response time is the right starting point for most relationships because it measures how long the customer sits in uncertainty; we go into how to bring it down in our article on first response time.
Who decides the priority?
In SLA negotiations the most contested clause is rarely the duration — it is the priority definition. A customer who accepts your response times and then opens every ticket as critical makes the commitment impossible by arithmetic alone. What prevents that is not goodwill but written definitions anchored to business impact. The tiers below cover most small and mid-sized providers and should be written to be read aloud with the customer.
- Critical: The service is fully down, no workaround exists and the customer's daily work cannot proceed. Keeping this tier narrow is the only thing that keeps it meaningful.
- High: An important function is broken but work continues; the workaround is partial or costly. Most genuine incidents land here.
- Medium: A usable workaround exists and the impact stays within a limited group of users or a single workflow.
- Low: Annoying but not blocking — display errors, isolated questions, gaps in documentation.
- Request: Not a fault but a new need. Requests should stay out of the SLA clock and run through a separate roadmap process.
- Dispute path: The provider makes the first classification; the customer can challenge it, and the challenge is reviewed jointly within a stated window.
The rule that works in practice is that priority follows business impact, not the tone of whoever filed the ticket. Write the right to downgrade into the contract, but never exercise it silently. Lowering a ticket's tier without telling anyone is the fastest way to win the clock and lose the customer.
When does the clock pause?
The mechanic that generates the most argument in any SLA is the pause. Stopping the clock while a ticket waits on the customer is reasonable; otherwise their delay shows up as your breach. But an invisible pause looks like manipulation. The rule is simple: a pause must be a status change the customer can see, it must state what is awaited and from whom, and a reminder should go out automatically if the wait drags on.
The opposite mistake is just as common. Teams that never pause the clock accumulate breaches on tickets where the customer went quiet for three days, and eventually stop trusting the report. A report nobody trusts is worse than no report, because it gets defended in meetings instead of corrected. When writing pause rules, also state how many times and for how long a ticket can be paused.
An SLA is not the number you can hit in an average week; it is the number you can hit in your worst one.
Write the exceptions honestly
Every SLA has exceptions: planned maintenance, force majeure, changes in the customer's own infrastructure, third-party outages. Writing them is not the problem — writing them broadly is. A phrase like anything outside the provider's control may be technically accurate and will still read as an escape hatch, dismantling on its own the trust the clause was meant to create. Exceptions should be enumerable and narrow.
The price of an exception is an obligation attached to it: advance notice. Announcing planned maintenance a stated period ahead is what earns the right to exclude it. Knowing where to look during an incident belongs to the same clause; we cover how to set that communication up in outage announcements and status pages. Unannounced planned maintenance feels like unplanned downtime no matter what the contract says.
No remedy, no promise
If the contract is silent on what happens when the threshold is missed, the clause is a wish. The common mechanism is a service credit: an offset against the next period, scaled to the severity of the breach, usually capped and usually requiring the customer to claim it. There are alternatives worth considering — extending the term, granting a penalty-free exit after repeated breaches, or committing to heightened reporting for a defined period.
Here is the unexpected but consistent part: credits rarely satisfy anyone. Nobody wants an offset in exchange for an outage; they want the incident not to repeat. Pairing the remedy with an obligation to deliver a root-cause report is worth more than the financial gesture. We wrote about how a bad incident can end up strengthening a relationship in service recovery. Have your own counsel finalize the wording for your situation; the mechanics described here do not substitute for that review.
Building the operation behind the promise
An SLA lives or dies not in the contract but in the first ten minutes after a ticket arrives. If the ticket does not reach the right queue, if the clock is not visible on anyone's screen, and if the priority field can be left blank, the signed window is decoration. The baseline requirement is that requests collect in one place and every record shows the time remaining; we walk through that setup in help desk and ticket management.
The second requirement is that escalation runs before the breach, not after it. A system that alerts you when the window closes only tells you that you lost. A rule that fires when a defined share of the remaining time is gone still leaves room to act. That is exactly what escalation automation is for: a ticket approaching its threshold becomes visible one level up before anyone has to notice it manually.
The same SLA for everyone?
Giving every customer the same window looks fair and is rarely economic. The condition for tiering is this: the difference between tiers has to be a real operational difference, not a smaller number in a table. An on-call rotation, a named contact, a separate queue, a wider maintenance window — those are concrete. A premium tier built only by shortening the clock collapses in its first busy week.
When designing tiers, cost to serve is a better guide than contract size; the calculation in cost to serve per customer makes that visible. It can also produce a surprising conclusion: the tightest SLA does not always belong to the largest account. Sometimes the workflow that suffers most from an outage sits inside a small customer, and that workflow is what deserves the narrowest commitment.
Tighter is not always better
This is where the usual advice needs contradicting: shortening the window does not automatically improve the service. A first response target set too tight pushes agents to send an empty message purely to stop the clock. The ticket opens, a we are looking into it reply goes out within a minute, the indicator looks excellent, and the customer has learned nothing. Total resolution time stretches, the team hits its metric, and the experience gets worse.
The remedy is to stop rewarding first response on its own. Reading it alongside reopen rate and one-touch resolution separates paper performance from the real thing; our article on first contact resolution works through that balance. Tying an SLA indicator to individual bonuses in isolation carries the same risk, for the same reason: what you measure gets shaped by people who know it is being measured.
Measure, report, revisit
Sending the SLA report before the customer asks for it is a small move that changes the tone of the relationship. A report produced on request is a defense; a report that arrives on its own is transparency, and even a bad month becomes discussable in that frame. The report should carry not only the number of breaches but their causes and the actions taken. That content is a natural part of the customer business review agenda.
An SLA is also a living clause. The product changes, the team grows, the customer's usage shifts, and a threshold written two years ago is now either far too loose or quietly dangerous. Use the renewal cycle to reread it. We collected ways to keep terms, renewals and amendments in one place in contract lifecycle management.
The safest way to start is to measure before you commit. Pick one metric, track it for a month without promising it to anyone, and look at the number from your worst week. The threshold you write into the contract should sit slightly above that figure, not above your average. This one discipline removes most breach conversations before they ever happen.
Writing an SLA without a place where tickets collect, priority is recorded and remaining time is visible to everyone means signing a promise you cannot keep. Rocketly brings support tickets, tasks and reminders, automation rules and reporting onto the same customer record — open a free account and build your own service level flow.