Key Takeaways

  • An SLA needs four elements: a precisely defined trigger, a measurable standard, an agreed measurement method, and a stated consequence. Missing any one turns it into an aspiration.
  • SLAs exist to remove ambiguity at a handoff. Their value is highest exactly where one team's output becomes another team's input, which is why the marketing-to-sales handoff is the most valuable internal SLA most companies never write.
  • The consequence does not have to be financial. For internal SLAs it is usually escalation and visibility — but if nothing at all happens on a breach, behaviour will not change.
  • Set targets from your own observed distribution, not from an aspiration or a competitor's published number. A target nobody has ever hit teaches the team that the SLA is decorative.
  • Measure the percentile, not the average. An average response time of two hours is consistent with most records answered in minutes and a tail waiting two days, and the tail is what customers and reps actually experience.
  • Edge cases are where SLAs fail: out-of-hours arrivals, paused clocks, reassignment and reopened records. An agreement that only documents the standard path will be argued about within a month.
  • Every SLA must be bidirectional. A one-way commitment where the receiving side has no obligations and no right to reject is not measurable and will not be honoured.

1. What Is an SLA? The Short Answer

A service level agreement is a written commitment that a defined service will be delivered to a measurable standard, within a stated time, with an agreed method of measurement and a defined consequence when the standard is missed.

Four elements, all required. A precise trigger, so both sides agree when the clock starts. A standard expressed as a number, so performance is a fact rather than an opinion. A measurement method both parties accept, including which system is authoritative. And a consequence, so that missing the standard produces something other than a conversation.

Remove any one and it stops being an agreement. Without a precise trigger, the two sides measure different things. Without a number, "promptly" means whatever each party needs it to mean. Without an agreed measurement method, every breach becomes a dispute about the data. Without a consequence, it is a statement of intent that will be ignored the first time it is inconvenient.

This is the same discipline as writing verifiable stage exit criteria in a pipeline — the standard has to be checkable by someone who was not in the room. We cover that parallel in the [funnel stages guide](/guides/funnel-stages-guide).

  • AEO Quick Answer: An SLA is a defined service, a measurable standard, an agreed measurement method and a stated consequence. All four are required.
  • Trigger precision decides whether both sides are measuring the same thing.
  • No consequence means no behaviour change — but the consequence need not be financial.

2. Three Different Things Are Called an SLA

The term covers three distinct instruments with different stakes and different failure modes. Establishing which one you are writing is the first decision.

The customer-facing SLA. A contractual commitment to your customers — uptime, support response, delivery windows — usually with financial remedies such as service credits. It is legally binding, it is negotiated, and it should be written with legal input because the remedy is real money.

The vendor SLA. The mirror image: what a supplier commits to you. Most businesses accept these as boilerplate, which is a mistake, because a vendor SLA covering a system your own customer-facing SLA depends on is a dependency you have not priced.

The internal or operational SLA. A commitment between two teams inside the same company — marketing to sales on lead volume and quality, sales to delivery on handover completeness, support to engineering on escalation. No contract, no money changes hands, and in our experience this is where the largest recoverable value sits, precisely because nobody treats it with the seriousness of the other two.

The mechanics are shared across all three. The differences are the remedy, the negotiating relationship and the consequence of failure. This guide covers the common mechanics and flags where the three diverge.

  • Customer-facing: contractual, financial remedies, needs legal input.
  • Vendor: usually accepted as boilerplate, which leaves an unpriced dependency.
  • Internal: no contract, largest recoverable value, most often absent entirely.

3. What Is an SLA For? The Problem It Solves

An SLA solves one problem: ambiguity at a handoff. Everything it delivers follows from that.

When work passes between parties — teams, companies, systems — there is a moment where responsibility transfers. If the expectations at that moment are unwritten, three predictable things happen. Each side assumes a different standard, and both believe they are meeting it. Failures are invisible until they aggregate into something large enough to notice. And when they are noticed, the conversation is about blame rather than about the mechanism, because there is no agreed standard to measure against.

An SLA makes the transfer explicit. It converts "we should respond quickly" into "first response within one business hour, measured from record creation, at the 90th percentile, reported weekly". That sentence can be verified by anyone, and disagreement about it is now a disagreement about data rather than about character.

The second thing it solves is the asymmetry of visible effort. In an unwritten handoff, the party who does more work gets no credit and the party who does less receives no signal. Marketing generates leads and hears nothing about what happened to them; sales receives leads of variable quality and has no mechanism to say so. Both conclude the other side is the problem. A bidirectional SLA with a rejection path gives each side a channel that produces data instead of resentment.

The third is prioritisation. When everything is urgent, nothing is. A stated standard tells a team what to do first when they cannot do everything, which is the normal condition of most teams.

4. Why SLAs Are Crucial: The Mechanisms

It is worth being specific about how an SLA produces improvement, because the mechanisms are what you are actually buying and each can be defeated.

Mechanism one: it makes a standard checkable. Before an SLA, performance at a handoff is an impression. After, it is a number with a denominator. This alone changes behaviour, because people respond to what is measured and visible far more reliably than to what is merely requested.

Mechanism two: it converts a tail into a target. Most handoff problems are not average problems; they are tail problems. The records that wait two days are the ones that produce lost deals and complaints. Expressing the SLA as a percentile — 90% within one hour — makes the tail the object of management rather than an anecdote.

Mechanism three: it forces definitions. You cannot write "response within one hour" without deciding what a response is, what starts the clock, and what happens outside business hours. That definitional work is most of the value, and teams routinely discover during it that the two sides had incompatible mental models all along.

Mechanism four: it creates a legitimate escalation path. Without an agreed standard, escalating is a political act. With one, escalation is procedural: the standard was missed, the documented step follows. This removes an enormous amount of friction from cross-team work.

Mechanism five: speed itself compounds. In lead handling specifically, response latency and contact rates are strongly related — the well-documented pattern is that contact and qualification rates fall sharply as time to first response extends, which is why first-response SLAs are the most common internal SLA and usually the one with the clearest return.

  • Makes a standard checkable — performance becomes a number with a denominator.
  • Targets the tail via percentiles rather than flattering averages.
  • Forces definitional work, which is where most of the value is created.
  • Makes escalation procedural rather than political.
  • In lead handling, response latency has a direct and well-documented effect on contact rates.

5. The Anatomy of a Complete SLA

A complete SLA has ten components. The first four are mandatory; the remaining six are what separate an agreement that survives from one that is argued about.

One — Scope. Which records, requests or services this covers, and explicitly which it does not. Exclusions written up front prevent the most common dispute, which is whether an incident was in scope at all.

Two — Trigger. The precise event that starts the clock, named as a system event where possible: record created in the CRM, ticket status set to New, order confirmed by the payment provider. "When we receive it" is not a trigger, because two systems will disagree about when that was.

Three — Standard. The number, with its statistical form. Not "one hour" but "90% of records within one business hour". State the percentile, the measurement window and the unit of time.

Four — Consequence. What happens on a breach. For customer-facing SLAs this is typically a service credit; for internal ones, escalation, reassignment and reporting. State it explicitly even when it is mild.

Five — Coverage calendar. Business hours, timezone, weekends, public holidays and the authoritative source for the holiday list. This single omission causes more SLA disputes than any other.

Six — Measurement method and system of record. Which system's timestamps are authoritative, how the metric is computed, and who publishes the report. When two systems disagree, the agreement must already say which one wins.

Seven — Obligations on both sides. What the receiving party must do for the commitment to hold — provide required information, accept or reject within a window, keep contact routes current. A one-directional SLA is not an agreement.

Eight — Exclusions and force majeure. Named circumstances where the clock does not apply, defined narrowly. Broad exclusions hollow out the agreement.

Nine — Review cadence. When the targets are re-examined against actual performance, and the process for changing them.

Ten — Edge case handling. Pauses, reassignment, reopening, bulk arrivals and duplicates — covered in detail below, because this is where agreements actually fail.

6. How an SLA Helps: What Changes in Practice

For the receiving team. Clear prioritisation when capacity is short, a defensible reason to push back on out-of-scope work, and protection from the expectation of instant response on everything. A well-written SLA is as much a shield as an obligation, and framing it that way is usually what converts resistance into adoption.

For the delivering team. Visible evidence of performance, and — through the rejection path — a feedback channel that produces data. Marketing teams operating under a bidirectional SLA learn which sources produce records sales rejects, which is information no volume report contains.

For management. A leading indicator. Response-time distributions degrade before revenue does, so an SLA dashboard gives warning months ahead of a pipeline shortfall. It also converts a recurring qualitative argument between two functions into a recurring quantitative review, which is a substantial saving in organisational energy.

For customers, where the SLA is external. A predictable experience and a credible commitment. The reputational value of a published SLA comes almost entirely from meeting it; publishing a target you routinely miss is worse than publishing nothing, because it converts a disappointment into a broken promise.

One effect worth naming because it is easy to miss: an SLA changes what gets built. Once first-response time is measured, the routing rules, notification design and queue tooling all become visibly consequential. Teams that have never measured it typically discover that a large share of their latency is mechanical — records routed to an absent owner, notifications nobody receives — rather than behavioural. Mechanical latency is cheap to fix, which is why the first months under a new SLA usually produce the largest gains.

7. Pros and Cons: An Honest Assessment

SLAs are widely recommended and less widely examined. Both columns are real.

The advantages. They remove ambiguity at handoffs, which is their whole purpose. They make performance visible and comparable over time. They give escalation a procedural basis. They force definitional clarity that has value independent of the agreement. They protect the receiving team as much as they bind it. And for lead handling specifically, they attack a latency problem with a well-established relationship to conversion.

The costs, which are usually understated. Measurement overhead is real — somebody instruments, reports and maintains this, and if that work is unowned the numbers degrade until nobody trusts them. Gaming is a genuine risk: any measured standard creates an incentive to satisfy the measure rather than the intent, and a first-response SLA met with an automated acknowledgement is the classic example. Rigidity can misallocate effort, since a single standard applied across records of very different value will push a team to answer a trivial request ahead of a significant one. Documentation drag means the agreement must be maintained as the business changes, and a stale SLA is actively misleading. And there is a real cultural risk: introduced as a compliance instrument rather than a coordination one, an SLA reads as surveillance and produces defensive minimum-compliance behaviour.

When an SLA is the wrong tool. When the underlying problem is capacity, not coordination — an SLA on a team that is structurally understaffed converts a resourcing problem into a performance problem, and the numbers will simply document the shortfall while morale falls. When volume is too low to produce a meaningful distribution, since a percentile over eleven records a month is noise. When the process is still changing weekly, because you will be rewriting the standard faster than you can measure against it. And when there is no genuine intent to act on a breach, in which case writing one teaches the organisation that stated standards are optional, which is worse than the original ambiguity.

  • Pros: removes handoff ambiguity, makes performance comparable, procedural escalation, forces definitions, protects the receiving team.
  • Cons: measurement overhead, gaming risk, rigidity across unequal records, maintenance drag, surveillance perception.
  • Wrong tool when: the constraint is capacity, volume is too low for percentiles, the process is still changing, or nobody intends to act on breaches.

8. How to Implement an SLA Properly

The SLA path, and what each edge case does to the clock

A six-stage SLA flow: record created, routed to owner, first response due, accepted or rejected, resolution clock, resolved. Four edge cases change how the clock behaves. Arriving outside business hours suspends the clock until the next opening, so a record created at 21:00 against a one-hour target is due one hour into the next business day. An unavailable owner triggers automatic reassignment while elapsed time carries across, so reassignment cannot be used to reset a breach. Waiting on the customer pauses the clock, which must be capped and reported or the pause becomes a way never to breach. A resolved record that is reopened within a stated window reattaches to the original record under a shortened secondary target, so repeated failures on one issue cannot each count as a fresh success.

The sequence below front-loads the cheap work and refuses to set a number until there is evidence for it.

Step one: measure before you commit. Instrument the handoff and observe it for a period long enough to see its distribution — commonly four to six weeks. You are looking for the median, the 90th percentile, the shape of the tail and the mechanical causes of the worst cases. Setting a target before this is guessing, and the guess is almost always either trivially easy or unreachable.

Step two: define the trigger and the response precisely, in terms of system events. Write down what specifically counts as a response. If an automated email counts, say so and expect to get automated emails; if it must be a human contact attempt logged against the record, say that instead.

Step three: set the target from the observed distribution. A defensible first target sits slightly beyond current performance — a common approach is to take roughly the current 75th percentile and make it the 90th percentile commitment. Ambitious enough to require change, close enough to be reached. Then improve it on a schedule rather than setting the eventual goal on day one.

Step four: write the coverage calendar. Hours, timezone, weekends, holidays, and the authoritative holiday source. If you support multiple regions, state whether coverage follows the customer's calendar or yours.

Step five: make it bidirectional. Write the receiving side's obligations, including the window in which they must accept or reject and the requirement that rejections carry a reason from a controlled list. Without this, the agreement cannot be measured in both directions and will be resented.

Step six: define the consequence and the escalation ladder. Who is notified at what point, what reassignment happens automatically, and what review is triggered by a pattern rather than by an incident. Distinguish the two: single breaches need mechanical handling, patterns need a review.

Step seven: instrument it in the system of record before launch, not after. If your CRM cannot compute the metric natively, decide now where it will be computed. Retrofitted measurement produces disputes about numbers in the first month, which is exactly when the agreement most needs credibility. Your platform's ability to do this is one of the selection criteria in the [CRM guide](/guides/choose-best-crm-guide).

Step eight: pilot on a segment. Run it on one source, region or product first. Pilots surface the edge cases cheaply, and edge cases discovered in a pilot are design inputs rather than disputes.

Step nine: publish the report on a fixed cadence, including the paused-time and exclusion counts. Reporting only the headline compliance percentage invites the gaming described above; reporting it alongside how much time was paused and how many records were excluded keeps it honest.

Step ten: review on schedule against actual performance, and change targets deliberately rather than quietly. A target that has been met at 99% for two quarters is no longer doing any work and should be tightened or retired.

9. Choosing the Numbers: Setting Defensible Targets

Target-setting is where most SLAs go wrong in a way that is hard to recover from, because a discredited target is difficult to re-legitimise.

Use percentiles, not averages. An average conceals the tail, and the tail is the experience that generates complaints and lost deals. State the percentile explicitly: 90% or 95% within the target. Reserve 99% and above for genuinely critical services, because the cost of the last few percent rises steeply and is usually paid in staffing.

Derive from observation, not aspiration. Your own measured distribution is the only defensible starting point. Competitor published SLAs are marketing artefacts and tell you nothing about their actual performance or their cost structure.

Segment by value where the records genuinely differ. A single standard across records of very different value forces a team to treat them identically, which is not what you want. Two or three tiers with different targets is usually enough; more than that becomes unmanageable.

Make the target reachable with the current mechanism. If the observed 90th percentile is eleven hours and mechanical fixes could plausibly bring it to four, do not commit to one. Commit to four, ship the mechanical fixes, then tighten.

Separate response from resolution. These are two different clocks with different drivers. Response time is almost entirely mechanical — routing, notification, capacity. Resolution time depends on the work itself. Combining them into one number produces a metric that cannot be acted on, because you cannot tell which half moved.

Finally, state the denominator. "90% of qualified inbound enquiries during published business hours, excluding duplicates and spam" is a target you can compute. "90% of leads" is not, because nobody agrees what counts.

  • Percentiles, never averages — the tail is the experience that matters.
  • Derive the target from your own observed distribution.
  • Segment by record value, but keep it to two or three tiers.
  • Separate the response clock from the resolution clock.
  • State the denominator precisely, including exclusions.

10. How to Handle Edge Cases in an SLA

Edge cases are where agreements fail. A document covering only the standard path will be argued about within the first month, and each unresolved argument costs the SLA credibility. Handle these explicitly, in writing, before launch.

Arrivals outside coverage hours. Decide whether the clock suspends outside business hours or runs continuously, and state the calendar, timezone and holiday source. Suspension is the usual choice for internal SLAs, because measuring a team against hours nobody staffed is meaningless. State the rule concretely: a record created at 21:00 against a one-hour target is due one hour into the next business day.

Paused clocks and customer-waiting states. This is the most abused mechanism in any SLA. Pausing while genuinely waiting on the customer is correct — you should be measured on time you controlled. But an uncapped pause state becomes a way never to breach. Three controls make it safe: cap the total paused time per record, require a structured reason to enter the state, and report paused duration alongside elapsed time so the pause is visible rather than silent.

Reassignment and absent owners. Reassign automatically after a stated idle period, and carry elapsed time across to the new owner. The common defect is restarting the clock on reassignment, which makes the metric look healthy while the customer waits twice as long. Never let an ownership change reset a measurement.

Reopened records. Define a reopen window during which a returning issue reattaches to the original record rather than creating a new one, and apply a shortened secondary target. Without this rule, closing and reopening resets the measurement, so repeated failures on one issue each count as fresh successes.

Bulk arrivals and spikes. A campaign, an outage or a data import can deliver a month of volume in an afternoon. Decide in advance whether spikes are excluded, whether the target degrades gracefully at defined volume thresholds, or whether an overflow path activates. Silence here means the SLA will be breached by a foreseeable event and then quietly ignored, which sets a precedent.

Duplicates, spam and misrouted records. State how they are identified and that they are excluded from the denominator, and report the exclusion count. Unreported exclusions are indistinguishable from gaming, even when they are entirely legitimate.

Partial and ambiguous responses. Define what counts. If an automated acknowledgement satisfies first response, you will get automated acknowledgements and your measured performance will improve while the customer experience does not. Most internal SLAs should require a logged human contact attempt.

Disputed timestamps. Name the authoritative system in the agreement. When two systems disagree — and with webhooks, retries and timezone handling they will — the agreement must already say which one wins, or every close call becomes a negotiation.

Dependency failures. Where meeting your commitment depends on a third party, either exclude their downtime explicitly or accept the risk knowingly. This is exactly where an unread vendor SLA becomes your problem, and where the two instruments need to be read together.

  • Out of hours: state the calendar, timezone and holiday source, and whether the clock suspends.
  • Pauses: cap total paused time, require a structured reason, report paused duration.
  • Reassignment: carry elapsed time across; never reset the clock on ownership change.
  • Reopening: define a window that reattaches to the original record with a shortened target.
  • Spikes: decide in advance between exclusion, graceful degradation, or an overflow path.
  • Exclusions: define duplicates and spam handling, and always report the exclusion count.
  • Timestamps: name the authoritative system before you need it.

11. Things to Keep in Mind Before You Implement

A short pre-flight list. Each item corresponds to a failure we have watched happen.

Confirm the constraint is coordination, not capacity. If the team cannot meet the target at any reasonable effort because there are not enough of them, an SLA will document that fact repeatedly and change nothing. Fix the resourcing or set a target the current team can meet.

Check you have the volume for the statistics. Percentile targets need enough records for the percentile to mean something. Below roughly thirty records per reporting period, report the raw distribution instead and revisit later.

Secure genuine agreement from both sides, in writing, from the people accountable. An SLA imposed by one function on another has a short life. The signature that matters is the one from the team that has to meet it.

Decide who owns the measurement before launch, by name. Unowned metrics decay, and a decayed SLA metric is worse than none because decisions continue to be made on it.

Plan the framing. Introduce it as a coordination mechanism that protects both sides, and be specific about what the receiving team gets — clearer priorities, a legitimate basis for pushing back, a rejection path that produces data. Introduced as monitoring, it produces minimum-compliance behaviour.

Write the edge cases before launch, not after the first dispute. Every one you handle in advance is a design decision; every one you handle afterwards is a negotiation conducted under pressure with someone's performance at stake.

Decide what you will do when it is met. A target consistently met at 99% has stopped doing work. Plan to tighten it, retire it, or move the measurement somewhere more useful.

Agree the review cadence and the change process up front, so that adjusting a target is a scheduled decision rather than an admission of failure.

12. The Marketing-to-Sales SLA: A Worked Example

The most valuable internal SLA in most commercial organisations, and the one most often missing. The figures below are an illustrative model, not client data.

The marketing commitment. Deliver an agreed volume of qualified records per month — say 500 — meeting a documented fit standard, with required fields populated and source attribution attached. Quality is measured by acceptance rate, with a floor: at least 70% of delivered records accepted by sales.

The sales commitment. First contact attempt within one business hour for 90% of records, with a documented minimum of five contact attempts across at least two channels over ten business days. Accept or reject within one business day, with rejections carrying a reason from a controlled list.

The shared definitions. What counts as qualified, using the fit-plus-behaviour standard rather than a single action. What counts as a contact attempt. What the rejection reasons are, as a short controlled vocabulary — out of territory, no budget, wrong role, duplicate, bad data, not in market. These definitions are the substance of the agreement; see the qualification section of the [funnel stages guide](/guides/funnel-stages-guide) for how to write them.

The consequences. Missed volume triggers a joint review of channel performance rather than a marketing-only post-mortem. Missed response time triggers automatic reassignment and appears on a weekly report. Acceptance rate falling below the floor for two consecutive periods triggers a joint review of the qualification standard — not a marketing problem or a sales problem, but a definition problem.

The edge cases this specific SLA needs. Records arriving Friday evening. Records arriving during a campaign spike well above the committed volume. Records that are duplicates of an existing open opportunity. Records rejected by sales that marketing disputes. Reassignment when the owning rep is on leave. Each needs a written rule.

What makes this work in practice is the rejection path. Marketing gets structured reasons instead of silence, and reasons aggregated by source become an acquisition decision — which is how an SLA quietly becomes an input to [channel budget allocation](/guides/channel-budget-allocation-guide) rather than merely a compliance report.

13. Customer-Facing SLAs: What Differs

External SLAs share the mechanics and add legal and commercial weight. Four differences matter.

The remedy is money. Service credits are the standard instrument, usually as a percentage of the period's fee scaled to severity. Two design points: credits should be meaningful enough to matter and bounded enough not to threaten the business, and the claim process should be documented — a credit that requires the customer to discover and claim it is a weaker commitment than one applied automatically.

Uptime arithmetic is unforgiving, and worth stating plainly because it is routinely underestimated. Over a thirty-day month, 99% allows roughly 7.2 hours of downtime, 99.9% allows about 43 minutes, and 99.99% allows about 4.3 minutes. Each additional nine costs substantially more in architecture and staffing than the last. Commit to the number you can actually operate, and define whether maintenance windows are excluded.

Definitions become contractual. What counts as downtime, how it is detected, who measures it, and whether partial degradation counts. These are negotiated terms with financial consequences, so they need legal review — the general principles here are not a substitute for it.

Your dependencies cap your commitment. You cannot credibly promise availability higher than the weakest dependency in your critical path. Read the vendor SLAs underneath you before publishing yours, and either exclude their failures explicitly or accept that you are absorbing their risk.

  • Credits should be meaningful, bounded, and ideally applied without requiring a claim.
  • Over 30 days: 99% ≈ 7.2 hours down, 99.9% ≈ 43 minutes, 99.99% ≈ 4.3 minutes.
  • Downtime definitions are contractual terms — get legal review.
  • You cannot promise more availability than your weakest dependency provides.

14. Common Mistakes, and What to Do Instead

Setting the target before measuring. Instead, instrument for four to six weeks and derive the target from the observed distribution.

Using averages. Instead, commit at a percentile and report the distribution, because the tail is what people experience.

Writing a one-way commitment. Instead, state both sides' obligations, including a rejection path with required reasons.

Leaving the calendar unstated. Instead, name hours, timezone, weekend treatment and the authoritative holiday source.

Allowing uncapped pause states. Instead, cap paused time, require a structured reason, and report paused duration next to elapsed time.

Restarting the clock on reassignment. Instead, carry elapsed time across owners so reassignment cannot mask a breach.

Accepting automated acknowledgements as responses. Instead, define response as a logged human contact attempt unless you genuinely intend otherwise.

Combining response and resolution into one number. Instead, run two clocks, because they have entirely different drivers.

Never acting on breaches. Instead, define mechanical handling for single breaches and a review trigger for patterns — and if you do not intend to act at all, do not write the SLA.

Never revisiting the target. Instead, review on a fixed cadence; a target met at 99% for two quarters has stopped doing any work.

15. An Implementation Checklist

Before launch, every line below should have an answer written down. Any that cannot be answered is a dispute waiting to happen.

Scope: which records are covered, and which are explicitly excluded. Trigger: the precise system event that starts each clock. Standard: the number, its percentile and its measurement window. Denominator: what is counted, including duplicate and spam handling.

Calendar: hours, timezone, weekends, holidays and the authoritative source. Clocks: whether response and resolution are separated, and the pause rules with their cap. Ownership: routing, automatic reassignment rules, and elapsed-time carry-over.

Both sides' obligations: what the receiving party must do, the accept-or-reject window, and the controlled list of rejection reasons. Consequences: mechanical handling for single breaches, review triggers for patterns.

Measurement: the authoritative system, how the metric is computed, who publishes it and on what cadence, and whether paused time and exclusion counts are published alongside the headline figure.

Edge cases: out-of-hours arrivals, spikes, reopened records, disputed timestamps and dependency failures. Governance: review cadence, change process, and the named owner of the measurement.

16. Putting It Together

An SLA is a promise with a clock, a measurement method and a consequence. The promise is the easy part. The clock requires you to define your trigger precisely, the measurement requires you to name an authoritative system and publish honestly, and the consequence requires you to decide in advance what happens when the standard is missed.

Most SLAs that fail do so for one of three reasons, all avoidable. The target was set from aspiration rather than observation, so it was never reached and became decorative. The edge cases were left unwritten, so the first month produced disputes that cost the agreement its credibility. Or nothing happened on a breach, so the organisation learned that stated standards are optional.

The highest-return SLA for most commercial teams is not a customer-facing one. It is the marketing-to-sales handoff, because that is where the largest unmeasured leak sits in most funnels, and because the definitional work it forces — what qualified means, what a contact attempt is, why records get rejected — improves the funnel whether or not the agreement is ever enforced.

If you want the handoff instrumented and the targets derived from your own data rather than from a benchmark, that diagnostic is where our [business operations](/solutions/business-ops) engagements start.

Frequently Asked Questions

What is an SLA?
A service level agreement is a written commitment that a defined service will be delivered to a measurable standard within a stated time, with an agreed measurement method and a defined consequence when the standard is missed. All four elements are required — remove any one and it becomes a statement of intent rather than an agreement.
Why is an SLA important?
An SLA removes ambiguity at a handoff, which is where most cross-team failure occurs. It makes performance checkable rather than impressionistic, targets the tail of the distribution rather than the flattering average, forces definitional clarity between the two sides, and turns escalation from a political act into a procedural one.
What are the disadvantages of an SLA?
Measurement overhead that decays if unowned; gaming risk, since any measured standard invites satisfying the measure rather than the intent; rigidity when one standard is applied to records of very different value; maintenance drag as the business changes; and a cultural risk that it reads as surveillance and produces defensive minimum-compliance behaviour.
How do I set SLA targets?
Instrument the handoff and observe it for four to six weeks first, then derive the target from the observed distribution. A defensible first target takes roughly the current 75th percentile and commits to it at the 90th percentile — demanding enough to require change, close enough to be reached. Tighten on a schedule rather than committing to the eventual goal on day one.
Should SLAs use averages or percentiles?
Percentiles. An average response time of two hours is consistent with most records answered in minutes and a tail waiting two days, and the tail is what customers and reps actually experience. Commit at the 90th or 95th percentile and reserve 99% and above for genuinely critical services, since the last few percent are paid for in staffing.
How do you handle SLA edge cases?
Write them before launch. State the coverage calendar and whether the clock suspends out of hours; cap and report paused time; carry elapsed time across reassignment rather than restarting; define a reopen window that reattaches to the original record; decide in advance how volume spikes are treated; exclude duplicates and spam but publish the exclusion count; and name the authoritative system for disputed timestamps.
What is a marketing to sales SLA?
A bidirectional internal agreement where marketing commits to a volume of records meeting a documented fit standard, and sales commits to a first-contact time, a minimum contact sequence, and accepting or rejecting within a window with a reason from a controlled list. It is usually the highest-return internal SLA and the one most often missing.
What happens when an SLA is breached?
Whatever the agreement says will happen — which is why it must say something. Distinguish single breaches, which need mechanical handling such as automatic reassignment and appearance on a report, from patterns, which need a joint review of the underlying mechanism. If nothing at all happens, the organisation learns that stated standards are optional.
Do internal SLAs need consequences if no money changes hands?
Yes, but the consequence does not have to be financial. Escalation, automatic reassignment and visibility on a published report are sufficient for most internal agreements. What matters is that a breach produces something other than a conversation, because a standard with no consequence will be ignored the first time meeting it is inconvenient.
When should you not use an SLA?
When the real constraint is capacity rather than coordination, since the SLA will document the shortfall and change nothing; when volume is too low for percentiles to be meaningful; when the process is still changing weekly; and when there is no genuine intent to act on breaches, in which case writing one teaches the organisation that stated standards are optional.
What uptime does 99.9% actually allow?
Over a thirty-day month, 99% allows roughly 7.2 hours of downtime, 99.9% about 43 minutes and 99.99% about 4.3 minutes. Each additional nine costs substantially more in architecture and staffing than the last, and you cannot credibly promise more availability than the weakest dependency in your critical path provides.