Insights  /  MSP SLA response times: setting targets you can actually hit

Insights

MSP SLA response times: setting targets you can actually hit

Insights By The Helios team  ·  12 min read

Most MSP SLA response times were copied from someone else's contract. They looked reassuring in the proposal, nobody has measured them since, and the first person to read them closely is an unhappy client with a lawyer. This guide covers what response actually means, sensible targets by priority, how to handle penalties and after-hours cover, and how to set numbers your service desk can hit every month, not just in a good week. It applies whether the machines belong to clients or to your own company.

Response is not resolution, and both need defining

The single most common SLA failure is not a missed target. It is a contract where "response" was never defined, so the MSP and the client are measuring two different things in good faith.

Pin down three separate clocks:

  • First response: a qualified human has read the ticket, triaged it and told the client what happens next. An automated "we got your email" receipt does not count, and clients notice when it is passed off as one.
  • Resolution: service is restored or a workaround is in place. Not "root cause fully understood", which can take much longer and belongs in a post-incident note, not the SLA clock.
  • Updates: how often the client hears from you while a ticket is open. On a P1 that might be every 30 minutes. Silence during an outage does more contract damage than the outage itself.

Write all three into the agreement, per priority level. If your PSA or service desk cannot report on all three, that is a tooling gap to fix before you sign anything.

Typical MSP SLA response times by priority

Numbers vary by market and price point, but most healthy MSP agreements land close to this shape:

  • P1, critical: a whole site or core system is down. First response in 15 to 30 minutes, work continues until service is restored, updates at least every 30 to 60 minutes.
  • P2, major: a team or key function is degraded. First response within 1 to 2 hours, resolution targeted the same business day.
  • P3, minor: one user affected, work continues with friction. First response within 4 business hours, resolution by the next business day.
  • P4, request: new starters, changes, questions. First response within 1 business day, fulfilment within 3 to 5.

Two caveats before you lift these into a contract. First, they assume business-hours cover; if you sell 24/7, the P1 line is the expensive one, so price it deliberately. Second, these are ceilings, not ambitions. If your desk routinely responds to P3s in 20 minutes, keep the 4-hour promise anyway. The gap between what you promise and what you deliver is your margin for the bad week, and every desk has bad weeks.

Rule of thumb: set the contractual target at roughly double your real-world average. If you consistently hit first response on P2s in 25 minutes, promise an hour. You will beat the SLA month after month, and the report that proves it becomes a retention tool rather than a liability.

The bad-week test: take your worst week from last quarter, the one with two technicians off and a P1 running, and ask whether the desk would still have hit the target you are about to sign. If the answer is no, it is not a target. It is a hope with a signature on it.

Is there an industry standard SLA? Not really, and that is fine

Prospects will ask what the industry standard response time is, usually because a competitor's proposal quoted one. There is no standard body, no benchmark and no regulation setting MSP response times. What exists is a market shape: the ranges above are what mid-market managed service agreements tend to converge on, because they balance what clients will pay against what a desk of realistic size can sustain.

That absence of a standard is useful to you. It means the honest answer to "is 15 minutes normal?" is "normal for a price point". A 15-minute first response on every priority is achievable, but only with staffing levels that roughly double the cost of the desk, and a client who wants it should see that number, not a discount. Anchor the conversation to your measured performance and your price, not to a folk standard that nobody can cite.

Priority must come from impact, not volume

An SLA is only as good as the triage behind it. If priority is set by whoever shouts loudest, every ticket from the shouty client becomes a P1, your team burns out chasing false urgency, and the genuinely critical ticket from a quiet client waits.

The fix is a simple, written impact and urgency matrix: how many people are affected, and can they work at all? A payroll server down on the 28th of the month is a P1 even if the ticket arrived politely. A director's second monitor is a P3 even if it arrived in capitals. Put the matrix in the client agreement too, so priority is a shared definition rather than a negotiation on every ticket.

A workable matrix needs only two axes:

  • Impact: who cannot work: the whole organisation, a team or function, or one person.
  • Urgency: is work stopped entirely, degraded with a workaround, or merely inconvenienced?

Cross them and the priorities fall out: organisation stopped is a P1, a team stopped or the organisation degraded is a P2, one person degraded is a P3, and anything that is a request rather than a fault is a P4 regardless of who asked. Two axes and four outcomes fit on half a page, and half a page that both sides have signed ends most priority arguments before they start.

Triage quality is also where alert-driven tickets earn their keep or destroy it. If your monitoring floods the desk with noise, technicians stop trusting priorities altogether, and the SLA clock runs on tickets nobody should be seeing. We covered how to fix that at the source in our guide to cutting MSP alert noise.

24/7 cover: price the P1 line like the commitment it is

Business hours are roughly 40 of the 168 hours in a week. A 24/7 SLA commits you to the other 128, and the arithmetic is unforgiving: keeping one seat genuinely staffed around the clock takes roughly five people once shift patterns, leave and sickness are counted. Signing a 24/7 P1 response without that maths done is how out-of-hours cover becomes a founder's phone on the bedside table.

Three honest ways to sell after-hours cover, in ascending order of cost:

  • Stacked calendars. P1 tickets get a 24/7 clock, everything else runs on business hours. A P3 logged at midnight starts its clock at 09:00. This is the most common shape because it covers the risk the client actually fears, the outage at 2am, without staffing for password resets at 3am.
  • On-call, not on-shift. Out of hours, a P1 gets a response from an on-call engineer within a longer window, say 60 minutes rather than 15. Write the different number into the contract explicitly. A single response time that silently assumes business hours is the clause a lawyer will read aloud.
  • A staffed overnight desk, your own or outsourced. Only viable at scale, and it should carry a price that reflects five-technicians-per-seat arithmetic, not a 20 per cent uplift picked to close the deal.

Skip this and your first out-of-hours P1 becomes a contractual breach, an exhausted engineer and a renewal conversation, all in the same week.

The mechanics that quietly break SLA reporting

Even sensible targets fall over on the details of how the clock runs. Four mechanics to get right in your tooling:

  1. Business-hours calendars. A P3 logged at 16:55 on Friday should not breach at 09:05 on Monday because the clock ran all weekend. Every SLA needs a calendar, including client-specific holidays if you serve multiple regions.
  2. Pause states. When a ticket is waiting on the client or on a vendor, the clock should pause, and the status change should be logged so the report can prove it. Without this, your worst-looking breaches are tickets where you were waiting three days for a reply.
  3. Reassignment resets. Escalating a ticket between technicians must not reset first response. It was responded to once; the client does not care about your internal routing.
  4. Breach visibility before the breach. A report that shows last month's misses is an autopsy. What changes behaviour is a queue that shows tickets at 75% of their target while there is still time to act.

Penalties and service credits: what happens when you miss

Many MSPs treat penalty clauses as something to resist at all costs. We think that instinct is wrong. A capped, well-defined service credit regime is cheaper and safer than a vague "reasonable endeavours" clause, because it converts an argument about damages into a line item both sides agreed in advance. The key is to define the shape yourself before the client's lawyer does.

  • Measure attainment, not single tickets. A credit that triggers on any individual miss makes one anomalous ticket expensive. A monthly attainment threshold, for example a percentage of P1 and P2 tickets meeting their targets, tolerates the odd bad ticket while still punishing a bad month. Set the threshold below 100, because 100 is not a target, it is a trap.
  • Cap the exposure. Credits should be a defined percentage of that month's fee, with a monthly cap, and they should be the sole remedy for the miss. Uncapped SLA liability is how a service contract becomes an insurance policy you did not price.
  • Write the exclusions down. Time waiting on the client, faults in third-party services outside your control, agreed maintenance windows and tickets raised outside the defined channels all pause or void the clock. Every one of these must appear in the agreement, because they will certainly appear in the dispute.
  • Consider earn-back. Some agreements let credits lapse if the following quarter's attainment recovers. It aligns incentives: the client wants sustained performance more than they want a discount.

Skip this and the first serious breach becomes a negotiation about the entire contract value, conducted at the worst possible moment, which is immediately after you let the client down.

Internal IT teams: same clocks, different consequences

If the machines belong to your own company rather than to clients, the lawyer disappears but nothing else does. Internal SLAs, sometimes dressed up as OLAs, fail in exactly the same ways: response never defined, priority set by seniority rather than impact, and a clock that nobody can pause when finance sits on a reply for a week.

The differences are political, not technical. There is no contract to breach, so the penalty for a bad month is quieter and worse: the sales director who stops raising tickets and starts emailing the CTO directly, the department that buys its own kit because the service desk is "slow". Publish the same priority matrix, run the same calendars and pause states, and put attainment in front of leadership monthly whether or not anyone asked. An internal desk that reports on itself before being challenged gets to set the terms of the conversation. One that does not gets its terms set for it, usually in a budget meeting.

Report on it before your client asks

An SLA you do not measure is a marketing line, and one you only measure when challenged is a liability. Put SLA attainment on the monthly report and the quarterly review: percentage of tickets that met first response and resolution by priority, the misses, and what changed as a result. Bringing your own misses to the table, with causes, is disarming and builds far more trust than a suspiciously perfect scorecard.

Expectations around response times should also be set on day one of a new contract, alongside how tickets are raised and what counts as urgent. It is one of the steps in our client onboarding checklist, because an SLA explained in week one prevents the escalation call in month six.

The two ways an SLA fails

Every broken SLA we have seen falls into one of two failure modes, and they sit at opposite ends of the same axis:

  • The vanity SLA. Sales promised 15 minutes on everything because the prospect flinched at an hour. The desk breaches weekly, technicians stop looking at the breach report because it is always red, and the attainment slide quietly vanishes from the quarterly review. This is failure by ambition: the numbers were a pitch, never a plan.
  • The wallpaper SLA. The targets are so loose that nothing ever breaches, so nobody ever looks, so nobody notices when real performance drifts. The document exists purely as decoration until the day a lawyer reads it, at which point it turns out even the loose targets were never actually measured. This is failure by neglect.

Both modes end in the same place: the contract and reality describing two different service desks. The cure is also the same: measure real performance first, set targets from the measurement with headroom for the bad week, and report on them every month whether or not anyone asks.

Where this fits with Helios

The Helios service desk was built around the mechanics above: per-client SLA policies with business-hours calendars, pause states for waiting-on-client time, and countdown timers on the queue so technicians see what is approaching breach rather than what already breached. Helio, the AI layer, auto-triages inbound tickets against your impact matrix, which keeps priorities consistent when they arrive at 2am or in capital letters. Attainment then flows into the client portal and monthly reports without anyone assembling a spreadsheet. To be clear about the limits: most of this article is contract design and triage discipline, not tooling. Helios makes the clocks honest; it cannot make the targets sensible. That part is yours.

Helios is the platform for MSPs and IT teams who want SLAs they can measure, hit and prove. 14-day trial and no feature gating.

Start free

Hold your own house to your clients' standard

Helios is an AI-native platform for MSPs and in-house IT teams: monitoring, patching, security and service desk in one place, with a 14-day trial and no feature gating.

See how Helios works

Read next

Insights RMM pricing in pounds: what UK MSPs actually pay at 100, 250, 500 and 1,000 endpoints Insights Leaving NinjaOne: how to migrate off it without losing scripts, policies or endpoints Insights RMM software in the UK: GBP pricing, VAT and support hours that match your day