Insights  /  AI service desk for MSPs: from ticket triage to verified resolution

Insights

AI service desk for MSPs: from ticket triage to verified resolution

Insights By The Helios team  ·  7 min read

An AI desk that only sorts tickets into categories is a filing clerk with a subscription fee. Useful, occasionally, but nowhere near the thing being sold. The phrase "AI service desk for MSPs" currently covers five different capabilities of wildly different value, and vendors have every incentive to blur them together, because the cheap rungs of the ladder demo just as well as the expensive ones. If you run a service desk and you are judging this tooling day to day, you need a way to tell them apart, a set of guardrails before any of it touches a client machine, and a measurement that survives contact with a real queue. Here is all three, rung by rung.

The maturity ladder: five rungs sold under one label

Rung one: categorise

The AI reads an inbound ticket, tags it, sets a priority and routes it to the right queue. This is real but modest. It saves a dispatcher a few seconds per ticket and it gets categories right more consistently than a rushed human. It resolves nothing. If a vendor's demo spends most of its time here, that tells you where the product spends most of its time too.

Rung two: draft

The AI writes a suggested reply for a technician to review and send. This saves typing, not thinking. The technician still has to know whether the answer is correct, which means the ticket still consumes a technician's attention. Drafting is worth having, but it moves the tickets-per-technician number by very little, because reviewing a wrong draft costs about as much as writing a right one.

Rung three: answer with sources

The AI answers the user directly, and, crucially, cites where the answer came from: your documentation, the client's runbook, a knowledge base article. This is the first rung with genuine deflection, and the citation is not decoration. An answer with a source can be checked in seconds. An answer without one has to be trusted, and trust is exactly what you cannot extend to a system that is sometimes confidently wrong. No sources, no rung three, whatever the marketing says.

Rung four: act on the device

The AI runs diagnostics on the endpoint, forms a hypothesis and applies a fix: restarts a service, clears a corrupt cache, repairs a broken update component. This is where the service desk stops being a conversation layer and starts being an operations tool, and it is where the risk arrives. We have written separately about what autonomous remediation really does and what it should never do; the short version is that action without constraint is not a feature, it is a liability with good presentation.

Rung five: verify the fix

After acting, the AI re-runs the check that failed, confirms the symptom is gone, and only then closes the ticket, with the evidence attached. This is the rung that separates a working system from a noisy one, because a system that closes tickets without verification is not resolving anything, it is optimising for a closure metric.

A closed ticket is a claim. A verified fix is evidence.

What an AI service desk for MSPs must prove before it acts

Rungs four and five involve an autonomous system executing changes on machines you are contractually responsible for, whether the machines belong to clients or to your own company. Before you allow that in a trial, let alone production, demand the following. We cover the full framework in AI guardrails for IT automation, but these are the non-negotiables:

  • Per-client approval scopes. The law firm and the skip hire company do not carry the same risk. You need to set, per client, which categories of action run automatically and which wait for a human. Skip this and one global setting will be tuned to your most nervous client, or your least, and both are wrong.
  • An allowlist, not a blocklist. The AI should only be able to run actions you have explicitly permitted. A system that can do anything except what you have forbidden will eventually find the gap in your imagination.
  • A visible plan before execution. For anything gated on approval, you should see what the AI intends to do, in plain steps, before it does it. Approving a summary you cannot inspect is a rubber stamp wearing a safety costume.
  • A complete audit trail. Every action, every command, every output, timestamped and attributable. When a client asks what happened on their server at 2am, "the AI handled it" is not an answer you can invoice against.
  • A kill switch that actually kills. One control that stops all autonomous action, per client and globally, that takes effect immediately. Test it during the trial, not during the incident.

Rule of thumb: the AI earns autonomy per action type, not per product. Start everything on approval, promote an action to automatic only after you have watched it succeed repeatedly, and never promote a category you have not personally reviewed.

Measuring whether it reduces tickets per technician

The only number that ultimately matters is tickets resolved per technician per month, measured against a baseline you record before the trial starts. The arithmetic is deliberately simple. If three technicians handle 600 tickets a month, your baseline is 200 per technician. If the AI genuinely resolves and verifies 90 of those without human touch, you are at roughly 170 per technician, and you can feel that difference in a fortnight.

Two distortions will try to flatter the number. The first is deflection theatre: a portal bot that makes tickets vanish by frustrating users until they ring instead, which moves the work off the dashboard and onto the phone. Measure inbound demand across every channel, not ticket count alone. The second is premature closure: tickets the AI marks resolved that reopen within days. Track the reopen rate on AI-closed tickets separately, and compare it against your human reopen rate. If the AI's is higher, rung five is missing, whatever the feature list says. Your SLA response times should improve as a side effect; if they do not, the AI is answering the easy tickets your team already answered quickly.

Acceptance criteria you can paste into a trial plan

Run the trial on a real queue for at least two weeks, with a baseline month measured first. Then hold the product to criteria like these:

  1. Categorisation accuracy: the AI's queue and priority assignment matches what your dispatcher would have chosen on at least nine tickets in ten, checked against a sample you review by hand.
  2. Sourced answers only: every direct answer to a user cites its source, and a spot check of twenty answers finds no confident fabrication. One invented fact fails the trial, because you will not catch the second one.
  3. Approval flow works under pressure: a gated action shows its full plan, waits for a named human, and logs who approved it. Test this on a Friday afternoon, not a quiet Tuesday.
  4. Verified resolutions: AI-closed tickets include evidence the original symptom is gone, and their reopen rate is no worse than your team's.
  5. The number moves: tickets requiring human touch drop measurably against baseline, with inbound demand across all channels flat or accounted for.

Near the end of any trial, the failure modes are more instructive than the successes. A high closure rate with a high reopen rate means the AI is grading its own homework. A pile of accurate drafts nobody sends means the team does not trust it, and they are usually right before they are wrong. A perfect record on an allowlist of two actions means the vendor tuned the demo, not the product.

Where this fits with Helios

Helios includes an AI agent, Helio, that works the full ladder: it triages inbound tickets, answers with sources from your documentation, investigates device alerts, and writes and verifies fixes under per-client approval guardrails, with every action logged. The verification rung is the one we built the rest around, because we run a live queue against it daily and unverified closures poison the metric that pays our wages too. Much of what this article describes is discipline rather than tooling: the baseline measurement and the promotion policy are yours to enforce, whatever platform you run.

Helios is a flat-rate RMM and PSA for small MSPs and internal IT teams, from £99 a month with every feature on every plan. 14-day trial, no card, no feature gating. Start free.

Hold your own house to your clients' standard

Helios is an AI-native platform for MSPs and in-house IT teams: monitoring, patching, security and service desk in one place, with a 14-day trial and no feature gating.

See how Helios works

Read next

Insights RMM pricing in pounds: what UK MSPs actually pay at 100, 250, 500 and 1,000 endpoints Insights Leaving NinjaOne: how to migrate off it without losing scripts, policies or endpoints Insights RMM software in the UK: GBP pricing, VAT and support hours that match your day