In the world of managed services, technical metrics are often treated like a scoreboard that only the engineers look at. We track uptime, ticket volumes, and response times, but there is one metric that sits at the very heart of your operational efficiency and your client’s perception of value: Mean Time to Resolve (MTTR).
Mean Time to Resolve (MTTR) is the average time it takes your team to fully resolve a technical issue, from the moment a ticket is opened or an alert is triggered until the service is fully restored and the ticket is closed. It is the definitive measure of how long a client is actually impacted by a problem.
Luis Navarro, who spent 15 years building Totality Services into a highly profitable MSP before a successful eight-figure exit, often noted that while "Mean Time to Respond" makes the SLA look good, Mean Time to Resolve is what keeps the client happy and your margins healthy.
Key Takeaways
- Definition: MTTR measures the full lifecycle of an incident, representing the true duration of downtime or decreased productivity for the client.
- Commercial Impact: Lowering MTTR directly increases labour profitability by reducing the "touches" required to fix a problem.
- Client Retention: High MTTR leads to "death by a thousand cuts"—clients feel unsupported even if your initial response was fast.
- Standardisation is Key: You cannot improve resolution times across a chaotic, non-standardised client stack.
- Beyond Technology: MTTR is often a communication and process problem, not just a technical one.
If you want to move from being a "firefighter" to a strategic partner, you have to master this metric. It isn't just about working faster; it’s about working smarter, documenting better, and ensuring your team has the right tools and authority to close issues without unnecessary escalations.
What is Mean Time to Resolve (MTTR)?
Mean Time to Resolve (MTTR) is a service level metric that calculates the average time elapsed between the initial report of a service failure and its complete resolution. Unlike response time, which only tracks how long it takes for a technician to say "hello," MTTR tracks how long it takes to actually fix the problem.
In a professional MSP environment, MTTR is calculated by taking the total time spent on resolutions during a specific period and dividing it by the number of resolved incidents.
The Core Components of the MTTR Clock:
- Identification: The time it takes for your system or the user to realise something is wrong.
- Triage & Assignment: The gap between the ticket appearing and the right resource starting work.
- Diagnosis: The investigation phase where the root cause is identified.
- Rectification: The actual technical work to fix the issue.
- Verification: Confirming with the user that the fix actually worked before closing the ticket.
MTTR vs. Other Common Metrics
It is easy to get lost in the "M-T-T-something" alphabet soup. To run a profitable MSP, you need to know which one actually drives the bottom line. While "Mean Time to Respond" is a legal requirement in most SLAs, it is often a vanity metric. If you respond in 5 minutes but take 5 days to fix the server, the client is still losing money.
| Metric | Focus | Client Perspective |
|---|---|---|
| Mean Time to Respond | Acknowledgment speed | "They heard me." |
| Mean Time to Resolve | Total downtime | "I can work again." |
| Mean Time to Recover | System availability | "The server is back up." |
| Mean Time Between Failures | Reliability | "Things rarely break." |
Why MTTR is a Commercial Metric, Not Just a Technical One
When Luis Navarro was growing Totality Services, he realised he wasn't just selling technical support; he was selling productivity and peace of mind. Every hour a ticket stays open is an hour of "work in progress" (WIP) for your team. WIP is the silent killer of MSP profitability.
High MTTR means your engineers are context-switching, juggling multiple open tasks, and likely touching the same ticket four or five times. That labour cost eats your fixed-fee recurring revenue alive.
The Impact on Profitability
If your average seat price is $150 and your cost to deliver service is $100, you have a $50 margin. If an engineer spends three extra hours on a basic printer issue because of poor documentation or lack of training, that $50 margin is gone. Mean Time to Resolve (MTTR) is the primary lever you have to protect your gross margins in a managed services model.
The Impact on Client Relationships
Clients don't remember the time you responded in two minutes. They remember the three days they couldn't access their finance software. High MTTR erodes trust. When trust is low, your Security Reviews and QBRs become defensive meetings where you justify your existence rather than proactive meetings where you sell projects.
The Structural Barriers to Low MTTR
If your MTTR is creeping up, it’s rarely because your engineers are "lazy." It is usually a symptom of structural issues within the MSP. After sitting in hundreds of client meetings and managing large technical teams in London and Johannesburg, Luis identified that the biggest delays come from a lack of clarity.
1. The "Information Gap"
If a technician has to spend thirty minutes hunting for a password or a network diagram, your MTTR is already failing. Documentation isn't just a chore; it is a direct contributor to your Mean Time to Resolve (MTTR). Poorly documented environments force every technician to reinvent the wheel every time a ticket is opened.
2. Escalation Bottlenecks
In many MSPs, Level 1 technicians act as expensive dispatchers. They hold onto a ticket for two hours, realise they can't fix it, and then escalate it. The clock has been ticking the whole time. A clear, time-based escalation policy—where a ticket must be moved if not resolved in 20 minutes—is essential for keeping MTTR under control.
3. Non-Standardised Stacks
It is impossible to have a low MTTR if every client has a different firewall, a different backup solution, and a different antivirus. Standardisation is the "secret sauce" of the most profitable MSPs. When your team knows one stack deeply, they don't have to "learn" the client's environment during an outage.
Practical Strategies to Reduce MTTR
Reducing your Mean Time to Resolve (MTTR) requires a combination of cultural shifts and process improvements. It isn't about cracking the whip; it’s about removing the friction that stops your team from closing tickets.
Empower the Front Line
The goal should be to resolve as much as possible at Level 1. This is often called "Shift Left." Give your L1 team the training and the permissions they need to actually fix things. If they have to wait for an L3 engineer to approve a password reset or a firewall change, your MTTR will skyrocket.
Automate the Identification Phase
If a client has to call you to tell you their server is down, you’ve already lost the MTTR battle. Proactive monitoring and RMM (Remote Monitoring and Management) tools should trigger alerts that create tickets automatically. The faster the identification, the faster the resolution cycle begins.
Improve the Triage Process
Triage is the most undervalued role in an MSP. A skilled dispatcher or triage lead can identify "quick wins" and route complex issues to the right specialist immediately. This prevents tickets from sitting in a generic queue where the MTTR clock just ticks away while nobody is looking at it.
The Relationship Between Security and MTTR
In the modern landscape, MTTR isn't just about broken printers; it's about security incidents. When a client is hit with a potential breach, MTTR becomes a measure of survival. A fast resolution can mean the difference between a minor cleanup and a total business collapse.
At MSP Agenda, we believe that security should never be discussed in isolation. When you present a Security Review to a client, you should frame it in terms of business continuity. If they invest in the recommended security stack, the Mean Time to Resolve (MTTR) for future incidents will be significantly lower because the environment is standardised and observable.
Luis Navarro’s experience building Totality Services taught him that clients value results, not technical effort. By linking security recommendations to faster recovery and less downtime, you make the commercial argument for the project much stronger. You aren't just selling a "firewall upgrade"; you are selling a reduction in the time the business stays dark during an emergency.
Common Pitfalls in Measuring MTTR
Tracking Mean Time to Resolve (MTTR) is only useful if the data is clean. Many MSPs accidentally "game" their own metrics, which leads to a false sense of security while profitability leaks out of the business.
- "Pending Client" Status: Many MSPs pause the clock when waiting for a client response. While fair, if a ticket stays in "Pending" for two weeks, it still represents a lingering issue that might result in a follow-up call.
- Re-opened Tickets: If a ticket is closed but the problem isn't fixed, and the client calls back an hour later, does your system count that as a new ticket? If it does, your MTTR looks better than it actually is. True MTTR should account for "First Contact Resolution" (FCR).
- Cherry-Picking: Avoid the temptation to exclude "difficult" clients or complex projects from your MTTR averages. The outliers are usually where your biggest process gaps are hidden.
How to Talk to Clients About MTTR
Your clients probably don't know what MTTR stands for, and they shouldn't have to. However, they care deeply about the outcome of a low MTTR. When you sit down for a QBR or an account management meeting, translate the metric into business language.
Don't say: "Our Mean Time to Resolve improved by 15% this quarter."
Do say: "On average, we are getting your staff back to work 45 minutes faster than we were last year. Across the 200 tickets we handled this month, that’s 150 hours of recovered productivity for your team."
This approach demonstrates the value of the MSP relationship. It moves the conversation away from "what did I pay you for?" to "look at how much time you saved me." This is how you build a highly profitable, eight-figure MSP like the one Luis Navarro exited.
Frequently Asked Questions
What is a "good" MTTR for an MSP?
There is no universal number, as it depends on the complexity of the clients. However, for standard Level 1 and Level 2 tickets (desktop support, password resets, basic software issues), an MTTR of under 4 to 8 business hours is generally considered healthy. For critical outages, the goal is obviously much shorter.
Does Mean Time to Resolve include weekends and off-hours?
Generally, MTTR should be calculated based on your "Business Hours" or the hours defined in your SLA. If a ticket is opened at 4
PM on a Friday and resolved at 9 AM on Monday, counting that as 64 hours of resolution time will skew your data. Most PSA tools allow you to filter by service hours.How does standardisation help with MTTR?
Standardisation reduces the "Diagnostic" phase of MTTR. When every client uses the same stack, your engineers don't have to spend time learning how a specific router is configured or where a specific backup agent stores its logs. They can move straight to the "Rectification" phase.
Can MTTR be too low?
Potentially. If technicians are rushing to close tickets just to keep the MTTR low, they may provide "band-aid" fixes that don't address the root cause. This leads to re-opened tickets and lower client satisfaction. MTTR should always be balanced with quality metrics like First Contact Resolution (FCR).
What is the difference between Mean Time to Resolve and Mean Time to Recovery?
Resolution usually implies the ticket is finished and the root cause is addressed. Recovery often just means the service is back up, even if a temporary workaround is in place. For a business owner, recovery is the priority; for the MSP's long-term efficiency, resolution is the goal.
Ultimately, Mean Time to Resolve (MTTR) is a reflection of your operational maturity. It shows how well your team communicates, how effectively you have documented your clients' environments, and how much you have standardised your offerings. By focusing on this metric, you aren't just improving a number on a dashboard; you are building a more resilient, more profitable, and more valuable managed service provider.