Dex
9 min readBy Dean Craftsman

The night shift you can't staff: what 24/7 IT coverage really costs an MSP

MSPs sell 24/7 SLAs they can't affordably staff. Here's the real cost of after-hours coverage - and what changes when L1-L3 resolves itself.

Almost every managed service provider sells some version of 24/7. Almost none of them staff it. The gap between the SLA in the contract and the coverage on the roster gets filled with an on-call rotation, a phone that rings at 2am, and a quiet hope that the night stays quiet. This post prices that gap: what a real night shift costs, what the on-call substitute costs once you count attrition and missed response targets, and what changes when the off-hours load resolves itself across every tenant without waking anyone. The conclusion is not that you can fire your on-call engineer. It is that most of the queue they have been staying up for was never work that needed a person.

What "24/7" means on the roster

A standard business day covers 50 hours of a 168-hour week. The other 118 hours - nights, weekends, holidays - are the coverage you sold.

Filling those 118 hours with one human present at all times takes about 3.5 full-time technicians once you build in PTO, sick days, and holiday relief. Night and weekend shifts also carry a differential; a technician competent enough to touch a client's Microsoft 365 tenant unsupervised at 4am is not taking that shift at day rate. At a fully loaded cost of roughly $95,000 per night-shift technician, the arithmetic is unforgiving:

118 hours/week of uncovered time
  ÷  40 hours per technician
  =  3.0 FTE  →  3.5 FTE with PTO and holiday relief

3.5 FTE  ×  $95,000 loaded  =  ~$332,500 / year

That buys exactly one person awake at a time. Two overlapping incidents in different tenants at 1am, and the second client waits. For most MSPs, $332,500 exceeds the entire annual gross margin on a meaningful slice of the client base. The night shift is not unstaffed because MSPs are careless. It is unstaffed because the math never closes.

The on-call substitute, priced honestly

So the roster gets replaced with a rotation: four engineers, one week each, a stipend for carrying the phone, and an hourly callout rate when it rings.

On the invoice that looks cheap. A $200 weekly stipend is $10,400 a year. Three hundred callouts at a two-hour minimum and $75 an hour is $45,000. Call it $55,000 in direct spend against $332,500 for real coverage. Finance signs it without reading twice.

The costs that make it expensive do not appear on that line.

Attrition. Rotating on-call through a small team is a reliable way to lose the people you can least afford to lose. Replacing a mid-level M365 engineer conservatively costs half a year of loaded salary once you count recruiting, ramp time, and the output the rest of the team loses covering the gap - call it $42,000 per departure. One extra resignation a year attributable to the rotation nearly doubles the true cost of the program, in a different budget line than the one you were watching.

Response time. An on-call engineer who is asleep is not a four-minute response. They are a 30-to-60-minute response on a good night, and only for the cases that clear the bar for waking them. Everything below that bar - lockouts, expired credentials, MFA resets, a broken SharePoint permission - waits for the morning queue. The client's user does not experience that as "outside SLA scope." They experience it as being blocked from 6am until someone gets to it.

SLA credits, and the client who stops renewing. Most MSP contracts credit five to ten percent of the monthly fee against a missed response target. Across a 25-tenant fleet at $6,000 average MRR, a few credited months a year is maybe $5,000 - trivial. What is not trivial is the client who quietly declines to renew after the third overnight incident that sat until 9am. One lost tenant is $72,000 in ARR, which dwarfs every other figure here.

Add it up and the on-call model runs near $100,000 a year in real cost, delivers response measured in hours rather than minutes, and puts your senior engineers' tenure at risk. It is the cheapest bad option, not a solution. We laid out the shape of this constraint in how MSPs scale IT support without a linear headcount curve; after-hours is where that curve gets steepest.

What actually arrives overnight

Before pricing a fix, be precise about the load. A 25-tenant MSP covering roughly 5,000 seats generates something like 20,000 support requests a year, and 10 to 15 percent of those arrive outside business hours - shift workers, distributed teams, Sunday-evening catch-up, and the early-morning wave from users a few time zones ahead of your office. That is roughly 2,400 requests a year, or six to seven a night.

The mix skews heavily routine: account lockouts after repeated failed logins, forgotten passwords, MFA device loss, expired sessions, access to a file or SharePoint site someone needs before a morning deadline, mailbox issues, single-device problems. Genuine emergencies - a tenant-wide outage, a suspected compromise, a mail flow failure - are a small fraction of the total, and they are the only fraction an on-call rotation is actually designed for.

Which makes the honest description of most MSP night shifts this: you are paying senior engineers a burnout premium to be interrupted by password resets, and calibrating your SLA around how long a person can reasonably be expected to sleep.

24-hour IT coverage timeline comparing staffed daytime hours to autonomous after-hours resolution

What changes when the off-hours load resolves itself

Dex is an autonomous IT engineer for Microsoft 365. It does not route the request or draft a reply for a human to send in the morning - it investigates the tenant, plans the change, executes it under explicit policy, and closes with a full audit trail. End users reach it inside Microsoft Teams and Slack through Dex Go; admins work through the Dex Pro console. Two properties matter specifically for after-hours coverage.

It resolves L1 through L3, not just L1. Routine password, MFA, and access work is the volume, but the deeper Tier 2 and Tier 3 troubleshooting and configuration work is the part that historically justified waking a senior person. Both are in scope. Only genuine architectural or judgment cases escalate.

It covers every tenant at once. Dex Pro operates per tenant on the admin's own delegated permissions, with per-org isolated databases and encryption keys and a separate policy set per client. There is no night-shift roster to divide across the fleet, and no second client waiting behind the first. Marginal coverage for tenant 26 is compute, not a rotation slot - the same mechanism we worked through in the MSP tenant math, applied to the hours nobody wants to work.

Run the model. Of 2,400 off-hours requests, the routine surface resolving end-to-end at Dex's 90%+ rate leaves roughly 240 for humans - about one every day and a half, instead of six or seven a night. On-call becomes real on-call: rare, genuinely urgent, worth waking up for.

The financial effect is not primarily "we cut the on-call budget." It is that response time for the overnight majority goes from a morning-queue wait to minutes, without buying a single hour of night labor - and the SLA you already sold becomes something you can defend in a QBR.

What still needs a human at 3am

Anyone telling you autonomous resolution retires the on-call phone is selling you something. It does not, and it should not.

Real incidents still page a person. Tenant-wide outages, mail flow failures, suspected account compromise, and anything with a security dimension go to a human. These are judgment calls with blast radius, and they belong to people who can be accountable.

Anything without a policy stops. Dex checks every action against an explicit, per-tenant policy at the code level - no policy, no action. The overnight consequence is that an uncovered case does not improvise at 3am; it escalates. Policy coverage becomes what determines autonomous coverage, which makes policy authoring the new night-shift investment.

Client-specific business decisions stay with the client. Provisioning a license that costs money, granting an exception to an access rule, approving something a manager would normally sign off on - Dex holds rather than guesses.

Your ITSM stays in place. Dex is not an ITSM and does not replace one. It removes the work before it becomes a ticket and uses your ITSM - SysAid, ServiceNow, Jira, Halo - as the system of record for what remains.

What changes is the quality of the escalation. When a human does get woken, the request arrives with the investigation already done: what was tried, what the tenant state looks like, what policy blocked the action. Being paged with context is a categorically different experience from being paged with "user says email is broken."

What to do with this on Monday

Three moves, in order.

  1. Price your current after-hours coverage properly. Stipends, plus callout hours, plus a realistic attrition allowance, plus the ARR of any client you have lost to overnight response. Most MSPs have only ever seen the first two, and the full figure usually lands near six figures for a mid-size fleet.

  2. Break your off-hours volume down by tier and category. Pull the last 12 months of requests that arrived outside business hours and sort them. If lockouts, credentials, MFA, and access dominate - and they almost always do - you are paying a burnout premium for work that never needed a person, which is the tax we described in why L1 volume is a structural tax on IT.

  3. Ask any vendor for end-to-end resolution across L1-L3, specifically overnight. Not deflection, not containment, not assisted triage. The question that matters is what percentage of requests the system resolves itself, against the real tenant, under policy, with an audit trail, when nobody is watching. If the answer is a deflection rate, the night shift stays on your payroll.

To see what this looks like against a live Microsoft 365 tenant - including what Dex does when it hits a case it is not allowed to touch - join the next Dex webinar. We run real IT work, end to end, and the escalation path gets the same airtime as the happy path.

The night shift was never a staffing problem. It was a structural one: MSPs sold continuous coverage in a market where labor is discontinuous, expensive, and needs to sleep. You cannot fix that by finding more disciplined humans. You fix it by removing the assumption that a password reset at 3am requires one.

Frequently asked

How much does 24/7 IT coverage cost an MSP?
Genuinely staffing the clock costs more than most MSPs assume. Covering the 118 hours a week that fall outside a standard business day takes roughly 3.5 full-time technicians once you account for holidays, sick days, and PTO relief, and night and weekend shifts carry a pay differential on top of base salary. At a fully loaded cost of around $95,000 per night-shift technician, that is roughly $330,000 a year to keep one person awake and available - before you have covered a second concurrent incident. That is why most MSPs sell 24/7 and deliver an on-call rotation instead.
Can an MSP offer 24/7 support without running a night shift?
Yes, but only if the off-hours work stops requiring a person. An on-call rotation does not solve it - it moves the cost from payroll to attrition and response time. The structural fix is to break the link between an after-hours request and human labor: an autonomous IT engineer that investigates, plans, and executes the routine L1 through L3 work overnight under explicit policy, and wakes a human only for the cases that genuinely need judgment.
Does autonomous IT handle after-hours tickets across multiple client tenants?
Yes. Dex Pro operates per tenant using the admin's own delegated permissions, with per-org isolated databases and encryption keys, so a single autonomous engineer covers every Microsoft 365 tenant in the fleet at once without co-mingling client data. Each tenant carries its own policy set, so the same engineer enforces different rules for different clients simultaneously - and 3am is no different from 3pm, because there is no queue to be at the back of.
Is autonomous resolution only for Tier 1 after-hours tickets?
No. Dex autonomously resolves L1 through L3 - password, MFA, lockout, and access work, and also deeper Tier 2 and Tier 3 troubleshooting, configuration, and engineering-adjacent tasks that used to require a senior technician. That matters more overnight than during the day, because the tier that historically justified waking someone up is exactly the tier a single on-call technician is least equipped to handle alone at 3am.
What still needs a human on call overnight?
Tenant-wide outages, suspected account compromise and other security incidents, anything requiring a business decision on the client's behalf, and anything with no policy covering it - Dex enforces no policy, no action at the code level, so an uncovered case stops rather than improvises. You still need an on-call rotation. The difference is that it gets called roughly once every few days on real incidents instead of several times a night on lockouts, and every escalation arrives with the full investigation attached.