What should an SLA actually include?

By Weapp · Updated

An SLA, or service-level agreement, defines how quickly and reliably a vendor handles hosting and maintenance. The key metrics are response time, resolution time, and uptime, set by priority. Match the levels to your business's actual needs – an extra nine in uptime costs disproportionately more – and tie in measurement, reporting, and penalties that actually get tracked.

An SLA, or service-level agreement, is the promise for how a system will be handled after it’s built – how quickly issues get fixed and how often the service is allowed to go down. It’s one of the most important attachments in a hosting agreement, and one of the most misunderstood. Many buyers either buy too little and get caught out at the next breakdown, or buy too much and pay for readiness the system never needs. The key is to set requirements that match your business’s actual needs.

Response time, resolution time, and ticket priority

The first thing to sort out is what these times actually measure, since vendors and buyers often mix up the terms.

  • Response time is how quickly the vendor acknowledges the ticket and starts working. It says little about when the issue is actually gone.
  • Resolution time is how long it can take until the problem is fixed. That’s the number that protects your business.

A fast response time without a clear resolution time is a trap: you get a friendly “we’re looking into it” and then sit with the system down for two days without the agreement being breached. So always require resolution times, not just response times.

Both should also be tied to priority. A total outage of the core system isn’t the same thing as a crooked icon. A common breakdown is critical (the business is down), high (an important function is down but there’s a workaround), and normal (annoying but not urgent). Define the levels concretely in the agreement – otherwise every ticket turns into a negotiation over how serious it really is.

Reasonable uptime levels and what they cost

Uptime is expressed as a percentage, and the difference between levels looks small but is enormous in both allowed downtime and price. Each extra nine requires more redundancy, more monitoring, and more on-call coverage.

Uptime levelAllowed downtime per month
99%About 7 hours
99.9%About 43 minutes
99.99%About 4 minutes

The step from 99.9 to 99.99 percent sounds like a minor detail but often multiplies the cost several times over, since it requires the system to tolerate parts breaking without any disruption. So ask the counter-question before setting requirements: what does an hour of downtime actually cost the business? For an internal tool, the answer is often “annoying but tolerable,” and then 99.9 percent is more than enough. For a payment flow in the middle of the holiday shopping rush, the answer is different.

Match the level to how critical the service is

The most expensive mistake is giving every system the same high level. A well-thought-out maintenance setup splits services apart and sets readiness according to how much they actually matter.

  • Business-critical (e.g., an e-commerce checkout): high uptime, short resolution time, round-the-clock on-call coverage.
  • Business-important (e.g., an internal ticketing system): good uptime, fixes during business hours, longer margins.
  • Supporting (e.g., a reporting tool): basic uptime, longer resolution times, no on-call coverage.

By tiering your services, you pay for high readiness only where it’s needed, instead of putting on-call costs on the reporting tool nobody misses on a Sunday night.

Penalties, measurement, and follow-up that actually happens

Penalty clauses – deductions when levels are missed – are the part everyone wants to negotiate and the one that matters least in kronor. The compensation rarely covers your actual loss from an outage. Their real value is that they force measurement: you can’t claim a penalty without someone measuring the outcome.

And that’s where most SLAs fall apart. A scenario: a company had an agreement with impressive numbers – 99.9 percent uptime, penalties for shortfalls – but no measurement and no monthly report. When the system kept acting up, there was no evidence to point to, and the agreement turned into a paper tiger. An SLA without measurement and reporting is just a statement of intent.

So require three things beyond the levels themselves: how the outcome is measured, a recurring report you actually receive, and a clear consequence for repeated breaches. Then the agreement becomes a management tool instead of decoration.

If you’d like to set service levels that fit your specific systems rather than a standard template, we at Weapp are happy to help – get in touch with a description of what you’re running.

Frequently asked questions

What's the difference between response time and resolution time?

Response time is how quickly the vendor acknowledges and starts working on a ticket. Resolution time is how long it can take before the issue is fixed. It's resolution time that protects your business – a fast acknowledgment is worthless if the fix takes days. So set clear resolution times per priority level, not just response times.

What does 99.9 percent uptime mean in practice?

99.9 percent allows around 43 minutes of downtime per month, while 99.99 percent allows just over four minutes. Each extra nine costs significantly more because it requires redundancy and on-call coverage. Ask what an outage actually costs your business before you pay for five nines – for many systems, 99.9 is more than enough.

Do all systems need the same service level?

No, and giving them all the same level is a costly mistake. A business-critical payment flow justifies high uptime and short resolution times around the clock, while an internal reporting tool gets by with business hours and longer resolution times. Split your services by how critical they are and set levels accordingly, so you only pay for the readiness you need.

Are penalty clauses worth the trouble of negotiating?

Penalty clauses mostly work as a steering signal, not as a source of income – the compensation rarely covers your actual loss. Their real value is that they force measurement and follow-up, and give you leverage for repeated failures. An SLA without measurement and reporting, on the other hand, is just a statement of intent, no matter how steep the penalties on paper.