Software Support SLA Checklist: Response, Resolution and Ownership
Define a software support SLA with clear scope, service windows, severity, response, communication, restoration, security and reporting.
By AUZtec Innovations

A software support SLA should define which systems and environments are covered, when service is available, how severity is determined, what response and restoration mean, who communicates and which dependencies are excluded. An aggressive response number without capable escalation or recovery evidence is not meaningful protection.
Use this checklist alongside the contract and operating runbook. Obtain legal review for contractual wording.
Define the service scope
List applications, APIs, environments, integrations, databases and infrastructure. State whether support includes cloud providers, third-party applications, end-user devices and data correction.
Separate production from staging and development. Define what counts as a defect, service request, security incident and change.
Set service windows
Specify time zone, business hours, weekends and public holidays. If 24/7 coverage is offered, identify the active responder and escalation—not merely an inbox.
Align coverage with the times users and transactions matter. An internal weekday tool may not need overnight response; an ecommerce checkout may.
Define severity by impact
A practical severity model considers users affected, critical journey, data/security risk, workaround and time sensitivity.
- Critical: widespread loss of a critical service, material data/security event or no viable workaround.
- High: major function unavailable or severely degraded for a significant group.
- Medium: limited impact with a workaround.
- Low: minor defect, information request or cosmetic issue.
Allow the provider and customer to reassess severity with evidence. Avoid calling every inconvenience critical.
Distinguish acknowledgement, response and restoration
Acknowledgement confirms receipt. Response means a capable person has begun diagnosis. Restoration returns an acceptable service, possibly through a workaround. Resolution corrects the root cause.
Set targets for the state you actually need. Resolution time may depend on third parties or careful testing; restoration is often the better urgent commitment.
Define communication
State channels, update frequency, required incident information and stakeholder contacts. Use a status page or agreed broadcast for widespread issues.
Protect sensitive details. Security incidents need a separate communication and escalation route agreed with relevant advisers.
Specify customer responsibilities
The customer may need to provide access, examples, approval, account ownership and timely decisions. Define what pauses an SLA clock and prevent unreasonable use of “waiting for customer”.
Maintain current authorised contacts and emergency decision-makers.
Handle third-party dependencies
List providers and how their outage affects commitments. The support team should still diagnose, communicate, apply workarounds and escalate even when it cannot repair a provider.
Monitor deprecation and account limits before failure. Vendor exclusion should not mean operational silence.
Include security support
Define vulnerability intake, triage, emergency patching, credential compromise and incident handoff. Severity should consider exploitability and data consequence, not only visible downtime.
Maintain least-privilege production access and audit privileged actions. No SLA can promise that an incident will never occur.
Define backups and recovery
State backup ownership, frequency, retention, encryption, test cadence, recovery objectives and restore authority. Read the SaaS Backup and Disaster Recovery Plan for evidence beyond job success.
Clarify maintenance and change
List routine dependency updates, certificate renewal, monitoring review and supported versions. Define maintenance windows and notice.
New features, significant configuration and provider migrations should follow a change process. The Software Maintenance Cost Guide separates operating work from improvement.
Reporting and governance
Monthly or quarterly reporting can include incidents by severity, response/restoration performance, recurring causes, availability where meaningful, vulnerabilities, change history and open risks.
Review exclusions and service levels as usage changes. A target achieved through repeated manual workarounds may still justify product improvement.
Exit and knowledge transfer
Define repository, account, documentation, runbook and credential handover. The business should retain controlled access to its production assets throughout the relationship.
Avoid a support model that becomes impossible to exit because only the provider can deploy or diagnose.
SLA checklist
- Covered systems and environments.
- Service hours and time zone.
- Severity and reassessment.
- Acknowledgement, response, restoration and resolution.
- Escalation and communications.
- Customer responsibilities.
- Third-party handling.
- Security and privacy incident route.
- Backup/recovery responsibilities.
- Maintenance and changes.
- Reporting and service review.
- Exit and knowledge transfer.
Rehearse the agreement
Before signature, walk through three realistic incidents: a complete outage, a degraded critical journey and a security concern. For each, confirm who detects it, how severity is assigned, which clock starts, who communicates, what evidence is required and how service is validated after restoration.
Check that monitoring covers the user journey named in the agreement. A server can be healthy while login, payment or document upload fails. Define planned maintenance, emergency change, data restoration and supplier escalation without hiding them inside broad exclusions.
Review performance with context
A monthly report should show incidents by impact and cause, achieved response and restoration times, recurring problems, maintenance work, recovery tests and open risks. Percentages need counts and service windows. One missed critical incident may matter more than hundreds of low-priority tickets answered quickly.
Use the review to agree prevention work and update runbooks. An SLA that only calculates credits after failure does not improve operational resilience.
Frequently asked questions
Is uptime a useful SLA?
It can be, when measurement point, exclusions and business journey are defined. An API can be technically reachable while checkout fails. Combine availability with journey monitoring and incident response.
Should service credits be the main protection?
Credits may create accountability but rarely compensate for business impact. Prioritise prevention, restoration capability, transparent communication and appropriate contractual advice.
What is a reasonable response time?
It depends on business impact, coverage and staffing. Require evidence that the proposed team can meet it, and distinguish a human acknowledgement from active technical diagnosis.
AUZtec Innovations provides software, cloud and security support matched to the real operating context. We can help define an SLA after establishing assets, critical journeys and recovery capability.
Review the SaaS backup and disaster recovery checklist alongside cloud and DevOps support so contractual recovery promises have a tested technical path.