Operate a reliable service

Reliability is a user need. A beautifully designed service that is unavailable when someone needs it has failed — often at the worst possible moment.

Plan for uptime, incident response, monitoring, and recovery. A service that is down when someone needs it fails the user regardless of how well it was designed.

Operational readiness includes monitoring, alerting, on-call rotation, incident management, backups, and recovery procedures. Beta is where many of these practices are tested under realistic load before full launch.

Reliability also means planning for dependency failures — third-party APIs, identity providers, payment gateways — and communicating clearly with users when something goes wrong.

How teams usually demonstrate this

  • Monitoring and alerting for critical user journeys
  • Incident response playbooks and tested recovery procedures
  • Uptime and performance targets aligned with user need
  • Support routes when the digital service is unavailable

Phases where this often matters

See also delivery phase guides and how frameworks fit together.

Related standard points

Point 14 of the Service Standard is published by the Government Digital Service on GOV.UK. Content is available under the Open Government Licence v3.0. UCD Services explains the standard in its own words and is not affiliated with GOV.UK or the Government Digital Service.