This ITIL core foundation video explains about availability management, key terms and incident lifecycle.
Purpose of Availability Management
To ensure that the level of service availability delivered in all services is matched to or exceeds the current and future business requirements, in a cost-effective and timely manner.
To ensure both the current and future availability need of the business are met.
Objectives of Availability Management
Provide a point of focus and management for all availability-related issues.
Produce and maintain an appropriate and up-to-date availability plan.
Ensure that proactive measures to improve the availability of services are implemented wherever it is cost-justifiable to do so.
Scope of Availability Management
Current business processes, their operation and requirements.
Future business plans and requirements.
Service targets and the current IT service operation and delivery.
Key terms in availability management.
Availability
The per cent time of agreed service hours the component or service is available.
Reliability
A measure of how long a component or IT service can perform its agreed operation without interruption.
Maintainability
A measure of how quickly and effectively a component or IT service can be restored to normal working after a failure.
Serviceability
The ability of a third-party supplier to meet the terms of its contract. This contract will include agreed levels of reliability, maintainability or availability for an IT service or component.
Vital Business Functions (VBF)
The business critical elements of the business process supported by an IT Service.
Typically this will be where more effort and investments will be spent to protect these vital business functions.
Service Availability
All aspects of service availability and unavailability and the impact of component availability, or the potential impact of component unavailability on service availability.
Component Availability
All aspects of component availability and unavailability.
Incident life cycle.
Any unplanned interruption to an IT Service or reduction in the quality of an IT service can be termed as incident.
As shown in the image, when ever an incident occurs, it goes through the following stages.
Detect
Record
Diagnose
Repair
Recover
Restore
The time taken from incident occurrence to restoring of the service is known as Mean Time to Restore Service MTRS.
The time between two incidents is called Mean Time Between System Incidents MTBSI.
The duration between the service restoring and a new incident is known as Mean Time Between Failures MTBF.