About the Role
This position supports the effective operation, coordination, control, and continual improvement of Incident Management, Major Incident Management, and Reactive and Proactive Problem Management across in-scope services. The role ensures that incidents are accurately recorded, prioritized, escalated, communicated, restored, validated, and closed within agreed targets while investigating and reducing recurring or high-impact issues through structured root-cause analysis, known-error management, and corrective action. Under the guidance of the Practice Manager or Senior Consultant, this role contributes to service reliability, customer experience, operational resilience, auditability, and continual learning.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
Responsibilities
- Support the day-to-day administration and consistent execution of Incident Management, Major Incident Management, and Problem Management procedures, workflows, templates, escalation models, and controls.
- Review incident records for accurate logging, categorization, impact, urgency, priority, assignment, timestamps, service and configuration-item links, diagnostic information, resolution details, user validation, and closure quality.
- Monitor high-priority, aging, breached, reopened, customer-sensitive, and supplier-dependent incidents; follow up with assigned teams and escalate risks or delays through approved paths.
- Support major-incident declaration, bridge-call coordination, participant activation, timeline and action logging, restoration tracking, stakeholder communications, and formal stand-down under the direction of the Major Incident Manager or Practice Lead.
- Coordinate service desk, technical, cloud, infrastructure, application, cybersecurity, supplier, customer, and business teams to support rapid and safe restoration, ensuring workarounds, emergency changes, failover, and recovery actions are authorized and recorded.
- Collect logs, monitoring information, decisions, impact details, timelines, and supporting evidence for major-incident reports, post-incident reviews, customer reporting, and audits; document lessons learned and track actions.
- Identify and register problem candidates from major and recurring incidents, trend and monitoring data, failed changes, post-release issues, supplier performance, availability or capacity concerns, and known operational risks.
- Support structured root-cause analysis using approved techniques, document causes and contributing factors, and track corrective and preventive actions to verified closure.
- Maintain problem backlogs, known errors, workarounds, temporary fixes, permanent-resolution actions, owners, target dates, risks, dependencies, and links among incident, problem, change, knowledge, supplier, service, and configuration-item records.
- Ensure verified workarounds and known errors are published to service desk and support teams, remain current, and are used to improve first-line resolution.
- Produce operational reports and dashboards covering incident trends, SLA performance, response and restoration times, backlog age, reopen rate, major and repeat incidents, problem status, known-error age, and action closure.
- Prepare meeting packs, minutes, decision records, and action logs for major-incident reviews, problem reviews, backlog meetings, supplier discussions, and service-improvement forums.
- Support monitoring, event-management, automation, runbook, knowledge, shift-left, swarming, and self-healing improvements that reduce detection, diagnosis, restoration, and recurrence times.
- Maintain procedures, priority models, communication templates, escalation lists, major-incident checklists, root-cause templates, known-error content, and knowledge articles under document control.
- Assist with record-quality reviews, process compliance checks, supplier follow-up, customer or audit evidence preparation, and closure of identified gaps or corrective actions.
- Identify recurring process weaknesses, data-quality issues, bottlenecks, supplier delays, and automation opportunities, and contribute to practical continual-improvement initiatives.
Qualifications
Education, Knowledge and Certifications
- Bachelor’s degree in information technology, computer science, software engineering, engineering, information systems, or a related discipline.
- ITIL 4 Foundation certification is required; relevant training in Incident Management, Major Incident Management, Problem Management, Monitor, Support and Fulfil, or Create, Deliver and Support is strongly preferred.
- Training in Site Reliability Engineering, SIAM, COBIT, ISO/IEC 20000, Lean Six Sigma, risk management, business continuity, project management, or information security is advantageous.
- Good working knowledge of IT service management, service operation and transition, service reliability, operational resilience, customer experience, service levels, supplier obligations, and continual improvement.
Hard Skills
- Working knowledge of Incident, Major Incident, and Problem Management lifecycles, roles, priority models, escalation paths, communications, governance forums, and record requirements.
- Ability to review incident and problem records, identify missing or inconsistent data, monitor SLA and aging status, maintain traceability, and follow approved escalation and closure procedures.
- Practical capability in major-incident coordination, bridge-call administration, timeline and action logging, stakeholder updates, restoration tracking, and post-incident documentation.
- Ability to support root-cause analysis, distinguish symptoms from causes, document contributing factors, and track corrective and preventive actions to verified closure.
- Working knowledge of enterprise ITSM platforms, monitoring and observability tools, knowledge bases, dashboards, workflow automation, CMDB, service mapping, and audit trails.
- Understanding of infrastructure, cloud, networks, telecommunications, cybersecurity, databases, middleware, applications, and end-user computing sufficient to coordinate multi-team support.
- Ability to prepare operational dashboards, trend reports, incident summaries, root-cause documents, meeting minutes, management presentations, and auditable supporting evidence.
- Working knowledge of analytical techniques such as Pareto analysis, 5 Whys, fishbone analysis, chronology analysis, causal mapping, and basic data analysis.
Experience
- 4–6 years of relevant experience in IT service operations, Incident Management, Major Incident Management, Problem Management, service desk, service reliability, or a comparable function.
- At least 2 years of practical experience coordinating priority incidents, supporting major incidents, maintaining problem records, assisting root-cause analysis, or tracking corrective actions.
- Experience using enterprise ITSM platforms, monitoring or observability tools, dashboards, knowledge bases, CMDB or service-mapping information, and workflow automation.
- Experience working with cross-functional infrastructure, cloud, application, network, cybersecurity, service desk, customer, and supplier teams in operational or 24x7 environments.
- Experience supporting operational reporting, customer communications, audits, post-incident reviews, problem backlogs, known-error management, or continual-improvement activities is advantageous.