About the Role / About the Job
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
As a Senior Site Reliability Engineer (SRE), you will be pivotal in enhancing the reliability and performance of our platforms. You will lead key reliability projects, improve observability, and respond to complex production incidents, ensuring seamless operations in a fast-paced environment.
Responsibilities:
- Maintain and evolve monitoring, alerting, and incident response systems.
- Proactively identify and mitigate performance bottlenecks and failure points.
- Contribute to infrastructure automation and deployment pipelines.
- Drive Service Level Objective (SLO) and Service Level Indicator (SLI) adoption in collaboration with engineering teams.
- Conduct root cause analysis and develop preventive solutions.
- Mentor junior engineers and foster operational excellence.
- Analyze service performance, identify bottlenecks, and create measurable improvement plans.
- Monitor technical and business metrics continuously to maintain environmental health, configuring alerts for potential issues and addressing risks proactively.
- Deploy application updates with minimal disruption to services.
- Identify, evaluate, and conduct proof-of-concepts for new technologies.
- Contribute to the knowledge base.
- Review and modify CI/CD principles and service maturity iteratively for continuous improvement.
- Comply with Quality Health Safety and Environment (QHSE), Business Continuity, Information Security, Privacy, Risk, Compliance Management, and Governance policies and related risk assessments.
Qualifications:
Requirements:
-
Educational Background:
- Bachelor’s Degree in Business Analytics, Data Science, Computer Science, Engineering, or a related field.
- Master’s Degree preferred.
-
Experience:
- 5+ years in a Site Reliability Engineer (SRE), DevOps, Sysadmin, or Platform Engineer role.
- 5+ years of experience managing Kubernetes clusters.
- 5+ years of experience configuring and utilizing monitoring/observability platforms.
- Familiarity with at least one type of database.
-
Skills:
- Solid experience with containerized environments (Docker, Kubernetes).
- Hands-on ability with CI/CD pipelines and automation tools.
- Proficiency in scripting languages (Python, Bash).
- Strong understanding of observability tools (Prometheus, Grafana, ELK, Sentry).
- Good knowledge of cloud platforms (Huawei Cloud, Azure preferred).
Mandatory Skills:
- Strong background in Linux/Unix Administration.
- Solid hands-on experience deploying and operating Kubernetes or Openshift clusters.
- Experience configuring and maintaining monitoring and observability solutions.
- Ability to troubleshoot and resolve complex production issues efficiently, performing root cause analysis and restoring services quickly during high-pressure incidents.
- Experience in backing up and restoring various systems.
- Collaborate with project managers and solution architects, serving as a subject matter expert.
- Implement basic network security measures (e.g., configuring VPCs, firewalls/security groups).
- Understand the dependencies of various GPU cards and upgrade container images as needed to ensure compatibility.
- Deploy and operate products provided by third-party vendors.
- Create and deploy release packages in coordination with the development team.
Bonus Skills:
- Good understanding of typical system architecture and interaction between its components.
- Experience automating tasks using infrastructure-as-code tools, like Ansible or Terraform.
- Thorough understanding of auxiliary components such as caching systems (e.g., Redis, Memcached) and message queues (e.g., RabbitMQ, Kafka).
- Good understanding of databases including Postgres, Elasticsearch, Clickhouse.
- Basic scripting abilities.
- Familiarity with OAuth 2.0, OpenID/OpenID-Connect, SAML 2.0, Kerberos, and LDAP.
What Working at the company Offers:
- Culture: Work in a fast-paced, varied environment that encourages initiative, surrounded by talented peers focused on solving transformational challenges.
- Career: Be part of a team that recognizes and rewards contributions while supporting you to operate at peak performance.
- Rewards: An attractive remuneration package including healthcare, education support for dependents, leave benefits, and more.