Company logo hidden

Senior Engineer - Site Reliability

Unlock employer Abu Dhabi, United Arab Emirates Direct to Company 1 hour ago · 16 Sep 2026

Financial

  • Estimate: $60k - $120k*
  • Zero income tax location

Accessibility

  • Apply from abroad
  • Visa Provided

Requirements

  • Experience: Senior
  • English: Professional

Position

Ready to apply for roles like this?

Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.

Unlock employer & apply directly

About the Role This position involves enhancing the reliability and performance of our platforms as a Senior Site Reliability Engineer (SRE). You will lead key reliability projects, improve observability, and respond to complex production incidents while maintaining our commitment to providing end-to-end AI solutions in the energy sector. Responsibilities - Maintain and evolve monitoring, alerting, and incident response systems. - Proactively find and fix performance bottlenecks and failure points. - Contribute to infrastructure automation and deployment pipelines. - Drive SLO/SLI adoption in collaboration with engineering teams. - Lead root cause analysis and build preventive solutions. - Mentor junior engineers and help scale operational excellence. - Analyze service performance, identify bottlenecks, and provide measurable improvement plans. - Maintain the environment’s health by continuously monitoring technical and business metrics, configuring alerts for potential issues, and proactively addressing risks to prevent disruption. - Deploy application updates with minimal disruption to services. - Identify, evaluate, and conduct proof-of-concepts for new technologies. - Contribute to the knowledge base. - Review and modify CI/CD principles and service maturity iteratively, striving for continuous improvement. - Comply with QHSE (Quality Health Safety and Environment), Business Continuity, Information Security, Privacy, Risk, Compliance Management, and Governance of Organizations policies, procedures, plans, and related risk assessments. Qualifications Requirements: - Bachelor’s Degree in Business Analytics, Data Science, Computer Science, Engineering, or a related field. - Master’s Degree is preferred. Experience: - 5+ years in an SRE/DevOps/Sysadmin/Platform Engineer role. - 5+ years of experience in managing Kubernetes clusters. - 5+ years of experience in configuring and using monitoring/observability platforms. - Familiarity with at least one type of database. Skills - Solid experience with containerized environments (Docker, Kubernetes). - Hands-on with CI/CD pipelines and automation tools. - Proficiency in scripting languages (Python, Bash). - Strong grasp of observability tools (Prometheus, Grafana, ELK, Sentry). - Good knowledge of cloud platforms (Huawei Cloud, Azure preferred). Mandatory Skills: - Strong background in Linux/Unix Administration. - Solid hands-on experience deploying and operating Kubernetes or OpenShift clusters. - Experience configuring and maintaining monitoring and observability solutions. - Ability to troubleshoot and resolve complex production issues efficiently, including performing root cause analysis and restoring services quickly during high-pressure incidents or critical outages. - Experience in backing up and restoring various systems. - Working together with project managers and solution architects while serving as subject matter experts. - Implementing basic network security (e.g., configuring VPCs, firewalls/security groups, etc.). - Understand the dependencies of various GPU cards, and upgrade container images as needed to ensure compatibility. - Deploy and operate products provided by third-party providers. - Creating releases together with the development team and deploying release packages to all required environments. Bonus Skills: - Good understanding of typical system architecture and interaction between its components. - Experience automating tasks using infrastructure-as-code tools, e.g., Ansible, Terraform. - Thorough understanding of a company's systems, including auxiliary components like caching systems (e.g., Redis, Memcached) and message queues (e.g., RabbitMQ, Kafka). - Good understanding of databases, e.g., Postgres, Elasticsearch, Clickhouse. - Basic scripting. - Working knowledge of OAuth 2.0, OpenID/OpenID-Connect, SAML 2.0, Kerberos, LDAP. What Working at the company Offers - Culture: Encouraging initiative in a fast-paced and varied environment, surrounded by talented peers from around the world who are similarly attracted to applying their skills to solve transformational challenges. - Career: Join a team where your contributions are recognized and rewarded while being supported to operate at your peak performance. - Rewards: An attractive remuneration package that includes healthcare, education support for dependents, leave benefits, and more.

Apply Direct

Jobs you might like   View all jobs

About IT Services and IT Consulting Company

Company details are hidden. Subscribe to view full company profile.

Ready to apply for this role?

Apply Direct