About the Role
This role is critical to the reliability and performance of our AI-native platform for customer experience (CX) intelligence, which manages entire customer lifecycles autonomously. As we scale globally, ensuring the reliability, performance, and resilience of our cloud infrastructure is mission-critical.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
As a Site Reliability Engineer, you will be at the heart of our platform's stability. You’ll own the reliability of our cloud infrastructure and ensure it scales seamlessly as we grow. This is not just about reacting to issues; it's about anticipating them, designing systems that prevent them, and building automation that eliminates them. If you enjoy solving complex infrastructure challenges, eliminating inefficiencies, and creating systems that “just work”, this is where you’ll thrive.
What You’ll Do
You’ll be responsible for outcomes, not just tasks. Here’s what success looks like in this role:
You’ll make reliability the default
- Design and maintain infrastructure that is highly available, fault-tolerant, and scalable.
- Proactively identify and eliminate single points of failure before they become incidents.
- Ensure our production systems remain stable, even under increasing scale and load.
You’ll own and optimize our cloud environments
- Manage and continuously improve workloads across AWS, GCP, or Azure.
- Use Infrastructure as Code (Terraform) to standardize and scale infrastructure.
- Optimize resource usage to balance performance and cost.
You’ll run and improve Kubernetes in production
- Operate and scale Kubernetes clusters (EKS, GKE, etc.) with confidence.
- Troubleshoot issues quickly and ensure smooth deployments and upgrades.
- Ensure our containerized workloads perform reliably at scale.
You’ll build strong observability and respond to incidents
- Implement and refine monitoring systems using tools like Prometheus, Grafana, Datadog, or ELK.
- Define meaningful alerting that is not noisy.
- Respond to incidents, lead root cause analysis, and ensure we learn from every failure.
You’ll automate everything that shouldn’t be manual
- Write scripts and build tooling to eliminate repetitive operational work.
- Continuously improve infrastructure efficiency through automation.
- Promote a culture where manual work is a temporary state, not the norm.
You’ll collaborate to improve the entire system
- Work closely with DevOps and engineering teams to solve performance bottlenecks.
- Contribute to CI/CD improvements and deployment reliability.
- Help shape reliability best practices across the organization.
What Success Looks Like (First 90 Days)
First 30 days:
- Build a strong understanding of our infrastructure, systems, and workflows.
- Contribute to day-to-day operations with support from the team.
- Start identifying areas for improvement in automation and reliability.
By 90 days:
- Independently manage infrastructure tasks and troubleshoot issues.
- Actively contribute to reliability and scalability improvements.
- Take ownership of parts of our infrastructure and make improvements.
Who You Are
This is what will make you successful in this role:
- Approximately 3 years of experience in SRE, DevOps, or infrastructure engineering; familiarity with what breaks at scale.
- Comfortable working in cloud environments like AWS, GCP, or Azure, with an understanding of distributed systems behavior.
- Hands-on experience with Kubernetes in production; confidence in troubleshooting issues.
- Proactive in fixing issues while understanding the root cause to prevent recurrence.
Technically, you likely:
- Use Terraform (or similar IaC tools) to manage infrastructure.
- Work confidently with Docker and Kubernetes.
- Write scripts in Python, Bash, or similar to automate workflows.
- Understand CI/CD pipelines (Jenkins, GitHub Actions, Bitbucket, etc.).
- Have a solid grasp of networking, load balancing, and high-availability design.
When it comes to monitoring:
- Implement tools like Prometheus, Grafana, Datadog, or ELK.
- Differentiate between useful alerts and noise.
- Focus on signals that actually drive action.
What Sets You Apart:
- Take ownership; you don’t wait for someone to inform you about issues.
- Maintain composure under pressure and methodical during incidents.
- Simplify complexity instead of contributing to it.
- Communicate clearly, even when discussing deeply technical issues.
- Care about building systems that enhance the effectiveness of other engineers.
Nice to Have (but not required)
- Experience with RabbitMQ or Redis in production.
- Familiarity with Ansible or AWX.
- Exposure to multi-cloud or hybrid environments.
- Cloud certifications (AWS, GCP) or Linux certifications.
- Background from ITI (Information Technology Institute).