The company is a specialist AI and data consultancy dedicated to transforming the regional corporate landscape through bespoke, high-impact intelligent systems. We are seeking a talented Site Reliability Engineer (SRE) to join our technical squad and own the ultimate production readiness, resilience, and operational health of our cutting-edge AI solutions.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
What You’ll Do
- Implement load testing, fault-injection techniques, and self-healing automation to ensure web/mobile backends and AI microservices easily sustain banking-scale traffic.
- Build Python-driven, CI/CD-integrated regression test suites that validate application logic, data pipelines, and Infrastructure as Code (IaC) states.
- Define and track critical Service Level Indicators (SLIs) and Objectives (SLOs). Set up real-time telemetry, dashboards, and proactive alerting to catch performance degradation before users do.
- Integrate specialized AI evaluation harnesses (e.g., DeepEval, Ragas) into deployment pipelines to benchmark model accuracy, hallucinations, and retrieval performance (RAG).
- Automate security scans and privacy checks to satisfy strict banking data residency standards, while monitoring compute efficiency to eliminate redundant cloud spend.
Requirements
- High proficiency in Python (mandatory) and Bash scripting, alongside automation test frameworks like PyTest, Selenium, or Robot Framework.
- Strong hands-on experience with Azure or AWS, focusing on cloud networking, auto-scaling topologies, containerized environments, and serverless reliability.
- Modern Tooling Command:
- CI/CD & IaC: GitHub Actions, Azure DevOps, and Terraform (reading/validating state).
- Performance & Load: K6 or JMeter.
- Observability: Azure Monitor, AWS CloudWatch, Grafana, or Prometheus.
- Solid grasp of prompt engineering, RAG architectures, and model evaluation metrics (ROUGE/BLEU scores, LLM benchmarks, LangSmith/Ragas).
- Experience operating in financial services, banking, or highly regulated enterprise environments where zero downtime and data privacy are non-negotiable.
It would also be great if you have:
- Experience managing production Kubernetes clusters (AKS/EKS), service meshes, and container security scanners.
- Familiarity with UAE Central Bank regulatory guidelines, NESA security compliance, and data residency mandates.
- Formal industry credentials such as AWS Certified SysOps Administrator, Microsoft Certified: Azure Administrator Associate, Certified Kubernetes Administrator (CKA), or SRE certifications.
Benefits
- Competitive salary
- Comprehensive personal health insurance
- Visa Sponsorship for the successful individual
- Professional development and certification support
- Subscription reimbursement relating to your role
- Opportunity to work on cutting-edge AI projects
- Monthly Employee Incentive program
- Career advancement opportunities in a rapidly growing AI company
This position offers a unique opportunity to shape the future of AI implementation while working with a talented team of professionals at the forefront of technological innovation. The successful candidate will play a crucial role in driving the company's success in delivering transformative AI solutions to our clients.