About the Role
As an MLOps Engineer, you will set the ML infrastructure and reliability strategy across the platform, including how large language models (LLMs) and other deep learning models are deployed, fine-tuned, and served at scale. You will own architecture decisions across both SaaS and on-prem operating models, mentor engineers across teams, and drive a multi-quarter ML infrastructure strategy.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
What You'll Do
- Define ML infrastructure architecture across the platform, including model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP).
- Set direction for ML system reliability, focusing on monitoring, latency, throughput, availability targets, and incident response across research and production environments.
- Mentor senior MLOps engineers and elevate the operational standards across multiple teams.
- Drive cross-team initiatives aimed at improving inference performance and cost-efficiency, including the use of distributed training frameworks (DeepSpeed, FSDP, Accelerate).
- Collaborate with ML researchers, product, and engineering leadership on a multi-quarter ML infrastructure strategy.
- Ensure that the ML infrastructure scales effectively for both managed SaaS and fully air-gapped on-prem deployments.
What You'll Bring
- 10+ years of experience in MLOps, ML infrastructure, or machine learning engineering, with a demonstrated history of architectural ownership.
- Proven track record in architecting large-scale model deployment, including LLMs and ML infrastructure at scale.
- Deep expertise in cloud technologies across major platforms (AWS, Azure, or GCP) and a strong proficiency in Python.
- A solid record of mentorship, with evidence of engineers you have developed now operating independently at higher levels.
- Comfort with architecting ML systems that function in both managed SaaS and on-premises, disconnected air-gapped environments.
- Kubernetes expertise at an architectural depth, including GPU scheduling, multi-tenancy, operators, and understanding the failure modes of distributed workloads on shared clusters.
- Excellent communication, stakeholder management, and decision-making skills, with a commitment to building diverse and inclusive engineering teams.
Strong Preference
- Experience with production reliability at the platform level, including SLO definition, incident command, postmortem practices, and driving reliability improvements across teams rather than individual services.
- Architectural experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
- In-depth knowledge of GPU systems: CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication issues across nodes.
- Experience with model optimization strategies at a portfolio level, such as quantization (FP8, AWQ, GPTQ) and speculative decoding, with measurable outcomes related to cost or latency.
- Previous experience in regulated or security-constrained environments involving compliance-driven architecture, model governance, lineage, audits, and secrets management.
- Experience with on-prem or air-gap ML delivery architecture at scale.
- Proven history of maturing MLOps practices in a growth-focused organization, with contributions to standards, platform abstractions, and paved paths that have a lasting impact.
- Knowledge of bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle management in customer-owned or data centers.
Nice to Have
- Public speaking at conferences, technical writing, or contributions to industry thought leadership.
- Involvement in open-source projects related to inference, serving, or ML infrastructure, especially at a maintainer level.
- Experience in C/C++ or CUDA kernel development for performance-critical paths.
- Proficiency in Arabic.
Why You Should Join
- Mission-Driven Work: Engage in cutting-edge AI applications with a skilled and passionate team, tackling real-world challenges in crucial sectors.
- Unparalleled Opportunity: Innovate and address significant issues using AI at a company with unique access to leading models and resources.
- Career Growth: Benefit from competitive compensation, bonuses, and substantial career advancement potential as a foundational member of the team.
- World-Class Environment: Enjoy a flexible work setting with access to the latest tools and technologies to optimize performance.