Company logo hidden

MLOps Engineer

Unlock employer Abu Dhabi, United Arab Emirates Direct to Company Under an hour ago · 21 Sep 2026

Financial

  • Estimate: $80k - $150k*
  • Zero income tax location

Accessibility

  • Office Only
  • Visa Provided

Requirements

  • Experience: Senior
  • English: Professional
  • Arabic: Preferred

Position

About the Role:
As an MLOps Engineer, you set the ML infrastructure and reliability strategy across the platform, focusing on the deployment, fine-tuning, and serving of LLMs and other deep learning models at scale. You will be responsible for architecture decisions within both SaaS and on-prem operating models, mentor engineers across teams, and drive a multi-quarter ML infrastructure strategy, functioning as a force multiplier.

Ready to apply for roles like this?

Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.

Unlock employer & apply directly

What You'll Do:

  • Define ML infrastructure architecture across the platform, including model deployment strategy (vLLM, Triton, or TGI), pipeline engineering (MLflow or Kubeflow), and cloud-native infrastructure across major cloud platforms (AWS, Azure, or GCP).
  • Set direction for ML system reliability, including monitoring, latency/throughput/availability targets, and incident response across research and production environments.
  • Mentor senior MLOps engineers; raise the operational bar across multiple teams.
  • Drive cross-team initiatives to improve inference performance and cost-efficiency, including distributed training frameworks (DeepSpeed, FSDP, Accelerate).
  • Partner with ML researchers, product, and engineering leadership on multi-quarter ML infrastructure strategy.
  • Ensure ML infrastructure scales across managed SaaS and fully air-gapped on-prem deployments.

What You'll Bring:

  • 10+ years of MLOps, ML infrastructure, or machine learning engineering with a history of architectural ownership.
  • Proven track record in architecting large-scale model deployment (including LLMs) and ML infrastructure at scale.
  • Deep cloud expertise across major cloud platforms (AWS, Azure, or GCP) and strong proficiency in Python.
  • A mentorship record where engineers you have grown now operate independently at higher levels.
  • Deep comfort architecting ML systems that run in both managed SaaS and on-premises/disconnected air-gapped environments.
  • Architectural depth in Kubernetes, including GPU scheduling, multi-tenancy, operators, and understanding of failure modes of distributed workloads on shared clusters.
  • Strong communication, stakeholder management, and decision-making skills, with a passion for building diverse, inclusive engineering teams.

Strong Preference:

  • Experience owning production reliability at the platform level, including SLO definition, incident command, postmortem practice, and driving reliability improvements across teams rather than services.
  • Architectural experience with distributed training and fine-tuning at scale (DeepSpeed, FSDP, Megatron-LM), including cluster design, checkpointing strategy, and failure recovery.
  • Deep knowledge of GPU systems, including CUDA, NCCL, interconnect topology (NVLink, InfiniBand/RoCE), and diagnosing performance and communication problems across nodes.
  • Model optimization strategy at the portfolio level, including quantization (FP8, AWQ, GPTQ) and speculative decoding with measurable cost or latency outcomes across multiple systems.
  • Experience in regulated or security-constrained environments with compliance-driven architecture, model governance, lineage, audit, and secrets management.
  • Experience in on-prem/air-gap ML delivery architecture at scale.
  • A track record of maturing MLOps practice in a growing organization with standards, platform abstractions, and paved paths that continue beyond your involvement.
  • Experience with bare-metal GPU cluster architecture, including scheduling (Slurm or Kubernetes) and hardware lifecycle in customer or owned data centers.

Nice to Have:

  • Experience with conference speaking, technical writing, or industry thought leadership.
  • Contributions to open-source projects related to inference, serving, or ML infrastructure, particularly at a maintainer level.
  • C/C++ or CUDA kernel experience for performance-critical paths.
  • Proficiency in Arabic language.

Why Join Us:

  • Mission-Driven Work: Engage on cutting-edge AI applications with a dedicated team, tackling real-world challenges in critical sectors.
  • Unparalleled Opportunity: Seize the chance to innovate and address real-world challenges using AI at a company with unique access to world-leading models and resources.
  • Career Growth: Benefit from competitive compensation, benefits, and significant career advancement opportunities as a foundational team member.
  • World-Class Environment: Experience a flexible working environment equipped with the latest tools and technologies necessary for your success.
Apply Direct

Jobs you might like   View all jobs

About Technology, Information and Internet Company

Company details are hidden. Subscribe to view full company profile.

Ready to apply for this role?

Apply Direct