Company logo hidden

Senior Engineer - HPC Operations

Unlock employer Abu Dhabi, United Arab Emirates Posted: 10 Aug 2026

Financial

  • Estimate: $90k - $120k*
  • Zero income tax location

Accessibility

  • Office Only
  • Apply from abroad
  • Visa Provided

Requirements

  • Experience: Senior
  • English: Professional

Position

The company, a leader in AI-powered cloud and digital infrastructure, is seeking a highly skilled Senior Engineer – HPC Operations to oversee the daily operations and support of high-performance computing clusters designed to power large-scale AI and ML workloads. This role ensures stable, secure, and high-performing infrastructure while leveraging technologies such as Slurm, Kubernetes, and modern MLOps platforms. The ideal candidate will possess deep technical expertise in HPC and a strong operational mindset to drive continuous improvement and automation across globally distributed environments. Responsibilities will include collaborating with multidisciplinary teams, leading complex projects, implementing cutting-edge technologies, and mentoring operations engineers.

Ready to apply for roles like this?

Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.

Unlock employer & apply directly

Key Responsibilities:

  • Lead daily operational support of HPC infrastructure including compute, storage, networking, and scheduling components (Slurm, Kubernetes, etc.).
  • Maximize the efficiency and performance of HPC systems, ensuring optimal resource utilization and minimal downtime.
  • Act as the primary technical escalation point for L2 support teams and ensure prompt resolution of incidents and service requests.
  • Monitor system health, performance, and utilization using advanced tools (e.g., Prometheus, Grafana, DCGM).
  • Manage user environments for AI/ML workloads, including container orchestration (e.g., Docker, Kubernetes) and workflow tools (e.g., MLflow, Kubeflow).
  • Implement and manage job scheduling policies, priorities, and partitions within Slurm and/or Kubernetes to ensure fairness and efficiency.
  • Conduct root cause analysis (RCA) of operational issues and contribute to post-mortem documentation and continuous improvement efforts.
  • Mentor junior engineers and participate in on-call rotation if required.
  • Ensure compliance with security and operational policies; assist in audits and documentation for change and incident management processes.

Qualifications:

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or related technical field.
  • 7+ years of experience in HPC operations, systems engineering, or DevOps roles.
  • Advanced knowledge and expertise in configuring, optimizing, and maintaining complex HPC environments, including hardware, software, and storage systems.
  • Hands-on experience managing Slurm clusters and/or Kubernetes-based environments for AI/ML workloads.
  • Expert knowledge of GPU resource management, workload schedulers, and performance tuning for AI/ML workloads.
  • Experience with monitoring and observability frameworks such as Prometheus, Grafana, and DCGM.
  • Strong scripting and automation skills (Python, Bash, Ansible, Terraform).
  • In-depth understanding of Linux (RHEL/CentOS/Ubuntu), networking concepts (RDMA, InfiniBand, RoCE), and storage technologies (NFS, Lustre, Ceph).

Language Requirements:

  • Proficiency in English is assumed, but other language skills may be advantageous.

What We Offer:

  • Competitive Salary: Attractive salary package based on skills and experience.
  • Yearly Bonus: Performance-based annual bonus recognition.
  • Exclusive Discount Cards: Access special benefits with Esaad and Fazaa cards across various services.
  • Premium Family Insurance: Comprehensive health coverage, including dental, vision, and life insurance for you and your family.
  • Learning & Development: Access to top-tier learning platforms for career growth with unlimited access to premium courses.
Apply Direct

Jobs you might like   View all jobs

About Artificial Intelligence Company

Company details are hidden. Subscribe to view full company profile.

Ready to apply for this role?

Apply Direct