Company logo hidden

Technical Lead - GPU Infrastructure (100% Remote - Worldwide)

Unlock employer Abu Dhabi, United Arab Emirates · Oman Direct to Company Under an hour ago · 12 Sep 2026

Financial

  • Estimate: $120k - $180k*
  • Zero income tax location

Accessibility

  • Fully Remote
  • Apply from abroad
  • No Relocation Support
  • Visa Provided

Requirements

  • Experience: Senior
  • English: Professional

Position

About the role
Cosmic AC is the company Data's GPU compute and managed inference platform, featuring GPU containers, managed inference endpoints, and platform observability, delivered as a self-hosted package on Kubernetes, with a control plane written in JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare-metal GPU infrastructure, starting with a managed Slurm scheduling layer for internal research and model-training teams and eventually transitioning to our own Kubernetes control plane for inference tenancy.

Ready to apply for roles like this?

Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.

Unlock employer & apply directly

The Technical Lead will own the architecture and delivery of this stack and lead an engineering team of about twelve engineers across backend, frontend, DevOps, QA, and documentation, who are distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing, and partner commitments; the Technical Lead will manage architecture, implementation and delivery plans, line-manage the engineers, and serve as the primary technical interface to our infrastructure partners.

This position is a hands-on infrastructure leadership role with a fixed delivery window in its first six months. It is not a research role, not a pure Kubernetes SRE role, and not solely a management role.

Responsibilities

  • Architecture: Own the platform architecture end to end, including architecture proposals, high-level and low-level designs, and ensuring they are kept current as the baseline through review.
  • Team leadership: Lead and line-manage a distributed team across backend (Node.js), frontend (React), DevOps, QA, and documentation, including engineering standards, code and design review, release gates, one-to-ones, and growth and performance input.
  • Bare-metal GPU scheduling layer: Design, build, and operate a managed Slurm service for research users, covering controller and accounting, partitions and login nodes, node onboarding and acceptance, driver and CUDA baseline upgrades, stalled-job and node-health detection, drainage and autohealing, storage visibility, identity, and isolation.
  • Kubernetes control plane and GPU enablement: Own cluster bootstrap and lifecycle management on partner-provided bare metal, manage NVIDIA GPU Operator and Network Operator, VM-based GPU isolation, and day-2 operations, including upgrades and backups.
  • Managed inference at scale: Develop serving architecture, multi-GPU and multi-node parallelism, autoscaling, request routing, and endpoint reliability along with confidential-compute-capable capacity for sensitive workloads.
  • Observability and operations: Implement metrics, logging, alerting, and SLOs across the control plane, GPU fleet, and application tiers, manage incident response and post-incident reviews, and create a sustainable on-call model.
  • Partners and vendors: Serve as the primary technical interface to infrastructure partners and vendors, translating requirements into specifications and acceptance tests, managing escalations, and providing input for capacity planning and hardware sourcing.
  • Internal consumers: Collaborate directly with research, model-training, and product teams to translate their workloads into platform requirements and broker capacity when short.
  • Hiring: Complete the platform team and set the technical bar for new engineers.

Must have

  • Experience: Eight or more years of hands-on engineering experience, including at least three years in leading teams that build and operate infrastructure platforms. Bachelor’s or Master’s degree in computer science or engineering, or equivalent practical experience.
  • Slurm at scale: Hand-on experience running slurmctld and slurmdbd for real users, including partitions, QoS and priority, accounting, prolog and epilog, and node health scripting. Experience operating an HPC or GPU training cluster is ideal.
  • GPU fleet operation on bare metal: Knowledge of NVIDIA driver and CUDA lifecycle, Fabric Manager and NVSwitch behaviour on SXM systems, health and utilisation via DCGM, and node burn-in and acceptance procedures.
  • High-performance interconnects: Familiarity with InfiniBand fabric and subnet configuration, RDMA, SR-IOV, and diagnosing multi-node NCCL performance problems.
  • Linux systems depth: Experience with kernel modules and drivers, PCIe passthrough and vfio-pci, and performance tuning for compute-heavy workloads.
  • Production Kubernetes operation: Demonstrated experience in Kubernetes control plane operations, including upgrades, CNI and CSI, operators, custom controllers, and multi-tenancy design.
  • HPC storage and data movement: Understanding of shared filesystems (VAST, Lustre, NFS), node-local NVMe caching, and distributed large model weights and datasets across nodes.
  • Observability and operations: Proficiency in using Prometheus, Grafana and Loki or equivalents, defining SLOs, and managing incident response and post-incident reviews.
  • JavaScript and Node.js: Working fluency sufficient to review a control plane, CLI, and worker services and make architecture decisions.
  • Shipped platform: Experience with a multi-tenant IaaS or PaaS, or a research computing service including resource isolation, quotas, usage metering, and user-facing API and CLI surfaces.
  • Leadership in code: People management across time zones, a collaborative review process, recording written architecture decisions with alternatives, and the ability to communicate limits to partners and executives.
  • English proficiency: Excellent written and spoken English, as most partner and leadership interactions occur in writing.
  • Location: This is a fully remote position. Candidates should be based between UTC and UTC+5:30 to ensure overlap of working hours with team members and partners in Europe and India. Occasional travel to partner sites and team events may be required.

Desirable

  • Experience with Slurm operators on Kubernetes (Soperator, Slinky) or Kubernetes-native schedulers (Kueue, Volcano, KAI, Kubeflow Trainer).
  • Knowledge of modern serving stacks (vLLM, SGLang, TensorRT-LLM) including strategies for parallelism, quantisation trade-offs, and GPU memory planning.
  • Familiarity with VM and container isolation techniques for multi-tenant GPU compute (KubeVirt, Kata Containers, QEMU, KVM, Firecracker) and confidential computing (Intel TDX, AMD SEV-SNP, NVIDIA confidential-compute mode).
  • Experience with Cluster API, kubeadm, Cilium, and tools for autohealing, infrastructure as code, and GitOps.
  • Background working on the operator side of a GPU cloud, at a national or university HPC center, or within an AI lab’s platform team.
  • Experience in peer-to-peer or distributed systems.
  • Experience with hardware provisioners who do not operate, and translating that relationship into a formal contract with acceptance tests.
Apply Direct

Jobs you might like   View all jobs

About Financial Services Company

Company details are hidden. Subscribe to view full company profile.

Ready to apply for this role?

Apply Direct