Develop and maintain production AI systems, ensuring delivery beyond research stages. Take ownership of the CI/CD process at scale, utilizing platforms like Azure DevOps and GitHub Actions. Lead a team in engineering standards across multiple squads, mentoring them in best practices.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
Location
Riyadh, Saudi Arabia
Requirements
- Bachelor's degree in Software Engineering, Computer Science, or a related field.
- 6–8+ years in software, DevOps, or platform engineering, including at least 2 years in an applied AI or ML engineering capacity.
- Proven delivery of production AI/LLM systems.
- Strong Python programming skills; comfortable with Bash and YAML.
- Hands-on experience with Kubernetes, Docker/Podman, and Terraform.
- Production experience with at least one major cloud service (Azure preferred; OCI or GCP acceptable).
- Experience leading a team and setting engineering standards across multiple squads.
- Arabic and English professional proficiency.
Preferred Qualifications
- Master's degree in Applied AI, Machine Learning, or a related discipline.
- Fine-tuning experience with QLoRA/LoRA on GPU clusters; familiarity with PyTorch and Transformers.
- Experience with vector databases (Milvus, Pinecone, or Weaviate) and RAG retrieval design.
- Prior experience delivering on Saudi government or large-scale national digital platforms, with knowledge of local compliance and standards.
Key Responsibilities
- Build, fine-tune, and evaluate LLM systems for domain-specific tasks (QLoRA / PEFT on open-weight models such as Llama-3 and Mistral).
- Design reproducible evaluation harnesses and A/B test frameworks.
- Architect multi-agent and RAG systems (LangGraph, FastAPI, vector databases) from prototype through production.
- Implement safety guardrails to reduce invalid or high-risk model actions.
- Translate business use cases into deployable prototypes with measurable acceptance criteria.
Platform & Infrastructure
- Design and operate cloud infrastructure and MLOps workspaces for AI workloads on Kubernetes.
- Build CI/CD pipelines and GitOps-based release promotion across development, test, and production environments.
- Implement end-to-end observability for defined detection and response targets.
- Own disaster recovery design with automated backups and failover.
Engineering Leadership
- Lead and mentor a cloud/AI operations team; define monitoring and incident response practices.
- Standardize SDLC practices to improve lead time and deployment frequency.
- Produce handover documentation and runbooks for operational transparency.
- Support vendor and licensing negotiations for cloud enterprise agreements.