About the Role
This is a senior technical leadership role requiring strong hands-on experience across modern data and AI platforms, DevOps/DataOps, production AI, observability, and Service Management integration. The successful candidate will be responsible for establishing and scaling the Data & AI ProductOps practice, implementing clear and reusable standards, and ensuring operational readiness and effective support models in client environments.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
Key Success Measures
- Data & AI ProductOps practice established, adopted and scaled.
- Clear and reusable standards established across DataOps, MLOps, LLMOps and AgentOps.
- ProductOps capabilities successfully implemented within clients’ environments.
- Clear production ownership, operational readiness and L0/L1/L2/L3 support models established.
- Improved product health and SLO compliance.
- Reduced incident detection and restoration time.
- Reduced recurring production issues and unnecessary engineering escalations.
- Increased automation, proactive monitoring and self-service.
- Strong integration between Data & AI ProductOps and Enterprise Service Management.
- Increased reuse of standards, patterns, playbooks and delivery accelerators.
- Strong internal capability development and effective client knowledge transfer.
Responsibilities
Practice Development
- Establish and lead the Data & AI ProductOps practice, including the operating model, standards, methods, governance, reusable patterns, playbooks and implementation approach.
- Define the practice across DataOps, MLOps, LLMOps and AgentOps, with clear standards for operating Data, BI, ML, GenAI and Agent-based products in production.
- Develop reusable ProductOps assets, including operational readiness standards, support models, SLO frameworks, monitoring patterns, runbooks, implementation templates and delivery accelerators.
Product Operations
- Define and implement operational readiness requirements covering product ownership, criticality, support levels, SLAs/SLOs, monitoring, alerting, runbooks, escalation, recovery, dependencies and rollback.
- Establish clear L0/L1/L2/L3 support models, with L0 focused on automation and self-service, L1 on Service Desk support and initial triage, L2 on ProductOps-led operational support and product-level troubleshooting, and L3 on complex issues requiring Engineering expertise.
- Establish product health and observability across availability, performance, data freshness and quality, pipeline and integration health, model and AI performance, usage, cost and other product-specific operational measures.
- Lead product-level incident investigation, restoration and root-cause analysis, coordinating across Engineering, Platform, Service Management, Security and vendors.
- Integrate ProductOps with enterprise Incident, Problem, Request, Change, Release, Knowledge, Service Level and Major Incident Management processes.
- Ensure recurring incidents and operational issues are converted into permanent corrective actions, automation opportunities and product improvement backlog items.
- Establish effective change and release practices supported by automated testing, CI/CD, versioning, controlled deployment, and rollback.
- Establish regular Product Operations reviews covering product health, SLOs, incidents, recurring problems, operational trends, technical debt, improvements and releases.
Client Delivery
- Assess clients Data & AI ProductOps maturity and define target operating models, support structures, service levels, observability, automation and implementation roadmaps.
- Design and implement ProductOps capabilities for Data, BI, ML, GenAI and Agent-based products within clients’ environments.
- Lead clients solutioning, architecture workshops and implementation engagements, defining roles, responsibilities, support models, operational processes, tooling and Service Management integration.
- Provide architecture and implementation oversight for complex Data and AI production environments, engaging senior stakeholders across Product, Engineering, Platform, Governance, Security and Service Management.
- Support proposals, technical solutioning and advisory engagements across Data & AI ProductOps, DataOps, MLOps, LLMOps and AgentOps.
Leadership Expectations
- Build and scale a new Data & AI ProductOps practice.
- Provide strong technical leadership while remaining hands-on where required.
- Lead multidisciplinary teams across Data, AI, Engineering, Platform, Product and Service Management.
- Provide architecture and implementation oversight across complex enterprise environments.
- Engage confidently with senior clients and internal stakeholders.
- Translate emerging Data and AI technologies into practical enterprise operating practices.
- Develop internal talent and build sustainable technical capability.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, Data/AI, Information Systems, or a related field.
- Relevant certifications across cloud, Data & AI platforms, DevOps/MLOps, Service Management, architecture, or data management are an advantage.
- 10+ years of experience across Data Engineering, AI/ML, DataOps, DevOps, SRE, Product Operations, Platform Engineering or related disciplines.
- Strong hands-on experience with modern enterprise data platforms and production Data/AI solutions.
- Demonstrated experience establishing or leading production operating capabilities for Data and AI products.
- Strong practical experience across DataOps, MLOps, LLMOps and AgentOps.
- Strong experience with operational readiness, observability, SLAs/SLOs, monitoring, support models, incident and problem management, CI/CD, automated testing, release management, recovery and rollback.
- Experience operating BI and data products, data pipelines, APIs, integrations, ML models, GenAI applications and AI Agents in production.
- Practical experience with model lifecycle management, model serving, drift monitoring, RAG, vector search, AI evaluation, AI observability, Agent tracing, tool execution and guardrails.
- Strong understanding of enterprise ITSM and Service Management, with experience integrating Data and AI product teams into established support, incident, problem, change and release processes.
- Strong incident leadership, troubleshooting, root-cause analysis and production problem-solving capabilities.
- Demonstrated experience developing technical standards, operating models, reusable patterns and implementation methods.
- Strong client-facing consulting, architecture, solutioning and stakeholder-management capabilities.
- Strong experience working with government agencies or large enterprises in Qatar or the GCC is highly preferred.
Preferred Technology Experience
Experience across a combination of the following technology areas:
- Data & Analytics Platforms: Databricks, Informatica IDMC, Microsoft Fabric, Power BI, Azure Data Services, Spark and Delta Lake.
- Data Integration & Streaming: Informatica, Kafka, APIs, batch integration, streaming and CDC patterns.
- AI, ML & GenAI: MLflow, Azure AI/ML, model serving and monitoring, RAG, vector search, AI evaluation and AI observability.
- GenAI & Agent Frameworks: LangChain, LangGraph, Semantic Kernel or equivalent GenAI and Agent orchestration frameworks.
- DevOps & Platform Engineering: Azure DevOps, GitHub, CI/CD, Docker, Kubernetes, Terraform, Python and SQL.
- Observability & Monitoring: OpenTelemetry, Azure Monitor, Application Insights or equivalent enterprise monitoring platforms.
- Service Management: ServiceNow or equivalent enterprise ITSM platforms.