Company logo hidden

LLM Research & Validation Specialist

Unlock employer Abu Dhabi, United Arab Emirates Direct to Company Under an hour ago · 22 Sep 2026

Financial

  • Estimate: $60k - $120k*
  • Zero income tax location

Accessibility

  • Office Only
  • Visa Provided

Requirements

  • Experience: Intermediate
  • English: Professional

Position

About the role:
Lead frontier research, quantitative evaluation, and independent validation of large language models, multimodal models, retrieval-augmented generation systems, and agentic AI used or proposed by the company. Translate mathematical and scientific methods into reproducible validation tests, challenger analyses, runtime controls, and decision-useful evidence for model governance. The role combines deep technical research with second-line effective challenge. It is expected to build validation toolkits and evaluation harnesses, independently assess conceptual soundness and production behaviour, and communicate material limitations clearly to technical teams, senior management, and governance forums. The role does not own model development or production approval.

Ready to apply for roles like this?

Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.

Unlock employer & apply directly

Key accountabilities / responsibilities:

  • Lead independent validation of LLM, multimodal, RAG, and agentic AI use cases across design, implementation, deployment, and ongoing monitoring.
  • Assess transformer architecture, tokenisation, embeddings, attention, context-window behaviour, decoding, fine-tuning, alignment, quantisation, and inference configuration.
  • Design reproducible evaluation harnesses, golden datasets, adversarial suites, counterfactual tests, canary sets, and statistically defensible acceptance criteria.
  • Evaluate task performance, hallucination and factuality, calibration, robustness, stability, long-context behaviour, retrieval quality, grounding, citation faithfulness, and uncertainty.
  • Perform deep testing of prompt injection, indirect injection, data leakage, tool-use safety, excessive agency, multi-step failure propagation, kill-switches, and human oversight.
  • Apply probability, statistics, optimisation, information theory, numerical methods, and experimental design to develop challenger tests and quantify uncertainty.
  • Review data provenance, representativeness, contamination, benchmark validity, leakage, drift, and limitations of synthetic or LLM-generated evaluation data.
  • Build and maintain reusable Python-based validation tooling, automated test pipelines, experiment tracking, results repositories, and technical documentation.
  • Conduct structured research on emerging model architectures, interpretability, mechanistic analysis, scalable oversight, model evaluation, and AI safety methods.
  • Independently challenge model owners, vendors, and developers, document findings, propose risk-based restrictions, and track remediation without assuming first-line ownership.
  • Prepare validation reports, research notes, standards, committee papers, and senior-management briefings that clearly distinguish evidence, judgement, and residual uncertainty.
  • Mentor junior validators, improve team methodology, and support knowledge transfer across Model Risk.

Education and experience:

  • Master's degree in Theoretical Physics, Applied Physics, Mathematics, Applied Mathematics, or a closely related quantitative discipline is required. A PhD or research-intensive master's is strongly preferred.
  • Typically, one to three years of relevant experience in AI research, machine learning, quantitative modelling, model validation, scientific computing, or a closely related field. Exceptional research profiles may be considered based on demonstrated capability.
  • Deep understanding of probability, statistics, linear algebra, optimisation, numerical computation, experimental design, and uncertainty quantification.
  • Strong understanding of transformers, LLM training and inference, embeddings, RAG, fine-tuning, alignment, evaluation, agentic systems, and AI safety failure modes.
  • Advanced Python proficiency and experience with scientific and ML libraries. Exposure to PyTorch, Hugging Face, evaluation frameworks, experiment tracking, SQL, Git, and cloud AI platforms is expected.
  • Ability to read research papers critically, reproduce methods, design controlled experiments, and convert findings into bank-grade validation evidence.
  • Experience with red teaming, adversarial testing, interpretability, calibration, robustness, privacy, security, or model risk management is strongly advantageous.
  • Excellent technical writing and communication, including the ability to explain mathematical concepts, assumptions, and limitations to non-specialist stakeholders.
  • Banking experience is advantageous but not mandatory. The role requires a willingness to develop knowledge of financial services, Islamic banking, CBUAE expectations, and the company governance.

Indicative success measures:

  • Validation conclusions are reproducible, evidence-based, and proportionate to use-case risk.
  • Reusable evaluation assets and automation measurably improve validation coverage, consistency, and efficiency.
  • Material LLM and agentic risks are identified early, clearly communicated, and translated into actionable controls or use restrictions.
  • Research outputs strengthen the company validation methodology and remain traceable to tested evidence rather than unsupported claims.
  • Stakeholders receive constructive, independent challenges while second-line ownership and decision rights remain clear.
Apply Direct

Jobs you might like   View all jobs

About Banking Company

Company details are hidden. Subscribe to view full company profile.

Ready to apply for this role?

Apply Direct