Company logo hidden

AI Product Evaluation Lead

Unlock employer Dubai, United Arab Emirates Direct to Company Under an hour ago · 03 Sep 2026

Financial

  • Estimate: $80k - $120k*
  • Zero income tax location

Accessibility

  • Apply from abroad
  • Visa Provided

Requirements

  • Experience: Intermediate
  • English: Professional

Position

About
We are the AI Vertical Scenarios Team under the Corporate Information Systems department. Our core mission is to serve the company's global internal office scenarios as the standard-setter for AI-native applications and the driver of globally consistent experiences. You will be deeply involved in the entire journey of an enterprise AI product — from idea to launch!

Ready to apply for roles like this?

Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.

Unlock employer & apply directly

Responsibilities

  • Lead evaluation product planning and system design for core scenarios such as AI Agents and Coding Agents, designing the evaluation framework, workflow, and product mechanisms, turning frontier capabilities into reusable and continuously evolving evaluation systems;
  • Uncover the real needs of enterprise users by distilling core pain points through interviews, data analysis, and behavioral insights, and drive them into the product;
  • Drive evaluation engineering and cross-functional collaboration by working closely with algorithm, engineering, and data teams to build evaluation platforms, automated pipelines, and result dashboards so evaluation can directly support model optimization and product decisions.

Requirements
Minimum Qualifications

  • Bachelor's degree or above in Computer Science, Artificial Intelligence, Information Management, Statistics, or related fields;
  • Solid structured thinking and sustained interest in the intersection of evaluation, data, and AI products;
  • Clear logical thinking, self-driven, and user-value oriented, with strong communication and cross-team collaboration skills;
  • Genuine interest in AI products, model evaluation, and benchmark systems, with a habit of following the latest developments in the LLM/Agent space;
  • Familiarity with AI/Agent tools such as Claude Code, Cursor, Codex.

Preferred Qualifications

  • Experience using, analyzing, or building around third-party benchmarks such as SWE-Bench is a plus;
  • Understanding of evaluation automation, data processing, or experiment pipelines is preferred; hands-on experience building agents, evaluation platforms, or AI applications is a plus.

Location
Dubai

Apply Direct

Jobs you might like   View all jobs

About Software Development Company

Company details are hidden. Subscribe to view full company profile.

Ready to apply for this role?

Apply Direct