Intern - International Services
Telecommunications Company Riyadh, Saudi Arabia
Unlock employer Dubai, United Arab Emirates Posted: 30 Jun 2026
As a member of the AI model team, you will drive innovation in multi-modal reinforcement learning to advance next-generation intelligent systems. Your work will focus on optimizing decision-making and adaptive behavior across integrated data modalities such as text, images, and audio to deliver enhanced intelligence, robust performance, and domain-specific capabilities for real-world challenges. You will develop and scale reinforcement learning techniques within complex multi-modal architectures, including diffusion-based generative models and autoregressive models for multimodal understanding, as well as resource-efficient models designed for constrained hardware environments. This includes conducting research on reinforcement learning algorithms for multimodal models, spanning diffusion models for image autoregressive models for multimodal reasoning, and unified multimodal frameworks.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directlyYou are expected to have deep expertise in designing multi-modal reinforcement learning systems and a strong background in advanced model architectures, with a hands-on, research-driven approach to building and deploying novel algorithms and training frameworks. You will design and develop RL infrastructure and reward modeling strategies to enable efficient large-scale training, improve training stability, and mitigate reward hacking and related failure modes. Your responsibilities also include curating multi-modal simulation environments and training datasets, improving baseline policy performance across modalities, and identifying and resolving bottlenecks in multi-modal learning and reward optimization. In addition, you will explore next-generation reinforcement learning paradigms that more directly and effectively learn from environment feedback, with the goal of unlocking superior, domain-adapted AI performance in dynamic, real-world environments.
Responsibilities
Requirements
Location
United Arab Emirates, Dubai
Telecommunications Company Riyadh, Saudi Arabia
IT Services and IT Consulting Company Abu Dhabi, United Arab Emirates
Fintech Company Abu Dhabi & Dubai, United Arab Emirates · Manama, Bahrain
Financial Services Company Abu Dhabi & Dubai, United Arab Emirates · Riyadh, Saudi Arabia
Financial Services Company Dubai, United Arab Emirates
Financial Services Company Dubai, United Arab Emirates
Ready to apply for this role?
Apply Direct