About the Role
You will build the internal AI systems that our engineers rely on daily. This includes developing agents that generate and maintain tests, creating pipelines to triage failures before they reach human eyes, and providing the necessary tooling to speed up code review and debugging. Your performance will be measured by the actual usage of the systems you deliver, making it a builder's role where you'll write production code and own systems in Continuous Integration (CI).
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
What Will You Do
Test automation across the stack
- Backend: Implement API and contract testing, service-level and integration coverage, create a stable data setup, and design tests that are resilient to schema changes.
- Frontend: Develop end-to-end (E2E) and component-level coverage using Playwright, conduct visual and RTL regression, and ensure that suites are quick enough to block merges instead of running overnight.
- Mobile: Create coverage for native and cross-platform applications using Appium or Maestro, establish a device-farm strategy, test offline and sync behavior, and ensure functionality with payment peripherals that only fail on real hardware.
- The connective tissue: Manage shared fixtures, environment and test-data, execute parallelization, and establish CI pipelines where a red build has significance.
- Experience in performance and load testing is also expected.
The AI layer on top of it
- Generate tests from specifications, code, and production traffic, while ensuring maintenance is addressed beyond initial drafts.
- Develop failure triage systems that classify red builds into categories (real bugs, flake, environment, or test rot) before human intervention.
- Create self-healing locators and suite health tools, including flake detection and coverage-gap analysis.
- Set up evaluation infrastructure for AI features across products, including datasets, scoring, and regression detection when prompts or models change, especially focused on Arabic and English behavior, RTL interfaces, and region-specific rules for POS, tax, and payments.
Agentic AI and orchestration
- Develop Agentic AI that performs tasks within pipelines: reading diffs, running relevant suites, reproducing failures, proposing fixes, and submitting pull requests.
- Create agents that manage quality workflows comprehensively, conducting exploratory testing, identifying coverage gaps, and assessing release risks, with escalation to humans when necessary.
- Ensure orchestration that withstands high load: execute multi-step plans, utilize tools, manage retries, maintain state across steps, handle sandboxed executions, and define boundaries between agentic and deterministic steps.
- Integrate with existing stacks (CI, Jira, observability) rather than operate in isolation.
- Apply sound judgement to determine when straightforward pipelines outshine agentic solutions.
The technical ground
You should be knowledgeable about the current practices in these areas, capable of discussing details rather than simply reciting:
- Test Automation: framework design, test pyramid applicability, flake economics, parallel execution, cross-browser realities, CI/CD gating, and familiarity with frameworks like Playwright, Appium, and Maestro.
- Agentic AI: orchestration methods, multi-step planning, management of memory and state, sandboxed execution, and integration of tools, as well as understanding the costs involved.
- Context Engineering: retrieval strategies, chunking, reranking, caching, and the management of long-context behavior.
- Evaluation: performance evaluations (offline and online), LLM-as-judge dynamics, human-in-the-loop reviews, and establishing statistical significance.
- Reliability: structured output management, design of guardrails and fallback processes, and handling non-determinism.
What Are We Looking For
- An engineer with a history of shipping production software and recent hands-on experience with LLM-backed systems relied upon by real users.
- Proficiency in Python, and comfortable with at least one of the following: .NET, Java, or TypeScript. Your experience should reflect tested and maintained code, rather than just notebooks.
- In-depth automation experience across multiple surfaces, demonstrated by owning a suite that manages backend and UI releases without compromising the quality of tests.
- Experience building evaluation systems with quantifiable improvement metrics.
- Practical depth in utilizing modern LLM toolkits: prompting, structured output, tool use, retrieval, and agentic orchestration, with a good understanding of trade-offs.
- Strong familiarity with testing principles, including test design and automation frameworks; a formal QA title is not necessary.
- A focus on adoption in your work; you value being measured by how much other engineers utilize what you produce, rather than just by your demonstrations.
What We Offer You
We believe you will love working here!
- An inclusive and diverse culture with a strong emphasis on innovation and flexible hybrid work setups.
- Highly competitive compensation packages, including bonuses and potential share options.
- A commitment to personal development with regular training sessions and an annual learning stipend to foster career growth in a rapidly evolving environment.
- The opportunity to join a talented team comprising over 30 nationalities operating in 14 countries, offering valuable experience in an exciting industry.
- Autonomy, mentorship, and challenging objectives that present substantial opportunities for both individual and company growth.