We are building the next generation of wargaming facilities purpose-built for the military to run massive-scale simulations of autonomous systems operating in contested environments. Those facilities are only useful if they are up, current, and trustworthy. A failed scenario run or a silent regression after a software release costs operators and engineers’ real time.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
As our founding Site Reliability Engineer, you will design, build, and operate the infrastructure that makes this possible. You'll work at the intersection of hardware, simulation software, and distributed systems. You will be the person who knows why the simulation broke, who catches it before anyone else notices, and who makes sure it doesn't break the same way twice. This is a hands-on role that blends software maintenance, infrastructure ownership, and release validation, with direct exposure to the development teams whose code you're keeping stable.
Responsibilities
- Maintain the simulation software stack - installation, configuration, updates, version management, and day-to-day functionality across the Simulation Center's tools and environments.
- Own the underlying infrastructure - compute, networking, storage, and environment configuration that the simulation depends on; keep it provisioned, patched, and performant.
- Build and maintain a post-release test suite - design, automate, and continually extend a regression and smoke-test process that runs after every software release or configuration change, so integration issues surface immediately rather than mid-exercise.
- Forecast, diagnose, and eliminate failure modes - root-cause errors and bugs in the system, drive them to permanent resolution, and implement the guardrails, monitoring, or process changes that prevent recurrence.
- Partner with development teams and stakeholders - review upcoming changes for reliability risk, surface concerns early, and implement mitigation strategies before releases land in the simulation environment.
- Monitor overall system health - instrument and watch the environment, triage issues within your scope, and escalate clearly and quickly with the context needed for others to act when an issue exceeds your ability to resolve it.
- Document what you learn - runbooks, known issues, environment configuration, and release validation results, so the Simulation Center's operational knowledge isn't held in one person's head.
Required Qualifications
- Proficiency in Python for automation, tooling, and test development.
- Working knowledge of C++; enough to read, debug, build, and trace issues in the simulation codebase.
- Solid general networking fundamentals: TCP/IP, UDP, multicast, DNS, routing, firewalls, and the ability to diagnose latency, packet loss, and connectivity problems across distributed systems.
- Experience with project management, issue tracking, bug triage, and coordinating work across engineering teams.
- Demonstrated experience maintaining production or production-adjacent systems, including troubleshooting under time pressure.
- Strong written and verbal communication; you can escalate an issue, explain a root cause, and write a runbook someone else can follow.
- Eligibility to pass the security and background check requirements for sensitive information systems.
Preferred Qualifications
- Experience with modeling and simulation, wargaming, or distributed simulation standards (DIS, HLA, TENA) and platforms such as AFSIM, VBS, or similar.
- Test automation and CI/CD experience; building automated validation pipelines, not just running them.
- Infrastructure-as-code and configuration management (Terraform, Ansible, Docker, Kubernetes).
- On-prem and cloud deployment experience.
- Observability tooling: Prometheus, Grafana, ELK, or equivalent.
- Linux systems administration depth; comfort in mixed Linux/Windows environments.
- Prior work in a defense, aerospace, or classified environment.
- Active security clearance.
Location
Abu Dhabi, United Arab Emirates; London, England, United Kingdom
Work Conditions
Candidate must be willing to relocate to the facility upon completion.