About the Role
We are looking for a Forward Deployed Engineer to join our team in Riyadh. In this hybrid role, you'll embed directly with enterprise and public-sector customers across the Kingdom to architect and ship production systems on our global network — voice, messaging, AI, and wireless — including self-hosted, open-weight LLM deployments that run inside customer environments. This isn't about demoing products; it's about building real solutions that work at scale.
This role would suit a technical professional who thrives at the intersection of engineering and customer success. You're equally comfortable debugging SIP traces as you are standing up a LiteLLM gateway with routing and failover across a customer's model providers, or leading a whiteboard session with a customer's engineering team. The ability to own outcomes from POC to production go-live is essential.
Saudi customers increasingly require Arabic-first conversational AI and in-Kingdom data residency, so a large part of this job is making open-weight models perform in production on infrastructure the customer controls. This is a customer-facing role conducted largely in Arabic — fluent spoken Gulf Arabic is a hard requirement.
This is a Riyadh-based role with 10–30% travel to customer sites for deployments, workshops, and escalations — primarily across the Kingdom (Riyadh, Jeddah, Dammam/Khobar, NEOM) with occasional travel elsewhere in the GCC.
Ready to apply for roles like this?
Unlock the company name and direct application link. Subscribers get instant access to fresh jobs across Dubai, Abu Dhabi and Riyadh, many with visa support.
Unlock employer & apply directly
Responsibilities
- Embed with enterprise and government customers to understand their communications workflows, AI use cases, and integration challenges firsthand
- Build and deploy custom implementations: AI Voice Assistants, APIs (Voice, Messaging, Fax, Wireless), WebRTC
- Deploy and operate open-weight LLMs (Llama, Qwen, Mistral, DeepSeek, gpt-oss, and Arabic-first models such as ALLaM, Fanar, and Jais) in customer environments, including air-gapped and in-Kingdom sovereign cloud deployments
- Deploy and operate LiteLLM as the model gateway in customer environments: a unified OpenAI-compatible interface across self-hosted open-weight models and hosted providers, with routing, load balancing, retries and fallbacks, rate limits, and per-team virtual keys
- Instrument and govern LLM usage through the gateway — cost tracking and budgets, caching, logging and observability (OpenTelemetry, Langfuse, or similar), and guardrails — so customers can see and control what their AI workloads are doing
- Design model routing strategies for real-time voice workloads, balancing latency, cost, and quality across Arabic-capable models, with sane fallback behavior when a provider degrades
- Make the build-vs-buy case between self-hosted open-weight models and hosted frontier APIs, and keep the customer's application code portable across both
- Adapt models to customer domains: prompt and RAG pipelines, LoRA/QLoRA fine-tuning, and evaluation harnesses for Arabic (including dialectal Arabic) and bilingual Arabic/English use cases
- Lead POCs, pilots, and production launches from whiteboard to go-live
- Own customer outcomes — stay engaged until the solution is live and stable
- Collaborate directly with Product and Engineering to shape the roadmap based on field insights from the Saudi and wider MENA market
- Create clear technical documentation, runbooks, and maintainable solutions for handoff, in English and where needed in Arabic
- Troubleshoot and resolve complex integration issues alongside customer teams
- Work with customers to meet local regulatory and data-residency requirements (CST/CITC, SDAIA, NCA, and PDPL obligations)
What We Are Looking For
- CS degree or equivalent experience
- 3+ years building software or doing technical consulting
- Proficiency in multiple languages: Python, Node.js, Go — you're more dangerous in some than others
- Hands-on experience running LiteLLM (or a comparable LLM gateway such as Portkey, Kong AI Gateway, or an in-house proxy) in production — config-driven model definitions, the proxy server, routing and fallback rules, and virtual key management
- Practical understanding of what breaks in front of a model in production: provider rate limits and quotas, timeout and retry behavior, streaming, token accounting and cost attribution, and the failure modes that only show up under concurrency
- Comfortable deploying containerized services on Kubernetes, with secrets management, config, and upgrades as part of the deployment story
- Strong API fluency, event-driven thinking, and cloud-native instincts
- Exposure to SIP, WebRTC, or real-time voice/messaging systems
- Ability to translate "it's not working" into root cause
- Comfortable working on customer sites and in high-stakes technical conversations
- Fluent spoken Gulf Arabic (Khaleeji) — required, not preferred. You'll run whiteboard sessions, live troubleshooting, and escalations in Arabic with customer engineering teams
- Professional working proficiency in English for internal collaboration, documentation, and work with Product and Engineering
- Based in Riyadh, or willing to relocate — this is a hybrid role with travel
- Legally authorized to work in Saudi Arabia, or eligible for sponsorship
Bonus Points For
- Experience with AI voice assistants, STT/TTS, or LLM-based conversational systems — especially Arabic ASR and speech synthesis
- Fine-tuning and post-training experience: LoRA/QLoRA, distillation, preference tuning, or building eval sets for a specific domain
- Familiarity with the Arabic open-weight model landscape and the trade-offs between Arabic-first and multilingual models
- Experience with an inference serving engine behind the gateway (vLLM, SGLang, TGI, Ollama) and sizing GPU capacity for self-hosted models
- SQL proficiency (Postgres, MySQL, Oracle)
- ETL and data wrangling experience
- DevOps fundamentals (Docker, Kubernetes, CI/CD)
- Background in telecom, CPaaS, or high-growth SaaS
- Experience with in-Kingdom sovereign or on-prem cloud deployments and PDPL-aligned architectures
- Security mindset (IAM, encryption, audit logging)