AI Solution/Principal Engineer
We are seeking an AI Solution/Principal Engineer to join a team building a sovereign, multi-tenant agentic AI platform together with the applied products running on top of it. Every component ships under strict data residency constraints and operates in both Arabic and English. Our squads are small and senior, and we need someone who owns how the organization knows its AI systems perform — evaluation across assistants, retrieval pipelines, agent workflows, voice and document systems, and the infrastructure that transforms "it seems better" into evidence someone can act on.
This is not a QA role with AI added. Test automation and performance testing are in scope, but the center of gravity is evaluating non-deterministic systems, where the same input produces different outputs, "correct" is often fuzzy, and failure modes such as hallucination, grounding failure, drift, and prompt injection are ones no assertion library catches.
Responsibilities
- Define evaluation strategy: what gets measured, at which layer, with what methodology, and how results feed product decisions
- Build the shared platform, covering evaluation harnesses, golden-set management, dataset versioning, automated grading, regression detection, and reporting leadership can read
- Ensure grading is trustworthy through judge model selection, rubric design, calibration against human labels, and knowing when automated grading cannot be trusted
- Assess retrieval and agents for grounding and citation correctness, tool-use validation, multi-step reasoning and failure recovery
- Create Arabic golden sets and judges calibrated for Arabic rather than assumed to transfer from English
- Execute online evaluation and drift detection
- Perform prompt injection, jailbreak and data-leakage red-teaming
- Integrate evaluation and test automation together as quality gates in CI/CD
Requirements
- 5+ years of experience building evaluation or test infrastructure others depend on, including harnesses, shared libraries and frameworks
- At least 1 year of relevant leadership experience
- Expertise in AI evaluation spanning output quality, retrieval and grounding, and regression detection on non-deterministic behavior
- Proficiency in LLM-as-judge methodology: rubric design, calibration against human labels, and understanding of its failure modes
- Proficiency in Python at an advanced level, deep enough to build shared libraries, with intermediate competency in TypeScript or Java
- Knowledge of statistics for non-deterministic systems, covering sampling, confidence intervals, inter-rater agreement and significance
- Skills in CI/CD framework design across API, web and data surfaces, with test selection, parallelization and flake management
- Background in golden sets and dataset versioning
- Proficiency in English at an Upper-Intermediate level (B2) or higher
Nice to have
- Familiarity with ragas, DeepEval, Promptfoo, Braintrust, LangSmith or Langfuse tools
- Expertise in Arabic evaluation, spanning golden sets, dialect coverage, right-to-left validation and judge calibration
- Background in adversarial and security testing, including red-teaming practices
- Capability to carry out voice and conversational evaluation, plus performance testing against AI services
- Showcase of delivery in a regulated or government environment
Benefits
CONTINUOUS UPSKILLING, LEARNING & DEVELOPMENT
- Diversity of tasks and projects
- Assessment center for objective review of competency level
- Personal development plan
- Mentoring programs and leadership development
- Certification and professional development support
- Access to learning platforms including more than 2,500 internal courses
- English courses taught by certified teachers
CORPORATE BENEFITS
- Extra leave days
- Referral bonuses
COMPENSATION PACKAGE
- Competitive compensation paid in USD
- Regular salary and performance reviews
MEDICAL & HEALTHCARE
- Private health insurance
- Well-being events
WORKING ENVIRONMENT
- Recreation areas and kitchens
- Tea, coffee and snacks
- Sports equipment and game consoles
- IT Equipment
- Microsoft’s Software Assurance Home Use Program (HUP)
About the Company
More jobs at EPAM Systems
-
Lead AI Engineer with Java 17
Turkiye · remote · Oct 9, 2026
-
Senior Software Engineer (Risk Batch)
· · Oct 9, 2026
-
Senior Full-Stack Engineer
Georgia, Kazakhstan, Kyrgyzstan, Uzbekistan · remote · Oct 9, 2026
-
AI/Machine Learning Engineer
Georgia, Armenia, Kazakhstan, Kyrgyzstan, Uzbekistan · remote · Oct 10, 2026
-
Senior Data Engineer
Netherlands · remote · Oct 10, 2026