Senior ML / Evaluation Engineer
EPAM Systems · Remote, Portugal
Apply directly with the employer or job board. Applications are never handled here.
We're looking for a Senior ML / Evaluation Engineer to join our team in Portugal in a fully remote working mode. In this role, you will own the design and implementation of advanced evaluation frameworks for an Enterprise Agent Development Platform—a production-grade, cloud-native ecosystem enabling scalable, secure AI agent deployment. You will create evaluation strategies that combine LLM-as-judge grading with deterministic checks, define enterprise evaluation standards, and implement CI/CD deployment gates to enforce quality metrics prior to release. This position requires strong expertise in ML system testing, evaluation design, and integration into automated pipelines for agentic environments.
Responsibilities
- Design and implement multi-layer evaluation frameworks for agentic workflows and AI-driven applications
- Build LLM-as-judge evaluators leveraging AWS AgentCore built-in modules and custom logic for correctness and helpfulness checks
- Develop deterministic evaluators as AWS Lambda functions for rule-based validation
- Define enterprise evaluation standards, including mandatory dimensions, scoring criteria, and pass/fail thresholds
- Implement CI/CD deployment gates using on-demand evaluation modes to enforce quality in automated pipelines
- Enable online evaluation in production by integrating sampling-based evaluation strategies and PII detection guardrails
- Incorporate observability signals (OpenTelemetry spans) from AWS AgentCore into grading frameworks for trace-level assessment
- Generate metrics, logs, and dashboards from evaluation outcomes via CloudWatch or equivalent monitoring platforms
- Collaborate with platform, orchestration, and DevOps teams to maintain evaluation reliability and scalability
Requirements
- 5+ years of experience in ML engineering, AI evaluation frameworks, or AI platform development
- Hands-on expertise designing LLM evaluation frameworks (LLM-as-judge and deterministic graders)
- Practical experience implementing CI/CD deployment gates for ML model or AI agent quality assurance
- Proficiency in Python for building evaluation logic (deterministic Lambda-based evaluators)
- Strong understanding of advanced validation dimensions, including multi-turn context integrity and workflow-level scoring
Nice to have
- Familiarity with AWS AgentCore Evaluations API (CreateEvaluation, GetEvaluationResult)
- Exposure to AWS Bedrock Guardrails for compliance and sensitive data validation
- Experience integrating evaluation metrics into AWS CloudWatch for monitoring and alerting
- Knowledge of OTel instrumentation and trace ingestion for quality scoring inputs
We offer
- Competitive compensation depending on experience and skills
- Variety of projects within one company
- Being a part of a project following engineering excellence standards
- Individual career path and professional growth opportunities
- Internal events and communities
- Flexible work hours
Keep looking
100+ English-friendly jobs in Remote
Every one checked for language requirements, updated daily.