Senior Site Reliability Engineer (SRE)
EPAM Systems · Remote, Spain
You apply on the site where the job is posted. I never handle applications.
We're looking for a Senior Site Reliability Engineer (SRE) to join our team in Spain in a remote working mode. In this role, you will collaborate with development, operations, security and quality teams to ensure highly reliable, scalable and efficient systems for business-critical applications in the financial domain. You will focus on implementing SRE practices, reducing toil through automation and driving operational excellence while meeting strict Service Level Objectives (SLOs).
This position offers the opportunity to influence system design for reliability and performance within a global delivery context, leveraging modern cloud technologies, observability tools and automation frameworks to maintain seamless user experiences.
Responsibilities
- Define and maintain Service Level Objectives (SLOs), SLIs and error budgets for critical services
- Collaborate with cross-functional teams to embed reliability into application and infrastructure design
- Automate operational tasks to reduce manual toil and improve service performance
- Troubleshoot and resolve infrastructure and application incidents quickly and effectively
- Implement robust monitoring and observability systems to detect and prevent outages
- Plan capacity and scaling strategies to ensure high availability and resiliency
- Contribute to incident postmortems and continuous improvement initiatives
- Support the adoption of SRE best practices across all SDLC stages
Requirements
- Bachelor’s degree in Computer Science, Engineering or related field
- Proven experience working in cloud environments (AWS, GCP or Azure)
- Practical knowledge of SRE principles (SLO/SLI design, error budgets, postmortems, automation)
- Proficiency in Python or other scripting language for automation tasks
- Strong understanding of monitoring tools and observability frameworks
- Experience with Infrastructure-as-Code and CI/CD tools (e.g., Terraform, Ansible, Jenkins, GitLab)
- Hands-on expertise with containerization and orchestration platforms such as Docker and Kubernetes
Nice to have
- Experience deploying and managing Large Language Models (LLMs), including RAG-based solutions
- Certifications in Kubernetes, AWS/GCP/Azure or related cloud technologies
- Background in DevOps practices and agile delivery frameworks
- Familiarity with AI/ML model operations: deployment, monitoring and optimization in production environments
We offer
- Private health insurance
- EPAM Employees Stock Purchase Plan
- 100% paid sick leave
- Referral Program
- Professional certification
- Language courses
EPAM is a leading digital transformation services and product engineering company with 61,700+ EPAMers in 55+ countries and regions. Since 1993, our multidisciplinary teams have been helping make the future real for our clients and communities around the world. In 2018, we opened an office in Spain that quickly grew to over 1,450 EPAMers distributed between the offices in Málaga, Madrid and Cáceres as well as remotely across the country. Here you will collaborate with multinational teams, contribute to numerous innovative projects, and have an opportunity to learn and grow continuously.
- Why Join EPAM
- WORK AND LIFE BALANCE. Enjoy more of your personal time with flexible work options, 24 working days of annual leave and paid time off for numerous public holidays.
- CONTINUOUS LEARNING CULTURE. Craft your personal Career Development Plan to align with your learning objectives. Take advantage of internal training, mentorship, sponsored certifications and LinkedIn courses.
- CLEAR AND DIFFERENT CAREER PATHS. Grow in engineering or managerial direction to become a People Manager, in-depth technical specialist, Solution Architect, or Project/Delivery Manager.
- STRONG PROFESSIONAL COMMUNITY. Join a global EPAM community of highly skilled experts and connect with them to solve challenges, exchange ideas, share expertise and make friends.