Senior AI Platform Engineer
InventYou · Stockholm, Sweden
You apply off-site, with the employer or the job board. I never handle applications.
We are looking for a Senior AI Platform Engineer to design, build, and operate a production-grade AI platform within a complex enterprise environment.
In this role, you will take end-to-end technical ownership of the AI platform, including AI gateway engineering, governance, agent management, access control, cost telemetry, observability, and developer enablement. This is a highly hands-on role for someone who has already built an AI governance/platform capability in production and can bring that experience into a new environment.
Key Responsibilities* Build and operate an enterprise AI Gateway, including SSO, request logging, data-classification tagging, policy enforcement, and model routing
- Work with Azure OpenAI / AI Foundry, Entra ID, and Azure API Management (APIM)
- Build and maintain a governed repository of approved AI agents, including ownership, permissions, and data scope
- Develop cost telemetry to attribute AI usage to specific cost centres and use cases
- Automate access provisioning and licence lifecycle management for enterprise AI tools
- Develop reusable developer templates and golden paths for RAG, AI agents, and evaluation
- Establish production observability, monitoring, and operational runbooks
- Build and operate containerised workloads on Kubernetes
- Implement identity and access controls including SSO, RBAC, and service principals
- Build internal platform capabilities and developer services that can be adopted across engineering teams
- Implement Infrastructure as Code and CI/CD practices for reliable platform delivery
Requirements
- Previous hands-on experience building an AI governance/platform capability end-to-end in production
- Ability to demonstrate a previous AI platform implementation, including what was built, what it governed, who used it, and lessons learned
- Strong production experience with Kubernetes and containerised workloads
- Hands-on experience with AI gateways / LLM proxies, such as APIM, Kong, LiteLLM, or equivalent
- Practical experience developing LLM applications, including API integration, token/cost behaviour, evaluation, and RAG patterns
- Strong production observability experience using OpenTelemetry, Grafana, Azure Log Analytics, or equivalent
- Strong backend engineering skills in Go (preferred), Python, or TypeScript
- Hands-on experience with Infrastructure as Code and CI/CD, including Terraform and GitHub Actions or Azure Pipelines
- Strong experience with identity and access management, including SSO, RBAC, and service principals
- Proven experience building an internal platform or developer service successfully adopted by other teams
- Strong ownership, problem-solving, and independent working capabilities
Tech StackAI & Azure: Azure OpenAI, Azure AI Foundry, APIM
Identity: Entra ID, SSO, RBAC, Service Principals
Backend: Go, Python, TypeScript
Platform: Kubernetes, Containers
AI Gateway: APIM, Kong, LiteLLM or equivalent
Observability: OpenTelemetry, Grafana, Azure Log Analytics
Infrastructure & CI/CD: Terraform, GitHub Actions, Azure Pipelines
AI Engineering: LLM APIs, RAG, Agents, Evaluation
You will be a great candidate for us if you* Have already built and operated an enterprise AI platform rather than only experimented with AI technologies
- Can take end-to-end technical ownership with a high degree of independence
- Understand both AI engineering and production platform engineering
- Build platforms and developer services that other engineering teams want to use
- Can balance governance, security, cost control, and developer experience
- Are comfortable making technical decisions in complex enterprise environments
Benefits
Why join inventYOU* Build and shape enterprise-scale AI platform capabilities
- Work hands-on with modern Azure and Generative AI technologies
- Take ownership of technically complex and business-critical solutions
- Work across AI engineering, cloud, platform engineering, security, and DevOps
- Contribute to the adoption of AI across large-scale enterprise environments