31 - Data Engineer
Tunga · Sweden
You apply on the site where the job is posted. I never handle applications.
Our client is looking for a Data Engineer to join a 10-person team building and operating the data platform behind a large European enterprise's analytical applications.
This is a five-month engagement.
The work sits on production systems. The platform is already live and feeding analytical applications that people rely on, and it is still evolving — new sources land, pipelines get reshaped, infrastructure gets rebuilt underneath you. You'll be designing and operating pipelines in that environment rather than building something greenfield in isolation.
This is an individual-contributor role. You won't be managing anyone, and you won't be handed fully specified tickets either. The team expects you to take a data problem, work out how the pipeline should be shaped, build it, and then own it in production.
Mid-level here means you've already done this work somewhere real. You've had a pipeline break at an inconvenient hour and you've fixed it. You know what a backfill costs. You don't need someone standing over you to get a PySpark job into production.
What you'll be doing
- Designing and building data pipelines that feed the client's analytical applications
- Operating and maintaining existing production pipelines — monitoring, debugging, fixing, improving
- Writing and optimising distributed processing jobs in PySpark
- Building and maintaining orchestration in Airflow: DAG design, dependencies, scheduling, retries, backfills
- Developing ETL workflows from source systems into the analytical layer
- Modelling and writing SQL against PostgreSQL for analytical workloads
- Contributing to the continuous evolution of the client's cloud-based (AWS) data infrastructure
- Working within existing CI/CD pipelines and Git workflows — reviewed code, tested changes, no direct-to-production edits
- Investigating data quality issues and making pipelines more resilient to the ones that keep coming back
- Collaborating with the wider platform and analytics teams on what the data needs to look like downstream
Required skills and experience
- Python — strong, production-level. This is the primary language of the role.
- PySpark — hands-on experience writing and tuning distributed processing jobs
- Airflow — you've built and operated DAGs, not just triggered someone else's
- AWS — practical experience with cloud-based data infrastructure
- PostgreSQL and SQL — comfortable writing and reasoning about analytical queries
- ETL design — you can take a source system and a target and work out the pipeline in between
- CI/CD and Git — you work inside a pipeline, with branches, reviews and automated checks
- Experience on live production data systems — not only development environments, academic projects or course work
- Comfortable in a fast-paced delivery environment where the platform changes while you're working on it
- Professional English, spoken and written, and the communication habits that make fully remote work function
Nice to have
- Hadoop ecosystem experience
- Data modelling for analytics and BI consumption
- Data quality, monitoring or observability tooling
- Exposure to infrastructure as code or containerised workloads
- Experience working with an enterprise client or in a regulated environment