Description
Data-engineering role focused on scalable batch and real-time pipelines, modern data platforms and production reliability.
Key responsibilities
- Design, develop, test and deploy ETL/ELT pipelines using Python and PySpark.
- Optimise data storage and query performance across data lakes, warehouses and lakehouse formats such as Delta Lake or Iceberg.
- Apply software-engineering best practices, including maintainable code, reviews, CI/CD and unit testing.
- Implement monitoring, logging and alerting and resolve pipeline failures.
- Partner with data scientists, analysts and product stakeholders and optimise Spark workloads for performance and cost
Requirements
- Advanced Python, including object-oriented programming, pandas and testing frameworks.
- Strong practical Apache Spark / PySpark experience.
- Advanced SQL and relational databases such as PostgreSQL or MySQL.
- Experience with a data warehouse such as Snowflake, BigQuery or Redshift.
- Experience with at least one major cloud platform: AWS, GCP or Azure.
- Workflow orchestration using Airflow, Prefect or Dagster.
- Excellent written and spoken English.
Benefits
The benefits will be determined once the contractual terms have been established.