Location: London Bridge – In office 4 days per week
Start Date: ASAP
Duration: 3-4 months with extension
Daily Rate: £400 - £450 per day
Summary
We require a high-agency Data Engineer who will sit directly alongside our clients Head of Data and Lead Data Scientist. You will actively shape their overarching data strategy, translate business objectives into high-value projects, and co-author Technical Design Documents (TDDs). Once the strategy is defined, you will personally execute at speed, writing production-grade Python, Dagster pipelines, and Terraform infrastructure to ship features end-to-end.
Requirements
- 4+ years commercial experience in Python distributed data engineering & systems backend development.
- Hands-on production expertise with Dagster or Airflow in a Medallion or data lakehouse environment.
- Production-grade Terraform experience managing AWS resources (EKS, Aurora PostgreSQL, IAM IRSA, VPCs).
- Security-first engineering mindset (familiarity with enterprise audit constraints such as SOC2 or ISO27001).
- Experience with multi agent frameworks/libraries such as Google ADK or Langchain etc.
- Experience building RAG pipelines, managing pgvector embeddings, or deploying LLM inference pipelines.
- Infrastructure technologies - EKS and terraform
Data Strategy & Architecture Leadership (Co-Ownership)
- Partner with the Head of Data and Lead Data Scientist/Engineer to define overarching data lakehouse strategy, schema evolution, and vector retrieval architectures.
- Translate top-down Objectives and Key Results (OKRs) into structured and detailed TDDs that satisfy Fortune 500 enterprise security and compliance standards.
- Establish data governance, lineage tracking, and multi-tenant isolation patterns for agentic workflows.
- Define JIRA Epics and user stories derived from approved TDDs to guide downstream implementation.
- Architect, code, and deploy idempotent, backfill-safe data pipelines using Python and Dagster
- Ability to utilise and work against AWS cloud environment
- Enforce data lineage, schema evolution, and optimal Parquet/Delta partitioning across S3 storage layers.
