About the role
From the employer’s listing · Wayve · posted 5 October 2026
Before the detail, here's the challenge you'd help us solve. We build the embodied intelligence that moves real vehicles safely, and the ecosystem a billion machines will run on in the future. Very few people in AI can say this. Every role here, whatever the team, plugs into that. Here’s what this particular role covers. 🛠️ About our Engineering Teams The Machine Learning team within Application Software, part of Product & Delivery, works on critical initiatives that push the frontier of model-based autonomous driving. That covers core driving performance as well as feature-level intelligence such as personalisation, comfort and collaboration. 🧠 Your day-to-day As a Data Engineer, you’ll design and deliver scalable data pipelines that turn vast amounts of data from diverse internal and external sources into structured, reliable and model-ready datasets. Your work spans data ingestion, data quality assurance, transformation, curation, evaluation and ML support. You’ll work closely with Wayve’s Data Corpus teams, customer programmes and ML engineers, bringing up ingestion pipelines for new vehicle platforms and sensor configurations, and tackling the highest-impact bottlenecks first. 🧩 What you’ll be working on: Building and improving scalable data pipelines that support model development, evaluation and production ML workflows for autonomous driving. Ingesting, transforming and curating large-scale real-world, synthetic and partner-provided datasets into structured, reliable, model-ready formats aligned with standardised taxonomies and coordinate systems. Developing data quality checks, validation processes and monitoring so that raw vehicle data and processed datasets are high-quality, complete, consistent, traceable and fit for ML use cases. Curating and mining real-world and synthetic data to drive scenario diversity, coverage and feature-specific development. Improving pipeline performance, reliability and usability, reducing bottlenecks and increasing iteration velocity across ML development. Collaborating closely with ML engineers, Data Corpus, AI Platform and external partners so data pipelines integrate effectively with production-scale learning systems. 🙌 You should apply if: Essential Proven experience building and operating scalable data pipelines or distributed data processing systems in production environments. Strong software engineering skills in Python, with a solid foundation in maintainable, reliable, and well-tested software development practices. Proficient in SQL and PySpark, with experience using warehouse/OLAP concepts, window functions, and Spark for distributed data processing. Experience with modern data pipeline architectures, including workflow orchestration and DAG-based systems such as Airflow, Flyte, Ray, or similar. Solid understanding of robotics and automated driving data concepts, including sensor characteristics, timestamping and clock synchronisation, coordinate transformations, calibration, and ego-motion
















