Overview
For Data Engineering Project II, our team built a complete data platform around mobility: how people, and students in particular, get around. We went through the whole chain, from collecting raw data to analysing it and applying machine learning.
The data
We brought together data from very different sources:
- traffic counts from counting stations along the road
- Blue-bike, the shared bikes at Belgian train stations
- train data
- a survey among students about how they travel to school
What we built
- Pipelines that fetch the data from every source, clean it and load it into the data warehouse.
- The data warehouse itself, designed and set up by us.
- Analyses on top of the warehouse, to see how the different ways of travelling relate to each other.
- Machine learning models on the combined data.
What I learned
The hardest part of data engineering is rarely the model. It’s getting data from sources that were never meant to work together to agree on dates, places and formats, and building pipelines that keep doing that reliably.