Raw hourly readings do not answer questions. A trustworthy pipeline and a few well chosen features do. This project builds the second thing from the first.
Open to data engineering & analytics rolesThe source is a decade of hourly meteorological readings from Szeged, Hungary, covering 2006 through 2016: temperature and apparent temperature, humidity, wind speed and bearing, visibility, pressure, and a text summary of conditions, ninety six thousand four hundred and fifty three rows in all. On its own it is an undifferentiated wall of numbers. The work is turning that wall into something queryable, trustworthy, and legible.
That means two disciplines working together. First, the engineering: a pipeline that cleans, validates, and lands the data somewhere it can be queried at scale. Second, the analysis: deriving the handful of features that let a decade of hourly noise resolve into the daily and seasonal rhythms that were there all along. The rest of this page follows the data through both.
Raw timestamps are almost useless for seeing a daily rhythm. So the transformation step derives three features from each reading: its Hour, its Weekday, and its Part of Day, morning, afternoon, evening, or night. That single engineered feature, Part of Day, is what turns a flat table into the shape below: the average temperature of one Szeged day, learned from ten years of them.
Roughly eight degrees separate the afternoon peak from the overnight floor, a swing you feel every day but rarely quantify. Feature engineering is what made it measurable: without Part of Day, this rhythm stays buried in timestamps. This is the quiet heart of the project, the point where data engineering becomes data analysis.
With the data clean and the features in place, the patterns surface quickly. Monthly averages breathe from roughly eight degrees in deep winter to the mid twenties at the height of summer, the long seasonal wave underneath the daily one.
Underneath the seasons, three sharper findings emerge from the engineered features, each one a small piece of atmospheric physics the data confirms on its own.
Every number above depends on the data being clean, complete, and correctly loaded. So the pipeline is built the way production data work actually gets built: five isolated phases, one orchestrator, structured logging at every step, and validation that fails loudly rather than silently.
The last one is the one worth defending in an interview: when the local infrastructure could not carry the plan, I changed the architecture instead of forcing it.
re.sub(r'[^a-zA-Z0-9_]', '_', name), so the schema is always legal without hand editing.pd.to_datetime(errors='coerce'), so malformed timestamps become handled nulls, never a crash mid run.A validated pipeline, a few deliberate features, and a live dashboard, the full path from raw hourly readings to a decade you can question in a glance.