← Back to Home
AI & automation

İzmir Hava Kalitesi Tahmini

PM10, NO2 and SO2 forecasts for a monitoring station in Bornova, and an audit of our own model

Dokuz Eylül University, Applied Artificial Intelligence course · 4-person group project, January 2026 · re-audited in September 2026

PythonpandasXGBoostscikit-learnSHAPStreamlit

Screenshots & Schematics

The course model's score by test method: 0.90 with shuffled rows, 0.31 on unseen days, -0.55 when forecasting the future
The course model's score by test method: 0.90 with shuffled rows, 0.31 on unseen days, -0.55 when forecasting the future

Problem & Challenge

Monitoring stations report pollution only after it happens; an early warning needs a forecast. The second problem was ours: a high test score did not show that the model could actually forecast.

How It Works

01

A year of hourly PM10, NO2 and SO2 readings from the Ministry of Environment's Bornova Eğitim station, plus temperature, wind, pressure and precipitation from Meteostat.

02

In the corrected version every hour is one row, weather timestamps are converted from UTC to Turkish time, and hours the station did not measure are left out of the score.

03

The model looks at the last 24 hours of all three pollutants and the weather at the target hour, and forecasts 1 and 24 hours ahead.

04

It is trained on December 2024 – October 2025 and tested on October – December 2025, which it has never seen; every score is compared with a "nothing changes" baseline.

05

Result: R² 0.79–0.85 one hour ahead; 0.20–0.33 a day ahead, with about 20% less error than the baseline.

Architecture & Technical Decisions

Python, pandas, XGBoost and scikit-learn; the course version also had a SHAP explainability analysis and a Streamlit interface.

As a course project we combined a year of hourly readings from the Bornova Eğitim station with Meteostat weather data and built an XGBoost model that forecasts PM10, NO2 and SO2, with a Streamlit interface; the submitted model scored R² 0.90. I later re-audited it. When the data was merged, every hour had been split into two rows by accident, and because the rows were shuffled before the train/test split, the model had seen near-copies of its test hours during training. On days it had never seen, the score fell to 0.31. The corrected version uses the latest readings, is tested on a period it has never seen, and makes about 20% less error than a simple baseline 24 hours ahead.