Week 3: Past data, measured values, checking the data
Goal of today: about 970 days of past forecasts and measured temperatures for both cities in weather.db,
the daily job also saves the measured values, and a short report tells you if the data is OK:
Budapest / models: 976 days (2024-02-05 .. 2026-10-07)
missing values: {'ecmwf': 0, 'gfs': 0, 'icon': 0, 'measured': 2}
complete days: 974
How to work: do the steps in order. After each step, run the code and compare with the expected output. Go on only when yours looks the same. (Your numbers will differ a little, because you run it on another day.)
Work in C:\weather-lab with the venv active ((.venv) at the start of the line).
What you need to know
Why do we need past data? A model learns from examples: "on this day the forecasts said 24, 25, 23 and in reality it was 23.5". Your live collection (week 2) gives one example per day. That is ~40 examples by week 8: enough to see something, but too little for a neural network. So today we also download a history set of almost 1000 days.
Weather model. A big computer program that simulates the atmosphere with physics and calculates the weather of the next days. Weather services run them several times a day. The three best-known ones:
| Name | Full name | Who runs it |
|---|---|---|
ecmwf |
ECMWF IFS | European Centre for Medium-Range Weather Forecasts (Europe) |
gfs |
Global Forecast System | NOAA, the US weather service |
icon |
ICON | DWD, the German weather service |
The providers you collect (Open-Meteo, MET Norway, wttr.in) take outputs of models like these and process them further. Only Open-Meteo keeps an archive of old forecasts, and it has one for exactly these three models.
Forecast vs. measurement. A forecast is made before the day (what the model expected). A measurement is known only after the day (what really happened). Our models will learn the difference between the two.
Day-ahead forecast. We must learn from the same kind of forecast we will use later: a forecast made
the day before. Open-Meteo's previous runs service gives exactly this: temperature_2m_previous_day1
is the temperature that the model predicted one day earlier.
Hourly → daily. That service gives hourly values. In week 1 you wrote daily_max_min(), which turns hourly values
into daily max/min. We use it again.
Measured values come from Open-Meteo's weather archive (ERA5 reanalysis: measurements from weather stations,
satellites and balloons put together on a map grid). Two traps, both handled in the code:
1. The newest days arrive late and may be corrected later. So the daily job asks for the last 10 days every time,
and INSERT OR REPLACE (week 2) overwrites the old values.
2. The archive also sends a value for today, but today is not over yet, so that value is only an estimate.
Learning from it would mean learning from wrong data. We never ask for anything after yesterday.
pandas is a Python library for tables. A table is called a DataFrame:
- it has columns (with names, like ecmwf, measured) and rows;
- every row has a label, the index (for us: the day);
- a missing value is shown as NaN ("not a number");
- table.dropna() keeps only the rows without missing values;
- pivot and join reshape and combine tables. You will try them in step 9 on a tiny example.
Step 1: Install pandas
Add pandas to requirements.txt. It should look like this:
requirements.txt
requests
tzdata
pandas
Install and check:
pip install -r requirements.txt
python -c "import pandas; print(pandas.__version__)"
You should see a version number, for example:
3.0.3
Step 2: Look at the history service in the browser
Open this address. It asks for the day-ahead forecasts of the three models for Eger, for 1 October:
You should see (shortened):
{"latitude":48.0,"longitude":20.5, ...
"hourly":{"time":["2026-10-01T00:00","2026-10-01T01:00", ... ,"2026-10-01T23:00"],
"temperature_2m_previous_day1_ecmwf_ifs025":[15.2,14.9,14.7, ...],
"temperature_2m_previous_day1_gfs_seamless":[...],
"temperature_2m_previous_day1_icon_seamless":[...]}}
models=...asks for three models at once. The answer has one list per model; the name of the list is the variable name +_+ the model name.start_dateandend_datechoose the days. For the history set we will ask for 2.5 years in one request.
Step 3: config.py: the two data sets
Add these lines at the end of config.py:
# History set: three weather models whose old day-ahead forecasts can be downloaded.
MODEL_SOURCES = ["ecmwf", "gfs", "icon"]
OPEN_METEO_MODELS = {"ecmwf": "ecmwf_ifs025", "gfs": "gfs_seamless", "icon": "icon_seamless"}
HISTORY_START = "2024-02-05"
SETS = {"models": MODEL_SOURCES, "providers": PROVIDER_SOURCES}
MODEL_SOURCES: our short names for the three models.OPEN_METEO_MODELS: short name → the name Open-Meteo uses.HISTORY_START: the first day for which all three models have archived forecasts.SETSgives a name to each group of three sources. Everything later (checks, web page, models) works with"models"or"providers"the same way.
Try it:
python
>>> from config import SETS, OPEN_METEO_MODELS
>>> SETS
{'models': ['ecmwf', 'gfs', 'icon'], 'providers': ['open_meteo', 'met_norway', 'wttr']}
>>> OPEN_METEO_MODELS
{'ecmwf': 'ecmwf_ifs025', 'gfs': 'gfs_seamless', 'icon': 'icon_seamless'}
>>> exit()
Step 4: sources.py: download the past forecasts
Open sources.py. First, add OPEN_METEO_MODELS to the import from config:
from config import OPEN_METEO_MODELS, TIMEZONE, USER_AGENT
Then add these two functions at the end of the file:
def _models_hourly(url, variable, city, extra_params):
"""Ask Open-Meteo for hourly temperatures of our three models -> {source: {day: (tmax, tmin)}}."""
params = {
"latitude": city["lat"], "longitude": city["lon"],
"hourly": variable, "models": ",".join(OPEN_METEO_MODELS.values()),
"timezone": TIMEZONE, **extra_params,
}
response = requests.get(url, params=params, timeout=120)
response.raise_for_status()
hourly = response.json()["hourly"]
return {source: daily_max_min(hourly["time"], hourly[f"{variable}_{model}"])
for source, model in OPEN_METEO_MODELS.items()}
def fetch_model_history(city, start, end):
"""Old forecasts that were made ONE DAY BEFORE each day (true day-ahead forecasts)."""
return _models_hourly("https://previous-runs-api.open-meteo.com/v1/forecast",
"temperature_2m_previous_day1", city,
{"start_date": start, "end_date": end})
What they do:
- _models_hourly() asks for hourly data of all three models in one request (step 2).
The answer has one list per model, e.g. temperature_2m_previous_day1_ecmwf_ifs025.
For each model, daily_max_min() (week 1) turns the hours into days. The result looks like
{"ecmwf": {day: (tmax, tmin)}, "gfs": {...}, "icon": {...}}.
- The _ at the start of the name means "helper, used only inside this file".
- **extra_params adds the extra parameters (here: the start and end date) to the dictionary.
- timeout=120: 2.5 years of hourly data is a bigger answer, so we wait longer.
- fetch_model_history() calls the helper with the previous runs address and the dates.
Try it for three days:
python
>>> from config import CITIES
>>> from sources import fetch_model_history
>>> history = fetch_model_history(CITIES["eger"], "2026-10-01", "2026-10-03")
>>> history.keys()
dict_keys(['ecmwf', 'gfs', 'icon'])
>>> history["ecmwf"]
{'2026-10-01': (22.1, 14.4), '2026-10-02': (22.0, 14.4), '2026-10-03': (22.1, 13.3)}
>>> history["icon"]
{'2026-10-01': (23.9, 7.3), '2026-10-02': (23.7, 8.3), '2026-10-03': (23.4, 5.1)}
>>> exit()
Look: for the same days, ICON expected nights 5–7 °C colder than ECMWF. The models really disagree.
Remember: after you change a file,
exit()and startpythonagain.
Step 5: sources.py: the models' forecast for tomorrow
The history set should keep growing every day, like the live set. So the daily job will also save the
three models' forecast for tomorrow. Add this function at the end of sources.py:
def fetch_models_tomorrow(city):
"""Today's forecast of the three models for tomorrow -> {source: (tmax, tmin)}."""
days = _models_hourly("https://api.open-meteo.com/v1/forecast", "temperature_2m", city,
{"forecast_days": 3})
return {source: values[tomorrow()] for source, values in days.items() if tomorrow() in values}
It uses the same helper, but with the normal forecast address and temperature_2m (today's forecast),
then keeps only tomorrow.
Try it:
python
>>> from config import CITIES
>>> from sources import fetch_models_tomorrow
>>> fetch_models_tomorrow(CITIES["eger"])
{'ecmwf': (22.2, 15.4), 'gfs': (24.2, 11.7), 'icon': (24.6, 7.0)}
>>> exit()
Step 6: sources.py: what really happened
Add this function at the end of sources.py:
def fetch_measured(city, start, end):
# Today is not over yet, so its max/min is not known. The archive still sends a
# value for today (an estimate), so we never ask for anything after yesterday.
end = min(end, (date.today() - timedelta(days=1)).isoformat())
params = {
"latitude": city["lat"], "longitude": city["lon"],
"daily": "temperature_2m_max,temperature_2m_min",
"timezone": TIMEZONE, "start_date": start, "end_date": end,
}
response = requests.get("https://archive-api.open-meteo.com/v1/archive", params=params, timeout=120)
response.raise_for_status()
daily = response.json()["daily"]
return {day: (tmax, tmin)
for day, tmax, tmin in zip(daily["time"], daily["temperature_2m_max"], daily["temperature_2m_min"])
if tmax is not None and tmin is not None}
- The first line is trap 2:
min(end, yesterday)makes sure we never ask for today or later. (Text dates like"2026-10-05"can be compared withminbecause they are written year-month-day.) zip()walks through the three lists together: day, max, min.if tmax is not None ...skips days that have no value yet.
Try it. Ask until today on purpose (use today's date):
python
>>> from config import CITIES
>>> from sources import fetch_measured
>>> fetch_measured(CITIES["eger"], "2026-10-01", "2026-10-06")
{'2026-10-01': (22.3, 12.0), '2026-10-02': (22.2, 11.4), '2026-10-03': (21.9, 10.9), '2026-10-04': (22.4, 10.6), '2026-10-05': (23.0, 11.4)}
>>> exit()
The last day is yesterday (10-05), not today (10-06). Trap 2 works.
Finally, change the first line of sources.py (the description) to:
"""Download forecasts and measured temperatures from the internet.
Step 7: db.py: a table for the measured values
Open db.py. Add a second table to SCHEMA, so it looks like this:
SCHEMA = """
CREATE TABLE IF NOT EXISTS forecasts (
city TEXT, source TEXT, day TEXT, tmax REAL, tmin REAL, saved_at TEXT,
PRIMARY KEY (city, source, day)
);
CREATE TABLE IF NOT EXISTS measured (
city TEXT, day TEXT, tmax REAL, tmin REAL,
PRIMARY KEY (city, day)
);
"""
One measured value per city and day: the primary key is (city, day).
Then add this function at the end of db.py:
def save_measured(city, values):
rows = [(city, day, tmax, tmin) for day, (tmax, tmin) in values.items()]
con = connect()
con.executemany("INSERT OR REPLACE INTO measured VALUES (?, ?, ?, ?)", rows)
con.commit()
con.close()
return len(rows)
It is the same idea as save_forecasts() from week 2.
Try it. Save the same 5 days twice:
python
>>> from config import CITIES
>>> from sources import fetch_measured
>>> from db import save_measured
>>> values = fetch_measured(CITIES["eger"], "2026-10-01", "2026-10-05")
>>> save_measured("eger", values)
5
>>> save_measured("eger", values)
5
>>> import sqlite3
>>> sqlite3.connect("weather.db").execute("SELECT COUNT(*) FROM measured").fetchone()
(5,)
>>> exit()
5 rows, not 10: INSERT OR REPLACE overwrote them the second time. No duplicates.
Step 8: load_history.py: download the past once
Create load_history.py:
load_history.py
"""Run ONCE: download the history set (old forecasts + measured values) into the database.
python load_history.py
"""
from datetime import date
from config import CITIES, HISTORY_START
from db import save_forecasts, save_measured
from sources import fetch_measured, fetch_model_history
end = date.today().isoformat()
for key, city in CITIES.items():
print(f"{city['name']}: downloading {HISTORY_START} .. {end}")
history = fetch_model_history(city, HISTORY_START, end)
for source, values in history.items():
print(f" {source:6s} {save_forecasts(key, source, values)} days")
measured = fetch_measured(city, HISTORY_START, end)
print(f" measured {save_measured(key, measured)} days")
print("Done.")
It goes through both cities, downloads the history from HISTORY_START until today, and saves it.
Run it:
python load_history.py
You should see:
Budapest: downloading 2024-02-05 .. 2026-10-06
ecmwf 975 days
gfs 975 days
icon 975 days
measured 974 days
Eger: downloading 2024-02-05 .. 2026-10-06
ecmwf 975 days
gfs 975 days
icon 975 days
measured 974 days
Done.
- Forecasts: 975 days, including today (the forecast for today was made yesterday, so it exists).
- Measured: 974 days, until yesterday (trap 2).
- You can run it again any time: thanks to
INSERT OR REPLACEit does not create duplicates.
Step 9: pandas in five minutes
Before we read the database into pandas, try the two important operations on a tiny, made-up table
(only 2 days and 2 sources, so you can follow every number; the real table has ~970 days and 3 sources).
Create try_pandas.py:
import pandas as pd
# One row per forecast, like in our database ("long" format)
long = pd.DataFrame({
"day": ["10-01", "10-01", "10-02", "10-02"],
"source": ["ecmwf", "gfs", "ecmwf", "gfs"],
"tmax": [22.1, 23.1, 22.0, 23.1],
})
print(long)
print("---")
# One row per day, one column per source ("wide" format)
wide = long.pivot(index="day", columns="source", values="tmax")
print(wide)
print("---")
# The measured values, with the day as index
measured = pd.DataFrame({"day": ["10-01"], "measured": [22.3]}).set_index("day")
# Put them next to the forecasts, matching by day
table = wide.join(measured, how="left")
print(table)
print("---")
print(table.dropna())
print("---")
print(table["ecmwf"].mean())
Run it:
python try_pandas.py
You should see:
day source tmax
0 10-01 ecmwf 22.1
1 10-01 gfs 23.1
2 10-02 ecmwf 22.0
3 10-02 gfs 23.1
---
source ecmwf gfs
day
10-01 22.1 23.1
10-02 22.0 23.1
---
ecmwf gfs measured
day
10-01 22.1 23.1 22.3
10-02 22.0 23.1 NaN
---
ecmwf gfs measured
day
10-01 22.1 23.1 22.3
---
22.05
pivotturned 4 rows (one per forecast) into 2 rows (one per day): the values ofsourcebecame column names.joinadded themeasuredcolumn, matching rows by the index (the day).how="left"keeps all days of the left table, even without a measured value: 10-02 gotNaN.dropna()kept only the complete row.table["ecmwf"]is one column;.mean()is its average.
This "one row per day: forecasts + measured" shape is exactly what machine learning needs later:
the inputs and the correct answer in one row. You can delete try_pandas.py now.
Step 10: db.py: read a city's data into one table
Add import pandas as pd at the top of db.py, after from datetime import datetime:
import sqlite3
from datetime import datetime
import pandas as pd
from config import DB_PATH
Then add this function at the end of db.py:
def load_table(city, sources):
"""One row per day: the max temperature forecast of each source + the measured value.
Days that have forecasts but no measurement yet (for example tomorrow) are kept,
their 'measured' value is empty (NaN).
"""
con = connect()
forecasts = pd.read_sql_query("SELECT day, source, tmax FROM forecasts WHERE city = ?", con, params=(city,))
measured = pd.read_sql_query("SELECT day, tmax AS measured FROM measured WHERE city = ?", con, params=(city,))
con.close()
forecasts = forecasts[forecasts["source"].isin(sources)]
table = forecasts.pivot(index="day", columns="source", values="tmax").reindex(columns=sources)
table = table.join(measured.set_index("day"), how="left")
return table.sort_index()
It is step 9, with real data:
1. pd.read_sql_query() runs a SQL query and gives the result as a DataFrame (long format, like long in step 9).
2. isin(sources) keeps only the sources we asked for (e.g. only the three models).
3. pivot() makes one row per day, one column per source. reindex(columns=sources) puts the columns in our order.
4. join(..., how="left") adds the measured value. Days without a measurement (today, tomorrow) get NaN.
5. sort_index() sorts by day.
This is the most important function of the project: everything from now on uses it.
Try it:
python
>>> from db import load_table
>>> table = load_table("eger", ["ecmwf", "gfs", "icon"])
>>> table.tail(4)
ecmwf gfs icon measured
day
2026-10-03 22.1 22.8 23.4 21.9
2026-10-04 22.5 23.2 24.2 22.4
2026-10-05 23.2 23.9 25.0 23.0
2026-10-06 23.2 24.3 25.2 NaN
>>> len(table)
975
>>> exit()
tail(4) shows the last 4 rows. Today (10-06) has forecasts, but no measured value yet: NaN. Correct.
Step 11: collect.py: save the models and the measured values every day
Replace collect.py with this version:
collect.py
"""Daily job: save tomorrow's forecasts and update the measured values.
python collect.py run once (this is what Windows Task Scheduler starts)
python collect.py --loop keep running, collect every day at 08:00
"""
import logging
import sys
import time
from datetime import date, datetime, timedelta
from config import BASE_DIR, CITIES
from db import save_forecasts, save_measured
from sources import PROVIDERS, fetch_measured, fetch_models_tomorrow, tomorrow
logging.basicConfig(
level=logging.INFO, format="%(asctime)s %(message)s",
handlers=[logging.FileHandler(BASE_DIR / "collect.log", encoding="utf-8"), logging.StreamHandler()],
)
log = logging.getLogger()
def collect_forecasts():
for key, city in CITIES.items():
# Live set: one provider failing must not stop the others.
for source, fetch in PROVIDERS.items():
try:
tmax, tmin = fetch(city)
save_forecasts(key, source, {tomorrow(): (tmax, tmin)})
log.info(f"{key:9s} {source:11s} {tomorrow()} max {tmax} min {tmin}")
except Exception as error:
log.error(f"{key:9s} {source:11s} FAILED: {error}")
# History set: keep it growing with today's model forecasts.
try:
for source, values in fetch_models_tomorrow(city).items():
save_forecasts(key, source, {tomorrow(): values})
log.info(f"{key:9s} models saved")
except Exception as error:
log.error(f"{key:9s} models FAILED: {error}")
def collect_measured():
# The measured values arrive a few days late, so always ask for the last 10 days again.
start = (date.today() - timedelta(days=10)).isoformat()
end = (date.today() - timedelta(days=1)).isoformat()
for key, city in CITIES.items():
try:
log.info(f"{key:9s} measured {save_measured(key, fetch_measured(city, start, end))} days")
except Exception as error:
log.error(f"{key:9s} measured FAILED: {error}")
def run_once():
collect_forecasts()
collect_measured()
def seconds_until(hour):
now = datetime.now()
next_run = now.replace(hour=hour, minute=0, second=0, microsecond=0)
if next_run <= now:
next_run += timedelta(days=1)
return (next_run - now).total_seconds()
if __name__ == "__main__":
run_once()
if "--loop" in sys.argv:
while True:
time.sleep(seconds_until(8))
run_once()
What is new compared with week 2:
- In collect_forecasts(), a second try block saves the three models' forecast for tomorrow (step 5).
So the history set keeps growing every day.
- The new function collect_measured() downloads the measured values of the last 10 days (trap 1) and saves them.
- run_once() calls both.
The Task Scheduler task from week 2 starts the same file, so you do not need to change it.
Run it:
python collect.py
You should see:
2026-10-06 14:34:58,063 budapest open_meteo 2026-10-07 max 25.0 min 12.3
2026-10-06 14:34:58,193 budapest met_norway 2026-10-07 max 22.8 min 9.0
2026-10-06 14:34:58,278 budapest wttr 2026-10-07 max 25.0 min 14.0
2026-10-06 14:34:58,356 budapest models saved
2026-10-06 14:34:58,434 eger open_meteo 2026-10-07 max 24.4 min 11.9
2026-10-06 14:34:58,561 eger met_norway 2026-10-07 max 22.0 min 10.4
2026-10-06 14:34:58,666 eger wttr 2026-10-07 max 23.0 min 10.0
2026-10-06 14:34:58,744 eger models saved
2026-10-06 14:34:58,824 budapest measured 10 days
2026-10-06 14:34:58,904 eger measured 10 days
Check the result in the interactive Python:
python
>>> from db import load_table
>>> load_table("eger", ["ecmwf", "gfs", "icon"]).tail(3)
ecmwf gfs icon measured
day
2026-10-05 23.2 23.9 25.0 23.0
2026-10-06 23.2 24.3 25.2 NaN
2026-10-07 22.2 24.2 24.6 NaN
>>> load_table("eger", ["open_meteo", "met_norway", "wttr"])
open_meteo met_norway wttr measured
day
2026-10-07 24.4 22.0 23.0 NaN
>>> exit()
Tomorrow (10-07) is now in the models set too. The providers set has as many days as you have collected since week 2.
Step 12: check_data.py: look before you learn
Bad data gives bad models. Before any machine learning, we check the data. Create check_data.py:
check_data.py
"""Week 3: look at the data before using it. Bad data -> bad models.
python check_data.py
"""
from config import CITIES, SETS
from db import load_table
for city in CITIES:
for data_set, sources in SETS.items():
table = load_table(city, sources)
name = CITIES[city]["name"]
if table.empty:
print(f"\n{name} / {data_set}: no data yet")
continue
print(f"\n{name} / {data_set}: {len(table)} days ({table.index.min()} .. {table.index.max()})")
# 1. Missing values per column
print(" missing values:", table.isna().sum().to_dict())
# 2. Complete days: every source AND the measured value present
print(" complete days: ", len(table.dropna()))
# 3. Suspicious values: a forecast more than 8 degrees away from the measured value
complete = table.dropna()
for source in sources:
wrong = complete[(complete[source] - complete["measured"]).abs() > 8]
if len(wrong):
print(f" {source}: {len(wrong)} days more than 8 °C wrong, e.g. {wrong.index[0]}")
For each city and each set, it prints:
1. how many values are missing in each column (isna().sum() counts the NaNs),
2. how many days are complete,
3. suspicious days: a forecast more than 8 °C away from the measured value (maybe a data error).
Run it:
python check_data.py
You should see (your numbers will be a bit bigger):
Budapest / models: 976 days (2024-02-05 .. 2026-10-07)
missing values: {'ecmwf': 0, 'gfs': 0, 'icon': 0, 'measured': 2}
complete days: 974
Budapest / providers: 1 days (2026-10-07 .. 2026-10-07)
missing values: {'open_meteo': 0, 'met_norway': 0, 'wttr': 0, 'measured': 1}
complete days: 0
Eger / models: 976 days (2024-02-05 .. 2026-10-07)
missing values: {'ecmwf': 0, 'gfs': 0, 'icon': 0, 'measured': 2}
complete days: 974
Eger / providers: 1 days (2026-10-07 .. 2026-10-07)
missing values: {'open_meteo': 0, 'met_norway': 0, 'wttr': 0, 'measured': 1}
complete days: 0
How to read it:
- The 2 missing measured values in models are today and tomorrow. That is correct: they have not happened yet.
- No "more than 8 °C wrong" lines: no suspicious values.
- providers will have more days: one more for every day collect.py has run since week 2.
(In this example the collection started today.)
Your complete changed files
Compare with yours. If something does not work, copy these.
config.py
"""Settings shared by all scripts."""
from pathlib import Path
# Files are always created next to this file, no matter where you start Python from.
BASE_DIR = Path(__file__).resolve().parent
DB_PATH = BASE_DIR / "weather.db"
TIMEZONE = "Europe/Budapest"
CITIES = {
"budapest": {"name": "Budapest", "lat": 47.4979, "lon": 19.0402},
"eger": {"name": "Eger", "lat": 47.9025, "lon": 20.3772},
}
# MET Norway wants to know who is asking. You may add your own e-mail address,
# but NOT a fake one like ...@example.com (that is blocked with error 403).
USER_AGENT = "weather-lab/1.0 (university lab project)"
# Live set: three independent forecast providers, collected by us every day.
PROVIDER_SOURCES = ["open_meteo", "met_norway", "wttr"]
# History set: three weather models whose old day-ahead forecasts can be downloaded.
MODEL_SOURCES = ["ecmwf", "gfs", "icon"]
OPEN_METEO_MODELS = {"ecmwf": "ecmwf_ifs025", "gfs": "gfs_seamless", "icon": "icon_seamless"}
HISTORY_START = "2024-02-05"
SETS = {"models": MODEL_SOURCES, "providers": PROVIDER_SOURCES}
sources.py
"""Download forecasts and measured temperatures from the internet.
Every function returns plain Python values: (tmax, tmin) or {day: (tmax, tmin)}.
A day is always a text like "2026-10-07" in local (Hungarian) time.
"""
from datetime import date, datetime, timedelta
from zoneinfo import ZoneInfo
import requests
from config import OPEN_METEO_MODELS, TIMEZONE, USER_AGENT
TIMEOUT = 30 # seconds to wait for an answer
LOCAL = ZoneInfo(TIMEZONE)
def tomorrow():
return (date.today() + timedelta(days=1)).isoformat()
def daily_max_min(times, temps):
"""Turn hourly temperatures into {day: (tmax, tmin)}.
times are local times like "2026-10-07T13:00". Days with fewer than
20 hourly values are skipped, because their max/min would be unreliable.
"""
by_day = {}
for time, temp in zip(times, temps):
if temp is not None:
by_day.setdefault(time[:10], []).append(temp)
return {day: (max(values), min(values)) for day, values in by_day.items() if len(values) >= 20}
# ---------- Live set: three providers, tomorrow only ----------
def fetch_open_meteo(city):
params = {
"latitude": city["lat"], "longitude": city["lon"],
"daily": "temperature_2m_max,temperature_2m_min",
"timezone": TIMEZONE, "forecast_days": 3,
}
response = requests.get("https://api.open-meteo.com/v1/forecast", params=params, timeout=TIMEOUT)
response.raise_for_status()
daily = response.json()["daily"]
i = daily["time"].index(tomorrow())
return daily["temperature_2m_max"][i], daily["temperature_2m_min"][i]
def fetch_met_norway(city):
params = {"lat": round(city["lat"], 4), "lon": round(city["lon"], 4)}
response = requests.get("https://api.met.no/weatherapi/locationforecast/2.0/compact",
params=params, headers={"User-Agent": USER_AGENT}, timeout=TIMEOUT)
response.raise_for_status()
times, temps = [], []
for step in response.json()["properties"]["timeseries"]:
# MET Norway gives UTC times ("...Z"); convert them to Hungarian time.
utc_time = datetime.fromisoformat(step["time"].replace("Z", "+00:00"))
times.append(utc_time.astimezone(LOCAL).strftime("%Y-%m-%dT%H:%M"))
temps.append(step["data"]["instant"]["details"]["air_temperature"])
return daily_max_min(times, temps)[tomorrow()]
def fetch_wttr(city):
url = f"https://wttr.in/{city['lat']},{city['lon']}"
response = requests.get(url, params={"format": "j1"}, timeout=TIMEOUT)
response.raise_for_status()
for day in response.json()["weather"]:
if day["date"] == tomorrow():
return float(day["maxtempC"]), float(day["mintempC"])
raise ValueError("wttr.in did not send tomorrow's forecast")
PROVIDERS = {"open_meteo": fetch_open_meteo, "met_norway": fetch_met_norway, "wttr": fetch_wttr}
# ---------- History set: three weather models, past and tomorrow ----------
def _models_hourly(url, variable, city, extra_params):
"""Ask Open-Meteo for hourly temperatures of our three models -> {source: {day: (tmax, tmin)}}."""
params = {
"latitude": city["lat"], "longitude": city["lon"],
"hourly": variable, "models": ",".join(OPEN_METEO_MODELS.values()),
"timezone": TIMEZONE, **extra_params,
}
response = requests.get(url, params=params, timeout=120)
response.raise_for_status()
hourly = response.json()["hourly"]
return {source: daily_max_min(hourly["time"], hourly[f"{variable}_{model}"])
for source, model in OPEN_METEO_MODELS.items()}
def fetch_model_history(city, start, end):
"""Old forecasts that were made ONE DAY BEFORE each day (true day-ahead forecasts)."""
return _models_hourly("https://previous-runs-api.open-meteo.com/v1/forecast",
"temperature_2m_previous_day1", city,
{"start_date": start, "end_date": end})
def fetch_models_tomorrow(city):
"""Today's forecast of the three models for tomorrow -> {source: (tmax, tmin)}."""
days = _models_hourly("https://api.open-meteo.com/v1/forecast", "temperature_2m", city,
{"forecast_days": 3})
return {source: values[tomorrow()] for source, values in days.items() if tomorrow() in values}
# ---------- What really happened ----------
def fetch_measured(city, start, end):
# Today is not over yet, so its max/min is not known. The archive still sends a
# value for today (an estimate), so we never ask for anything after yesterday.
end = min(end, (date.today() - timedelta(days=1)).isoformat())
params = {
"latitude": city["lat"], "longitude": city["lon"],
"daily": "temperature_2m_max,temperature_2m_min",
"timezone": TIMEZONE, "start_date": start, "end_date": end,
}
response = requests.get("https://archive-api.open-meteo.com/v1/archive", params=params, timeout=120)
response.raise_for_status()
daily = response.json()["daily"]
return {day: (tmax, tmin)
for day, tmax, tmin in zip(daily["time"], daily["temperature_2m_max"], daily["temperature_2m_min"])
if tmax is not None and tmin is not None}
db.py
"""Everything that touches the SQLite database file (weather.db)."""
import sqlite3
from datetime import datetime
import pandas as pd
from config import DB_PATH
SCHEMA = """
CREATE TABLE IF NOT EXISTS forecasts (
city TEXT, source TEXT, day TEXT, tmax REAL, tmin REAL, saved_at TEXT,
PRIMARY KEY (city, source, day)
);
CREATE TABLE IF NOT EXISTS measured (
city TEXT, day TEXT, tmax REAL, tmin REAL,
PRIMARY KEY (city, day)
);
"""
def connect():
con = sqlite3.connect(DB_PATH)
con.executescript(SCHEMA) # creates the tables the first time, does nothing later
return con
def now():
return datetime.now().isoformat(timespec="seconds")
def save_forecasts(city, source, values):
"""values: {day: (tmax, tmin)}. Saving the same day again overwrites it (no duplicates)."""
rows = [(city, source, day, tmax, tmin, now()) for day, (tmax, tmin) in values.items()]
con = connect()
con.executemany("INSERT OR REPLACE INTO forecasts VALUES (?, ?, ?, ?, ?, ?)", rows)
con.commit()
con.close()
return len(rows)
def save_measured(city, values):
rows = [(city, day, tmax, tmin) for day, (tmax, tmin) in values.items()]
con = connect()
con.executemany("INSERT OR REPLACE INTO measured VALUES (?, ?, ?, ?)", rows)
con.commit()
con.close()
return len(rows)
def load_table(city, sources):
"""One row per day: the max temperature forecast of each source + the measured value.
Days that have forecasts but no measurement yet (for example tomorrow) are kept,
their 'measured' value is empty (NaN).
"""
con = connect()
forecasts = pd.read_sql_query("SELECT day, source, tmax FROM forecasts WHERE city = ?", con, params=(city,))
measured = pd.read_sql_query("SELECT day, tmax AS measured FROM measured WHERE city = ?", con, params=(city,))
con.close()
forecasts = forecasts[forecasts["source"].isin(sources)]
table = forecasts.pivot(index="day", columns="source", values="tmax").reindex(columns=sources)
table = table.join(measured.set_index("day"), how="left")
return table.sort_index()
New files: load_history.py (step 8), collect.py (step 11) and check_data.py (step 12) are shown in full above.
Check
- ☐
check_data.pyshows ~970+ complete days formodelsin both cities. - ☐ Running
load_history.pyagain does not change the number of days. - ☐
collect.pynow printsmodels savedandmeasured 10 dayslines too. - ☐ There is no measured value for today. Check it:
python -c "import sqlite3; print(sqlite3.connect('weather.db').execute('SELECT city, MAX(day) FROM measured GROUP BY city').fetchall())"You should see yesterday's date for both cities, e.g.[('budapest', '2026-10-05'), ('eger', '2026-10-05')].
Extra
- Which model has the most days with an error over 3 °C? Change the limit in
check_data.py. - Add a check: are there days where
tmax < tmin? (There should not be.) - Think: why is a random split of days into train/test dangerous for this kind of data? (Answer in week 5.)
If something goes wrong
| Problem | Reason |
|---|---|
ModuleNotFoundError: No module named 'pandas' |
Run pip install -r requirements.txt with the venv active. |
ImportError: cannot import name 'OPEN_METEO_MODELS' |
You did not add the new lines to config.py (step 3). |
KeyError: 'hourly' |
The service answered with an error message instead of data. Print response.json() to see it. |
ReadTimeout in load_history.py |
Many students at the same time. Wait a minute and run it again. |
providers has 0 complete days |
Normal at the beginning: tomorrow's forecast has no measured value yet. Wait a few days. |
| A function "does not exist" in the interactive Python | You changed the file: exit() and start python again. |