Week 3: Past data, measured values, checking the data

Goal of today: about 970 days of past forecasts and measured temperatures for both cities in weather.db, the daily job also saves the measured values, and a short report tells you if the data is OK:

Budapest / models: 976 days (2024-02-05 .. 2026-10-07)
  missing values: {'ecmwf': 0, 'gfs': 0, 'icon': 0, 'measured': 2}
  complete days:  974

How to work: do the steps in order. After each step, run the code and compare with the expected output. Go on only when yours looks the same. (Your numbers will differ a little, because you run it on another day.)

Work in C:\weather-lab with the venv active ((.venv) at the start of the line).


What you need to know

Why do we need past data? A model learns from examples: "on this day the forecasts said 24, 25, 23 and in reality it was 23.5". Your live collection (week 2) gives one example per day. That is ~40 examples by week 8: enough to see something, but too little for a neural network. So today we also download a history set of almost 1000 days.

Weather model. A big computer program that simulates the atmosphere with physics and calculates the weather of the next days. Weather services run them several times a day. The three best-known ones:

Name Full name Who runs it
ecmwf ECMWF IFS European Centre for Medium-Range Weather Forecasts (Europe)
gfs Global Forecast System NOAA, the US weather service
icon ICON DWD, the German weather service

The providers you collect (Open-Meteo, MET Norway, wttr.in) take outputs of models like these and process them further. Only Open-Meteo keeps an archive of old forecasts, and it has one for exactly these three models.

Forecast vs. measurement. A forecast is made before the day (what the model expected). A measurement is known only after the day (what really happened). Our models will learn the difference between the two.

Day-ahead forecast. We must learn from the same kind of forecast we will use later: a forecast made the day before. Open-Meteo's previous runs service gives exactly this: temperature_2m_previous_day1 is the temperature that the model predicted one day earlier.

Hourly → daily. That service gives hourly values. In week 1 you wrote daily_max_min(), which turns hourly values into daily max/min. We use it again.

Measured values come from Open-Meteo's weather archive (ERA5 reanalysis: measurements from weather stations, satellites and balloons put together on a map grid). Two traps, both handled in the code: 1. The newest days arrive late and may be corrected later. So the daily job asks for the last 10 days every time, and INSERT OR REPLACE (week 2) overwrites the old values. 2. The archive also sends a value for today, but today is not over yet, so that value is only an estimate. Learning from it would mean learning from wrong data. We never ask for anything after yesterday.

pandas is a Python library for tables. A table is called a DataFrame: - it has columns (with names, like ecmwf, measured) and rows; - every row has a label, the index (for us: the day); - a missing value is shown as NaN ("not a number"); - table.dropna() keeps only the rows without missing values; - pivot and join reshape and combine tables. You will try them in step 9 on a tiny example.


Step 1: Install pandas

Add pandas to requirements.txt. It should look like this:

requirements.txt

requests
tzdata
pandas

Install and check:

pip install -r requirements.txt
python -c "import pandas; print(pandas.__version__)"

You should see a version number, for example:

3.0.3

Step 2: Look at the history service in the browser

Open this address. It asks for the day-ahead forecasts of the three models for Eger, for 1 October:

https://previous-runs-api.open-meteo.com/v1/forecast?latitude=47.9025&longitude=20.3772&hourly=temperature_2m_previous_day1&models=ecmwf_ifs025,gfs_seamless,icon_seamless&timezone=Europe/Budapest&start_date=2026-10-01&end_date=2026-10-01

You should see (shortened):

{"latitude":48.0,"longitude":20.5, ...
 "hourly":{"time":["2026-10-01T00:00","2026-10-01T01:00", ... ,"2026-10-01T23:00"],
           "temperature_2m_previous_day1_ecmwf_ifs025":[15.2,14.9,14.7, ...],
           "temperature_2m_previous_day1_gfs_seamless":[...],
           "temperature_2m_previous_day1_icon_seamless":[...]}}

Step 3: config.py: the two data sets

Add these lines at the end of config.py:

# History set: three weather models whose old day-ahead forecasts can be downloaded.
MODEL_SOURCES = ["ecmwf", "gfs", "icon"]
OPEN_METEO_MODELS = {"ecmwf": "ecmwf_ifs025", "gfs": "gfs_seamless", "icon": "icon_seamless"}
HISTORY_START = "2024-02-05"

SETS = {"models": MODEL_SOURCES, "providers": PROVIDER_SOURCES}

Try it:

python
>>> from config import SETS, OPEN_METEO_MODELS
>>> SETS
{'models': ['ecmwf', 'gfs', 'icon'], 'providers': ['open_meteo', 'met_norway', 'wttr']}
>>> OPEN_METEO_MODELS
{'ecmwf': 'ecmwf_ifs025', 'gfs': 'gfs_seamless', 'icon': 'icon_seamless'}
>>> exit()

Step 4: sources.py: download the past forecasts

Open sources.py. First, add OPEN_METEO_MODELS to the import from config:

from config import OPEN_METEO_MODELS, TIMEZONE, USER_AGENT

Then add these two functions at the end of the file:

def _models_hourly(url, variable, city, extra_params):
    """Ask Open-Meteo for hourly temperatures of our three models -> {source: {day: (tmax, tmin)}}."""
    params = {
        "latitude": city["lat"], "longitude": city["lon"],
        "hourly": variable, "models": ",".join(OPEN_METEO_MODELS.values()),
        "timezone": TIMEZONE, **extra_params,
    }
    response = requests.get(url, params=params, timeout=120)
    response.raise_for_status()
    hourly = response.json()["hourly"]
    return {source: daily_max_min(hourly["time"], hourly[f"{variable}_{model}"])
            for source, model in OPEN_METEO_MODELS.items()}


def fetch_model_history(city, start, end):
    """Old forecasts that were made ONE DAY BEFORE each day (true day-ahead forecasts)."""
    return _models_hourly("https://previous-runs-api.open-meteo.com/v1/forecast",
                          "temperature_2m_previous_day1", city,
                          {"start_date": start, "end_date": end})

What they do: - _models_hourly() asks for hourly data of all three models in one request (step 2). The answer has one list per model, e.g. temperature_2m_previous_day1_ecmwf_ifs025. For each model, daily_max_min() (week 1) turns the hours into days. The result looks like {"ecmwf": {day: (tmax, tmin)}, "gfs": {...}, "icon": {...}}. - The _ at the start of the name means "helper, used only inside this file". - **extra_params adds the extra parameters (here: the start and end date) to the dictionary. - timeout=120: 2.5 years of hourly data is a bigger answer, so we wait longer. - fetch_model_history() calls the helper with the previous runs address and the dates.

Try it for three days:

python
>>> from config import CITIES
>>> from sources import fetch_model_history
>>> history = fetch_model_history(CITIES["eger"], "2026-10-01", "2026-10-03")
>>> history.keys()
dict_keys(['ecmwf', 'gfs', 'icon'])
>>> history["ecmwf"]
{'2026-10-01': (22.1, 14.4), '2026-10-02': (22.0, 14.4), '2026-10-03': (22.1, 13.3)}
>>> history["icon"]
{'2026-10-01': (23.9, 7.3), '2026-10-02': (23.7, 8.3), '2026-10-03': (23.4, 5.1)}
>>> exit()

Look: for the same days, ICON expected nights 5–7 °C colder than ECMWF. The models really disagree.

Remember: after you change a file, exit() and start python again.


Step 5: sources.py: the models' forecast for tomorrow

The history set should keep growing every day, like the live set. So the daily job will also save the three models' forecast for tomorrow. Add this function at the end of sources.py:

def fetch_models_tomorrow(city):
    """Today's forecast of the three models for tomorrow -> {source: (tmax, tmin)}."""
    days = _models_hourly("https://api.open-meteo.com/v1/forecast", "temperature_2m", city,
                          {"forecast_days": 3})
    return {source: values[tomorrow()] for source, values in days.items() if tomorrow() in values}

It uses the same helper, but with the normal forecast address and temperature_2m (today's forecast), then keeps only tomorrow.

Try it:

python
>>> from config import CITIES
>>> from sources import fetch_models_tomorrow
>>> fetch_models_tomorrow(CITIES["eger"])
{'ecmwf': (22.2, 15.4), 'gfs': (24.2, 11.7), 'icon': (24.6, 7.0)}
>>> exit()

Step 6: sources.py: what really happened

Add this function at the end of sources.py:

def fetch_measured(city, start, end):
    # Today is not over yet, so its max/min is not known. The archive still sends a
    # value for today (an estimate), so we never ask for anything after yesterday.
    end = min(end, (date.today() - timedelta(days=1)).isoformat())
    params = {
        "latitude": city["lat"], "longitude": city["lon"],
        "daily": "temperature_2m_max,temperature_2m_min",
        "timezone": TIMEZONE, "start_date": start, "end_date": end,
    }
    response = requests.get("https://archive-api.open-meteo.com/v1/archive", params=params, timeout=120)
    response.raise_for_status()
    daily = response.json()["daily"]
    return {day: (tmax, tmin)
            for day, tmax, tmin in zip(daily["time"], daily["temperature_2m_max"], daily["temperature_2m_min"])
            if tmax is not None and tmin is not None}

Try it. Ask until today on purpose (use today's date):

python
>>> from config import CITIES
>>> from sources import fetch_measured
>>> fetch_measured(CITIES["eger"], "2026-10-01", "2026-10-06")
{'2026-10-01': (22.3, 12.0), '2026-10-02': (22.2, 11.4), '2026-10-03': (21.9, 10.9), '2026-10-04': (22.4, 10.6), '2026-10-05': (23.0, 11.4)}
>>> exit()

The last day is yesterday (10-05), not today (10-06). Trap 2 works.

Finally, change the first line of sources.py (the description) to:

"""Download forecasts and measured temperatures from the internet.

Step 7: db.py: a table for the measured values

Open db.py. Add a second table to SCHEMA, so it looks like this:

SCHEMA = """
CREATE TABLE IF NOT EXISTS forecasts (
    city TEXT, source TEXT, day TEXT, tmax REAL, tmin REAL, saved_at TEXT,
    PRIMARY KEY (city, source, day)
);
CREATE TABLE IF NOT EXISTS measured (
    city TEXT, day TEXT, tmax REAL, tmin REAL,
    PRIMARY KEY (city, day)
);
"""

One measured value per city and day: the primary key is (city, day).

Then add this function at the end of db.py:

def save_measured(city, values):
    rows = [(city, day, tmax, tmin) for day, (tmax, tmin) in values.items()]
    con = connect()
    con.executemany("INSERT OR REPLACE INTO measured VALUES (?, ?, ?, ?)", rows)
    con.commit()
    con.close()
    return len(rows)

It is the same idea as save_forecasts() from week 2.

Try it. Save the same 5 days twice:

python
>>> from config import CITIES
>>> from sources import fetch_measured
>>> from db import save_measured
>>> values = fetch_measured(CITIES["eger"], "2026-10-01", "2026-10-05")
>>> save_measured("eger", values)
5
>>> save_measured("eger", values)
5
>>> import sqlite3
>>> sqlite3.connect("weather.db").execute("SELECT COUNT(*) FROM measured").fetchone()
(5,)
>>> exit()

5 rows, not 10: INSERT OR REPLACE overwrote them the second time. No duplicates.


Step 8: load_history.py: download the past once

Create load_history.py:

load_history.py

"""Run ONCE: download the history set (old forecasts + measured values) into the database.

    python load_history.py
"""
from datetime import date

from config import CITIES, HISTORY_START
from db import save_forecasts, save_measured
from sources import fetch_measured, fetch_model_history

end = date.today().isoformat()

for key, city in CITIES.items():
    print(f"{city['name']}: downloading {HISTORY_START} .. {end}")

    history = fetch_model_history(city, HISTORY_START, end)
    for source, values in history.items():
        print(f"  {source:6s} {save_forecasts(key, source, values)} days")

    measured = fetch_measured(city, HISTORY_START, end)
    print(f"  measured {save_measured(key, measured)} days")

print("Done.")

It goes through both cities, downloads the history from HISTORY_START until today, and saves it.

Run it:

python load_history.py

You should see:

Budapest: downloading 2024-02-05 .. 2026-10-06
  ecmwf  975 days
  gfs    975 days
  icon   975 days
  measured 974 days
Eger: downloading 2024-02-05 .. 2026-10-06
  ecmwf  975 days
  gfs    975 days
  icon   975 days
  measured 974 days
Done.

Step 9: pandas in five minutes

Before we read the database into pandas, try the two important operations on a tiny, made-up table (only 2 days and 2 sources, so you can follow every number; the real table has ~970 days and 3 sources). Create try_pandas.py:

import pandas as pd

# One row per forecast, like in our database ("long" format)
long = pd.DataFrame({
    "day":    ["10-01", "10-01", "10-02", "10-02"],
    "source": ["ecmwf", "gfs",   "ecmwf", "gfs"],
    "tmax":   [22.1,    23.1,    22.0,    23.1],
})
print(long)
print("---")

# One row per day, one column per source ("wide" format)
wide = long.pivot(index="day", columns="source", values="tmax")
print(wide)
print("---")

# The measured values, with the day as index
measured = pd.DataFrame({"day": ["10-01"], "measured": [22.3]}).set_index("day")

# Put them next to the forecasts, matching by day
table = wide.join(measured, how="left")
print(table)
print("---")

print(table.dropna())
print("---")
print(table["ecmwf"].mean())

Run it:

python try_pandas.py

You should see:

     day source  tmax
0  10-01  ecmwf  22.1
1  10-01    gfs  23.1
2  10-02  ecmwf  22.0
3  10-02    gfs  23.1
---
source  ecmwf   gfs
day
10-01    22.1  23.1
10-02    22.0  23.1
---
       ecmwf   gfs  measured
day
10-01   22.1  23.1      22.3
10-02   22.0  23.1       NaN
---
       ecmwf   gfs  measured
day
10-01   22.1  23.1      22.3
---
22.05

This "one row per day: forecasts + measured" shape is exactly what machine learning needs later: the inputs and the correct answer in one row. You can delete try_pandas.py now.


Step 10: db.py: read a city's data into one table

Add import pandas as pd at the top of db.py, after from datetime import datetime:

import sqlite3
from datetime import datetime

import pandas as pd

from config import DB_PATH

Then add this function at the end of db.py:

def load_table(city, sources):
    """One row per day: the max temperature forecast of each source + the measured value.

    Days that have forecasts but no measurement yet (for example tomorrow) are kept,
    their 'measured' value is empty (NaN).
    """
    con = connect()
    forecasts = pd.read_sql_query("SELECT day, source, tmax FROM forecasts WHERE city = ?", con, params=(city,))
    measured = pd.read_sql_query("SELECT day, tmax AS measured FROM measured WHERE city = ?", con, params=(city,))
    con.close()
    forecasts = forecasts[forecasts["source"].isin(sources)]
    table = forecasts.pivot(index="day", columns="source", values="tmax").reindex(columns=sources)
    table = table.join(measured.set_index("day"), how="left")
    return table.sort_index()

It is step 9, with real data: 1. pd.read_sql_query() runs a SQL query and gives the result as a DataFrame (long format, like long in step 9). 2. isin(sources) keeps only the sources we asked for (e.g. only the three models). 3. pivot() makes one row per day, one column per source. reindex(columns=sources) puts the columns in our order. 4. join(..., how="left") adds the measured value. Days without a measurement (today, tomorrow) get NaN. 5. sort_index() sorts by day.

This is the most important function of the project: everything from now on uses it.

Try it:

python
>>> from db import load_table
>>> table = load_table("eger", ["ecmwf", "gfs", "icon"])
>>> table.tail(4)
            ecmwf   gfs  icon  measured
day
2026-10-03   22.1  22.8  23.4      21.9
2026-10-04   22.5  23.2  24.2      22.4
2026-10-05   23.2  23.9  25.0      23.0
2026-10-06   23.2  24.3  25.2       NaN
>>> len(table)
975
>>> exit()

tail(4) shows the last 4 rows. Today (10-06) has forecasts, but no measured value yet: NaN. Correct.


Step 11: collect.py: save the models and the measured values every day

Replace collect.py with this version:

collect.py

"""Daily job: save tomorrow's forecasts and update the measured values.

    python collect.py          run once (this is what Windows Task Scheduler starts)
    python collect.py --loop   keep running, collect every day at 08:00
"""
import logging
import sys
import time
from datetime import date, datetime, timedelta

from config import BASE_DIR, CITIES
from db import save_forecasts, save_measured
from sources import PROVIDERS, fetch_measured, fetch_models_tomorrow, tomorrow

logging.basicConfig(
    level=logging.INFO, format="%(asctime)s %(message)s",
    handlers=[logging.FileHandler(BASE_DIR / "collect.log", encoding="utf-8"), logging.StreamHandler()],
)
log = logging.getLogger()


def collect_forecasts():
    for key, city in CITIES.items():
        # Live set: one provider failing must not stop the others.
        for source, fetch in PROVIDERS.items():
            try:
                tmax, tmin = fetch(city)
                save_forecasts(key, source, {tomorrow(): (tmax, tmin)})
                log.info(f"{key:9s} {source:11s} {tomorrow()} max {tmax} min {tmin}")
            except Exception as error:
                log.error(f"{key:9s} {source:11s} FAILED: {error}")
        # History set: keep it growing with today's model forecasts.
        try:
            for source, values in fetch_models_tomorrow(city).items():
                save_forecasts(key, source, {tomorrow(): values})
            log.info(f"{key:9s} models      saved")
        except Exception as error:
            log.error(f"{key:9s} models      FAILED: {error}")


def collect_measured():
    # The measured values arrive a few days late, so always ask for the last 10 days again.
    start = (date.today() - timedelta(days=10)).isoformat()
    end = (date.today() - timedelta(days=1)).isoformat()
    for key, city in CITIES.items():
        try:
            log.info(f"{key:9s} measured    {save_measured(key, fetch_measured(city, start, end))} days")
        except Exception as error:
            log.error(f"{key:9s} measured    FAILED: {error}")


def run_once():
    collect_forecasts()
    collect_measured()


def seconds_until(hour):
    now = datetime.now()
    next_run = now.replace(hour=hour, minute=0, second=0, microsecond=0)
    if next_run <= now:
        next_run += timedelta(days=1)
    return (next_run - now).total_seconds()


if __name__ == "__main__":
    run_once()
    if "--loop" in sys.argv:
        while True:
            time.sleep(seconds_until(8))
            run_once()

What is new compared with week 2: - In collect_forecasts(), a second try block saves the three models' forecast for tomorrow (step 5). So the history set keeps growing every day. - The new function collect_measured() downloads the measured values of the last 10 days (trap 1) and saves them. - run_once() calls both.

The Task Scheduler task from week 2 starts the same file, so you do not need to change it.

Run it:

python collect.py

You should see:

2026-10-06 14:34:58,063 budapest  open_meteo  2026-10-07 max 25.0 min 12.3
2026-10-06 14:34:58,193 budapest  met_norway  2026-10-07 max 22.8 min 9.0
2026-10-06 14:34:58,278 budapest  wttr        2026-10-07 max 25.0 min 14.0
2026-10-06 14:34:58,356 budapest  models      saved
2026-10-06 14:34:58,434 eger      open_meteo  2026-10-07 max 24.4 min 11.9
2026-10-06 14:34:58,561 eger      met_norway  2026-10-07 max 22.0 min 10.4
2026-10-06 14:34:58,666 eger      wttr        2026-10-07 max 23.0 min 10.0
2026-10-06 14:34:58,744 eger      models      saved
2026-10-06 14:34:58,824 budapest  measured    10 days
2026-10-06 14:34:58,904 eger      measured    10 days

Check the result in the interactive Python:

python
>>> from db import load_table
>>> load_table("eger", ["ecmwf", "gfs", "icon"]).tail(3)
            ecmwf   gfs  icon  measured
day
2026-10-05   23.2  23.9  25.0      23.0
2026-10-06   23.2  24.3  25.2       NaN
2026-10-07   22.2  24.2  24.6       NaN
>>> load_table("eger", ["open_meteo", "met_norway", "wttr"])
            open_meteo  met_norway  wttr  measured
day
2026-10-07        24.4        22.0  23.0       NaN
>>> exit()

Tomorrow (10-07) is now in the models set too. The providers set has as many days as you have collected since week 2.


Step 12: check_data.py: look before you learn

Bad data gives bad models. Before any machine learning, we check the data. Create check_data.py:

check_data.py

"""Week 3: look at the data before using it. Bad data -> bad models.

    python check_data.py
"""
from config import CITIES, SETS
from db import load_table

for city in CITIES:
    for data_set, sources in SETS.items():
        table = load_table(city, sources)
        name = CITIES[city]["name"]
        if table.empty:
            print(f"\n{name} / {data_set}: no data yet")
            continue
        print(f"\n{name} / {data_set}: {len(table)} days ({table.index.min()} .. {table.index.max()})")
        # 1. Missing values per column
        print("  missing values:", table.isna().sum().to_dict())
        # 2. Complete days: every source AND the measured value present
        print("  complete days: ", len(table.dropna()))
        # 3. Suspicious values: a forecast more than 8 degrees away from the measured value
        complete = table.dropna()
        for source in sources:
            wrong = complete[(complete[source] - complete["measured"]).abs() > 8]
            if len(wrong):
                print(f"  {source}: {len(wrong)} days more than 8 °C wrong, e.g. {wrong.index[0]}")

For each city and each set, it prints: 1. how many values are missing in each column (isna().sum() counts the NaNs), 2. how many days are complete, 3. suspicious days: a forecast more than 8 °C away from the measured value (maybe a data error).

Run it:

python check_data.py

You should see (your numbers will be a bit bigger):

Budapest / models: 976 days (2024-02-05 .. 2026-10-07)
  missing values: {'ecmwf': 0, 'gfs': 0, 'icon': 0, 'measured': 2}
  complete days:  974

Budapest / providers: 1 days (2026-10-07 .. 2026-10-07)
  missing values: {'open_meteo': 0, 'met_norway': 0, 'wttr': 0, 'measured': 1}
  complete days:  0

Eger / models: 976 days (2024-02-05 .. 2026-10-07)
  missing values: {'ecmwf': 0, 'gfs': 0, 'icon': 0, 'measured': 2}
  complete days:  974

Eger / providers: 1 days (2026-10-07 .. 2026-10-07)
  missing values: {'open_meteo': 0, 'met_norway': 0, 'wttr': 0, 'measured': 1}
  complete days:  0

How to read it: - The 2 missing measured values in models are today and tomorrow. That is correct: they have not happened yet. - No "more than 8 °C wrong" lines: no suspicious values. - providers will have more days: one more for every day collect.py has run since week 2. (In this example the collection started today.)


Your complete changed files

Compare with yours. If something does not work, copy these.

config.py

"""Settings shared by all scripts."""
from pathlib import Path

# Files are always created next to this file, no matter where you start Python from.
BASE_DIR = Path(__file__).resolve().parent
DB_PATH = BASE_DIR / "weather.db"

TIMEZONE = "Europe/Budapest"

CITIES = {
    "budapest": {"name": "Budapest", "lat": 47.4979, "lon": 19.0402},
    "eger": {"name": "Eger", "lat": 47.9025, "lon": 20.3772},
}

# MET Norway wants to know who is asking. You may add your own e-mail address,
# but NOT a fake one like ...@example.com (that is blocked with error 403).
USER_AGENT = "weather-lab/1.0 (university lab project)"

# Live set: three independent forecast providers, collected by us every day.
PROVIDER_SOURCES = ["open_meteo", "met_norway", "wttr"]

# History set: three weather models whose old day-ahead forecasts can be downloaded.
MODEL_SOURCES = ["ecmwf", "gfs", "icon"]
OPEN_METEO_MODELS = {"ecmwf": "ecmwf_ifs025", "gfs": "gfs_seamless", "icon": "icon_seamless"}
HISTORY_START = "2024-02-05"

SETS = {"models": MODEL_SOURCES, "providers": PROVIDER_SOURCES}

sources.py

"""Download forecasts and measured temperatures from the internet.

Every function returns plain Python values: (tmax, tmin) or {day: (tmax, tmin)}.
A day is always a text like "2026-10-07" in local (Hungarian) time.
"""
from datetime import date, datetime, timedelta
from zoneinfo import ZoneInfo

import requests

from config import OPEN_METEO_MODELS, TIMEZONE, USER_AGENT

TIMEOUT = 30  # seconds to wait for an answer
LOCAL = ZoneInfo(TIMEZONE)


def tomorrow():
    return (date.today() + timedelta(days=1)).isoformat()


def daily_max_min(times, temps):
    """Turn hourly temperatures into {day: (tmax, tmin)}.

    times are local times like "2026-10-07T13:00". Days with fewer than
    20 hourly values are skipped, because their max/min would be unreliable.
    """
    by_day = {}
    for time, temp in zip(times, temps):
        if temp is not None:
            by_day.setdefault(time[:10], []).append(temp)
    return {day: (max(values), min(values)) for day, values in by_day.items() if len(values) >= 20}


# ---------- Live set: three providers, tomorrow only ----------

def fetch_open_meteo(city):
    params = {
        "latitude": city["lat"], "longitude": city["lon"],
        "daily": "temperature_2m_max,temperature_2m_min",
        "timezone": TIMEZONE, "forecast_days": 3,
    }
    response = requests.get("https://api.open-meteo.com/v1/forecast", params=params, timeout=TIMEOUT)
    response.raise_for_status()
    daily = response.json()["daily"]
    i = daily["time"].index(tomorrow())
    return daily["temperature_2m_max"][i], daily["temperature_2m_min"][i]


def fetch_met_norway(city):
    params = {"lat": round(city["lat"], 4), "lon": round(city["lon"], 4)}
    response = requests.get("https://api.met.no/weatherapi/locationforecast/2.0/compact",
                            params=params, headers={"User-Agent": USER_AGENT}, timeout=TIMEOUT)
    response.raise_for_status()
    times, temps = [], []
    for step in response.json()["properties"]["timeseries"]:
        # MET Norway gives UTC times ("...Z"); convert them to Hungarian time.
        utc_time = datetime.fromisoformat(step["time"].replace("Z", "+00:00"))
        times.append(utc_time.astimezone(LOCAL).strftime("%Y-%m-%dT%H:%M"))
        temps.append(step["data"]["instant"]["details"]["air_temperature"])
    return daily_max_min(times, temps)[tomorrow()]


def fetch_wttr(city):
    url = f"https://wttr.in/{city['lat']},{city['lon']}"
    response = requests.get(url, params={"format": "j1"}, timeout=TIMEOUT)
    response.raise_for_status()
    for day in response.json()["weather"]:
        if day["date"] == tomorrow():
            return float(day["maxtempC"]), float(day["mintempC"])
    raise ValueError("wttr.in did not send tomorrow's forecast")


PROVIDERS = {"open_meteo": fetch_open_meteo, "met_norway": fetch_met_norway, "wttr": fetch_wttr}


# ---------- History set: three weather models, past and tomorrow ----------

def _models_hourly(url, variable, city, extra_params):
    """Ask Open-Meteo for hourly temperatures of our three models -> {source: {day: (tmax, tmin)}}."""
    params = {
        "latitude": city["lat"], "longitude": city["lon"],
        "hourly": variable, "models": ",".join(OPEN_METEO_MODELS.values()),
        "timezone": TIMEZONE, **extra_params,
    }
    response = requests.get(url, params=params, timeout=120)
    response.raise_for_status()
    hourly = response.json()["hourly"]
    return {source: daily_max_min(hourly["time"], hourly[f"{variable}_{model}"])
            for source, model in OPEN_METEO_MODELS.items()}


def fetch_model_history(city, start, end):
    """Old forecasts that were made ONE DAY BEFORE each day (true day-ahead forecasts)."""
    return _models_hourly("https://previous-runs-api.open-meteo.com/v1/forecast",
                          "temperature_2m_previous_day1", city,
                          {"start_date": start, "end_date": end})


def fetch_models_tomorrow(city):
    """Today's forecast of the three models for tomorrow -> {source: (tmax, tmin)}."""
    days = _models_hourly("https://api.open-meteo.com/v1/forecast", "temperature_2m", city,
                          {"forecast_days": 3})
    return {source: values[tomorrow()] for source, values in days.items() if tomorrow() in values}


# ---------- What really happened ----------

def fetch_measured(city, start, end):
    # Today is not over yet, so its max/min is not known. The archive still sends a
    # value for today (an estimate), so we never ask for anything after yesterday.
    end = min(end, (date.today() - timedelta(days=1)).isoformat())
    params = {
        "latitude": city["lat"], "longitude": city["lon"],
        "daily": "temperature_2m_max,temperature_2m_min",
        "timezone": TIMEZONE, "start_date": start, "end_date": end,
    }
    response = requests.get("https://archive-api.open-meteo.com/v1/archive", params=params, timeout=120)
    response.raise_for_status()
    daily = response.json()["daily"]
    return {day: (tmax, tmin)
            for day, tmax, tmin in zip(daily["time"], daily["temperature_2m_max"], daily["temperature_2m_min"])
            if tmax is not None and tmin is not None}

db.py

"""Everything that touches the SQLite database file (weather.db)."""
import sqlite3
from datetime import datetime

import pandas as pd

from config import DB_PATH

SCHEMA = """
CREATE TABLE IF NOT EXISTS forecasts (
    city TEXT, source TEXT, day TEXT, tmax REAL, tmin REAL, saved_at TEXT,
    PRIMARY KEY (city, source, day)
);
CREATE TABLE IF NOT EXISTS measured (
    city TEXT, day TEXT, tmax REAL, tmin REAL,
    PRIMARY KEY (city, day)
);
"""


def connect():
    con = sqlite3.connect(DB_PATH)
    con.executescript(SCHEMA)  # creates the tables the first time, does nothing later
    return con


def now():
    return datetime.now().isoformat(timespec="seconds")


def save_forecasts(city, source, values):
    """values: {day: (tmax, tmin)}. Saving the same day again overwrites it (no duplicates)."""
    rows = [(city, source, day, tmax, tmin, now()) for day, (tmax, tmin) in values.items()]
    con = connect()
    con.executemany("INSERT OR REPLACE INTO forecasts VALUES (?, ?, ?, ?, ?, ?)", rows)
    con.commit()
    con.close()
    return len(rows)


def save_measured(city, values):
    rows = [(city, day, tmax, tmin) for day, (tmax, tmin) in values.items()]
    con = connect()
    con.executemany("INSERT OR REPLACE INTO measured VALUES (?, ?, ?, ?)", rows)
    con.commit()
    con.close()
    return len(rows)


def load_table(city, sources):
    """One row per day: the max temperature forecast of each source + the measured value.

    Days that have forecasts but no measurement yet (for example tomorrow) are kept,
    their 'measured' value is empty (NaN).
    """
    con = connect()
    forecasts = pd.read_sql_query("SELECT day, source, tmax FROM forecasts WHERE city = ?", con, params=(city,))
    measured = pd.read_sql_query("SELECT day, tmax AS measured FROM measured WHERE city = ?", con, params=(city,))
    con.close()
    forecasts = forecasts[forecasts["source"].isin(sources)]
    table = forecasts.pivot(index="day", columns="source", values="tmax").reindex(columns=sources)
    table = table.join(measured.set_index("day"), how="left")
    return table.sort_index()

New files: load_history.py (step 8), collect.py (step 11) and check_data.py (step 12) are shown in full above.

Check

Extra

  1. Which model has the most days with an error over 3 °C? Change the limit in check_data.py.
  2. Add a check: are there days where tmax < tmin? (There should not be.)
  3. Think: why is a random split of days into train/test dangerous for this kind of data? (Answer in week 5.)

If something goes wrong

Problem Reason
ModuleNotFoundError: No module named 'pandas' Run pip install -r requirements.txt with the venv active.
ImportError: cannot import name 'OPEN_METEO_MODELS' You did not add the new lines to config.py (step 3).
KeyError: 'hourly' The service answered with an error message instead of data. Print response.json() to see it.
ReadTimeout in load_history.py Many students at the same time. Wait a minute and run it again.
providers has 0 complete days Normal at the beginning: tomorrow's forecast has no measured value yet. Wait a few days.
A function "does not exist" in the interactive Python You changed the file: exit() and start python again.

← Week 2Week 4 →