Week 8: Automation, live evaluation, tests
Goal of today: the system runs by itself. Every morning it collects the forecasts, retrains the models, forecasts tomorrow and saves our forecast. The web page shows how good our forecasts were in reality (the live score, with a red line on the chart). Automatic tests check the most important parts of the code. And finally you train on your own live data from the three providers.
How to work: do the steps in order. After each step, run the code and compare with the expected output. Go on only when yours looks the same. (Your temperatures and dates will differ, because you run it on another day.)
Start from your week 7 project (or the week 7 download), with the venv active.
What you need to know
Live evaluation. In weeks 5–7 we measured the error on past days that the model did not see while learning. That is good, but it is still the past: we chose the data, the split, the settings, and we may have tuned them (even by accident) until the numbers looked nice. The real test is the future: we save our forecast every morning, and a day later, when the measured value arrives, we compare. Nobody can cheat this test.
Automation. collect.py already runs every morning (week 2). Now it will also retrain the models (a few seconds)
and save tomorrow's forecast. So the model always knows the latest days, and nobody has to press a button.
SQL JOIN. We will have two tables: our predictions, and the measured values. A JOIN puts together rows of two
tables that belong together, here: the same city and the same day.
predictions measured predictions LEFT JOIN measured
day predicted day measured day predicted measured
2026-10-01 23.0 2026-10-01 22.3 2026-10-01 23.0 22.3
2026-10-09 20.0 2026-10-09 20.0 (empty)
LEFT JOIN keeps every row of the left table (predictions), even if the right table has nothing for it yet.
(A plain JOIN would drop 2026-10-09, because its measured value does not exist yet.)
Automated test. A small function that runs a piece of your code with a known input and checks the result. If somebody (you, next month) breaks that piece, the test fails immediately, instead of the web page showing nonsense.
assert checks that something is true. If it is not, the program stops with an AssertionError:
assert 1 + 1 == 2 # nothing happens
assert 1 + 1 == 3 # AssertionError
pytest is a tool that finds tests and runs them. It looks for files called test_*.py and, inside them,
functions called test_.... You do not call them yourself.
Fixture. Something a test needs prepared before it runs (for example: an empty database). You write it once, and every test that names it as a parameter gets it.
monkeypatch temporarily replaces a value during one test. We use it to point DB_PATH to a temporary file,
so the tests never touch your real weather.db.
Part 1: Saving and showing our predictions
Step 1: A table for our predictions
Open db.py and add a third table at the end of SCHEMA (before the closing """):
CREATE TABLE IF NOT EXISTS predictions (
city TEXT, data_set TEXT, day TEXT, tmax REAL, saved_at TEXT,
PRIMARY KEY (city, data_set, day)
);
One row = our forecast for one city, one data set (models or providers) and one day.
The primary key makes sure there is only one forecast per city + data set + day.
Try it. The table is created the next time anything connects to the database:
python
>>> from db import connect
>>> con = connect()
>>> con.execute("SELECT name FROM sqlite_master WHERE type = 'table'").fetchall()
[('forecasts',), ('measured',), ('predictions',)]
>>> con.close()
>>> exit()
sqlite_master is a built-in table that lists all tables of the database file. The new one is there.
Step 2: Saving a prediction
Add this function to db.py, after save_measured():
def save_prediction(city, data_set, day, tmax):
con = connect()
con.execute("INSERT OR REPLACE INTO predictions VALUES (?, ?, ?, ?, ?)",
(city, data_set, day, tmax, now()))
con.commit()
con.close()
It works like save_forecasts() from week 2: INSERT OR REPLACE with ? placeholders, then commit().
Try it with three made-up predictions for three days of last week (use your own dates: any days from the last week that already have a measured value):
python
>>> from db import connect, save_prediction
>>> save_prediction("eger", "models", "2026-10-01", 23.0)
>>> save_prediction("eger", "models", "2026-10-02", 23.4)
>>> save_prediction("eger", "models", "2026-10-03", 22.9)
>>> con = connect()
>>> con.execute("SELECT * FROM predictions").fetchall()
[('eger', 'models', '2026-10-01', 23.0, '2026-10-06T14:35:14'), ('eger', 'models', '2026-10-02', 23.4, '2026-10-06T14:35:14'), ('eger', 'models', '2026-10-03', 22.9, '2026-10-06T14:35:14')]
>>> con.close()
>>> exit()
These rows are only for trying the next steps. We delete them in step 7.
Step 3: Reading predictions together with the measured values
Add this function at the end of db.py:
def load_predictions(city, data_set):
con = connect()
table = pd.read_sql_query(
"SELECT p.day, p.tmax AS predicted, m.tmax AS measured FROM predictions p "
"LEFT JOIN measured m ON m.city = p.city AND m.day = p.day "
"WHERE p.city = ? AND p.data_set = ? ORDER BY p.day", con, params=(city, data_set))
con.close()
return table.set_index("day")
predictions pandmeasured mgive the two tables short names, so we can writep.dayandm.tmax.LEFT JOIN measured m ON m.city = p.city AND m.day = p.day: for every prediction, find the measured value of the same city and day (see What you need to know).p.tmax AS predicted,m.tmax AS measured: both tables have atmaxcolumn, so we rename them.
Try it. Add one more made-up prediction, for a day in the future, and read everything back:
python
>>> from db import load_predictions, save_prediction
>>> save_prediction("eger", "models", "2026-10-09", 20.0)
>>> load_predictions("eger", "models")
predicted measured
day
2026-10-01 23.0 22.3
2026-10-02 23.4 22.2
2026-10-03 22.9 21.9
2026-10-09 20.0 NaN
>>> exit()
The past days got their measured value. The future day is still in the table, with NaN (empty):
this is what LEFT JOIN does.
Your db.py is now complete. It must look like the complete file at the end of this page.
Step 4: collect.py: train and forecast every morning
Open collect.py. First change the imports:
from config import BASE_DIR, CITIES, SETS
from db import save_forecasts, save_measured, save_prediction
from ml import predict_tomorrow, train_and_save
Add this function after collect_measured():
def train_and_predict():
for key in CITIES:
for data_set in SETS:
try:
scores, days = train_and_save(key, data_set)
if scores is None:
log.info(f"{key:9s} {data_set:9s} only {days} complete days, not trained yet")
continue
prediction = predict_tomorrow(key, data_set)
if prediction is not None:
save_prediction(key, data_set, tomorrow(), prediction)
log.info(f"{key:9s} {data_set:9s} trained on {days} days, tomorrow: {prediction}")
except Exception as error:
log.error(f"{key:9s} {data_set:9s} FAILED: {error}")
And call it at the end of run_once():
def run_once():
collect_forecasts()
collect_measured()
train_and_predict()
What train_and_predict() does, for both cities and both data sets:
1. train_and_save() (from week 7) trains the combined model on all complete days and saves it.
If there are fewer than 30 complete days, it returns None, and we only write a note into the log.
2. predict_tomorrow() (from week 7) forecasts tomorrow.
3. save_prediction() saves the forecast, so that we can compare it with reality later.
4. try / except: if one city or data set fails, the others still run (same idea as in week 2).
Also change the first line (the description) of the file to:
"""Daily job: save tomorrow's forecasts, update measured values, retrain, predict.
Run it:
python collect.py
You should see (it takes 10–30 seconds, because the models are trained now):
2026-10-06 14:35:26,590 budapest open_meteo 2026-10-07 max 25.0 min 12.3
2026-10-06 14:35:26,726 budapest met_norway 2026-10-07 max 22.8 min 9.0
2026-10-06 14:35:26,832 budapest wttr 2026-10-07 max 25.0 min 14.0
2026-10-06 14:35:26,910 budapest models saved
2026-10-06 14:35:26,987 eger open_meteo 2026-10-07 max 24.4 min 11.9
2026-10-06 14:35:27,116 eger met_norway 2026-10-07 max 22.0 min 10.4
2026-10-06 14:35:27,200 eger wttr 2026-10-07 max 23.0 min 10.0
2026-10-06 14:35:27,279 eger models saved
2026-10-06 14:35:27,360 budapest measured 10 days
2026-10-06 14:35:27,442 eger measured 10 days
2026-10-06 14:35:34,241 budapest models trained on 974 days, tomorrow: 23.8
2026-10-06 14:35:34,246 budapest providers only 0 complete days, not trained yet
2026-10-06 14:35:40,711 eger models trained on 974 days, tomorrow: 23.0
2026-10-06 14:35:40,717 eger providers only 0 complete days, not trained yet
The last four lines are new. (In your case providers will have more days, see step 13.)
The predictions table now has a row for tomorrow for each city.
The Task Scheduler task from week 2 needs no change: it starts the same collect.py, so from tomorrow
morning this all happens by itself.
Step 5: The live score
Now we show on the web page how good our saved predictions were. Open app.py.
Change the db import:
from db import load_predictions, load_table
Add this function after source_errors():
def live_score(city, data_set):
"""How good were OUR daily predictions, compared with the sources on the same days?"""
predictions = load_predictions(city, data_set).dropna()
if predictions.empty:
return None
table = load_table(city, SETS[data_set]).loc[predictions.index]
score = {"our prediction": round(mae(predictions["predicted"], predictions["measured"]), 2)}
score.update(source_errors(table, SETS[data_set]))
return {"days": len(predictions), "errors": score}
Step by step:
1. Load our predictions and keep only the days that have a measured value (dropna()).
2. If there are none yet, return None (the page will not show the live score).
3. Take the forecasts of the three sources for the same days (.loc[predictions.index]).
This is important for a fair comparison: we compare everybody on exactly the same days.
4. Compute our error and the error of each source with mae() and source_errors().
Then give it to the template: in city_page(), add one line after best_forecast=...:
live=live_score(city, data_set),
Try it (the made-up rows from step 2 are still there):
python
>>> from app import live_score
>>> live_score("eger", "models")
{'days': 3, 'errors': {'our prediction': 0.97, 'ecmwf': 0.2, 'gfs': 0.87, 'icon': 1.53}}
>>> live_score("budapest", "models")
>>> exit()
Eger: 3 days with a measured value. (Our "prediction" is bad here because we made it up.)
Budapest prints nothing: the function returned None, because there are no predictions with a measured value yet.
Now show it on the page. In templates/city.html, add this before <h3>Latest days</h3>:
{% if live %}
<h3>Live score: our daily predictions ({{ live.days }} days)</h3>
<table>
<tr>{% for name in live.errors %}<th>{{ name }}</th>{% endfor %}</tr>
<tr>{% for value in live.errors.values() %}<td>{{ value }} °C</td>{% endfor %}</tr>
</table>
{% endif %}
{% if live %}: the table appears only if live_score() did not return None.
Run it:
python app.py
Open http://127.0.0.1:5000/city/eger.
You should see a new table above Latest days:
Live score: our daily predictions (3 days)
our prediction ecmwf gfs icon
0.97 °C 0.2 °C 0.87 °C 1.53 °C
Budapest has no live score table yet. Leave the server running.
Step 6: Our predictions on the chart
In app.py, in city_page(), load the predictions right after table = load_table(...).tail(60):
# Predictions saved by collect.py, shown on the chart next to the sources.
predicted = load_predictions(city, data_set)["predicted"]
and add one more entry to the chart dictionary:
"predicted": [predicted.get(day) for day in table.index],
predicted.get(day) gives our prediction for that day, or None if we have none. None becomes a gap in the line.
In templates/city.html, add one line in the <script> part, just before new Chart(...):
datasets.push({ label: "our prediction", data: chart.predicted, borderColor: "red", borderWidth: 2 });
Save, and reload the page in the browser (debug=True restarts the server by itself).
You should see a new red line our prediction in the legend. On the chart it has three points on your made-up days, and one point at the right end: tomorrow's real prediction from step 4.
Your app.py and templates/city.html are now complete (see the complete files).
Step 7: Delete the made-up predictions
The rows from steps 2 and 3 were only for trying. They must not count in the live score. Delete them (use your own dates):
python
>>> from db import connect
>>> con = connect()
>>> con.execute("DELETE FROM predictions WHERE day IN ('2026-10-01', '2026-10-02', '2026-10-03', '2026-10-09')")
<sqlite3.Cursor object at 0x744a2cf46540>
>>> con.commit()
>>> con.execute("SELECT city, data_set, day, tmax FROM predictions").fetchall()
[('budapest', 'models', '2026-10-07', 23.8), ('eger', 'models', '2026-10-07', 23.0)]
>>> con.close()
>>> exit()
Only the two real predictions for tomorrow remain. Reload the page: the live score table is gone, and it will come back by itself in a day or two, with real numbers.
Part 2: Automated tests
Step 8: Install pytest and write the first test
Add pytest to requirements.txt:
requirements.txt
requests
pandas
scikit-learn
flask
tzdata
pytest
pip install -r requirements.txt
Create a folder tests in your project folder, and in it a file tests/test_weather.py:
"""Tests that run without internet. Start them with: python -m pytest"""
import sys
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) # so "import db" works
from sources import daily_max_min
def test_daily_max_min():
times = [f"2026-10-07T{h:02d}:00" for h in range(24)]
temps = list(range(24))
assert daily_max_min(times, temps) == {"2026-10-07": (23, 0)}
- The
sys.path.insertline lets the test file importdb,mlandsourcesfrom the folder above it. - The test gives 24 hourly values (0, 1, 2, ... 23) for one day. The correct answer is max 23, min 0. We know it in advance, so we can check it.
Run it from the project folder:
python -m pytest
You should see:
. [100%]
1 passed in 0.19s
Every dot is a passed test.
Step 9: A second test, and what a failing test looks like
Add this at the end of tests/test_weather.py:
def test_daily_max_min_skips_incomplete_days():
times = [f"2026-10-07T{h:02d}:00" for h in range(10)]
assert daily_max_min(times, [5] * 10) == {}
Only 10 hourly values for the day: daily_max_min() must drop it (week 1: fewer than 20 values = unreliable).
python -m pytest
You should see: 2 passed.
Now let's break the code on purpose, to see what a test is good for. In sources.py, in daily_max_min(),
change >= 20 to >= 5, and run the tests again.
You should see:
.F [100%]
=================================== FAILURES ===================================
___________________ test_daily_max_min_skips_incomplete_days ___________________
def test_daily_max_min_skips_incomplete_days():
times = [f"2026-10-07T{h:02d}:00" for h in range(10)]
> assert daily_max_min(times, [5] * 10) == {}
E AssertionError: assert {'2026-10-07': (5, 5)} == {}
...
1 failed, 1 passed in 0.16s
F means failed. pytest shows which test, which line (>), and what the function returned instead of the expected
value (E). Change >= 5 back to >= 20 and check that both tests pass again.
Step 10: A test with a temporary database (fixture)
Testing the database is harder: a test must not write into your real weather.db.
Add two imports to the top of the test file. import pytest goes after from pathlib import Path;
import db and import ml go after the sys.path line (they need it):
import pytest
import db
import ml
from sources import daily_max_min
Then add the fixture after the imports:
@pytest.fixture
def empty_db(tmp_path, monkeypatch):
"""Every test gets its own, empty database and model folder."""
monkeypatch.setattr(db, "DB_PATH", tmp_path / "test.db")
monkeypatch.setattr(ml, "MODELS_DIR", tmp_path / "models")
@pytest.fixturemarks the function as a fixture.tmp_pathandmonkeypatchare fixtures that pytest already has: a new empty temporary folder for every test, and the tool that replaces values temporarily.- During the test,
db.DB_PATHpoints into the temporary folder. After the test, everything is restored.
And a test that uses it. Add it at the end of the file:
def test_saving_twice_does_not_duplicate(empty_db):
db.save_forecasts("eger", "wttr", {"2026-10-07": (20, 10)})
db.save_forecasts("eger", "wttr", {"2026-10-07": (21, 11)})
table = db.load_table("eger", ["wttr"])
assert len(table) == 1
assert table.loc["2026-10-07", "wttr"] == 21
The parameter empty_db is how the test asks for the fixture. The test saves the same day twice and checks
what week 2 promised: one row, with the newer value.
python -m pytest
You should see: 3 passed.
Your real weather.db is untouched: look at Eger / wttr for tomorrow in DB Browser. It still has the real value, not 21.
Step 11: Testing the date split
Add import pandas as pd before import pytest, and this test at the end:
def test_split_keeps_time_order():
days = [f"2026-01-{d:02d}" for d in range(1, 11)]
table = pd.DataFrame({"x": range(10)}, index=days)
train, test = ml.split_by_date(table)
assert train.index.max() < test.index.min()
assert len(test) == 2
10 days; the test part must be the newest 20% (2 days), and every training day must be older than every test day. This is the rule from week 5: never test on the past and train on the future.
python -m pytest
You should see: 4 passed.
Step 12: Testing the machine learning with fake data
The last two tests check training and predicting, without internet. They build fake sources with known errors around a "true" temperature: one source is 1.5 °C too cold, one is right, one is 1.5 °C too warm, all with random noise. A working model must be better than the worst of them.
Add from datetime import date, timedelta after import sys, and import numpy as np before import pandas as pd.
Then add at the end:
def fill_fake_data(sources, days=60):
"""Sources that are wrong in different ways around a 'true' temperature."""
rng = np.random.default_rng(0)
start = date.today() - timedelta(days=days)
for i in range(days + 2): # includes today and tomorrow
day = (start + timedelta(days=i)).isoformat()
true = 15 + 8 * np.sin(i / 10)
if day < date.today().isoformat():
db.save_measured("eger", {day: (true, true - 10)})
for k, source in enumerate(sources):
value = true + (k - 1) * 1.5 + rng.normal(0, 1) # bias -1.5, 0, +1.5 plus noise
db.save_forecasts("eger", source, {day: (value, value - 10)})
def test_train_and_predict(empty_db):
fill_fake_data(ml.SETS["providers"])
scores, days = ml.train_and_save("eger", "providers")
assert days == 60
# The combined model must be better than the worst single source.
assert scores["combined"] < max(scores[s] for s in ml.SETS["providers"])
prediction = ml.predict_tomorrow("eger", "providers")
assert prediction is not None
assert 0 < prediction < 30
def test_not_enough_data(empty_db):
fill_fake_data(ml.SETS["providers"], days=10)
scores, days = ml.train_and_save("eger", "providers")
assert scores is None
assert ml.predict_tomorrow("eger", "providers") is None
fill_fake_data()savesdaysdays of fake forecasts and measured values, plus forecasts for today and tomorrow (without measured values, like in real life).rng = np.random.default_rng(0): random numbers, but the same ones at every run, so the test always gives the same result.test_not_enough_datachecks the safety rule: with only 10 days, there is no training and no prediction.
Your test file must now look like the complete file. Run all tests, with -v (verbose) to see their names:
python -m pytest -v
You should see:
tests/test_weather.py::test_daily_max_min PASSED [ 16%]
tests/test_weather.py::test_daily_max_min_skips_incomplete_days PASSED [ 33%]
tests/test_weather.py::test_saving_twice_does_not_duplicate PASSED [ 50%]
tests/test_weather.py::test_split_keeps_time_order PASSED [ 66%]
tests/test_weather.py::test_train_and_predict PASSED [ 83%]
tests/test_weather.py::test_not_enough_data PASSED [100%]
============================== 6 passed in 4.61s ===============================
From now on: run python -m pytest after every change. It takes 5 seconds.
Part 3: Your own live data
Step 13: How much live data do you have?
You have been collecting the three providers since week 2. Check:
python check_data.py
Look at the providers lines. You should see something like complete days: 38. (Right after setting up a new
project you see complete days: 0, like this:)
Budapest / providers: 1 days (2026-10-07 .. 2026-10-07)
missing values: {'open_meteo': 0, 'met_norway': 0, 'wttr': 0, 'measured': 1}
complete days: 0
- 30 or more complete days: go on to step 14.
- Fewer, or big gaps (the computer was off, Task Scheduler did not run): download the class database
(see Downloads), close DB Browser and
app.py, and save it asweather.dbin your project folder (keep your own file under another name). Then runcheck_data.pyagain.
Step 14: Train on the live set and compare
python ml.py
Now the providers lines show results too, in the same format as models:
the three providers, their average, linear, random forest, neural net and combined.
(Below 30 complete days you see not enough data yet (need 30).)
What to expect: - The errors are less reliable than on the history set: the test part is only the newest 20%, that is 7–9 days. One unusual day changes the result a lot. - The simple models (linear) usually do better than the complex ones (random forest, neural net), which need more data. You saw this in the small data experiment in week 6.
Write the results of both sets and both cities into your results table: you need it for the presentation.
Then run python collect.py once more: with 30+ complete days, the log now says
providers trained on ... days, tomorrow: ..., and the providers page gets its own Our best forecast.
Check
- ☐
python -m pytest→6 passed. - ☐
collect.pylogstrained on ... days, tomorrow: ...formodels, and forprovidersif you have 30+ complete days. - ☐ The
predictionstable has a row for tomorrow, and the made-up rows are deleted. - ☐ After a few days: the page shows the Live score table and the red line.
- ☐
python ml.pyshows results for both sets. They are in your results table.
Final presentation (10 minutes per team)
- Live demo: the web page, both cities, both sets.
- The data: sources, history vs. live, how many days, problems you found (week 3).
- Results table: single sources, average, linear, random forest, neural net, combined, for both sets and cities.
- Live score: how good were your real daily predictions?
- Honest conclusions: What worked? What did not? Why did the neural net not win? What would you do with one more year?
Extra
- Predict
tmintoo. (Hint: most functions only need acolumnparameter instead of the fixed"tmax".) - Add a fourth provider. Everything else (database, web page, models) works without change. Only
sources.pyandconfig.pychange. - Write a test for
live_score(). (Hint: use theempty_dbfixture, save a few forecasts, measured values and predictions.) - Train one model on both cities together, with the city as an extra input (0 = Budapest, 1 = Eger). Does more data help?
If something goes wrong
| Problem | Reason |
|---|---|
no such table: predictions |
The predictions table is missing from SCHEMA in db.py (step 1). |
ImportError: cannot import name 'save_prediction' |
save_prediction() is missing from db.py (step 2), or there is a typo in its name. |
ModuleNotFoundError: No module named 'db' in pytest |
Start pytest from the project folder: python -m pytest. Check the sys.path.insert line. |
fixture 'empty_db' not found |
The @pytest.fixture line is missing above def empty_db. |
| The live score never appears | Predictions need 1–2 days until their measured value arrives. Check that collect.py runs every day (collect.log). |
providers still not trained |
Fewer than 30 complete days. Check python check_data.py, or use the class database. |
database is locked |
DB Browser or another program has weather.db open. Close it. |
Complete files
Compare with yours. If something does not work, copy these.
db.py
"""Everything that touches the SQLite database file (weather.db)."""
import sqlite3
from datetime import datetime
import pandas as pd
from config import DB_PATH
SCHEMA = """
CREATE TABLE IF NOT EXISTS forecasts (
city TEXT, source TEXT, day TEXT, tmax REAL, tmin REAL, saved_at TEXT,
PRIMARY KEY (city, source, day)
);
CREATE TABLE IF NOT EXISTS measured (
city TEXT, day TEXT, tmax REAL, tmin REAL,
PRIMARY KEY (city, day)
);
CREATE TABLE IF NOT EXISTS predictions (
city TEXT, data_set TEXT, day TEXT, tmax REAL, saved_at TEXT,
PRIMARY KEY (city, data_set, day)
);
"""
def connect():
con = sqlite3.connect(DB_PATH)
con.executescript(SCHEMA) # creates the tables the first time, does nothing later
return con
def now():
return datetime.now().isoformat(timespec="seconds")
def save_forecasts(city, source, values):
"""values: {day: (tmax, tmin)}. Saving the same day again overwrites it (no duplicates)."""
rows = [(city, source, day, tmax, tmin, now()) for day, (tmax, tmin) in values.items()]
con = connect()
con.executemany("INSERT OR REPLACE INTO forecasts VALUES (?, ?, ?, ?, ?, ?)", rows)
con.commit()
con.close()
return len(rows)
def save_measured(city, values):
rows = [(city, day, tmax, tmin) for day, (tmax, tmin) in values.items()]
con = connect()
con.executemany("INSERT OR REPLACE INTO measured VALUES (?, ?, ?, ?)", rows)
con.commit()
con.close()
return len(rows)
def save_prediction(city, data_set, day, tmax):
con = connect()
con.execute("INSERT OR REPLACE INTO predictions VALUES (?, ?, ?, ?, ?)",
(city, data_set, day, tmax, now()))
con.commit()
con.close()
def load_table(city, sources):
"""One row per day: the max temperature forecast of each source + the measured value.
Days that have forecasts but no measurement yet (for example tomorrow) are kept,
their 'measured' value is empty (NaN).
"""
con = connect()
forecasts = pd.read_sql_query("SELECT day, source, tmax FROM forecasts WHERE city = ?", con, params=(city,))
measured = pd.read_sql_query("SELECT day, tmax AS measured FROM measured WHERE city = ?", con, params=(city,))
con.close()
forecasts = forecasts[forecasts["source"].isin(sources)]
table = forecasts.pivot(index="day", columns="source", values="tmax").reindex(columns=sources)
table = table.join(measured.set_index("day"), how="left")
return table.sort_index()
def load_predictions(city, data_set):
con = connect()
table = pd.read_sql_query(
"SELECT p.day, p.tmax AS predicted, m.tmax AS measured FROM predictions p "
"LEFT JOIN measured m ON m.city = p.city AND m.day = p.day "
"WHERE p.city = ? AND p.data_set = ? ORDER BY p.day", con, params=(city, data_set))
con.close()
return table.set_index("day")
collect.py
"""Daily job: save tomorrow's forecasts, update measured values, retrain, predict.
python collect.py run once (this is what Windows Task Scheduler starts)
python collect.py --loop keep running, collect every day at 08:00
"""
import logging
import sys
import time
from datetime import date, datetime, timedelta
from config import BASE_DIR, CITIES, SETS
from db import save_forecasts, save_measured, save_prediction
from ml import predict_tomorrow, train_and_save
from sources import PROVIDERS, fetch_measured, fetch_models_tomorrow, tomorrow
logging.basicConfig(
level=logging.INFO, format="%(asctime)s %(message)s",
handlers=[logging.FileHandler(BASE_DIR / "collect.log", encoding="utf-8"), logging.StreamHandler()],
)
log = logging.getLogger()
def collect_forecasts():
for key, city in CITIES.items():
# Live set: one provider failing must not stop the others.
for source, fetch in PROVIDERS.items():
try:
tmax, tmin = fetch(city)
save_forecasts(key, source, {tomorrow(): (tmax, tmin)})
log.info(f"{key:9s} {source:11s} {tomorrow()} max {tmax} min {tmin}")
except Exception as error:
log.error(f"{key:9s} {source:11s} FAILED: {error}")
# History set: keep it growing with today's model forecasts.
try:
for source, values in fetch_models_tomorrow(city).items():
save_forecasts(key, source, {tomorrow(): values})
log.info(f"{key:9s} models saved")
except Exception as error:
log.error(f"{key:9s} models FAILED: {error}")
def collect_measured():
# The measured values arrive a few days late, so always ask for the last 10 days again.
start = (date.today() - timedelta(days=10)).isoformat()
end = (date.today() - timedelta(days=1)).isoformat()
for key, city in CITIES.items():
try:
log.info(f"{key:9s} measured {save_measured(key, fetch_measured(city, start, end))} days")
except Exception as error:
log.error(f"{key:9s} measured FAILED: {error}")
def train_and_predict():
for key in CITIES:
for data_set in SETS:
try:
scores, days = train_and_save(key, data_set)
if scores is None:
log.info(f"{key:9s} {data_set:9s} only {days} complete days, not trained yet")
continue
prediction = predict_tomorrow(key, data_set)
if prediction is not None:
save_prediction(key, data_set, tomorrow(), prediction)
log.info(f"{key:9s} {data_set:9s} trained on {days} days, tomorrow: {prediction}")
except Exception as error:
log.error(f"{key:9s} {data_set:9s} FAILED: {error}")
def run_once():
collect_forecasts()
collect_measured()
train_and_predict()
def seconds_until(hour):
now = datetime.now()
next_run = now.replace(hour=hour, minute=0, second=0, microsecond=0)
if next_run <= now:
next_run += timedelta(days=1)
return (next_run - now).total_seconds()
if __name__ == "__main__":
run_once()
if "--loop" in sys.argv:
while True:
time.sleep(seconds_until(8))
run_once()
app.py
"""The web app. python app.py then open http://127.0.0.1:5000"""
from flask import Flask, abort, redirect, render_template, request, url_for
from config import CITIES, SETS
from db import load_predictions, load_table
from ml import mae, predict_tomorrow, print_scores, train_and_save
from sources import tomorrow
app = Flask(__name__)
def source_errors(table, sources):
"""Average error of each source on the days where we know the measured value."""
known = table.dropna()
if known.empty:
return {}
return {source: round(mae(known[source], known["measured"]), 2) for source in sources}
def live_score(city, data_set):
"""How good were OUR daily predictions, compared with the sources on the same days?"""
predictions = load_predictions(city, data_set).dropna()
if predictions.empty:
return None
table = load_table(city, SETS[data_set]).loc[predictions.index]
score = {"our prediction": round(mae(predictions["predicted"], predictions["measured"]), 2)}
score.update(source_errors(table, SETS[data_set]))
return {"days": len(predictions), "errors": score}
@app.route("/")
def home():
return redirect(url_for("city_page", city="budapest"))
@app.route("/city/<city>")
def city_page(city):
if city not in CITIES:
abort(404)
data_set = request.args.get("set", "models")
if data_set not in SETS:
abort(404)
sources = SETS[data_set]
table = load_table(city, sources).tail(60)
# Predictions saved by collect.py, shown on the chart next to the sources.
predicted = load_predictions(city, data_set)["predicted"]
chart = {
"days": list(table.index),
"measured": [None if v != v else v for v in table["measured"]], # NaN -> None (empty in the chart)
"sources": {s: [None if v != v else v for v in table[s]] for s in sources},
"predicted": [predicted.get(day) for day in table.index],
}
tomorrow_forecasts = table.loc[tomorrow()].drop("measured").to_dict() if tomorrow() in table.index else {}
return render_template(
"city.html", cities=CITIES, city=city, data_set=data_set, sets=SETS,
chart=chart, errors=source_errors(table, sources),
tomorrow=tomorrow(), tomorrow_forecasts=tomorrow_forecasts,
best_forecast=predict_tomorrow(city, data_set),
live=live_score(city, data_set),
rows=table.iloc[::-1].head(15).round(1).to_dict("index"),
)
@app.route("/train/<city>/<data_set>", methods=["POST"])
def train(city, data_set):
if city not in CITIES or data_set not in SETS:
abort(404)
scores, days = train_and_save(city, data_set)
if scores:
print_scores(scores)
best = min(scores, key=scores.get) if scores else None
return render_template("train.html", cities=CITIES, city=city, data_set=data_set, sets=SETS,
scores=scores, days=days, best=best)
if __name__ == "__main__":
app.run(debug=True)
templates/city.html
{% extends "base.html" %}
{% block content %}
<h2>{{ cities[city].name }} – {{ data_set }}</h2>
<div class="box">
<b>Tomorrow ({{ tomorrow }}), max temperature</b><br>
{% for source, value in tomorrow_forecasts.items() %}
{{ source }}: {{ value }} °C
{% else %}
No forecasts for tomorrow yet. Run collect.py.
{% endfor %}
<br>
<b>Our best forecast: {{ best_forecast if best_forecast is not none else "no trained model yet" }}{% if best_forecast is not none %} °C{% endif %}</b>
<form method="post" action="{{ url_for('train', city=city, data_set=data_set) }}" style="display:inline">
<button>Train now</button>
</form>
</div>
<canvas id="chart" height="110"></canvas>
<h3>Average error of each source (last 60 days)</h3>
<table>
<tr>{% for source in errors %}<th>{{ source }}</th>{% endfor %}</tr>
<tr>{% for value in errors.values() %}<td>{{ value }} °C</td>{% endfor %}</tr>
</table>
{% if live %}
<h3>Live score: our daily predictions ({{ live.days }} days)</h3>
<table>
<tr>{% for name in live.errors %}<th>{{ name }}</th>{% endfor %}</tr>
<tr>{% for value in live.errors.values() %}<td>{{ value }} °C</td>{% endfor %}</tr>
</table>
{% endif %}
<h3>Latest days</h3>
<table>
<tr><th>day</th>{% for source in sets[data_set] %}<th>{{ source }}</th>{% endfor %}<th>measured</th></tr>
{% for day, row in rows.items() %}
<tr><td>{{ day }}</td>
{% for source in sets[data_set] %}<td>{{ row[source] if row[source] == row[source] else "" }}</td>{% endfor %}
<td><b>{{ row.measured if row.measured == row.measured else "" }}</b></td></tr>
{% endfor %}
</table>
<script>
const chart = {{ chart | tojson }};
const datasets = [{ label: "measured", data: chart.measured, borderColor: "black", borderWidth: 3 }];
const colors = ["#1f77b4", "#ff7f0e", "#2ca02c"];
Object.entries(chart.sources).forEach(([name, values], i) => {
datasets.push({ label: name, data: values, borderColor: colors[i], borderWidth: 1 });
});
datasets.push({ label: "our prediction", data: chart.predicted, borderColor: "red", borderWidth: 2 });
new Chart(document.getElementById("chart"), {
type: "line",
data: { labels: chart.days, datasets: datasets },
options: { spanGaps: false, pointRadius: 1 },
});
</script>
{% endblock %}
tests/test_weather.py
"""Tests that run without internet. Start them with: python -m pytest"""
import sys
from datetime import date, timedelta
from pathlib import Path
import numpy as np
import pandas as pd
import pytest
sys.path.insert(0, str(Path(__file__).resolve().parent.parent)) # so "import db" works
import db
import ml
from sources import daily_max_min
@pytest.fixture
def empty_db(tmp_path, monkeypatch):
"""Every test gets its own, empty database and model folder."""
monkeypatch.setattr(db, "DB_PATH", tmp_path / "test.db")
monkeypatch.setattr(ml, "MODELS_DIR", tmp_path / "models")
def test_daily_max_min():
times = [f"2026-10-07T{h:02d}:00" for h in range(24)]
temps = list(range(24))
assert daily_max_min(times, temps) == {"2026-10-07": (23, 0)}
def test_daily_max_min_skips_incomplete_days():
times = [f"2026-10-07T{h:02d}:00" for h in range(10)]
assert daily_max_min(times, [5] * 10) == {}
def test_saving_twice_does_not_duplicate(empty_db):
db.save_forecasts("eger", "wttr", {"2026-10-07": (20, 10)})
db.save_forecasts("eger", "wttr", {"2026-10-07": (21, 11)})
table = db.load_table("eger", ["wttr"])
assert len(table) == 1
assert table.loc["2026-10-07", "wttr"] == 21
def test_split_keeps_time_order():
days = [f"2026-01-{d:02d}" for d in range(1, 11)]
table = pd.DataFrame({"x": range(10)}, index=days)
train, test = ml.split_by_date(table)
assert train.index.max() < test.index.min()
assert len(test) == 2
def fill_fake_data(sources, days=60):
"""Sources that are wrong in different ways around a 'true' temperature."""
rng = np.random.default_rng(0)
start = date.today() - timedelta(days=days)
for i in range(days + 2): # includes today and tomorrow
day = (start + timedelta(days=i)).isoformat()
true = 15 + 8 * np.sin(i / 10)
if day < date.today().isoformat():
db.save_measured("eger", {day: (true, true - 10)})
for k, source in enumerate(sources):
value = true + (k - 1) * 1.5 + rng.normal(0, 1) # bias -1.5, 0, +1.5 plus noise
db.save_forecasts("eger", source, {day: (value, value - 10)})
def test_train_and_predict(empty_db):
fill_fake_data(ml.SETS["providers"])
scores, days = ml.train_and_save("eger", "providers")
assert days == 60
# The combined model must be better than the worst single source.
assert scores["combined"] < max(scores[s] for s in ml.SETS["providers"])
prediction = ml.predict_tomorrow("eger", "providers")
assert prediction is not None
assert 0 < prediction < 30
def test_not_enough_data(empty_db):
fill_fake_data(ml.SETS["providers"], days=10)
scores, days = ml.train_and_save("eger", "providers")
assert scores is None
assert ml.predict_tomorrow("eger", "providers") is None