Week 6: A neural network
Goal of today: add a neural network to the comparison of week 5, and understand when it helps and when it does not:
Eger / models: 974 complete days
ecmwf 0.82 °C
gfs 1.56 °C
icon 1.28 °C
average of sources 0.93 °C
linear 0.68 °C <-- best
random_forest 0.82 °C
neural_net 0.69 °C
How to work: do the steps in order. After each step, run the code and compare with the expected output. Go on only when yours looks the same. (Your numbers can differ a little: you have more days than when this page was written.)
What you need to know
A neuron is a tiny calculator. It multiplies each input by a weight, adds them up, and adds a constant (the bias). That is exactly what linear regression does. Example with three forecasts as inputs:
inputs: ecmwf = 22, gfs = 24, icon = 23
weights: 0.5 0.2 0.3 bias: -0.4
22·0.5 + 24·0.2 + 23·0.3 - 0.4 = 11 + 4.8 + 6.9 - 0.4 = 22.3
Activation. Then the neuron passes the result through a simple function. We use ReLU:
negative numbers become 0, positive numbers stay the same (ReLU(3.5) = 3.5, ReLU(-2) = 0).
This small "bend" is what lets a network learn non-linear rules (curves), not only straight lines.
Layers. Neurons are arranged in layers. Every neuron of a layer gets all outputs of the previous layer. Ours:
3 inputs 16 neurons 8 neurons 1 output
(ecmwf, gfs, icon) ───▶ layer 1 ───▶ layer 2 ───▶ tomorrow's max
In scikit-learn this is MLPRegressor(hidden_layer_sizes=(16, 8)). MLP means multi-layer perceptron, the classic neural network.
It has about 200 weights; linear regression has 4.
Training = learning from errors. At the start the weights are random, so the predictions are nonsense. Then training repeats:
- predict all training days,
- measure the error,
- change every weight a tiny bit in the direction that makes the error smaller.
One round over all training days is an iteration (or epoch). It takes hundreds of rounds. max_iter=3000 is the upper limit.
Because the start is random, two trainings can give slightly different results. random_state=0 fixes the random start, so you get the same result every time.
Scaling. Neural networks learn badly if the inputs have very different sizes (for example temperature around 20 and air pressure around 1000). A scaler shifts and stretches each input so that its average is 0 and its typical spread is 1.
Overfitting and regularization. A network with many weights can learn the training days by heart, including their random noise.
Then it is great on the training days and bad on new days. This is overfitting. alpha=0.01 punishes very big weights a little,
which keeps the network simpler. This is regularization.
Honest expectation. In week 5 the linear model was the best: our problem is almost a straight line. A neural network can learn a straight line too, but it needs more data and more care. So do not expect it to win here. Expect it to be about as good as linear regression. That is a typical real-life result: a bigger model is not automatically better. Next week we let the models work together.
Steps
Work in your project folder with the venv active ((.venv) at the start of the line).
Step 1: A neural network on a toy example
Before the real data, let us see a neural net learn a very simple rule: double the number.
Start the interactive Python and type the lines after >>>:
python
>>> from sklearn.neural_network import MLPRegressor
>>> X = [[1], [2], [3], [4], [5], [6]]
>>> y = [2, 4, 6, 8, 10, 12]
>>> net = MLPRegressor(hidden_layer_sizes=(8,), max_iter=5000, random_state=0)
>>> net.fit(X, y)
MLPRegressor(hidden_layer_sizes=(8,), max_iter=5000, random_state=0)
>>> net.predict([[7], [2.5]])
array([13.63308002, 5.1416677 ])
>>> net.n_iter_
357
>>> exit()
What happened:
- X are the inputs (a list of rows, here one number per row), y the answers. The same shape as in week 5.
- hidden_layer_sizes=(8,): one layer with 8 neurons. Always write hidden_layer_sizes=: without the name it is an error in new scikit-learn versions.
- fit trained it (the interactive Python then prints the model with its settings; that is normal). n_iter_ shows it needed 357 rounds.
- For 7 it says 13.6 (correct: 14), for 2.5 it says 5.1 (correct: 5). Close, but not perfect.
It never saw 7, and it had only 6 examples. Linear regression would give exactly 14 here, because the rule is a straight line.
Keep this in mind: it is the same story with our weather data.
Step 2: What the scaler does
Try StandardScaler on three rows of made-up data: temperature and air pressure.
python
>>> from sklearn.preprocessing import StandardScaler
>>> X = [[20, 1010], [25, 1020], [30, 1000]]
>>> scaler = StandardScaler()
>>> scaler.fit_transform(X)
array([[-1.22474487, 0. ],
[ 0. , 1.22474487],
[ 1.22474487, -1.22474487]])
>>> scaler.mean_
array([ 25., 1010.])
>>> exit()
fitlearns the average (25 and 1010) and the spread of each column.transformapplies it.fit_transformdoes both at once.- After scaling, both columns are small numbers around 0. The pressure (around 1000) no longer looks 50 times more important than the temperature.
- The 25 °C day is exactly average, so it becomes 0.
Step 3: A neural network on the real data
Now the real thing, still in the interactive Python, so you see every piece. Use the functions from week 5:
python
>>> from ml import build_dataset, split_by_date, mae
>>> from config import SETS
>>> sources = SETS["models"]
>>> table = build_dataset("eger", sources)
>>> train, test = split_by_date(table)
>>> len(train), len(test)
(779, 195)
Now build the network, with the scaler in front of it:
>>> from sklearn.neural_network import MLPRegressor
>>> from sklearn.pipeline import make_pipeline
>>> from sklearn.preprocessing import StandardScaler
>>> net = make_pipeline(StandardScaler(), MLPRegressor(hidden_layer_sizes=(16, 8), alpha=0.01, max_iter=3000, random_state=0))
>>> net.fit(train[sources], train["measured"])
Pipeline(steps=[('standardscaler', StandardScaler()),
('mlpregressor',
MLPRegressor(alpha=0.01, hidden_layer_sizes=(16, 8),
max_iter=3000, random_state=0))])
>>> mae(net.predict(test[sources]), test["measured"])
0.6916833211567296
About 0.69 °C, very close to linear regression (0.68) and much better than the best single model (ECMWF, 0.82).
Look at the first three test days:
>>> net.predict(test[sources].head(3)).round(1)
array([15.3, 13.8, 11.8])
>>> test.head(3)
ecmwf gfs icon measured
day
2026-03-25 15.4 15.3 16.4 15.4
2026-03-26 14.3 15.7 13.9 15.3
2026-03-27 12.1 10.8 12.7 11.2
>>> exit()
make_pipeline(A, B)glues two steps into one model.fitfirst fits the scaler on the training days, scales them, then trains the network on the scaled data.predictscales the new data with the same numbers and then predicts. So you can never forget the scaling, and you cannot accidentally scale the test days with their own averages.- The pipeline has the same
fit/predictas every other model, so the rest of our code does not need to know it is two steps.
Step 4: ml.py: the imports
Now put it into the program. Open ml.py and add three lines to the imports at the top, after the LinearRegression import:
from sklearn.neural_network import MLPRegressor
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
The top of the file now looks like this:
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import LinearRegression
from sklearn.neural_network import MLPRegressor
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from config import CITIES, SETS
from db import load_table
Step 5: ml.py: a function that makes the network
Add this function before make_models():
def make_neural_net():
# The scaler moves every input to a similar range; neural nets learn badly without it.
return make_pipeline(
StandardScaler(),
MLPRegressor(hidden_layer_sizes=(16, 8), alpha=0.01, max_iter=3000, random_state=0),
)
It is the same network as in step 3, in a function, so we can create a fresh, untrained copy any time we need one.
Try it:
python
>>> from ml import make_neural_net
>>> make_neural_net()
Pipeline(steps=[('standardscaler', StandardScaler()),
('mlpregressor',
MLPRegressor(alpha=0.01, hidden_layer_sizes=(16, 8),
max_iter=3000, random_state=0))])
>>> exit()
You see the two steps of the pipeline and the settings.
Step 6: ml.py: add it to the models
Change make_models() so that it also returns the neural network:
def make_models():
linear = LinearRegression()
forest = RandomForestRegressor(n_estimators=200, min_samples_leaf=5, random_state=0)
neural = make_neural_net()
return {"linear": linear, "random_forest": forest, "neural_net": neural}
evaluate() from week 5 goes through every model in this dictionary, so nothing else needs to change: the new model is trained, tested and printed automatically.
Step 7: Run the comparison
python ml.py
It takes a few seconds longer than last week: that is the neural net learning.
You should see:
Budapest / models: 974 complete days
ecmwf 0.64 °C
gfs 1.18 °C
icon 0.70 °C
average of sources 0.67 °C
linear 0.58 °C <-- best
random_forest 0.73 °C
neural_net 0.59 °C
Budapest / providers: 0 complete days
not enough data yet (need 30)
Eger / models: 974 complete days
ecmwf 0.82 °C
gfs 1.56 °C
icon 1.28 °C
average of sources 0.93 °C
linear 0.68 °C <-- best
random_forest 0.82 °C
neural_net 0.69 °C
Eger / providers: 0 complete days
not enough data yet (need 30)
The neural net is 0.01 °C behind linear regression in both cities, and far better than every single weather model. Exactly the "honest expectation" from the top of the page.
Step 8: Write down the results
Add the neural_net column to your results table from week 5:
| ecmwf | gfs | icon | average | linear | random forest | neural net | |
|---|---|---|---|---|---|---|---|
| Budapest | 0.64 | 1.18 | 0.70 | 0.67 | 0.58 | 0.73 | 0.59 |
| Eger | 0.82 | 1.56 | 1.28 | 0.93 | 0.68 | 0.82 | 0.69 |
You need it for week 7 and for the final presentation.
Your complete ml.py
Compare with yours. If something does not work, copy this one.
ml.py
"""Machine learning: learn how to combine several forecasts into a better one.
Run it directly to compare all models: python ml.py
"""
import numpy as np
from sklearn.ensemble import RandomForestRegressor
from sklearn.linear_model import LinearRegression
from sklearn.neural_network import MLPRegressor
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from config import CITIES, SETS
from db import load_table
MIN_DAYS = 30 # do not train on fewer complete days than this
def build_dataset(city, sources):
"""Complete days only: every forecast AND the measured value must be present."""
return load_table(city, sources).dropna()
def split_by_date(table, test_share=0.2):
"""Older days for training, the newest days for testing. Never shuffle time series!"""
cut = int(len(table) * (1 - test_share))
return table.iloc[:cut], table.iloc[cut:]
def mae(predicted, measured):
"""Mean absolute error: on average, how many degrees we are wrong."""
return float(np.mean(np.abs(np.asarray(predicted) - np.asarray(measured))))
def baseline_scores(test, sources):
"""How good are the forecasts without any machine learning?"""
scores = {source: mae(test[source], test["measured"]) for source in sources}
scores["average of sources"] = mae(test[sources].mean(axis=1), test["measured"])
return scores
def make_neural_net():
# The scaler moves every input to a similar range; neural nets learn badly without it.
return make_pipeline(
StandardScaler(),
MLPRegressor(hidden_layer_sizes=(16, 8), alpha=0.01, max_iter=3000, random_state=0),
)
def make_models():
linear = LinearRegression()
forest = RandomForestRegressor(n_estimators=200, min_samples_leaf=5, random_state=0)
neural = make_neural_net()
return {"linear": linear, "random_forest": forest, "neural_net": neural}
def evaluate(city, data_set):
"""Train every model on the older days, measure the error on the newest days."""
sources = SETS[data_set]
table = build_dataset(city, sources)
if len(table) < MIN_DAYS:
return None, len(table)
train, test = split_by_date(table)
scores = baseline_scores(test, sources)
for name, model in make_models().items():
model.fit(train[sources], train["measured"])
scores[name] = mae(model.predict(test[sources]), test["measured"])
return scores, len(table)
def print_scores(scores):
best = min(scores, key=scores.get)
for name, value in scores.items():
mark = " <-- best" if name == best else ""
print(f" {name:20s} {value:5.2f} °C{mark}")
if __name__ == "__main__":
for city in CITIES:
for data_set in SETS:
scores, days = evaluate(city, data_set)
print(f"\n{CITIES[city]['name']} / {data_set}: {days} complete days")
if scores is None:
print(f" not enough data yet (need {MIN_DAYS})")
else:
print_scores(scores)
Check
- ☐
python ml.pyshows aneural_netline for both cities. It is clearly better than every single model, and close tolinear. - ☐ Change
random_state=0to1,2,3inmake_neural_net()and run again. In Eger you get about 0.67–0.70. (Neural nets depend on their random start. Linear regression always gives the same answer.) Set it back to0.
Extra
- Size experiment. Try
hidden_layer_sizes=(4,),(16, 8),(64, 32),(128, 64, 32). Make a table: size → error. Does a bigger network help? - Without the scaler. Use only
MLPRegressor(...)withoutStandardScaler(). What happens to the error? (Here: almost nothing, because all three inputs are temperatures of similar size. The scaler becomes essential when inputs have very different sizes, e.g. air pressure around 1000 hPa next to temperature around 20 °C.) - Small data. Use only the last 60 days of the history set (
table = table.tail(60)inbuild_dataset). Compare linear regression and the neural net. Which one suffers more from little data? (This is what will happen with your live set in week 8.)
If something goes wrong
| Problem | Reason |
|---|---|
InvalidParameterError: The 'loss' parameter of MLPRegressor must be ... |
You wrote MLPRegressor((16, 8)). Write the name: MLPRegressor(hidden_layer_sizes=(16, 8)). |
ConvergenceWarning: Maximum iterations reached |
The network was still improving when it hit max_iter. It is a warning, not an error. Raise max_iter, or ignore it. |
NameError: name 'make_pipeline' is not defined |
The imports of step 4 are missing. |
| The neural net is much worse than linear (e.g. 2 °C) | Check that you train on train and test on test, and that alpha=0.01 and max_iter=3000 are there. |
ModuleNotFoundError: No module named 'ml' in the interactive Python |
Start python in the project folder, where ml.py is. |