← back to the tape

LINER NOTES · RESEARCH · CNRS LAMIH · 2025–2026

Two tests that lowered my numbers

I model hourly EV-charging load at a CNRS lab and built a simulation that shows it live. Along the way two tests made my results worse on paper and better in fact: one caught a train/test split that leaked time, the other caught a feature that drifted between training and serving.

the data

Public charging-session logs from seven places: Palo Alto, Caltech’s ACN, Perth, SAP, Boulder, Dundee and Paris. Each session has a plug-in time, an unplug time and the energy delivered. I split every session’s kWh across the hours it spanned, in proportion to the minutes it overlapped each hour, then reindexed every site to a full hourly timeline so idle hours appear as zeros instead of missing rows. That gives 261,090 station-hours, about 65% of them with no charging at all.

the model

A hurdle model asks two questions separately. An XGBoost classifier predicts whether a site is charging in a given hour. A second XGBoost regressor with a Tweedie objective, trained only on the non-zero hours, predicts how many kWh. The output is the classifier’s yes/no times the regressor’s magnitude.

Both use 500 trees, depth 6, learning rate 0.05, with early stopping. The 44 features are local time (each region keeps its own timezone), day-of-week and holiday signals, sessions started that hour, lags from 1 hour back to 168 hours, and rolling statistics over those lags.

test one: the split that lied

The first experiment used a 70/15/15 split drawn per site. It scored R² 0.92 on test with MAE 4.1 kWh. We were pleased for about a day.

The problem: a random split of hourly rows puts Tuesday 3 pm in training and Tuesday 4 pm from the same site in test. With lag features reaching back 168 hours, the model is effectively shown the answer. The score measured interpolation, not prediction.

The fix: sort everything by time, hold out the last 15% as a final test set nobody touches, run 5-fold TimeSeriesSplit on the remaining 85% so every fold trains on the past and validates on the future, then retrain on the 85% and score the untouched 15% once.

PROTOCOLTEST R²MAE (kWh)WAPE
Original 70/15/15 per-site split0.9194.1021.5%
Chronological 85% CV + 15% untouched test0.8891.9326.4%

The fold scores tell the story a random split hides: early folds trained on little history sit around R² 0.70, later ones climb past 0.95. MAE fell while R² fell, because the final test window holds calmer sites with smaller loads. Two metrics moving in opposite directions is a reminder to report both, with the data they were measured on.

the simulation

The lab wanted to see the model work, not read a table. So I built a FastAPI service and a MapLibre front end on the IGN basemap of Île-de-France: six charging sites, 71 charge points, and simulated vehicles that drive to a site, plug in, charge, and leave. The cars follow the real OpenStreetMap road graph for Paris, about 4,500 junctions after contraction, routed locally in around 10 ms per trip. Every simulated hour the model predicts each site’s load and is scored live against a persistence baseline.

One label sits in the header on purpose: the model on that machine was trained on synthetic sessions, because the real corpus isn’t there, so the error figures demonstrate the pipeline rather than report results. Saying so first means nobody gets to catch you with it later.

test two: the feature that drifted

Training builds features from each site’s entire history. The simulation builds them from a ring buffer of the last 336 hours. Any feature that isn’t a pure function of a bounded trailing window will disagree between the two, and the model will be served inputs it never saw in training, silently.

$ python3 test_parity.py
Window: offline = full history (23,040 rows), live = last 336h
Comparing 38 features across 6 entities

PASS  37 features identical to 1e-09
FAIL  1 features differ:
   expanding_mean           max abs diff 0.4127
These are not pure functions of a bounded trailing window.

The parity test replays the same data through both paths and asserts every feature matches to 1e-9. It caught expanding_mean: a running average over all history, which a 336-hour buffer cannot reproduce. The live buffer now carries that feature as explicit state (running sum and count) instead of recomputing it, and the test passes on all 38 features. The tolerance is a tripwire, not a budget: identical computations agree far below 1e-9, so anything larger is a different computation.

what I’d tell my past self

— Gaurav, research intern at LAMIH (CNRS UMR 8201), Université Polytechnique Hauts-de-France, with Prof. Alaa Daoud.