← notebook

What 671 days of sleep logs taught me

2026-08-19

The engine inside this app was fitted and validated on 671 days of one child's real sleep — 2,181 sessions, month 1 to month 22.

My partner and I logged every one of them. Wake, down, up, night waking, for the better part of two years, mostly at the hours when writing anything down is the last thing a person wants to do. What you are left with, eventually, is a CSV nobody else has and some firm opinions about what a sleep tracker owes you in return.

Before I let that engine predict anything for anyone, I made it earn the right. This is the scorecard, including the part where it loses, because a scorecard without that part is marketing.

The physics is real

Sleep pressure — Process S in the literature — rises while a child is awake and drains while they sleep. I fitted its build-up rate month by month, and it grew about 3.5× between month 1 and month 22: exactly the direction the developmental literature predicts, arrived at from the logs alone. Then I let the fitted model free-run a whole day knowing nothing but the morning wake time. It reproduced the actual nap count in 18 of 22 months, and all four misses landed inside the child's real nap-transition windows — the weeks when he himself couldn't decide between two naps and one.

That is the good half. Here is the other one.

A dumb average beats my physics

On raw prediction error, a personal baseline with no biology in it whatsoever — this nap usually starts around when it started this week — wins. Roughly 17.5 minutes median error against 21–23 for the model.

I could have left that sentence out. Nobody was going to audit me.

The logged schedule is largely the schedule my partner and I enacted. Imitating the family is close to unbeatable by construction.

But the number stops being embarrassing the moment you understand what it measures, and understanding it is the entire point. A rolling average is an excellent mirror — it is very good at telling me what we already did. It cannot do a single thing I actually need: a live pressure gauge between naps, a sensible answer on a strange day, cold-start guidance for a child with no history, or a principled way to say this timing drifted from the biology rather than this timing is unusual for you. The average is better at predicting the family. The physics is better at knowing what is going on.

I would rather ship the second one and tell you it loses the first race.

Late costs more than early

Days where a nap started more than 20 minutes past the model's predicted window showed more fragmented nights in every age band, and shorter naps in most.

Early costs little. Late compounds. That asymmetry now runs through the whole app — the scoring, the guidance corridors, the direction I nudge — and it is the most useful thing 671 days gave me.

What I would actually trust