Skip to content
aadya

Accuracy that is measured, not asserted.

Every method sees the same users, the same splits and the same scoring rule. Differences come with paired confidence intervals, and a tie is called a tie.

How we measure

One score for the whole forecast.

CRPS (continuous ranked probability score) rewards forecasts that put probability near the real outcome and are neither over- nor under-confident. It is in days, and lower is better.

Proper scoring

CRPS can only be minimised by honest probabilities. A method cannot win by shouting a single date or by hedging everywhere.

Paired intervals

Differences between two methods are bootstrapped per user. If the 95% interval includes zero, we call it a tie.

Calibration too

We also report coverage: an 80% window should hold the real day about 80% of the time, and how wide that window is.

Want to feel the score? Try the interactive demo: slide the real cycle length and watch the error change.

Simulated users

The leaderboard.

539 simulated validation users across six scenarios (seeds 101 to 105, decision point 0, at least one past cycle). Simulated data does not say how a method does on real people, so the real-data tracks are below.

AadyaBench, simulated users

AadyaBench official v2

Mean CRPS, lower is better
View the numbers
MethodValue95% interval
aadya-m12.9252.774 to 3.089
LSTM (same data)2.9652.807 to 3.136
Bayes reference3.0262.863 to 3.200
Poisson with skips3.1092.959 to 3.277
Shrinkage3.5083.328 to 3.701
Rolling median3.7023.509 to 3.911
Exp. smoothing3.7023.480 to 3.942
28-day constant3.7123.529 to 3.913
Robust AR3.7323.523 to 3.952
Rolling mean3.7783.559 to 4.016
Last cycle4.3824.127 to 4.666
#MethodCRPSvs Bayes ref.80% coverWidth
1aadya-m12.925-0.101 [-0.142, -0.065]0.8513.9
2LSTM (same data)2.965-0.061 [-0.100, -0.024]0.8413.2
3Bayes reference3.026reference0.8514.0
4Poisson with skips3.109+0.083 [0.051, 0.115]0.9117.5
5Shrinkage3.508+0.482 [0.418, 0.550]0.8116.7
6Rolling median3.702+0.676 [0.602, 0.757]0.8216.7
7Exp. smoothing3.702+0.676 [0.582, 0.782]0.8116.7
828-day constant3.712+0.686 [0.574, 0.799]0.8216.7
9Robust AR3.732+0.706 [0.633, 0.784]0.8116.7
10Rolling mean3.778+0.752 [0.665, 0.849]0.8016.7
11Last cycle4.382+1.356 [1.239, 1.488]0.7716.7

Simulated users stand in for real people; this board is the least conclusive of the tracks. On it, aadya-m1 has the lowest error of the runnable open methods.

Paired difference against the Bayes reference

Negative is better. Intervals are 95% paired bootstrap.

Bayes reference← betterworse →aadya-m1-0.101LSTM (same data)-0.061Poisson with skips+0.083Shrinkage+0.482Rolling median+0.676Exp. smoothing+0.67628-day constant+0.686Robust AR+0.706Rolling mean+0.752Last cycle+1.356
View the numbers
MethodMean difference95% interval
aadya-m1-0.101-0.142 to -0.065
LSTM (same data)-0.061-0.100 to -0.024
Poisson with skips0.0830.051 to 0.115
Shrinkage0.4820.418 to 0.550
Rolling median0.6760.602 to 0.757
Exp. smoothing0.6760.582 to 0.782
28-day constant0.6860.574 to 0.799
Robust AR0.7060.633 to 0.784
Rolling mean0.7520.665 to 0.849
Last cycle1.3561.239 to 1.488

aadya-m1 and the LSTM both sit below zero; the heuristic baselines are well above it.

By scenario

Where the lead comes from.

The six simulated scenarios stress different things: well-specified cycles, drift, slow adaptation, irregularity, a mix of life stages, and a blend of populations.

Error by scenario

Mean CRPS, lower is better.

0246S0Well-specifiedS1Mild driftS2Slow adaptationS3IrregularS4Life-stage mixS5Population blend
28-day baselineBayes referenceaadya-m1
View the numbers
Scenarioaadya-m1Bayes reference28-dayBest
S0 Well-specified2.582.563.44hierarchical-baseline
S1 Mild drift2.712.733.11aadya-m1
S2 Slow adaptation1.811.822.58aadya-m1
S3 Irregular2.992.994.2B8
S4 Life-stage mix4.785.335.35aadya-m1
S5 Population blend2.672.723.59aadya-m1

The lead comes mainly from the hard scenarios. On a larger board of 5,400 simulated users aadya-m1 is first overall, but slightly behind the Bayes reference on the well-specified (S0) and irregular (S3) scenarios, and clearly ahead on mixed life stages (S4) and the population blend (S5).

Real users

What happens on real data.

Two public cohorts: Creighton (251 users, cross-fitted so no user is scored by a model that saw them) and Marquette (113 users). 364 real people in total.

Real cohorts

Mean CRPS, lower is better.

View the numbers
MethodValue95% interval
Creighton: aadya-m12.091-
Creighton: Bayes ref.2.119-
Creighton: LSTM2.139-
Marquette: aadya-m11.572-
Marquette: Bayes ref.1.580-
Marquette: LSTM1.639-

Against the usual tracker method

Real-data tracks only.

Usual tracker method (rolling median; 28-day for cold start)aadya-m1
View the numbers
TrackDataUsualaadya-m1Lower byAgainst the strongest comparison
Real data: all public users (pooled)364 users2.351.9318%slightly ahead of the Bayes reference (small but clear)
Real data: Creighton (cross-fitted)251 users2.542.0917.6%slightly ahead of the Bayes reference (small but clear)
Real data: Marquette113 users1.951.5719.4%on par with the Bayes reference

On pooled real users the lead over the Bayes reference is about 1% of the error: clear on Creighton and on the pooled 364 users, within noise on Marquette alone.

Every track

The full headline table.

The usual tracker method is a rolling median of the last three cycles with an empirical spread; for cold start, a 28-day baseline.

aadya-m1 against the usual tracker method

Lower is better. Includes the ties and the one case where it is level.

Usual tracker method (rolling median; 28-day for cold start)aadya-m1
View the numbers
TrackDataUsualaadya-m1Lower byAgainst the strongest comparison
Simulated users, all cases539 users3.702.9221%slightly ahead of the best competitor (LSTM)
Hard case: life-stage mix (S4)simulated6.834.7830%better than the Bayes reference
Real data: all public users (pooled)364 users2.351.9318%slightly ahead of the Bayes reference (small but clear)
Real data: Creighton (cross-fitted)251 users2.542.0917.6%slightly ahead of the Bayes reference (small but clear)
Real data: Marquette113 users1.951.5719.4%on par with the Bayes reference
Mid-cycle on real data, no marker (mcPHASES)110 cycles3.492.4330.2%Poisson-with-skips slightly ahead (2.34 vs 2.43; significance not tested)
With an ovulation-test marker (mcPHASES)110 cycles, 41 people3.491.7051.2%better than the best competitor (Poisson-with-skips)
20% of period logs missed (simulated)simulated12.459.5523.3%on par with the Bayes reference
No history, Marquette (cold start)113 users2.132.18-2.3%on par with the 28-day baseline

Does more history help?

Mean CRPS by number of logged past cycles (simulated users).

View the numbers
MethodValue95% interval
1 past cycle3.29-
2 past cycles3.27-
3 to 53.02-
6 or more2.91-

Diminishing returns. On real data, forecasts with six or more past cycles are only about 8% better than with one: cycle dates alone have a noise floor.

Honesty

What we do not claim.

No multiple-times lead

On clean real data, the best methods tie within about 1%. With cycle dates alone there is little room left; bigger gains need new information (like an ovulation test) or far more users.

No comparison to closed apps

We compare to open, runnable methods. We do not claim anything about closed commercial trackers.

Evidence level

Simulated users plus two public cohorts and one small study for the optional ovulation-test input. Not a clinical trial. Not a medical device.

Reproduce or challenge it

The protocol and sources are in the paper. If you think a result is wrong, open an issue on GitHub.