Accuracy that is measured, not asserted.
Every method sees the same users, the same splits and the same scoring rule. Differences come with paired confidence intervals, and a tie is called a tie.
One score for the whole forecast.
CRPS (continuous ranked probability score) rewards forecasts that put probability near the real outcome and are neither over- nor under-confident. It is in days, and lower is better.
Proper scoring
CRPS can only be minimised by honest probabilities. A method cannot win by shouting a single date or by hedging everywhere.
Paired intervals
Differences between two methods are bootstrapped per user. If the 95% interval includes zero, we call it a tie.
Calibration too
We also report coverage: an 80% window should hold the real day about 80% of the time, and how wide that window is.
Want to feel the score? Try the interactive demo: slide the real cycle length and watch the error change.
The leaderboard.
539 simulated validation users across six scenarios (seeds 101 to 105, decision point 0, at least one past cycle). Simulated data does not say how a method does on real people, so the real-data tracks are below.
AadyaBench, simulated users
AadyaBench official v2
View the numbers
| Method | Value | 95% interval |
|---|---|---|
| aadya-m1 | 2.925 | 2.774 to 3.089 |
| LSTM (same data) | 2.965 | 2.807 to 3.136 |
| Bayes reference | 3.026 | 2.863 to 3.200 |
| Poisson with skips | 3.109 | 2.959 to 3.277 |
| Shrinkage | 3.508 | 3.328 to 3.701 |
| Rolling median | 3.702 | 3.509 to 3.911 |
| Exp. smoothing | 3.702 | 3.480 to 3.942 |
| 28-day constant | 3.712 | 3.529 to 3.913 |
| Robust AR | 3.732 | 3.523 to 3.952 |
| Rolling mean | 3.778 | 3.559 to 4.016 |
| Last cycle | 4.382 | 4.127 to 4.666 |
| # | Method | CRPS | vs Bayes ref. | 80% cover | Width |
|---|---|---|---|---|---|
| 1 | aadya-m1 | 2.925 | -0.101 [-0.142, -0.065] | 0.85 | 13.9 |
| 2 | LSTM (same data) | 2.965 | -0.061 [-0.100, -0.024] | 0.84 | 13.2 |
| 3 | Bayes reference | 3.026 | reference | 0.85 | 14.0 |
| 4 | Poisson with skips | 3.109 | +0.083 [0.051, 0.115] | 0.91 | 17.5 |
| 5 | Shrinkage | 3.508 | +0.482 [0.418, 0.550] | 0.81 | 16.7 |
| 6 | Rolling median | 3.702 | +0.676 [0.602, 0.757] | 0.82 | 16.7 |
| 7 | Exp. smoothing | 3.702 | +0.676 [0.582, 0.782] | 0.81 | 16.7 |
| 8 | 28-day constant | 3.712 | +0.686 [0.574, 0.799] | 0.82 | 16.7 |
| 9 | Robust AR | 3.732 | +0.706 [0.633, 0.784] | 0.81 | 16.7 |
| 10 | Rolling mean | 3.778 | +0.752 [0.665, 0.849] | 0.80 | 16.7 |
| 11 | Last cycle | 4.382 | +1.356 [1.239, 1.488] | 0.77 | 16.7 |
Simulated users stand in for real people; this board is the least conclusive of the tracks. On it, aadya-m1 has the lowest error of the runnable open methods.
Paired difference against the Bayes reference
Negative is better. Intervals are 95% paired bootstrap.
View the numbers
| Method | Mean difference | 95% interval |
|---|---|---|
| aadya-m1 | -0.101 | -0.142 to -0.065 |
| LSTM (same data) | -0.061 | -0.100 to -0.024 |
| Poisson with skips | 0.083 | 0.051 to 0.115 |
| Shrinkage | 0.482 | 0.418 to 0.550 |
| Rolling median | 0.676 | 0.602 to 0.757 |
| Exp. smoothing | 0.676 | 0.582 to 0.782 |
| 28-day constant | 0.686 | 0.574 to 0.799 |
| Robust AR | 0.706 | 0.633 to 0.784 |
| Rolling mean | 0.752 | 0.665 to 0.849 |
| Last cycle | 1.356 | 1.239 to 1.488 |
aadya-m1 and the LSTM both sit below zero; the heuristic baselines are well above it.
Where the lead comes from.
The six simulated scenarios stress different things: well-specified cycles, drift, slow adaptation, irregularity, a mix of life stages, and a blend of populations.
Error by scenario
Mean CRPS, lower is better.
View the numbers
| Scenario | aadya-m1 | Bayes reference | 28-day | Best |
|---|---|---|---|---|
| S0 Well-specified | 2.58 | 2.56 | 3.44 | hierarchical-baseline |
| S1 Mild drift | 2.71 | 2.73 | 3.11 | aadya-m1 |
| S2 Slow adaptation | 1.81 | 1.82 | 2.58 | aadya-m1 |
| S3 Irregular | 2.99 | 2.99 | 4.2 | B8 |
| S4 Life-stage mix | 4.78 | 5.33 | 5.35 | aadya-m1 |
| S5 Population blend | 2.67 | 2.72 | 3.59 | aadya-m1 |
The lead comes mainly from the hard scenarios. On a larger board of 5,400 simulated users aadya-m1 is first overall, but slightly behind the Bayes reference on the well-specified (S0) and irregular (S3) scenarios, and clearly ahead on mixed life stages (S4) and the population blend (S5).
What happens on real data.
Two public cohorts: Creighton (251 users, cross-fitted so no user is scored by a model that saw them) and Marquette (113 users). 364 real people in total.
Real cohorts
Mean CRPS, lower is better.
View the numbers
| Method | Value | 95% interval |
|---|---|---|
| Creighton: aadya-m1 | 2.091 | - |
| Creighton: Bayes ref. | 2.119 | - |
| Creighton: LSTM | 2.139 | - |
| Marquette: aadya-m1 | 1.572 | - |
| Marquette: Bayes ref. | 1.580 | - |
| Marquette: LSTM | 1.639 | - |
Against the usual tracker method
Real-data tracks only.
View the numbers
| Track | Data | Usual | aadya-m1 | Lower by | Against the strongest comparison |
|---|---|---|---|---|---|
| Real data: all public users (pooled) | 364 users | 2.35 | 1.93 | 18% | slightly ahead of the Bayes reference (small but clear) |
| Real data: Creighton (cross-fitted) | 251 users | 2.54 | 2.09 | 17.6% | slightly ahead of the Bayes reference (small but clear) |
| Real data: Marquette | 113 users | 1.95 | 1.57 | 19.4% | on par with the Bayes reference |
On pooled real users the lead over the Bayes reference is about 1% of the error: clear on Creighton and on the pooled 364 users, within noise on Marquette alone.
The full headline table.
The usual tracker method is a rolling median of the last three cycles with an empirical spread; for cold start, a 28-day baseline.
aadya-m1 against the usual tracker method
Lower is better. Includes the ties and the one case where it is level.
View the numbers
| Track | Data | Usual | aadya-m1 | Lower by | Against the strongest comparison |
|---|---|---|---|---|---|
| Simulated users, all cases | 539 users | 3.70 | 2.92 | 21% | slightly ahead of the best competitor (LSTM) |
| Hard case: life-stage mix (S4) | simulated | 6.83 | 4.78 | 30% | better than the Bayes reference |
| Real data: all public users (pooled) | 364 users | 2.35 | 1.93 | 18% | slightly ahead of the Bayes reference (small but clear) |
| Real data: Creighton (cross-fitted) | 251 users | 2.54 | 2.09 | 17.6% | slightly ahead of the Bayes reference (small but clear) |
| Real data: Marquette | 113 users | 1.95 | 1.57 | 19.4% | on par with the Bayes reference |
| Mid-cycle on real data, no marker (mcPHASES) | 110 cycles | 3.49 | 2.43 | 30.2% | Poisson-with-skips slightly ahead (2.34 vs 2.43; significance not tested) |
| With an ovulation-test marker (mcPHASES) | 110 cycles, 41 people | 3.49 | 1.70 | 51.2% | better than the best competitor (Poisson-with-skips) |
| 20% of period logs missed (simulated) | simulated | 12.45 | 9.55 | 23.3% | on par with the Bayes reference |
| No history, Marquette (cold start) | 113 users | 2.13 | 2.18 | -2.3% | on par with the 28-day baseline |
Does more history help?
Mean CRPS by number of logged past cycles (simulated users).
View the numbers
| Method | Value | 95% interval |
|---|---|---|
| 1 past cycle | 3.29 | - |
| 2 past cycles | 3.27 | - |
| 3 to 5 | 3.02 | - |
| 6 or more | 2.91 | - |
Diminishing returns. On real data, forecasts with six or more past cycles are only about 8% better than with one: cycle dates alone have a noise floor.
What we do not claim.
No multiple-times lead
On clean real data, the best methods tie within about 1%. With cycle dates alone there is little room left; bigger gains need new information (like an ovulation test) or far more users.
No comparison to closed apps
We compare to open, runnable methods. We do not claim anything about closed commercial trackers.
Evidence level
Simulated users plus two public cohorts and one small study for the optional ovulation-test input. Not a clinical trial. Not a medical device.