Skip to content
aadya

Wins, ties and falling short.

Each study asks one concrete question, shows the chart, and states the verdict. A result that does not favour aadya-m1 stays on this page.

Win

Irregular and changing cycles

Cycles change with age, stress and life stage. We built a scenario that mixes life stages and asked whether a learned model beats a classical one when the rules shift.

It does, clearly. This is the biggest margin over the Bayes reference we measured. It is on simulated users, so treat it as a stress test rather than a promise about real people.

30%
less error than the usual method on the hardest group
4.78
CRPS against 5.33 for the Bayes reference

Error by simulated scenario

S4 is a mix of life stages (cycles that change character over time).

0246S0Well-specifiedS1Mild driftS2Slow adaptationS3IrregularS4Life-stage mixS5Population blend
28-day baselineBayes referenceaadya-m1
View the numbers
Scenarioaadya-m1Bayes reference28-dayBest
S0 Well-specified2.582.563.44hierarchical-baseline
S1 Mild drift2.712.733.11aadya-m1
S2 Slow adaptation1.811.822.58aadya-m1
S3 Irregular2.992.994.2B8
S4 Life-stage mix4.785.335.35aadya-m1
S5 Population blend2.672.723.59aadya-m1

Where cycles are not stable, the neural model pulls ahead of the classical reference. Where cycles are already well specified (S0), the classical model is a hair ahead.

Hard for everyone

When people forget to log

Real logs are messy. People forget, or log late. We simulated 5%, 10% and 20% of periods going unlogged and re-scored everything.

The headline here is not a win. All methods lose a lot of accuracy, and the best response is in the app, not the model: prompt the person to check for a missed period.

2.99 → 9.55
aadya-m1 error when 20% of period logs are missed
23%
still lower than the usual method at 20% missed

Error as logs go missing

Share of periods that were never logged, simulated users.

048123.73.0All logged6.44.95% missed7.96.410% missed12.49.520% missed
Usual tracker methodaadya-m1
View the numbers
LevelUsual methodaadya-m1Bayes reference
All logged3.732.993.06
5% missed6.364.874.91
10% missed7.856.426.47
20% missed12.459.559.61

Every method gets worse, by a similar factor. Missing data is a data problem, not a model problem. An app should ask “did you forget to log a period?” when a cycle looks too long.

What a long gap looks like

One cycle in this history is 58 days, which could be a real long cycle or a forgotten log.

Probability of each cycle length3%6%9%12%5101520253035404550556065707580859095100105110115120days from the start of the last periodmost likely day 29 · 80% window days 24–38

History: 28, 29, 58, 30, 27 days. A 58-day cycle could be real or a forgotten log, which is why an app should ask.

Tie

364 real users

Real cycles are more regular than a stress test. Two public cohorts give 364 people. Against the rolling-median method many trackers use, aadya-m1 removes 18% of the error.

Against a strong classical model the story is a close race. That is useful in itself: a small learned model matches the best classical reference and adds calibrated uncertainty plus an optional marker.

18%
less error than the usual method on real users
~1%
lead over the Bayes reference, small but clear on the pooled set

Creighton (251) and Marquette (113)

Mean CRPS, lower is better. Creighton is cross-fitted: no user is scored by a model that saw them.

View the numbers
MethodValue95% interval
Creighton: aadya-m12.091-
Creighton: Bayes reference2.119-
Marquette: aadya-m11.572-
Marquette: Bayes reference1.580-

On Marquette the model is level with the Bayes reference (1.572 against 1.580); on Creighton and pooled it is ahead by a small margin. No method beats the reference by a multiple on clean real data, and we measured that this is a noise floor, not a missed opportunity.

Win (experimental)

Adding an ovulation test

Cycle dates alone have a noise floor. New information can lower it, and a positive ovulation test is information many people already collect.

On an independent dataset (mcPHASES, used for evaluation only), adding the test input narrowed the window from about 12 to about 7 days. A personalised version of the marker was not validated, so it is not included.

12 → 7 days
80% window, without and with the test
51%
less error than the usual method (1.70 against 3.49)

One history, with and without the test

Cycles 26, 31, 28, 33, 27, 30 days, 21 days since the last period. Bars: cycle dates only. Line: with the test.

Probability of each cycle length4%8%12%16%202530354045505560days from the start of the last periodmost likely day 29 · 80% window days 25–33 · with test: 25–32

Real aadya-m1 output for one history. The gain varies by case; across the study the average 80% window went from about 12 to about 7 days. It is not shown as an ovulation date.

Error on 110 cycles from 41 people

Mean CRPS with 95% interval. Lower is better.

View the numbers
MethodValue95% interval
28-day constant2.582.12 to 3.09
Last cycle3.562.91 to 4.29
Rolling median3.492.85 to 4.23
Poisson with skips2.342.00 to 2.73
Bayes reference2.442.06 to 2.87
aadya-m12.432.05 to 2.87
aadya-m1 + ovulation test1.701.33 to 2.13

Without the test, aadya-m1 is level with the Bayes reference and slightly behind a Poisson-with-skips model (2.34 against 2.43; significance not tested). With the test it beats every method we ran. This is one small study; treat it as promising, not proven.

On par

A brand-new user

What does it say on day one, before any period is logged? A population prior is the best anyone can do, and we show it as a wide window.

On real users with no history, aadya-m1 ties a plain 28-day baseline whose spread was fitted to data. A cold-start win would be easy to claim and wrong, so we report it as a tie.

2.18
aadya-m1 with no history, Marquette
2.13
a plain 28-day baseline with a fitted spread

No history versus one cycle

Real model output for a user who has logged nothing, and for one who has logged a single 31-day cycle.

No history

Probability of each cycle length1015202530354045505560days from the start of the last periodmost likely day 28 · 80% window days 23–38

One 31-day cycle

Probability of each cycle length1015202530354045505560days from the start of the last periodmost likely day 30 · 80% window days 26–37

With nothing logged, the model returns a broad population-level distribution. One logged cycle already changes it. We do not claim a cold-start win: it is level with a good 28-day baseline.

Diminishing returns

How much history do you need?

How long until a forecast is good? Quickly. A model that waits for a year of data would be of no use to a new user.

The practical lesson is for the product: start forecasting from the first cycle, show a wide window, and let it tighten.

3.29 → 2.91
error with one past cycle against six or more
~8%
real-data gain from 1 to 6+ cycles

Error by history length

Mean CRPS on simulated users.

View the numbers
MethodValue95% interval
1 past cycle3.29-
2 past cycles3.27-
3 to 53.02-
6 or more2.91-

The first cycles matter most. After a handful, extra history adds little, because cycle dates alone carry limited information.

Falls short

A generator it never saw

If a model is trained only on one simulator, how will it do on a different one? We wrote a fresh generator, never used for training or tuning, and re-ran the board.

Slightly behind the two strongest comparisons, and by a small margin. This is why we say the lead on real data is small, and why we keep collecting real data.

+0.012
aadya-m1 minus the Bayes reference (positive = worse)
+0.032
aadya-m1 minus the LSTM, interval [0.009, 0.055]

Simulator S6, never used for training

A different way of generating cycle histories (Markov regimes, Gamma lengths). Mean CRPS with 95% interval.

View the numbers
MethodValue95% interval
LSTM (same data)2.6722.541 to 2.801
Bayes reference2.6922.554 to 2.828
aadya-m12.7042.565 to 2.839
Rolling median2.8802.757 to 3.001
28-day constant3.6043.429 to 3.792

aadya-m1 is clearly ahead of every heuristic here, but slightly behind the Bayes reference and the LSTM. Trained only on synthetic data, it can inherit the quirks of its simulator. We kept this result in.

Did not work

Ideas we tried and did not ship

Progress is also knowing where the ceiling is. With cycle dates alone, the best methods tie within about 1%. Each lever above was tested against that floor.

We publish the list so that nobody repeats the work, and so you can judge how much of a claim is evidence and how much is hope. See the benchmark.

5
levers measured and closed
0
added to the released model

The closed-door list

Every one of these was pre-declared, measured, and rejected by its own rule.

  • Fine-tuning on real data: it did not beat the shipped model, so it was not installed.
  • A learned blend of the neural and classical parts: no gain; the hand-tuned blend stayed.
  • An ensemble with an LSTM: no clear gain on simulated users and significantly worse on Marquette, so no second network ships.
  • A skip-aware layer: a probe showed at most 1%, 3% and 7% gain even if the true skip rate were known, so it was not built.
  • A personalised marker: underpowered on the available data; the public API uses the fixed marker only.