Read it, cite it, challenge it.
An open benchmark with proper scoring and paired confidence intervals, an open calibrated probability model, honest negative results, and an optional validated marker.
The contribution
AadyaBench, a protocol for next-period forecasting scored with CRPS and paired bootstrap intervals; aadya-m1, an open model that outputs a calibrated probability for each cycle length; and a record of what did not work.
Evaluation data
Simulated users from AadyaBench; the Marquette FedCycleData cohort; the Creighton Model cycle-length dataset (Stanford and Najmabadi, Hive, CC BY-NC 4.0); and the mcPHASES dataset from PhysioNet for the ovulation-test study. Restricted data stays on the evaluator’s machine and only aggregates are reported.
Reproducibility
Fixed seeds, fixed splits, and a published protocol. Every number on this site comes from a committed benchmark report.
Citation.
@misc{dutta2026aadyam1,
title = {aadya-m1: An Open, On-Device Probabilistic Model for Calibrated Next-Period Forecasting},
author = {Dutta, Manas},
year = {2026},
doi = {10.5281/zenodo.23236800},
url = {https://doi.org/10.5281/zenodo.23236800}
}The preprint is archived on Zenodo with DOI 10.5281/zenodo.23236800.
Beat us on the benchmark.
Every method gets the same users and the same scoring rule. Negative results are welcome, and rows are never removed for scoring badly.
A submission is a class with fit(train, calib) and forecast(past_lengths, elapsed, cohort), returning a probability for each cycle length. To propose a method, a dataset or a correction, open an issue on GitHub.
We are also looking for partners with larger real-world cycle datasets who can evaluate on-site without moving data. That is the most valuable thing that could change the real-data results.