✦ Astrology Academy

Backtest: measure your chart against events that already happened

Most astrology tools only look forward, and nothing checks them. This one runs backward. You enter real events with a date and a life-area category, and it compares the score your chart produced for that area on that day against other days inside the same Firdaria major period. Then it reports how often the score cleared the high-signal threshold, what the average was, and whether that average is distinguishable from chance.

Written by: AstroStarCheck Astrological Research Team · Review: Internal editorial review · Last Revised: 2026-08-19
Positions computed with the Swiss Ephemeris

Written and reviewed in house by the AstroStarCheck research team. Astronomical positions come from the Swiss Ephemeris; interpretive rules are checked against the classical texts listed under Sources.

What gets measured, and against what baseline

Each event carries a date and one of twelve life areas, such as partnership, vocation or home. The engine rebuilds the day forecast for that date and reads the score for the matching area.

The comparison baseline is not every other day of your life. It is the other days inside the same Firdaria major period. Periods differ in baseline level, so pooling them would let a naturally high period pass itself off as accuracy.

A day counts as a hit when its score reaches the 70th percentile of that period or higher. 70 is a convention, not a finding. The underlying percentile is always shown so you can move the line yourself.

The three numbers that come back

Mean percentile is the average score across your events. Lift is that average minus 50, the midpoint of the period. Lift near zero means your events landed on ordinary days.

Hit rate at 70 is the share of events that cleared the threshold. It answers a different question than the mean: a handful of very high days can pull the mean up while most events sit flat.

Both appear together because either one alone is easy to misread.

Why two p-values appear

One is a one-sample t-test asking whether the mean percentile differs from 50. The other is a binomial test asking whether the hit rate exceeds 30 percent, the share you would expect if events scattered at random across the percentile scale.

Neither p-value proves astrology works. A small one says the pattern is unlikely under one specific null model, and that model assumes you chose your events without looking at the chart first. If you picked dates because they felt astrologically dramatic, the null model is already broken.

Below five events the page returns no p-value at all. Printing one there would dress noise up as evidence.

Sample size decides what you may conclude

Under twenty events: exploratory. Read it as a description of the days you happened to record, nothing more.

Twenty to forty-nine: preliminary. Patterns can be noticed here and then tested against events you add later.

Fifty or more: usable for calibration. At that point a stable hit rate can inform how much weight you give the timing layer.

The tier is printed on every result. It is the single line most worth reading.

How to record events so the test stays honest

Write the event down before you look at any score. Recording after you have seen the chart is the fastest way to manufacture agreement.

Use dates you actually remember. An event rounded to the nearest month cannot test a day-level score.

Choose the category you would have chosen at the time. Reclassifying an event so it lands in whatever area scored highest is the same error as moving the date.

Dull events matter as much as dramatic ones. A list of nothing but peaks says nothing about whether the score separates peaks from ordinary days.

Interpretive limits

Astrology is a symbolic interpretive tradition, and it has not been shown to predict specific events reliably under controlled conditions. A favourable backtest says your recorded events lined up with one scoring model on one set of days. That establishes neither causation nor a basis for decisions about health, money or law. Treat the output as a reason to look more closely at how you remember those periods.

Frequently asked questions

How many events do I need?

Below twenty the result is exploratory and is not worth quoting to anyone. Around thirty a stable picture usually starts to form. Fifty or more is where the hit rate becomes something you can reasonably weigh.

Why compare within one Firdaria period?

Major periods differ in baseline level. Comparing across periods would credit the period itself for the result. Within-period comparison asks the narrower and more useful question: on this kind of day, did the event stand out from its neighbours?

Can I use this to predict what happens next?

Not directly. The backtest says whether the scoring model separated your recorded events from ordinary days in the past. Carrying that into a forecast is a separate step, and the reliability tier from your own sample is the best guide to how much weight it deserves.

Where are my birth data and events stored?

On this site's own infrastructure. Birth data and event records are processed on our servers to run the calculation and are not sent to third parties. AI interpretation and PDF export are the two features that hand text to an external model, and both are optional.

My hit rate is low. Does that mean the chart is wrong?

It means the scoring model did not separate your recorded events from ordinary days. Common causes are a small sample, vague dates, categories assigned after the fact, or a genuine mismatch between the model and your history. Rule out the first three before drawing any conclusion about the fourth.

Run a backtest on your own events

Related on this site

Sources and further reading

Astronomical positions and interpretive rules on this page are checked against these works.