What a Base Rate Is
— Why "a 70% Hit Rate" Is a Meaningless Number
You often see a performance claim that just prints a big hit rate. But without a comparison baseline, that number proves nothing. We explain what a base rate is, and why it's the starting point of every signal validation, using real measured data.
What a Base Rate Is
A base rate is the average rate at which an event happens with no specific condition applied. Translated to a stock scanner, it becomes this:
"If you held the entire universe of eligible stocks without looking at any signal at all, what percent of them reach +15% within 20 trading days?"
That value is the base rate. And for a signal's hit rate to mean anything, it needs to be clearly higher than this base rate.
Why Looking at the Hit Rate Alone Is Misleading
The same "70% hit rate" can mean something completely different depending on the base rate.
| Situation | Signal Hit Rate | Base Rate | Lift | Interpretation |
|---|---|---|---|---|
| Bear market | 70% | 20% | +50%p | An excellent signal |
| Bull market | 70% | 68% | +2%p | Effectively no contribution |
| Overheated market | 70% | 75% | −5%p | Actually a net loss |
All three cases could equally be marketed with the exact same phrase, "70% hit rate." This is exactly why a hit rate published without its base rate is dangerous. In the worst case, you could be bragging about a result that's actually below the base rate.
A Real Example — the DawnScan Universe
DawnScan actually measures this baseline and publishes it on the base-rate proof page. The measurement standard is as follows.
- Denominator — not a set of pre-selected candidates, but the entire universe of stocks that pass the liquidity filter
- Hit determination — whether the high within 20 trading days reaches +15% versus the entry candle's close
- Deduplication — if the same stock triggers repeatedly, count it only 1x per episode, using a 21-day window
- Delisted stocks included — excluding them would inflate the result (survivorship bias), so they stay in the sample
Measured this way, the recent-sample universe base rate was in the low 30%s. In other words, even holding stocks completely at random, with no signal at all, roughly 3 out of every 10 stocks hit +15% within 20 days.
3 Common Mistakes When Comparing to a Base Rate
① Computing the Base Rate From Candidates Alone
If you compute the base rate using only the stocks a scanner has already filtered down to, your denominator is already a set of good stocks. That either inflates the base rate and understates the signal's lift, or mixes the selection logic into the denominator, creating circular logic. The denominator must always be the full universe, before any filtering.
② Counting Only Survivors
Dropping delisted or trading-halted stocks from your sample creates survivorship bias. A sample with the failures removed looks better than reality. That's why cases with a bad outcome need to stay in the sample all the way through.
③ Locking In a Conclusion From a Single Market Regime's Sample
If you conclude "this signal is valid" using data gathered only during a bull market, it can fall apart in a downturn. Because the base rate itself shifts dramatically with the market regime, you need to split the data by regime (cohort) and look at each one separately.
The Uncomfortable Fact the Base Rate Reveals
Once you measure the base rate properly, it becomes clear that many "famous signals" don't differ much from the baseline at all. That isn't a disappointing result — it's a normal one. Most technical signals are already widely known, and known information tends to already be priced in.
So what matters isn't "the signal was right" — it's "how much better than the baseline, and how consistently." Measuring that difference is the whole point of the base-rate proof page.