🏠 Home 🔍 Today's Scan 📰 Daily Report 📈 Hit Rate 📊 Base Rate ❓ Methodology 📚 Learn 🪙 Crypto Scanner 📋 All Tools ⏪ Investment Simulator 🧾 Tax Calculator🧮 Pension vs. Direct ⚡ Leverage ⚖️ Rebalancing 💹 DCA 📉 Averaging-Down B/E 💰 Dividend Calendar
← Back to the Learn hub
📐 Statistics · Decision Criteria

Lift and Confidence Intervals
— How to Tell If a Signal Is Real

Which is more trustworthy: "60% hit rate, a sample of 20" or "42% hit rate, a sample of 1,400"? Understanding lift and confidence intervals lets you answer this. We break it down with no formulas.

Written by Dawn · IT Engineer · Published
💡 Key takeawayLift tells you how much better a signal is than the baseline; a confidence interval tells you whether that number is just sample luck. Without both, a hit rate can't be interpreted.

Lift — How Much Better Than the Baseline

Lift is simple.

Lift = Hit rate with the signal − the base rate

If the base rate is 30% and a signal's hit rate is 42%, the lift is +12%p. That +12%p is the signal's actual contribution — not the raw number 42% itself.

If the lift is close to 0, that signal is giving you essentially no information. If it's negative, you'd actually be better off without it.

Confidence Interval — Can You Trust That Number?

The problem is that with a small sample, lift can come out large from luck alone.

If you flip a coin 10 times and get heads 7 times, you don't conclude "this coin lands heads 70% of the time." With only 10 flips, that kind of spread is common. A hit rate works exactly the same way.

A confidence interval tells you "accounting for sampling error, the true value is likely somewhere in this range." The smaller the sample, the wider the range.

ObservationHit Rate95% CI (approx.)Verdict
12 hits out of a 20-stock sample60%39% ~ 78%Could overlap with the 30% base rate → hold judgment
588 hits out of a 1,400-stock sample42%39% ~ 45%Clearly beats the 30% base rate

Looking at the hit rate alone, 60% looks better — but the trustworthy one is 42%. A 60% from a 20-stock sample wouldn't be strange to see land at 35% over the next 20.

DawnScan's Decision Criteria

The base-rate proof page displays these three things together.

  • Sample size — under 100 gets flagged as "insufficient sample" and grayed out. It's shown for reference only and isn't used as grounds for a judgment.
  • 95% CI lower bound — a conservative lower bound computed using the Wilson method. This value has to be above the base rate for it to be treated as more than chance.
  • ✓BH — shows whether it also passed the multiple-testing correction.

All three conditions have to pass for something to be treated as a "significant signal." Passing just one makes it only a candidate.

⚠️ Why We Use the CI Lower Bound — Using the lower bound instead of the median (the hit rate itself) is a conservative choice. It means deciding based on "how much, even in the worst case," not "how much, if things go well."

How Large a Sample Is Enough?

There's no single right answer, but you can get a feel for it. To reliably distinguish a difference of roughly +10%p at a 30% base-rate level, you generally need a sample in the hundreds.

That's why DawnScan sets its minimum bar at 100, and even that isn't used to mean "enough" — it's used as the line meaning "you at least need to clear this before the discussion even starts." Once you also apply deduplication of repeat observations on top of that, the actual effective sample shrinks further.

📌 Summary — When you look at a hit rate, check three numbers as a set. ① the base rate (what to compare against) ② the lift (the difference) ③ the sample size and confidence interval (whether it's trustworthy). A hit rate missing even one of these can't be interpreted.
📮 Daily US Market Morning Brief — We send an analysis of the previous day's top 10 US gainers (TOP10) and what they had in common, every day at 8am (KST). Telegram @dawnbrief · Free · No ads · Not stock recommendations.

Frequently Asked Questions

What is lift?

It's the hit rate with the signal applied, minus the base rate. In a market with a 30% base rate, if the hit rate is 42%, the lift is +12%p, and that 12%p is the signal's actual contribution. If the lift is close to 0, the signal gives essentially no information, and if it's negative, you'd be better off without it.

Why is a confidence interval necessary?

Because a small sample can produce a high hit rate from pure chance alone. Flipping a coin 10 times and getting heads 7 times doesn't make it a 70% coin. A confidence interval gives you a range that accounts for sampling error, and the smaller the sample, the wider that range gets.

Which is better: a 60% hit rate (20 stocks) or 42% (1,400 stocks)?

The 42% side is the trustworthy one. The 60% from a 20-stock sample has a roughly 39–78% 95% confidence interval — very wide, with room to overlap the 30% base rate. The 42% from a 1,400-stock sample, by contrast, has a narrow 39–45% interval that clearly beats the base rate. You have to look at the interval and the sample size together, not the raw number.

How large a sample is enough?

There's no fixed standard, but reliably distinguishing a +10%p difference at a 30% base-rate level generally takes a sample in the hundreds. DawnScan sets 100 as the minimum bar and flags anything below that as "insufficient sample." 100 doesn't mean enough — it's the minimum bar for starting the discussion at all.

Related Reading