Building Your Own Forecasting Model
A useful forecasting model makes its assumptions inspectable. The work begins with defining an outcome, collecting information available before that outcome, and testing predictions on later data. A disagreement with a price is a research question; it does not establish a betting advantage.
Collect data without hindsight
Record results, competition, season, venue, availability reports, and the price and timestamp at which a selection could actually have been made. Keep a data dictionary for every feature. End-of-season ratings, corrected injury reports, or a closing price published after your decision cannot be treated as information you had earlier.
Start with an interpretable baseline such as logistic regression for a binary outcome. Choose features with a stated rationale: opponent-adjusted performance, venue, rest, and confirmed availability are candidates to test. More features do not automatically improve forecasts. Compare each addition on held-out data, including missing-data handling and the cost of collecting it.
Backtest and check calibration
An illustrative walk-forward schedule trains on 2019–2021 and tests on 2022, then trains through 2022 and tests on 2023. Reserve a final later period that was not used to choose features, parameters, or selection thresholds. Repeatedly choosing the best historical strategy can overfit the evaluation itself.
Calibration asks whether forecasts near 70% occur about 70% of the time across a suitable sample. Show the number of predictions in each bin and uncertainty around observed frequencies. Use probability scores, calibration plots, and realized returns as distinct diagnostics. A negative return does not identify its cause: sampling noise, poor probabilities, prices, and costs can all matter.
The scikit-learn calibration documentation describes reliability diagrams and calibration methods. A calibrator needs data separate from model fitting; it can introduce its own estimation error. Do not assume every model is overconfident or that recalibration guarantees improvement.
Probability points, expected return, and Kelly
Hypothetical example: assume p = 0.60 and an exact offered break-even probability q = 0.52. The gap is 8 percentage points. Decimal return d = 1/q = 1.9230769; net win payout b = d − 1 = 0.9230769 per unit staked. Expected net return is p × b − (1 − p) = p × d − 1 = 0.1538462, or 15.3846% of stake. That is different from the probability-point gap.
For a binary win/loss wager, full Kelly is f = (b × p − (1 − p))/b. Here f = 1/6, or 16.6667% of bankroll. Quarter Kelly would be 4.1667%. These are model illustrations, not stake recommendations: Kelly assumes accurate probabilities, repeatable opportunities, divisibility, and the stated payout without extra costs. Correlated simultaneous wagers need a joint model. A negative fraction means no positive stake under this single-wager rule.
There is no universal 3–5% selection threshold. Set any research threshold using uncertainty, costs, and validation outside the tuning sample. CLV compares entry and closing prices; it cannot prove expected value, profitability, or model correctness. Keep the probability forecast and its uncertainty alongside that benchmark.
Continue reading: Statistical Variance · Kelly Criterion.
Frequently Asked Questions
Does 60% versus 52% mean an 8% expected return?
No. It is an 8 percentage-point probability gap. With exact decimal odds 1/0.52 and an assumed 60% win probability, expected net return is 15.3846% of stake.
Can CLV validate the model by itself?
No. CLV is a price benchmark. Evaluate forecasts on held-out outcomes, calibration, uncertainty, executable prices, and costs; a closing-price comparison alone proves none of those.
Should every model be recalibrated?
No. Measure calibration on separate data first. Calibration methods can help some models but also require validation and enough relevant observations.