Skip to content
OwnTheLinesOwnTheLines

Building Proprietary Power Rankings

Power ratings put assumptions about team strength on a numerical scale. They help explain why a projection differs from a market, but a disagreement does not establish value. Define the rating units and test predictions on later games before using a rating difference as a forecast.

From Ratings to an Expected Margin

A point-based rating can represent expected neutral-venue scoring margin against a reference team. Its scale must be defined: an arbitrary ranking score cannot simply be subtracted and interpreted as points. Keep opponent strength, roster information, venue, and the forecast timestamp in the record.

A practical research workflow is to collect dated results, choose initial ratings, state how each result updates them, and reserve later games for evaluation. Select features and tuning parameters without using the final test period. Report whether forecasts are means, medians, or another quantity.

Home-field and rest adjustments are hypotheses to estimate. Neither a universal three-point home-field value nor diminishing returns to rest is established here. A non-linear rest function is a candidate model, not a known causal effect. Team-specific estimates can be especially uncertain when based on few games.

A Transparent Hypothetical Update

Use the illustrative rule:

New rating = old rating + K × (observed margin − expected margin).

Assume ratings are measured in points, the old rating is 90, K = 0.05, observed margin is +17, and the earlier expected margin is +14.5. The residual is 2.5 points, the update is 0.125 points, and the new rating is 90.125. K is dimensionless here and is an arbitrary example, not a recommended or historically typical range.

If both opponents are updated, specify whether they receive equal opposite updates and how ratings are centered. Choose a consistent policy rather than adjusting one team opportunistically after seeing its result. Larger K makes this rule react more strongly to the same residual; it does not guarantee better predictions.

Choose Parameters Without Hindsight

Compare candidate K values and context adjustments using chronological training and validation periods. Keep a later test set untouched during selection. Evaluate error and probability calibration where applicable, not just a favorable handful of returns. Roster change or a new rules environment can weaken the relevance of old observations.

Worked Matchup and Its Limits

Assume Team A has a rating of 90 and Team B 80 on the point scale. For this example only, assign A a 3.5-point home adjustment and a 1-point advantage for the assumed rest difference. The projected mean margin is 90 − 80 + 3.5 + 1 = 14.5 points.

A market spread of A -10.5 differs by 4 points from that mean. This is a projection gap, not an expected-value edge. A model must also supply a margin distribution to estimate P(margin > 10.5), plus the offered payout and uncertainty. Many distributions with the same 14.5 mean can give different cover probabilities.

Do a sensitivity check before interpreting the discrepancy. If the home adjustment were 1.5 and the rest effect zero, the mean would become 11.5. Both scenarios are invented assumptions; neither is a measured NFL team effect. Changes in venue or availability should prompt a documented forecast update, not an assertion that the market must be wrong.

Power-Rating Questions

Use Building Your Own Model for validation and Closing Line Value for a price benchmark with explicit limitations.

Q: How often should I update a rating?

A: Define the update schedule in advance and test it. Updating after each completed game is one possible policy, but more frequent changes do not automatically improve predictions.

Q: How should I choose K?

A: Tune it on chronological training and validation data, then evaluate on a separate later sample. The illustrative K = 0.05 is not a universal optimum.

Q: Does a 14.5-point mean establish value at -10.5?

A: No. Estimate the margin distribution, cover probability, offered payout, and uncertainty. The four-point projection gap alone is insufficient.

Q: Must rest have diminishing returns?

A: No universal rest function is established here. Compare explicit candidate assumptions on relevant data, including uncertainty and possible confounding factors.