Calibration · Kalshi temperature ladders and games

When the price says 20%, does it happen 20% of the time?

A price is a probability with money behind it. If it is honest, the brackets priced near 20c should pay about one time in five, the ones near 60c about three in five, and so on up the scale. This page checks that for every settled bracket of Kalshi's daily-high and daily-low ladders the desk has read, at four moments of the measured day, and for the games at their last price before the start. It counts; it does not recommend.

loading the record…

The price, bin by bin

Each dot is one price band. Across, the band's average price; up, how often its brackets paid. On the dashed line a price pays exactly as often as it says. The whisker is the 95% interval on how often it paid: a band whose whisker crosses the line is in line with its price, and one whose whisker misses it paid more or less often than it was priced, by more than chance usually allows.

when
ladders

As the measured day begins

every settled bracket, priced at its midpoint · dashed line: a price that pays exactly as often as it says
Source: Kalshi's public market data, read hourly by the desk's logger; results are Kalshi's own settlements. Hollow dots have fewer than 30 brackets and are shown, not read.
price bandbracketsladdersaverage pricepaidhow often (95%)no bidagainst the price
loading…

"in line": the band's average price sits inside the 95% interval on how often it paid. "ladders": at most one bracket per ladder can pay, so this is the honest count beside the brackets. "no bid": the share of the band that nobody was bidding on. The intervals treat brackets as independent draws, which they are not, in two ways that pull against each other: the brackets of one ladder are mutually exclusive, which makes an interval a little too wide, and cities share a day's weather, which makes it too narrow. Which wins is not known, so read a band that only just misses as a lean, not a finding.

The score, from the start of the day to its last hour

The Brier score is the average squared gap between the price and what happened, 1 if the bracket paid and 0 if not: 0 is perfect. It means little alone, so it sits beside two forecasters that know nothing. Always No prices every bracket at zero; because one bracket in six pays, it scores about 0.17. Even split prices each bracket at one over the ladder's size. Both follow from the ladder's shape alone, six brackets of which one pays, so they read the same at every moment; only the price's score moves. Most of what a low score rewards is knowing which brackets are hopeless, so the bin table above, not this number, is the test of whether the prices in between are right; the score puts the four moments on one yardstick. A coin's 0.25 is the wrong yardstick for a six-bracket ladder, and is not used.

gradedbracketsmarket daysthe price95% (days resampled)always Noeven split
loading…

The interval resamples whole market days, a thousand times, because brackets that settle on the same day share its weather. The four rows are not four looks at one forecast. The first is the forecast, taken as the measured day begins. Noon is still mostly a forecast for the highs, since the afternoon has not happened, and mostly a reading for the lows, since the overnight low has; the highs and lows buttons show the two apart. At 6 pm and 11 pm most highs and many lows have already been recorded by the station, so those rows mostly measure how fast the price absorbs a reading it can already see. A moment counts only where the logger read it, so the rows are not quite the same ladders: the first row starts a day later than the others.

The games, at the line

The same question for every settled game on the games page: the favourite at its last price before the start, grouped by how strong a favourite it was, against how often it won. For the games the line is the two asks normalised to sum to one at the last read before the start, as the games page defines it, not a midpoint; a game whose book was wider than 25c has no line and is not here, and ties and voids are left out. The sample is young, so the whiskers are wide.

Favourites at the line

every settled game with a line
Source: Kalshi's public game markets, read hourly; the line is the last read before the start; results are Kalshi's settlements.
favourite atgamesaverage linewonhow often (95%)against the line
loading…

How this is measured

The price is the midpoint of the best bid and the best ask, as a probability. It is nobody's trading price: a buyer pays the ask plus Kalshi's fee, which is why the favourite rule on the front page loses money on brackets that mostly pay. A band that paid more often than its midpoint is a measurement about the price, not a trade.

The moments. Each temperature market has a close, Kalshi's end of the measured day: midnight local standard time, the end of the National Weather Service's climate day. So 24 hours before the close is the moment the measured day begins, not the afternoon before; 12 hours before is noon, 6 hours is 6 pm and 1 hour is 11 pm, local standard time. For each market the desk takes its last read at or before each moment, and only if the read falls within the ninety minutes before it; a market the logger did not read in that window is left out rather than filled in. A read is kept once its moment has passed and is never changed.

The bands are fixed in advance: 0–5c, 5–10c, then tens to 90c, then 90–95c and 95–100c, finer at the ends where most of a ladder's brackets sit. The intervals are Wilson 95% intervals per band and a day-resampled interval for the Brier score. Only settled markets count; the record starts with the logger on 2026-09-18.

Where it lives. retention.sql copies each read, its midpoint and the shape of its book, not the quotes, into a permanent table (calib_reads) before its raw row leaves the rolling window, and the views calibration_bins, calibration_scores and calibration_days grade them against the outcomes; the games come from the games_public view. The method covers the logger and its corrections.