Every probability the desk publishes is computed from public National Weather Service data by a fixed, documented pipeline — and every published call is scored after settlement, in public. This page is the pipeline, end to end.
The desk holds no private data and no inside information. Three classes of government-published weather data feed the model, each used for a different job:
The National Blend of Models (NBM) is NOAA's flagship statistical blend of dozens of weather models. Its probabilistic text bulletins (NBP) publish not just an expected high but a full spread — percentiles of tomorrow's temperature distribution per station. New guidance is produced four times a day (01, 07, 13 and 19 UTC), and the desk re-reads each cycle shortly after it lands.
Each market settles against a specific airport station (Central Park for New York, Midway for Chicago, and so on). The desk polls that station's METAR observations through the day, including the 6-hour maximum and minimum temperature groups embedded in synoptic-hour reports — which catch short spikes that hourly readings can miss.
The official daily climate report (CLI) that NWS offices issue each morning states yesterday's final maximum and minimum temperature. This is the same value Kalshi's weather markets settle on — so the desk grades itself against the identical source of truth the market uses, not its own opinion of what happened.
The other half of every gap is the market itself: Kalshi's public quote for each contract — best bid, best ask, last trade, volume and open interest — snapshotted with a timestamp at the moment of comparison. Market data is read-only and public; the desk holds no trading account and places no orders.
A Kalshi contract like NYC high ≥ 90°F is a yes/no question about an integer. The desk answers it by fitting a smooth distribution to the NBM's published percentiles for that station and date, then reading the tail probability above the strike — with a continuity correction, because settlement is a whole number: "≥ 90" means the rounded official high reaches 90, so the model integrates from 89.5.
Bracket contracts ("high between 85–86°F") are the same computation on both edges: the probability mass that falls inside the bracket after rounding.
What the model knows depends on the time of day, so every read is tagged with the regime it was computed under:
Day-ahead. Built purely from NBM guidance — the widest uncertainty, and the regime where the desk is most conservative about firing.
Settlement day. The forecast distribution is conditioned on what the station has already recorded — a morning running cool caps how high the afternoon can plausibly reach. Intraday claims must rest on a recent observation; stale data disqualifies the read.
The station has already touched the strike. Once an observed max prints at or above the threshold, "≥" is effectively decided — newer observations can only raise a maximum, never lower it. These are the desk's highest-confidence reads.
Day-ahead (forecast-regime) signals carry a Provisional label until that city's own settled track record earns full status — roughly: enough settled calls (30+), a strong standalone Brier score, and a head-to-head win over the market price on the same markets. Until then they are published, labeled, and graded — never silently hidden, never oversold.
A raw difference between the model's probability and the market's price is not a signal — it might be noise, an illiquid book, or a spread wide enough to swallow the whole edge. Every candidate must clear all of these guards at once:
Suppressed candidates aren't discarded — they're recorded with the reason they didn't fire, so the desk's restraint is auditable too.
Every number in a rendered brief line is checked against the underlying fetched record. If a figure can't be traced to a real fetched value — an NWS bulletin, a station observation, a Kalshi snapshot with its timestamp — the line is blocked. No unsourced numbers, ever.
The full text is scanned against a banned-language list: no "buy", no "sell", no "guaranteed", no "can't lose", no advice framing, no hype. The desk describes a divergence; it never recommends a position. Lines that fail are blocked, not softened.
Only what clears both gates is published — to the member desk, the daily brief, and the public scorecard.
The morning after a market settles, the desk pulls the official CLI climate report and records the outcome. Scoring rules are fixed in advance:
One prediction of record per market. For each contract and date, the read that counts is the forecast-regime probability subscribers actually saw in the morning brief — the desk can't quietly re-time its predictions after watching the day unfold.
Brier score, head to head. Both the desk's probability and the market's price at the same moment are scored against the outcome with the Brier score (mean squared error of probabilities — lower is better, 0.25 is coin-flip guessing). The comparison uses the same settled markets for both sides. If the desk's Brier isn't lower than the market's, the desk added nothing over just reading the price — and the public scorecard will say so.
Wins and losses post with equal prominence. The scorecard is written by the pipeline, not the marketing department. It can get worse as easily as better.
This same settled history feeds back into the desk: per-city calibration decides which cities' day-ahead signals have earned full status, and which stay Provisional. The model has to prove itself city by city, on settled outcomes, before its labels strengthen.
It also bends the numbers themselves. Once a city and regime have accumulated at least 30 settled outcomes, an isotonic (monotone, non-parametric) calibration layer fit on that history corrects the raw model probability before any gap is measured — if the desk has historically said 80% when reality delivered 70%, the published read becomes 70%. Every corrected row on the members' desk carries a small CAL marker whose tooltip shows the raw read and the sample it was fit on. Below 30 outcomes the raw probability stands untouched, and no correction ever pushes a probability past the 1–99% guardrails.
The Gap Desk is an information product. It surfaces where a market price and public data disagree, shows both numbers with their sources and timestamps, and keeps a public record of how those disagreements resolved. It does not recommend trades, size positions, manage risk, execute orders, or hold funds — and its language gate is built to keep it that way mechanically, not just as policy.
A gap is a divergence, not a promise. Markets can know things models don't: a stalled front, a sea-breeze the blend missed, or simply information that hasn't reached the guidance yet. The desk shows you the disagreement and its receipts. What you do with it is entirely your call.
Model estimates are exactly that — estimates. Verify independently. Prediction-market trading carries risk and may be restricted in your jurisdiction.
The sky doesn't care what the market thinks. That disagreement is the product — and the record of how it resolves is the receipt.
The Gap Desk · why every call is graded in public
A desk that publishes its scorecard shouldn't hide its disclaimers. This is the complete legal posture, in plain language:
Everything The Gap Desk publishes — this site, the daily brief, gap alerts, the scorecard — is information and education only. Nothing here is investment, financial, trading, legal, or tax advice, and nothing is a recommendation or solicitation to buy or sell any contract. The desk is not a registered investment adviser, broker-dealer, commodity trading advisor, or fiduciary with any regulator, anywhere.
Event contracts are all-or-nothing: a position can, and regularly does, go to zero. If you choose to trade anywhere, on anything, only ever risk money you can bear to lose entirely — money whose total loss would not change your life. If losing it would hurt, it doesn't belong at risk.
A model probability is an estimate built from public data, and public data can be wrong, late, or incomplete. The desk warrants nothing about accuracy, completeness, or availability. The public scorecard is history, and history is not a promise: past performance does not guarantee future results.
Prediction-market trading is regulated and may be unavailable or restricted where you live. It is your responsibility to know and follow the rules that apply to you; if you trade on Kalshi, Kalshi's own terms and eligibility rules govern — not anything written here.
The Gap Desk is independent. It is not affiliated with, endorsed by, or sponsored by Kalshi, NOAA, or the National Weather Service. Their names appear only to identify the markets and the public data sources being described.
The desk holds no trading account, never touches your funds, and never executes an order. It shows you a divergence and its receipts. Every decision you make — including the decision to do nothing — is yours alone.