Methodology

No black box. This is exactly how the desk works.

Every probability the desk publishes is computed from public National Weather Service data by a fixed, documented pipeline — and every published call is scored after settlement, in public. This page is the pipeline, end to end.

01 · The data

Everything starts from public NWS sources.

The desk holds no private data and no inside information. Three classes of government-published weather data feed the model, each used for a different job:

Forecast guidance — the NBM

The National Blend of Models (NBM) is NOAA's flagship statistical blend of dozens of weather models. Its probabilistic text bulletins (NBP) publish not just an expected high but a full spread — percentiles of tomorrow's temperature distribution per station. New guidance is produced four times a day (01, 07, 13 and 19 UTC), and the desk re-reads each cycle shortly after it lands.

Live observations — METAR

Each market settles against a specific airport station (Central Park for New York, Midway for Chicago, and so on). The desk polls that station's METAR observations through the day, including the 6-hour maximum and minimum temperature groups embedded in synoptic-hour reports — which catch short spikes that hourly readings can miss.

Settlement truth — the CLI report

The official daily climate report (CLI) that NWS offices issue each morning states yesterday's final maximum and minimum temperature. This is the same value Kalshi's weather markets settle on — so the desk grades itself against the identical source of truth the market uses, not its own opinion of what happened.

The other half of every gap is the market itself: Kalshi's public quote for each contract — best bid, best ask, last trade, volume and open interest — snapshotted with a timestamp at the moment of comparison. Market data is read-only and public; the desk holds no trading account and places no orders.


02 · The model

From percentiles to a probability.

A Kalshi contract like NYC high ≥ 90°F is a yes/no question about an integer. The desk answers it by fitting a smooth distribution to the NBM's published percentiles for that station and date, then reading the tail probability above the strike — with a continuity correction, because settlement is a whole number: "≥ 90" means the rounded official high reaches 90, so the model integrates from 89.5.

Bracket contracts ("high between 85–86°F") are the same computation on both edges: the probability mass that falls inside the bracket after rounding.

Three regimes, three levels of trust

What the model knows depends on the time of day, so every read is tagged with the regime it was computed under:

Forecast

Day-ahead. Built purely from NBM guidance — the widest uncertainty, and the regime where the desk is most conservative about firing.

Intraday

Settlement day. The forecast distribution is conditioned on what the station has already recorded — a morning running cool caps how high the afternoon can plausibly reach. Intraday claims must rest on a recent observation; stale data disqualifies the read.

Locked

The station has already touched the strike. Once an observed max prints at or above the threshold, "≥" is effectively decided — newer observations can only raise a maximum, never lower it. These are the desk's highest-confidence reads.

Day-ahead (forecast-regime) signals carry a Provisional label until that city's own settled track record earns full status — roughly: enough settled calls (30+), a strong standalone Brier score, and a head-to-head win over the market price on the same markets. Until then they are published, labeled, and graded — never silently hidden, never oversold.


03 · When a gap fires

Most divergences never make the brief.

A raw difference between the model's probability and the market's price is not a signal — it might be noise, an illiquid book, or a spread wide enough to swallow the whole edge. Every candidate must clear all of these guards at once:

minimum gap
|model − market| must be at least 8 points.Small disagreements are within model error; the desk doesn't publish coin-flip differences.
spread
The bid/ask spread must be 10¢ or tighter.A wide spread means the "market price" is a guess between two distant numbers — too noisy to compare against.
activity
The book must show real participation (volume / open interest).A dead book is not a market opinion. A gap against nobody is not a gap.
net edge
The gap must survive spread and trading fees with at least 2 points left.A divergence that frictions would fully consume is trivia, not information.
observation freshness
Unlocked intraday claims require a recent station observation.A claim about "so far today" resting on hours-old data is not a claim about today.
conviction
A blended 0–100 quality score must reach Tier C (40) or better; fired signals carry their tier (A/B/C).Gap size, book quality, regime trust and freshness rolled into one rankable number.

Suppressed candidates aren't discarded — they're recorded with the reason they didn't fire, so the desk's restraint is auditable too.


04 · Before anything publishes

Two gates stand between the model and you.

Source-traceability gate

Every number in a rendered brief line is checked against the underlying fetched record. If a figure can't be traced to a real fetched value — an NWS bulletin, a station observation, a Kalshi snapshot with its timestamp — the line is blocked. No unsourced numbers, ever.

Language gate

The full text is scanned against a banned-language list: no "buy", no "sell", no "guaranteed", no "can't lose", no advice framing, no hype. The desk describes a divergence; it never recommends a position. Lines that fail are blocked, not softened.

Only what clears both gates is published — to the member desk, the daily brief, and the public scorecard.


05 · Keeping score

Every call is graded. In public. Against the market.

The morning after a market settles, the desk pulls the official CLI climate report and records the outcome. Scoring rules are fixed in advance:

One prediction of record per market. For each contract and date, the read that counts is the forecast-regime probability subscribers actually saw in the morning brief — the desk can't quietly re-time its predictions after watching the day unfold.

Brier score, head to head. Both the desk's probability and the market's price at the same moment are scored against the outcome with the Brier score (mean squared error of probabilities — lower is better, 0.25 is coin-flip guessing). The comparison uses the same settled markets for both sides. If the desk's Brier isn't lower than the market's, the desk added nothing over just reading the price — and the public scorecard will say so.

Wins and losses post with equal prominence. The scorecard is written by the pipeline, not the marketing department. It can get worse as easily as better.

This same settled history feeds back into the desk: per-city calibration decides which cities' day-ahead signals have earned full status, and which stay Provisional. The model has to prove itself city by city, on settled outcomes, before its labels strengthen.

It also bends the numbers themselves. Once a city and regime have accumulated at least 30 settled outcomes, an isotonic (monotone, non-parametric) calibration layer fit on that history corrects the raw model probability before any gap is measured — if the desk has historically said 80% when reality delivered 70%, the published read becomes 70%. Every corrected row on the members' desk carries a small CAL marker whose tooltip shows the raw read and the sample it was fit on. Below 30 outcomes the raw probability stands untouched, and no correction ever pushes a probability past the 1–99% guardrails.


06 · What the desk is not

Descriptive, never prescriptive.

The Gap Desk is an information product. It surfaces where a market price and public data disagree, shows both numbers with their sources and timestamps, and keeps a public record of how those disagreements resolved. It does not recommend trades, size positions, manage risk, execute orders, or hold funds — and its language gate is built to keep it that way mechanically, not just as policy.

A gap is a divergence, not a promise. Markets can know things models don't: a stalled front, a sea-breeze the blend missed, or simply information that hasn't reached the guidance yet. The desk shows you the disagreement and its receipts. What you do with it is entirely your call.

Model estimates are exactly that — estimates. Verify independently. Prediction-market trading carries risk and may be restricted in your jurisdiction.

The sky doesn't care what the market thinks. That disagreement is the product — and the record of how it resolves is the receipt.

The Gap Desk · why every call is graded in public