Research · September 2026

Can an AI model make a better rain forecast?

That's the experiment behind Ninety Sky. We give Jev, an AI model, the same radar and weather-model evidence a forecaster would look at, and ask it whether it will rain in the next 15, 30, 45, 60 and 90 minutes. Then we check every answer against the radar. Here's what we've found so far.

9,233 forecasts45,522 verified windows37 places2026-09-23 – 2026-09-26

What we're testing

Radar is very good at the next few minutes. It shows where rain is falling and which way it's moving. It struggles with rain that hasn't formed yet, and with storms that grow or fall apart on the way to you. Weather models see that bigger picture, but they're coarse and update slowly. A good forecaster looks at both and makes a call.

Jev is our attempt to make that call automatically, every five minutes, for every place someone is looking at. It reads a short summary of the radar, how the rain is moving, what's nearby and what NOAA's models expect. For each window, it answers with a probability of rain. We blend that answer with the raw numbers, and the blend is what you see in the app.

The question we're testing is whether the forecast gets better with Jev in it.

How a forecast is made

For each place, we ask: will it rain here in the next 0–15 minutes? 15–30? 30–45? 45–60? 60–90? “Here” means an area about 8 km across.

MRMS RADARevery 2 minutes · 2 km MOTIONslides the rain forward NEIGHBOURHOODrain within reach, 50 km NOAA MODELSHRRR · NWS hourly chance JEVan AI modelreads the evidence andanswers each windowwith a probability BLENDweighs Jev againstthe raw inputs,separately foreach window THE RIBBON height = chancecolour = how hard the raw inputs also go straight into the blend CHECK EVERY WINDOWcompare each forecast with what radar saw,then refit the blend on what we learn new weights

Radar comes from NOAA's Multi-Radar Multi-Sensor mosaic, which updates every two minutes. Motion works out which way the rain is heading and slides it forward. The neighbourhood measures how much rain is close enough to reach you, and how active the surrounding 50 km are. That way a storm sitting just off to the side still counts. The models are HRRR, NOAA's hourly-updating 3 km model, and the National Weather Service's hourly chance of rain.

Jev reads all of that and gives a probability for each window. It doesn't write the forecast. It weighs the evidence. The blend is a simple statistical formula that combines Jev's answer with the raw inputs. Each window gets its own weights, which we fit on past days and check on days the formula hasn't seen.

Afterwards, we check every window against what the radar actually saw. That's where the results below come from.

Results

A window counts as rainy if radar showed at least 0.2 mm/h anywhere in the area during it. We compare Ninety Sky with three simpler forecasts built from the same data. Persistence says it will keep raining if it's raining now. Radar motion slides the current rain forward along its track. The weather model turns NOAA's model guidance into a probability.

The score is skill. Zero means no better than always guessing the average chance of rain, and 1 is perfect.

In the first 15 minutes, anything built on radar does well. Ninety Sky scores 0.75 and radar motion 0.67. The difference shows up further out. By 60–90 minutes, radar motion drops to 0.25, and Ninety Sky still scores 0.52. That later stretch, where radar alone loses track of what the weather is doing, is the part we're trying to improve.

Figure 1 · Skill by how far ahead we're forecasting (higher is better)
All four forecasts are scored on the same windows. Ninety Sky is shown before the calibration step described below.
WindowCasesRainedNinety SkyWeather modelRadar motionPersistenceNinety Sky AUC

What Jev adds

The blend uses plenty of inputs besides Jev, so a good forecast doesn't prove Jev is doing the work. To test that, we fit the same formula twice, once with Jev's answer and once without. We trained both on earlier days and tested them on later days neither had seen. For this test we replayed archived radar and HRRR data, because we need more history than the live log has.

With Jev, error was lower by 1–4% in four of the five windows, and about even in the 45–60 minute window. That's a real improvement, but a modest one. Jev is picking up something the numbers alone miss. It doesn't transform the forecast. Part of the gain may also come from the fuller picture of the radar that Jev gets, not only from its judgment.

Trained on 2026-08-01 to 2026-08-22, with a 24-hour gap before the test days: 2026-08-26, 2026-08-27, 2026-09-12, 2026-09-13, 2026-09-14.

WindowTest casesError with JevError without JevImprovement
0-15 min1,6350.02580.0263+1.9%
15-30 min1,5880.04250.0444+4.3%
30-45 min1,5540.05820.0588+1.0%
45-60 min1,5070.06280.0627-0.2%
60-90 min1,4240.07670.0792+3.2%

Jev on its own

If Jev is the interesting part, why not show its answer directly? We compared the two on the windows where Jev was asked, which is 59% of them. When radar and the models agree it's dry, we skip Jev. The blend's error was 14–32% lower in every window.

Jev isn't worse at telling rainy windows from dry ones. On that measure (AUC in the table), Jev and the blend are nearly identical. The gap is in the percentages. Jev hedges toward the middle. When it said about 11%, it rained 1% of the time. When it said about 86%, it rained 96% of the time. The blend pulls those numbers toward what actually happens.

So the two do different jobs. Jev judges which windows are riskier, and the blend turns that into a chance you can take at face value.

WindowCasesError, Jev aloneError, blendImprovementAUC, JevAUC, blend
0-15 min5,4120.0410.028+32%0.9830.985
15-30 min5,3940.0480.037+23%0.9730.975
30-45 min5,3780.0560.045+20%0.9570.961
45-60 min5,3630.0620.052+17%0.9480.951
60-90 min5,3350.0740.063+14%0.9340.935

Results by region

Besides forecasting for people using the app, we run forecasts around the clock for 24 cities chosen to cover ten climates. The regions below had enough rain in this period to score.

Ninety Sky did best in the Northeast, where rain tends to arrive in large, organized systems. It did worst in Denver and Asheville, where mountains set off storms that radar and models both struggle to anticipate. The weather model had its hardest time with the Southwest monsoon's pop-up afternoon storms, scoring below zero there.

ClimateCitiesCasesRainedNinety SkyWeather modelRadar motion
Southwest monsoonPhoenix, Tucson, Albuquerque5,8778%0.56-0.020.43
NortheastNew York, Boston, Washington5,81511%0.740.580.57
PlainsDallas, Oklahoma City, Kansas City5,9729%0.650.400.52
Pacific NorthwestSeattle, Portland3,97315%0.600.130.48
Southeast stormsCharlotte, Atlanta, Orlando, Nashville6,33713%0.590.210.38
Appalachians & Front RangeDenver, Asheville3,7737%0.43-0.500.18

Do the percentages match what happened?

A forecast can tell rainy from dry reliably and still get the odds wrong. This chart compares what we said with what happened. Points on the dotted line are exactly right, and bigger dots mean more cases.

Further out, our chances line up well with what happened. In the first half hour we undersell low chances. When we said about 14% for the next 15 minutes, it rained about a third of the time.

Figure 2 · What we said vs. how often it rained
Points above the line mean it rained more often than we said; points below mean less often. Groups with fewer than eight cases are left out.

The app now applies a correction to the 0-15, 15-30, 30-45, 45-60 minute windows to fix this. The chart shows forecasts from before that correction.

The hourly forecast

Past 90 minutes, the app gives an hourly chance of at least 0.1 mm of rain. An earlier version of this page compared that with the National Weather Service, but it scored the wrong event, so we took it down. We'll publish a fair comparison once we're measuring the same thing the NWS forecasts.

What's next

This report covers 4 days of live forecasts. That's enough to see a pattern but not enough to cover every season. Forecasts are for an area about 8 km across, so conditions on your street can differ. The comparisons are with simple forecasts we built from the same data, not with other weather apps. We'll keep checking every forecast and update this page as the record grows.

Rainy window
Radar showed at least 0.2 mm/h in the 8 km area on any scan during the window. We only score windows where the radar record is complete.
Error
The Brier score: the average squared gap between the forecast chance and what happened (1 for rain, 0 for dry). Lower is better.
Skill
How much lower our error is than always forecasting the average chance of rain. 0 is no better than that, 1 is perfect, and below 0 is worse.
AUC
How well the forecast separates rainy windows from dry ones. If you pick one of each at random, AUC is the chance the rainy one got the higher number. 0.5 is a coin flip, 1 is perfect.
Weather model
Our conversion of HRRR's rain amount and type into a probability, with the NWS hourly chance as a fallback. It isn't a probability that HRRR produces itself.
Data
Production forecast snapshot retrieved September 26, 2026. 9,233 forecasts at 37 places, 45,522 scored windows, 3,385 of them rainy, 2026-09-23 – 2026-09-26.