← All posts

Detecting and Understanding Weather Anomalies

Sunil Nagaraj · · 7 minute read

We had touched upon the current state of tracking and detecting weather patterns in an earlier post. In this post, we will outline our approach of systematically tracking unusual weather patterns or anomalies. We currently track anomalies for every urbanized area of the earth across different time granularities. We also clearly distinguish between those that are based on observations from those that are based on forecasts. We have written about data source types and time granularities earlier , but it’s intuitive to think of the anomalies as unusual patterns occurring on a daily, multi-day, weekly, monthly or even yearly basis. Each of these have different socio-economic implications and we have tried to provide a uniform and consistent way to act on them. 

Relevant Local Context

There are resources that track weather anomalies and while they are based on global gridded data, it's often impossible to zoom into details that provides relevant local context (e.g. the city, parameter measurements, related extremes). Providing such rich context improves the actionability of the signal. For instance, media will be able to communicate with appropriate level of local context to its audience. Business that use statistical models will benefit from being able to easily plugin data at the right level of detail and granularity.  An example shows a quick summary of an unusually hot day in March in Seville, Spain. In a glance, we are able to get an idea of the severity. (The maximum temperature is ~5C above normal)  and clicking through yields more information. 

The comparisons provide complete transparency regarding the baselines, the quality of data (210 points from 30 years for the week of March 26) and the observations (based on 47 observations).

 

 Observed Anomalies

A bigger impact is felt on society when anomalies persist. We track streaks of anomalies based on observations at a regional level.  This shows that Sevilla and indeed surrounding areas have experienced above normal daily high temperatures for most of March. Clicking through provides more detail about locations that are experiencing the most severe effects.

 

 

It is possible to verify this streak through our analytics functionality that maximum temperatures were indeed significantly above the 1991-2020 normal during the last three weeks of March.

 

Forecast Anomalies

In operational scenarios like estimating power demand or merely predicting if people are more likely to shop outdoors,  it is useful to detect anomalies that will be persistent. We make use of forecasts to build models of impending streaks across regions. Going back to our example. we can view the recent weather conditions of Seville where streaks, both present and future are summarized. 

The output suggests that temperatures are forecast to return to less severe levels (within 1-sigma) while it will continue to be drier than normal for the region. Of course, since we know that forecasts are not always accurate it is possible to see differences between observed streaks and streaks that are based on forecasts.

Methodology

We used a fairly straightforward way to detect anomalies. We sampled distributions of weather parameters across at least 30 years for the time periods of a week, month and year.  We observed that daily temperature, dew point, mean wind  were approximately normal but daily precipitation (rain, snow) totals were not. Daily rainfall totals were exponentially distributed (for rainy days). We also observed that means (and p10, p50, p90) of weekly, monthly and yearly temperature, dew point, max_wind, average wind, conditions (number of rainy days, thunder etc) were approximately normally distributed while accumulated total of precipitation were not. At the time of writing, we have observed that yearly total precipitation is close to normal distribution and weekly totals are more close to exponential. Literature suggests that attempts have been made to fit precipitation totals to Weibull distributions with appropriate shape and scale parameters. 

Distributions of Daily Observations 

Heathrow, 1991-2020. (Source: GHCND)

 Precipitation

 

Weekly                                          Monthly                                 Yearly

Temperature

Weekly                                     Monthly                                      Yearly

Distribution of Means and Accumulations 

Heathrow, 1993-2022 (Source: GSOD)

Temperature (TMAX) 

Temperature means are approximately normally distributed.

 

Weekly Mean                              Monthly Mean                       Yearly Mean

 

Precipitation (Rain)

Only yearly total appears to be normally distributed, weekly and monthly are NOT normally distributed.       

 

     Weekly Total Monthly Total                         Yearly Total

To enable us to learn and keep things simple, we started by detecting anomalies based on percentiles. For example, daily observations have to exceed p95(p5) of the corresponding distributions of the baseline.  We estimated the severity of the anomaly using standardized anomaly measures (e.g. z-scores for temperature).  We are also experimenting with detection of anomalies based entirely on models as we gain more understanding of the distributions of parameters across locations.

Comparisons Summary

The table summarizes the comparisons performed on observations (and forecasts) with the corresponding baseline. 

Anomaly Type

Observed

Baseline

Daily,  Streaks

Daily 

Daily across week and month (-30yrs)  

Weekly

Weekly Means

Weekly Means (-30 yrs) 

Monthly

Monthly Means

Monthly Means (-30 yrs)

Yearly

Yearly Means

Yearly Means (-30 yrs)

In addition, weekly , monthly and yearly anomalies also check if extremes have been exceeded on a given day for the corresponding year and/or month as applicable. For instance 40.2C at Heathrow on Jul 19 2022, would get flagged in weekly, monthly and yearly anomaly data as an extreme for month of July and for the year , in addition to outlier detection of weekly means, monthly means and yearly means respectively.

We are looking for help and feedback. If you are a climatologist or data scientist and can think of more effective methods or find glaring gaps please reach out to us.

 

Anomaly APIs

Almost all of the functionality that are described above are powered by APIs . Anomalies and streak data are also available as data files upon request. We believe that this makes it possible for customers to build custom visualizations or directly plug in the data to their models , whether it is a spreadsheet or a software system.  We expect the APIs and methodology to evolve as we get more feedback from potential customers. If you want to try our API or sample our data please get in touch.

Beyond Anomalies

Anomaly detection is based on comparing observed or forecast values with the latest climatology (past 30 years). While it is an important part of understanding weather patterns, there are other questions that follow. How has the climate itself changed for a city or a region? We will take a closer look at some of our efforts in trying to understand both shifts in climate and extremes over a period longer than the past 30 years. Stay tuned!