What’s available#
We are publishing statistically downscaled, daily climate model output to support the regional evaluation of the potential impacts of stratospheric aerosol injection (SAI). The outputs cover four scenarios, two global climate models (GCMs), all available ensemble members, and five variables. We bias-correct GCM output against ERA5, and downscale it to a global 0.25° grid with two methods.
Property |
Value |
|---|---|
Release |
|
GCMs |
|
Downscaling methods |
|
Observations |
ERA5 |
Variables |
|
Temporal resolution |
Daily |
Spatial extent |
Global |
In addition to the historical period (historical) and the baseline climate scenario (ssp245),
we downscale two stratospheric aerosol injection scenarios. In the scenario g6_1p5k, greenhouse gas
emissions continue at ssp245 levels while sulfate aerosols are injected into the stratosphere to
hold warming to 1.5 °C. The termination shock scenario (g6_1p5k_end) extends g6_1p5k ensemble
member 002 from CESM2-WACCM6 to 2100, simulating an abrupt end of aerosol injection at the end of 2084.
Scenario |
Group name |
Years |
GCMs |
|---|---|---|---|
Historical |
|
1978 to 2014 |
|
SSP2-4.5 |
|
2015 to 2099 (some |
|
G6-1.5K |
|
2035 to 2084 |
|
G6-1.5K termination |
|
2085 to 2100 |
|
At a high level, the release includes two output data products, along with the processed GCM input data they were built from. The table below summarizes all three.
Product |
Description |
Grid |
Variables |
|---|---|---|---|
Downscaled |
Bias-corrected and spatially disaggregated. This is the main product. |
0.25° |
|
Coarse bias-corrected |
Bias-corrected, but not spatially disaggregated |
Native GCM grid, about 1° to 2° |
The same five, plus |
Processed input |
Daily GCM output that the pipeline started from, before bias correction |
Native GCM grid, about 1° to 2° |
The same five |
Data location#
Important
By downloading, copying, or using this data, you agree to the Terms of Data Access.
All data lives in CarbonPlan’s
Source Cooperative repository. The bucket is
public in AWS us-west-2, so you don’t need AWS credentials to read it. Each store is an
Icechunk repository under
s3://us-west-2.opendata.source.coop/carbonplan/srm-downscaling/, at the paths below:
Data |
Path |
Branch |
|---|---|---|
|
|
|
|
|
|
|
|
|
|
|
|
A branch is a version of a store. We publish one branch per release, and so far that is
only v1.0.0, which is the branch the example below opens. The input stores keep their data on
main.
Data access#
We offer two ways to access the data from the Source Cooperative repository. Which one fits best depends on how much data you need, and whether you want a local copy.
Download a local copy. If you want to work with a small amount of data on your own machine, or you prefer netCDF files, we built a set of access utilities. You can use them to subset, transform, and export the downscaled data from the cloud to your local environment without writing code yourself.
Stream data from the cloud. If you’re comfortable working with data in the cloud without keeping a local copy, you can use tools like Zarr-Python, Icechunk, or Xarray. The following example opens one downscaled group from the current release.
import icechunk
import xarray as xr
storage = icechunk.s3_storage(
bucket="us-west-2.opendata.source.coop",
prefix="carbonplan/srm-downscaling/output/production/CESM2-WACCM6-ERA5-global.icechunk",
anonymous=True,
region="us-west-2",
)
repo = icechunk.Repository.open(storage)
session = repo.readonly_session(branch="v1.0.0")
ds = xr.open_zarr(session.store, group="bcsd/g6_1p5k/tas/001", consolidated=False)
Data shape#
Group layout#
Zarr organizes data in a nested structure that you can think of as a file system. The key building block is the group, a bundle of arrays and metadata at a particular path that you can open as a single labeled dataset.
Each output store holds both data products for one GCM, and each release of that store is
a branch. Within a branch, a group’s path is built from the method, scenario, variable, and
ensemble member, with an extra debiased_coarse level for the coarse bias-corrected product.
Opening one group gives you a dataset with one climate variable on (time, lat, lon), plus any
quality flags. Both methods publish the same set of groups.
{method}/{scenario}/{variable}/{member} # downscaled
{method}/debiased_coarse/{scenario}/{variable}/{member} # coarse bias-corrected
Input data stores are organized differently. Each group holds one scenario, so opening a single
group gives you an array for each variable with the dimensions (ensemble_member, time, lat, lon).
Ensemble members#
Available ensemble members depend on the GCM, scenario, and variable. Both output data products and
both methods publish the same members, and under debiased_coarse, dtr has the same members as
tasmax and tasmin.
CESM2-WACCM6
Scenario |
Variables |
Members |
Years |
|---|---|---|---|
|
|
|
1978 to 2014 |
|
|
|
1978 to 2014 |
|
|
|
2015 to 2099 |
|
All five |
|
2015 to 2068 |
|
All five |
|
2015 to 2069 |
|
All five |
|
2035 to 2084 |
|
All five |
|
2085 to 2100 |
For the model CESM2-WACCM6, no historical ensemble member carries all five variables, and in
ssp245 only members 006 to 010 do. To avoid mixing realizations, we advise you to pick an
ensemble member that carries every variable you need.
UKESM1-1-LL
Scenario |
Variables |
Members |
Years |
|---|---|---|---|
|
All five |
|
1978 to 2014 |
|
All five |
|
2015 to 2099 |
|
All five |
|
2035 to 2084 |
For UKESM1-1-LL we downscaled one historical member, u-by791. That name is a Met Office suite
ID rather than a variant label like r2i1p1f2, and every scenario member above was bias-corrected
against it.
Grid, time, and chunks#
The groups described above tell you which part of the data you’re reading. Within a group, each variable is an array that is physically split into chunks: fixed-size blocks that are compressed and read as a single unit. A read fetches whole chunks, so how much data a request moves depends on how many chunks it touches, rather than how many values it returns. For this dataset, a time series at a single point reads about one chunk per year, while a wide region reads many chunks for every year.
Chunks are bundled into larger files called shards. Shards don’t change which chunks a request reads, so you can mostly ignore them when estimating how much a request will move. The table below summarizes the grid, time axis, and chunk layout of both output products.
Property |
Downscaled |
Coarse bias-corrected |
|---|---|---|
Dimensions |
|
|
Grid |
0.25°, 721 × 1440 cells |
|
Longitude convention |
-180 to 180 |
-180 to 180 |
Calendar |
Proleptic Gregorian |
Proleptic Gregorian |
Data type |
|
|
Chunk size ( |
365 × 36 × 72 (1 year × 9° × 18°) |
365 × 30 × 60 |
Shard size ( |
1095 × 180 × 360 (3 years × 45° × 90°) |
1095 × 90 × 180 |
Quality flags#
Groups carry quality flags alongside their variable, in both products. Each flag is a uint8
array where 0 means no known issue and 1 means a known issue. The dtr groups under
debiased_coarse carry no flags.
Flag |
Dimensions |
Present on |
Marks |
|---|---|---|---|
|
|
Every group except |
Pixel-days that fail a quality check: outlier screening, a variable-specific plausible range, or a physical relationship such as |
|
|
Every group except |
Pixels where debiasing or downscaling distorts how scenarios compare with each other or with historical, relative to raw GCM output. It’s evaluated on the ensemble mean. |
The flags summarize checks we run after downscaling. For the details of each check, see the output integrity checks and trend distortion check notebooks.