dapper package¶
Subpackages¶
- dapper.config package
- dapper.domains package
- Submodules
- dapper.domains.domain module
DomainDomain.cellsDomain.copy()Domain.domain_ncDomain.elm_latlon_layout()Domain.ensure_cells_lon_lat()Domain.ensure_output_dirs()Domain.export_domain()Domain.export_landuse()Domain.export_met()Domain.export_surface()Domain.from_elm_domain()Domain.from_file()Domain.from_gdf()Domain.from_geometry()Domain.from_provided()Domain.gdfDomain.gidsDomain.group_nameDomain.has_topounits()Domain.iter_runs()Domain.make_topounits()Domain.met_dirDomain.met_supportDomain.modeDomain.nameDomain.path_domain_nc()Domain.path_landuse_nc()Domain.path_met_dir()Domain.path_outDomain.path_surface_nc()Domain.path_zone_mappings()Domain.providedDomain.rep_points()Domain.run_dirDomain.run_groupDomain.simplify_support()Domain.supportDomain.support_for()Domain.to_df_loc()Domain.topo_supportDomain.topounitsDomain.topounits_dim_nameDomain.topounits_for_gid()Domain.topounits_gid_colDomain.topounits_id_colDomain.with_step_support()Domain.with_topounits()
- dapper.domains.elm_domain module
- Module contents
DomainDomain.cellsDomain.copy()Domain.domain_ncDomain.elm_latlon_layout()Domain.ensure_cells_lon_lat()Domain.ensure_output_dirs()Domain.export_domain()Domain.export_landuse()Domain.export_met()Domain.export_surface()Domain.from_elm_domain()Domain.from_file()Domain.from_gdf()Domain.from_geometry()Domain.from_provided()Domain.gdfDomain.gidsDomain.group_nameDomain.has_topounits()Domain.iter_runs()Domain.make_topounits()Domain.met_dirDomain.met_supportDomain.modeDomain.nameDomain.path_domain_nc()Domain.path_landuse_nc()Domain.path_met_dir()Domain.path_outDomain.path_surface_nc()Domain.path_zone_mappings()Domain.providedDomain.rep_points()Domain.run_dirDomain.run_groupDomain.simplify_support()Domain.supportDomain.support_for()Domain.to_df_loc()Domain.topo_supportDomain.topounitsDomain.topounits_dim_nameDomain.topounits_for_gid()Domain.topounits_gid_colDomain.topounits_id_colDomain.with_step_support()Domain.with_topounits()
- dapper.elm package
- dapper.elmtest package
- dapper.geo package
- Submodules
- dapper.geo.constants module
- dapper.geo.lonwrap module
- dapper.geo.sampling module
- dapper.geo.zonal module
- Module contents
- dapper.integrations package
- dapper.io package
- dapper.landuse package
- dapper.met package
- dapper.schemas package
- dapper.surf package
- Submodules
- dapper.surf.fraction_closure module
- dapper.surf.sample module
- dapper.surf.schema module
- dapper.surf.sfile module
CustomizeErrorSurfaceFileSurfaceFile.add_params_from_df()SurfaceFile.add_topounits_from_domain()SurfaceFile.basic_registry_check()SurfaceFile.drop_params()SurfaceFile.export()SurfaceFile.from_domain()SurfaceFile.from_halfdegree_point()SurfaceFile.from_netcdf()SurfaceFile.resize_dim()SurfaceFile.set_global_attrs()SurfaceFile.set_scalar()SurfaceFile.to_netcdf()SurfaceFile.validate()
build_surface_dataset()build_surface_dataset_cellset()customize_surface()write_surface_nc()
- dapper.surf.surface_var_specs module
- dapper.surf.validate module
- Module contents
SurfaceFileSurfaceFile.add_params_from_df()SurfaceFile.add_topounits_from_domain()SurfaceFile.basic_registry_check()SurfaceFile.drop_params()SurfaceFile.export()SurfaceFile.from_domain()SurfaceFile.from_halfdegree_point()SurfaceFile.from_netcdf()SurfaceFile.resize_dim()SurfaceFile.set_global_attrs()SurfaceFile.set_scalar()SurfaceFile.to_netcdf()SurfaceFile.validate()
- dapper.topounit package
Module contents¶
dapper public package interface.
The supported public import surface is:
Other submodules may be importable, but are not considered part of the stable “from dapper import …” API.
- class dapper.Domain(name, mode, provided, support, cells, topounits=None, topounits_dim_name='topounit', topounits_id_col='topounit_id', topounits_gid_col='gid', met_support=None, topo_support=None, domain_nc=None, path_out=None, run_group=None)[source]¶
Bases:
objectCanonical spatial object passed through the dapper pipeline.
- Geometry views:
provided: exactly what the user supplied (provenance/plotting)
support : what sampling SHOULD use (may be simplified/processed later)
cells : what ELM RUNS on (site points now; per-cell geometries for cellset)
- Prepared sampling views (set during the relevant pipeline step; not on init):
met_support
topo_support
- Mode:
sites : one set of outputs per row (exporters loop internally)
cellset : one set of outputs total, including all rows
- Parameters:
name (str)
mode (DomainMode)
provided (gpd.GeoDataFrame)
support (gpd.GeoDataFrame)
cells (gpd.GeoDataFrame)
topounits (gpd.GeoDataFrame | None)
topounits_dim_name (str)
topounits_id_col (str)
topounits_gid_col (str)
met_support (Optional[gpd.GeoDataFrame])
topo_support (Optional[gpd.GeoDataFrame])
domain_nc (Optional[Path])
path_out (Optional[Path])
run_group (Optional[str])
- cells: gpd.GeoDataFrame¶
- domain_nc: Optional[Path] = None¶
- elm_latlon_layout(decimals=6, use_lon_0360=True)[source]¶
Compute lat/lon axes and a gid -> (iy, ix) index map for ELM-style lat/lon layouts. Useful mainly when your cells actually lie on a lat/lon lattice (dense or sparse).
- Return type:
tuple[ndarray,ndarray,dict[str,tuple[int,int]]]- Parameters:
decimals (int)
use_lon_0360 (bool)
- ensure_cells_lon_lat()[source]¶
Ensure
cellscontainslonandlatcolumns, derived from geometry if needed.- Return type:
- ensure_output_dirs(*, met=True)[source]¶
Create output directories implied by this Domain (and its runs, if mode=’sites’). Does not write any files.
- Return type:
None- Parameters:
met (bool)
- export_domain(*, filename='domain.nc', out_dir=None, overwrite=False, append_attrs=None, **kwargs)[source]¶
Export ELM domain NetCDF(s) for this Domain.
- Output layout:
mode=’cellset’: <run_dir>/domain.nc
mode=’sites’ : <run_dir>/<gid>/domain.nc
- Return type:
dict[str,Path]- Returns:
dict[run_id, output_path]
- Parameters:
filename (str)
out_dir (str | Path | None)
overwrite (bool)
append_attrs (dict | None)
- export_landuse(src_path, *, filename='landuse_timeseries.nc', out_dir=None, overwrite=False, append_attrs=None, **kwargs)[source]¶
Export landuse timeseries NetCDF(s) for this Domain.
- Return type:
dict[str,Path]- Parameters:
src_path (str | Path)
filename (str)
out_dir (str | Path | None)
overwrite (bool)
append_attrs (dict | None)
- export_met(src_path, *, adapter, out_dir=None, filename=None, overwrite=False, append_attrs=None, pack_scope=None, clip_to_full_years=None, **kwargs)[source]¶
Export meteorological forcing NetCDF(s) for this Domain.
- Parameters:
src_path (Path-like) – Directory containing input CSV(s) for the adapter.
adapter (object) – Adapter instance implementing the met adapter protocol.
out_dir (Path-like, optional) – Override output root. Defaults to Domain.run_dir.
filename (str, optional) – Optional filename template for output NetCDFs. Supports two formats: - Template with {var} placeholder: ‘ERA5_{var}_1950-2025_z01’ generates ‘ERA5_TBOT_1950-2025_z01.nc’ - Simple prefix (legacy): ‘prefix’ generates ‘prefix_TBOT.nc’ If omitted, files use ‘<driver>_<var>_<start>-<end>_z01.nc’. The ‘.nc’ extension is added automatically.
overwrite (bool) – If False, raises if MET output(s) already exist.
clip_to_full_years (bool or None) – Controls whether the export year range is clipped to complete calendar years. Raises if clipping is enabled but no complete year exists. None preserves the adapter default (True for ERA5).
append_attrs (dict | None)
- Return type:
dict[run_id, met_dir]
- export_surface(src_path, *, filename='surfdata.nc', out_dir=None, overwrite=False, append_attrs=None, **kwargs)[source]¶
Export surface NetCDF(s) for this Domain.
- Return type:
dict[str,Path]- Parameters:
src_path (str | Path)
filename (str)
out_dir (str | Path | None)
overwrite (bool)
append_attrs (dict | None)
- classmethod from_elm_domain(path_nc, *, name=None, mask_name='mask', frac_name='frac', frac_threshold=0.0, path_out=None, run_group=None)[source]¶
Build a Domain from an ELM domain NetCDF. This naturally produces a ‘cellset’.
- Return type:
- Parameters:
path_nc (str | Path)
name (str | None)
mask_name (str)
frac_name (str)
frac_threshold (float)
path_out (str | Path | None)
run_group (str | None)
- classmethod from_file(path, *, name=None, layer=None, id_col='gid', mode=None, cell_kind=None, path_out=None, run_group=None)[source]¶
Load a geospatial file (e.g., GeoPackage, Shapefile) and construct a Domain.
- Return type:
- Parameters:
path (str | Path)
name (str | None)
layer (str | None)
id_col (str)
mode (Literal['sites', 'cellset'] | None)
cell_kind (Literal['site_points', 'as_provided'] | None)
path_out (str | Path | None)
run_group (str | None)
- classmethod from_gdf(gdf, **kwargs)[source]¶
Alias for
Domain.from_provided().- Return type:
- Parameters:
gdf (geopandas.GeoDataFrame | DataFrame)
- classmethod from_geometry(geometry, *, gid='site', name='domain', mode='cellset', cell_kind='site_points', path_out=None, run_group=None)[source]¶
Construct a single-feature Domain from a shapely geometry.
- Return type:
- Parameters:
geometry (shapely.geometry.base.BaseGeometry)
gid (str)
name (str)
mode (Literal['sites', 'cellset'])
cell_kind (Literal['site_points', 'as_provided'])
path_out (str | Path | None)
run_group (str | None)
- classmethod from_provided(provided, *, name='domain', mode=None, id_col='gid', support=None, cells=None, cell_kind=None, domain_nc=None, path_out=None, run_group=None)[source]¶
Construct a Domain from a provided geometry table.
The input can be a GeoDataFrame (preferred) or a DataFrame with a geometry column. Use mode to choose between a single-run cellset or a multi-run sites container. If cells is not provided, cells are derived from the provided/support geometry depending on cell_kind.
- Return type:
- Parameters:
provided (geopandas.GeoDataFrame | DataFrame)
name (str)
mode (Literal['sites', 'cellset'] | None)
id_col (str)
support (geopandas.GeoDataFrame | DataFrame | None)
cells (geopandas.GeoDataFrame | DataFrame | None)
cell_kind (Literal['site_points', 'as_provided'] | None)
domain_nc (str | Path | None)
path_out (str | Path | None)
run_group (str | None)
- property gdf: geopandas.GeoDataFrame¶
Backwards-compat alias for older code.
Historically dapper exposed a single GeoDataFrame on the domain object (often a
df_loc-style table with columns like gid/lon/lat/geometry). The closest equivalent iscells(the run-level geometry table).
- property gids: list[str]¶
List of
gidvalues (as strings) for the current cell table.
- property group_name: str¶
Name of the output group directory for this Domain.
- has_topounits()[source]¶
Return True if this Domain has a non-empty topounits table attached.
- Return type:
bool
- iter_runs()[source]¶
Yield (run_id, run_domain) where run_domain is always a single-run ‘cellset’ Domain. - mode=’cellset’ -> yields exactly one run (self) - mode=’sites’ -> yields one run per gid (single-row Domains)
- make_topounits(*, binning, sources=None, combine='cartesian', combine_order=None, max_topounits=256, dem_source='arcticdem', export_scale='native', min_patch_pixels=None, target_pixels_per_topounit=500, target_scale=None, verbose=False, allow_slow_ncells=25)[source]¶
Convenience wrapper that computes topounits for this Domain and returns a new Domain with domain.topounits attached.
If this Domain has multiple rows (cellset/sites), it computes topounits per-row (per gid) using dapper.topounit.topomake.make_topounits_for_domain.
If this Domain has one row, it still goes through the same path (safe + consistent).
Users should not need to deal with ee.Geometry vs ee.Feature vs FeatureCollection here.
- Return type:
- Parameters:
binning (dict)
sources (list[str] | None)
combine (str)
max_topounits (int)
dem_source (str)
export_scale (str)
target_pixels_per_topounit (int)
target_scale (float | None)
verbose (bool)
allow_slow_ncells (int)
- property met_dir: Path¶
Convenience property for
Domain.path_met_dir().
- met_support: Optional[gpd.GeoDataFrame] = None¶
- mode: DomainMode¶
- name: str¶
- path_domain_nc(filename='domain.nc', run_id=None)[source]¶
Default output path for the domain NetCDF for this Domain (or a specific run_id).
- Return type:
Path- Parameters:
filename (str)
run_id (str | None)
- path_landuse_nc(filename='landuse_timeseries.nc', run_id=None)[source]¶
Default output path for the landuse NetCDF for this Domain (or a specific run_id).
- Return type:
Path- Parameters:
filename (str)
run_id (str | None)
- path_met_dir(run_id=None)[source]¶
Path to the MET output directory for this Domain (or a specific run_id).
- Return type:
Path- Parameters:
run_id (str | None)
- path_out: Optional[Path] = None¶
- path_surface_nc(filename='surfdata.nc', run_id=None)[source]¶
Default output path for the surface NetCDF for this Domain (or a specific run_id).
- Return type:
Path- Parameters:
filename (str)
run_id (str | None)
- path_zone_mappings(filename='zone_mappings.txt', run_id=None)[source]¶
Default output path for
zone_mappings.txt(under the MET directory).- Return type:
Path- Parameters:
filename (str)
run_id (str | None)
- provided: gpd.GeoDataFrame¶
- rep_points(*, source='support', step=None)[source]¶
Representative points for a given geometry view. If source=’support’ and step is provided, uses the prepared support for that step.
- Return type:
GeoDataFrame- Parameters:
source (Literal['provided', 'support', 'cells'])
step (Literal['met', 'topounits'] | None)
- property run_dir: Path¶
Directory holding the main run outputs for this Domain instance.
- For top-level cellset or top-level sites container:
path_out/<group_name>- For per-site/per-cellset run domains (created by
iter_runsin sites mode): path_out/<group_name>/<domain.name>
- run_group: Optional[str] = None¶
- simplify_support(tolerance_m, *, step, preserve_topology=True, equal_area_epsg=6933)[source]¶
Simplify the support geometry for a step and store it as met_support/topo_support. Does NOT modify provided/support/cells.
- Return type:
- Parameters:
tolerance_m (float)
step (Literal['met', 'topounits'])
preserve_topology (bool)
equal_area_epsg (int)
- support: gpd.GeoDataFrame¶
- support_for(*, step=None)[source]¶
Return the geometry set that should be used for the given step. - step=None -> support - step=”met” -> met_support if set else support - step=”topounits” -> topo_support if set else support
- Return type:
GeoDataFrame- Parameters:
step (Literal['met', 'topounits'] | None)
- to_df_loc(*, lon_col='lon', lat_col='lat', weight_col='weight', frac_col='frac', default_weight=1.0)[source]¶
Derived location/weight table from cells (internal glue; users shouldn’t need to touch).
- Return type:
DataFrame- Parameters:
lon_col (str)
lat_col (str)
weight_col (str)
frac_col (str)
default_weight (float)
- topo_support: Optional[gpd.GeoDataFrame] = None¶
- topounits: gpd.GeoDataFrame | None = None¶
- topounits_dim_name: str = 'topounit'¶
- topounits_for_gid(gid)[source]¶
Return the topounits subset for a single gid (or None if no topounits).
- Parameters:
gid (str)
- topounits_gid_col: str = 'gid'¶
- topounits_id_col: str = 'topounit_id'¶
- with_step_support(step, gdf)[source]¶
Attach a step-specific support GeoDataFrame (for “met” or “topounits”).
- Return type:
- Parameters:
step (Literal['met', 'topounits'])
gdf (geopandas.GeoDataFrame)
- with_topounits(topounits, *, id_col='band_name', gid_col='gid', dim_name='topounit')[source]¶
Attach topounits GeoDataFrame to this Domain. - Ensures a stable id column name (self.topounits_id_col == ‘topounit_id’) - Ensures gid linkage column exists (self.topounits_gid_col)
- Return type:
- Parameters:
topounits (geopandas.GeoDataFrame)
id_col (str)
gid_col (str)
dim_name (str)
- class dapper.ERA5Adapter[source]¶
Bases:
BaseAdapterERA5-Land → ELM adapter.
This adapter implements the
BaseAdapterinterface for ERA5-Land hourly data. It handles source-specific details—file discovery, unit conversions, humidity diagnostics, renaming to ELM short names, and nonnegativity enforcement, so the upstreamExportercan remain source-agnostic.Responsibilities¶
discover_files: Find CSV shards in a directory and infer the overall (start_year, end_year) using their date coverage.
normalize_locations: Validate and normalize the locations table (adds
lon_0-360, ensures/createszone, stable sorting).id_column_for_csv: Declare the identifier column name in the input CSVs. For ERA5 we require
gid.preprocess_shard: Convert one merged shard (CSV rows joined to locations) into canonical ELM columns. Steps include:
time filtering and optional “noleap” removal of Feb 29
ERA5→ELM unit conversions (e.g., J/hr/m² → W/m², m/hr → mm/s)
optional humidity computation (RH/Q) if temperature, dewpoint, and surface pressure are available
renaming raw ERA5 fields to ELM short names via a mapping
clipping canonical nonnegative variables
returning only required columns in a deterministic order
required_vars: Report the canonical ELM variable names required for the requested output format.
pack_params: Provide robust
(add_offset, scale_factor)for a canonical ELM variable, given optional data to tune ranges.
Notes
Humidity computation is performed only when
temperature_2m,dewpoint_temperature_2m, andsurface_pressureare present.Precipitation conversion uses
m/hr → mm/svia division by3.6.
- DRIVER_TAG = 'ERA5'¶
- GEE_COLLECTION_START = Timestamp('1950-01-01 01:00:00')¶
- INTERVAL_END_VARS = ('FSDS', 'FLDS', 'PRECTmms')¶
- SOURCE_INTERVAL_HOURS = 1.0¶
- SOURCE_NAME = 'ERA5-Land hourly reanalysis'¶
- discover_files(csv_directory, calendar, *, clip_to_full_years=None)[source]¶
Discover ERA5 CSV shards in a directory and infer the inclusive year range.
- id_column_for_csv(df_csv, id_col)[source]¶
Return the required identifier column name expected in ERA5 CSV shards (“gid”).
- pack_params(elm_var, data=None)[source]¶
Return (add_offset, scale_factor) used to pack a variable for NetCDF output.
- preprocess_shard(df_merged, start_year, end_year, calendar, dformat)[source]¶
Filter time & handle no-leap
Apply ERA5 → ELM unit conversions
Compute humidities (if columns available)
Rename columns to canonical ELM names using RAW_TO_ELM
Clip canonical nonnegative variables
Return only the canonical vars required by elm_required_vars(dformat), plus LONGXY/LATIXY/time/gid/zone (coords/meta).
- class dapper.ERA5SamplingPlan(requested_backend, backend, sampling_method, feature_count, grid_cell_count, estimated_seconds, estimated_download_mb, reason)[source]¶
Bases:
objectSummary of Dapper’s ERA5-Land backend decision.
- Parameters:
requested_backend (str)
backend (str)
sampling_method (str)
feature_count (int)
grid_cell_count (int)
estimated_seconds (float)
estimated_download_mb (float)
reason (str)
-
backend:
str¶
-
estimated_download_mb:
float¶
-
estimated_seconds:
float¶
-
feature_count:
int¶
-
grid_cell_count:
int¶
-
reason:
str¶
-
requested_backend:
str¶
-
sampling_method:
str¶
- class dapper.Exporter(adapter, src_path, *, domain, out_dir=None, calendar='noleap', dtime_resolution_hrs=1, dtime_units='days', dformat='BYPASS', append_attrs=None, chunks=None, include_vars=None, exclude_vars=None, clip_to_full_years=None)[source]¶
Bases:
objectSource-agnostic meteorological exporter.
This class orchestrates a two-pass pipeline that ingests time-sharded CSVs for many sites/cells, preprocesses them via a pluggable adapter, and writes ELM-ready NetCDF outputs in two layouts:
"cellset"– one NetCDF per variable with dims('DTIME','lat','lon')(global packing; sparse lat/lon axes are OK)."sites"– one directory per site; each directory contains one NetCDF per variable with dims('n','DTIME')wheren=1(per-site packing).
Exporter is source-agnostic: all dataset-specific logic (file discovery, unit conversions, renaming to ELM short names, etc.) lives in an adapter that implements the BaseAdapter interface (e.g., an
ERA5Adapter). The exporter handles staging (CSV → per-site parquet), global DTIME axis creation, packing scans, chunking, and NetCDF I/O.- Parameters:
adapter (BaseAdapter) – Implements:
discover_files,normalize_locations,preprocess_shard,required_vars, andpack_params.csv_directory (str or pathlib.Path) – Directory containing time-sharded CSV files for all sites/cells.
out_dir (str or pathlib.Path) – Destination directory for NetCDF outputs and temporary parquet shards.
df_loc (pandas.DataFrame) – Locations table with at least columns
["gid","lat","lon"]; optional"zone". The adapter’snormalize_locations: - validates columns, - adds"lon_0-360", - fills/validates"zone", - sorts for stable site order.id_col (str, optional) – Kept for backward compatibility (unused when
"gid"is assumed).calendar ({"noleap","standard"}, default "noleap") – Calendar for numeric DTIME coordinate; Feb 29 filtered for “noleap”.
dtime_resolution_hrs (int, default 1) – Target time resolution in hours for the DTIME axis.
dtime_units ({"days","hours"}, default "days") – Units of the numeric DTIME coordinate (e.g.,
"days since YYYY-MM-DD HH:MM:SS").domain (Domain)
dformat (str)
append_attrs (dict | None)
clip_to_full_years (bool | None)
- dformat{“BYPASS”,”DATM_MODE”}, default “BYPASS”
Target ELM format selector passed through to the adapter.
- append_attrsdict, optional
Extra global NetCDF attributes to include in every file. The exporter also adds:
export_mode("cellset"or"sites") andpack_scope("global"or"per-site").- chunkstuple[int,…], optional
Explicit NetCDF chunk sizes.
- include_vars / exclude_varsIterable[str], optional
Allow-/block-lists of ELM short names applied after preprocess. Meta columns
{"gid","time","LATIXY","LONGXY","zone"}are always kept.- clip_to_full_yearsbool or None, optional
Controls whether the discovered export year range is clipped to full calendar years. Clipping raises if no complete year exists.
Nonepreserves the adapter default (True for ERA5).
Side Effects¶
Creates a temporary directory of per-site parquet shards under
out_dir.Writes NetCDF files to
out_dirin the chosen layout.Writes a
zone_mappings.txtfile either at the root (cellset) or inside each site directory (sites).
Notes
Packing: global packing for
cellset; per-site packing forsites.Required columns: CSV shards and
df_locboth use"gid"; CSVs include the adapter’s date/time column (renamed to"time"during preprocess).Combined (lat/lon) layout: does not enforce regular grids; axes are the unique sorted lat/lon from
df_loc(sparse OK).
- run(*, pack_scope=None, filename=None, overwrite=False)[source]¶
Run the MET export for this exporter’s Domain.
- The output layout is derived from
Domain.mode: sites: writes<run_dir>/<gid>/MET/{prefix_}{var}.ncand a per-sitezone_mappings.txt(always zone=01, id=1).cellset: writes<run_dir>/MET/{prefix_}{var}.ncand a singlezone_mappings.txtcovering all locations (zones taken from df_loc, default 1).
- Parameters:
pack_scope – Optional packing strategy override. Defaults to
per-sitefor sites andglobalfor cellset outputs.filename (
str|None) – Optional filename template for output NetCDF files. Supports two formats: - Template with {var} placeholder: ‘ERA5_{var}_1950-2025_z01’ generates ‘ERA5_TBOT_1950-2025_z01.nc’ - Simple prefix (legacy): ‘prefix’ generates ‘prefix_TBOT.nc’ If omitted, the default is ‘<driver>_<var>_<start>-<end>_z01.nc’. The ‘.nc’ extension is added automatically.overwrite (
bool) – If True, clears existing MET outputs before writing.
- Return type:
None
- The output layout is derived from
- dapper.plan_era5_land_sampling(domain, start_date, end_date='latest', *, backend='auto', sampling_method='auto', variables='elm', arco_max_grid_cells=25, arco_max_estimated_seconds=300.0)[source]¶
Plan an ERA5-Land request without contacting ECMWF or GEE.
- Return type:
- Parameters:
domain (Domain)
backend (str)
sampling_method (str)
arco_max_grid_cells (int)
arco_max_estimated_seconds (float)
- dapper.sample_e5lh(params, domain_name=None, skip_tasks=False)[source]¶
Submit Google Earth Engine (GEE) export tasks for ERA5-Land Hourly time series.
This prepares the ERA5-Land Hourly ImageCollection (
"ECMWF/ERA5_LAND/HOURLY"), validates bands, ensures each geometry samples at least one pixel center (falling back to points when needed), batches the requested date range into N-year chunks, and (unlessskip_tasks=True) starts one Drive export task per batch.- Parameters:
params (dict) –
Configuration dictionary. Expected keys (case-sensitive):
start_date (str): Inclusive output start date or datetime.
end_date (str): Exclusive output end date/datetime, or
"latest".geometries: One of the following:
str: GEE asset ID for a FeatureCollection (e.g.,
"users/me/my_fc").ee.FeatureCollection: a pre-constructed collection.
GeoDataFrame: must contain geometry and an ID column (see
geometry_id_field).AOI:
dapper.domains.aoi.AOIinstance; uses its internal GeoDataFrame.Domain:
dapper.domains.domain.Domaininstance; usesDomain.to_geometries().
geometry_id_field (str, optional): ID column in provided geometries. Defaults to
"gid". Values are copied into the"gid"property on each feature.gee_bands (str or list[str]): Which ERA5-Land bands to export. One of:
"all": all available bands (fromera5.ALL_BANDS)"elm": bands required to derive ELM variables (fromera5.REQUIRED_RAW_BANDS)a list of band names validated against the collection
gdrive_folder (str): Google Drive folder name where CSV chunks are written.
job_name (str): Base name used to build per-batch export descriptions/filenames.
gee_scale (str or int or float): Sampling scale in meters. If
"native"(or a value < 11132), the native ERA5-Land scale of 11132 m is used.gee_years_per_task (int, optional): Years per export batch (default:
5).
The function sets
params["gee_ic"] = "ECMWF/ERA5_LAND/HOURLY"internally.domain_name (str, optional) – Optional name for the returned Domain.
skip_tasks (bool, default False) – If True, do everything except starting the GEE export tasks.
- Returns:
Domain describing the sampling locations. The underlying GeoDataFrame contains at least
"gid","lon", and"lat".- Return type:
Notes
Call
ee.Initialize()before using this function.CSV selectors include
["gid", "date"] + params["gee_bands"].Dates are derived from
system:time_startand formatted in UTC.Sampling includes the image at
end_dateas a one-hour lookahead for ERA5-Land interval fields. That image is not part of the requested output time axis when the CSVs are converted to ELM forcing files.
- Raises:
KeyError – If required keys are missing from
params.ValueError – If dates are malformed or
geometriesis an unsupported type.TypeError – If
gee_scaleis not"native"and not numeric.ee.EEException – Propagated Earth Engine errors (e.g., authentication, export quota).
Examples
params = { "start_date": "1950-01-01", "end_date": "1951-12-31", "geometries": "users/me/my_sites_fc", "geometry_id_field": "gid", "gee_bands": "elm", "gee_scale": "native", "gee_years_per_task": 5, "gdrive_folder": "era5_exports", "job_name": "era5l_sites", } domain = sample_e5lh(params) domain.gdf.head()
- dapper.sample_era5_land(domain, start_date, end_date='latest', *, backend='auto', sampling_method='auto', variables='elm', output_dir=None, overwrite=False, keep_downloads=False, arco_max_grid_cells=25, arco_max_estimated_seconds=300.0, cds_client=None, gdrive_folder='dapper_era5_land', job_name=None, gee_scale='native', gee_years_per_task=5, skip_tasks=False)[source]¶
Sample ERA5-Land through ECMWF ARCO or Google Earth Engine.
Dates use an inclusive start and exclusive output end. ARCO downloads are synchronous and written as GEE-compatible CSV files under
output_dir. GEE exports remain asynchronous Drive tasks.backend='auto'chooses ARCO only when every feature fits its service limits and the estimated request is below the configured cell/time thresholds.ECMWF omits the incomplete 1950-01-01 source day from its ERA5-Land ARCO archive. When ARCO is selected for a request beginning on that date, Dapper warns and begins sampling on 1950-01-02 instead of silently switching the entire request to GEE. Sampling metadata records both dates. Use
export_met(..., clip_to_full_years=True)to omit partial boundary years, or selectbackend='gee'explicitly when 1950-01-01 is required.For polygon supports,
sampling_method='auto'uses a local area-weighted mean of intersecting ERA5-Land cells with ARCO and GEE’s spatial mean with GEE. Points use nearest-cell sampling. Usesampling_method='nearest'to sample polygon representative points instead.Sampling provenance is attached without changing the input Domain’s model coordinates or weights. GEE geometry-reference coordinates are recorded as
sampling_reference_lonandsampling_reference_latwhen available.