ids#
ids.py
Functions for computing geographic identifiers.
geo_id: to link identical polygons securely through time (within a small spatial tolerance, to minimize corrections).
openlocationcode: for point locations
unique building IDs (UBID) for building footprints
Functions#
|
Generate stable, unique parcel IDs from polygon geometry. |
|
Return the GeoDataFrame using geo_id as the index |
|
Assign a location index based on centroid (Open Location Code). |
|
Return the GeoDataFrame using geo_id as the index |
|
Return the GeoDataFrame using geo_id as the index |
|
Decode level-11 Open Location Codes to a GeoDataFrame of points. |
|
Decode a Series of UBIDs to a GeoDataFrame of bounding boxes. |
|
Standardize raw parcel ids into a matching key via a conversion code. |
|
Return the dominant extraction pattern of raw parcel ids. |
|
Compute the standardized |
Module Contents#
- openplaces.geo.ids.get_geo_ids(gdf, grid_degrees=1e-06, hash_length=24, handle_duplicates=True, verbose=False)#
Generate stable, unique parcel IDs from polygon geometry.
Uses fixed degree grid in EPSG:4326 for full Earth coverage.
- Parameters:
gdf (GeoDataFrame) – GeoDataFrame with parcel geometries
grid_degrees (float) – Grid size in degrees
hash_length (int) – Number of hex characters in output
handle_duplicates (bool) – Add unique numeric suffix to duplicate geo_ids (default True)
verbose (bool) – Print information on duplicates (default False)
- Returns:
Series of geo_ids with same index as input GeoDataFrame
- Return type:
pd.Series
- openplaces.geo.ids.add_geo_id_index(gdf, name='geo_id', handle_duplicates=True, verbose=False)#
Return the GeoDataFrame using geo_id as the index
- Parameters:
gdf (GeoDataFrame) – Polygon data
name (str) – Name of index column
handle_duplicates (bool) – If True, adds numeric suffix to duplicate GIDs (default True)
verbose (bool) – If True, prints information on duplicates
- openplaces.geo.ids.get_openlocationcodes(gdf: geopandas.GeoDataFrame, name='openlocationcode', codelength=11, handle_duplicates=True)#
Assign a location index based on centroid (Open Location Code).
- Parameters:
gdf (GeoDataFrame) – Vector data with polygon geometries in any CRS.
name (str) – Name of index
codelength (int) – openlocationcode code length
handle_duplicates (bool) – If True, adds numeric suffix to duplicate OLCs (default True)
- openplaces.geo.ids.add_openlocationcode_index(gdf, name='openlocationcode', codelength=11, handle_duplicates=True)#
Return the GeoDataFrame using geo_id as the index
- Parameters:
gdf (GeoDataFrame) – Polygon data
name (str) – Name of index column
codelength (int) – openlocationcode code length
handle_duplicates (bool) – If True, adds numeric suffix to duplicate GIDs (default True)
- openplaces.geo.ids.add_ubid_index(gdf, name='ubid', duplicates='raise')#
Return the GeoDataFrame using geo_id as the index
- Parameters:
gdf (GeoDataFrame) – Polygon data
name (str) – Name of index column
duplicates (str) – ‘raise’ or ‘drop’. Duplicate UBID indices are not permitted.
- openplaces.geo.ids.decode_openlocationcodes(codes: list[str] | pandas.Series) geopandas.GeoDataFrame#
Decode level-11 Open Location Codes to a GeoDataFrame of points.
Returns the center of each OLC cell as a Point geometry.
- Parameters:
codes (list[str] or pd.Series) – Sequence of OLC strings with codelength=11 (e.g. ‘85G8Q23G+CFM’). If a Series, its index is preserved in the output.
name (str) – Name for the geometry column (default ‘geometry’).
- Returns:
Points at OLC cell centers, CRS EPSG:4326.
- Return type:
GeoDataFrame
- openplaces.geo.ids.decode_ubids(ubids: pandas.Series, outer: bool = True) geopandas.GeoDataFrame#
Decode a Series of UBIDs to a GeoDataFrame of bounding boxes.
- Parameters:
ubids (pd.Series) – Series of UBID strings in the format ‘OLC-N-E-S-W’.
outer (bool) – If True (default), return the outermost plausible bounding box. If False, return the innermost plausible bounding box (shrinks each side by 0.5 OLC cells to account for rounding).
- openplaces.geo.ids.convert_parcel_id(series: pandas.Series, pattern=None, conv_code: str = 'simple')#
Standardize raw parcel ids into a matching key via a conversion code.
Implements only the operations seen in the auto-selected best solutions:
simple(keep alphanumerics),no_conv,string_lengths,split_groups,drop_cols,keep_length,fill_zeros,switch,merge_after,max_length,join_char,skip_empty.- Parameters:
series (pd.Series) – Raw parcel identifiers (
parcel_id_assessor).pattern (str, optional) – Pattern name (in
parcel_id_patterns.csv) or a raw^...$regex with capture groups. Ignored whenstring_lengthsis in the code.conv_code (str) – Conversion code, e.g.
'string_lengths: 2 2 3 & skip_empty: 1'or the bare'simple'/'no_conv'.
- Returns:
Standardized matching key (
parcel_id_local); NA where extraction failed or the result is empty.- Return type:
pd.Series
- openplaces.geo.ids.dominant_parcel_id_pattern(series: pandas.Series, min_match_ratio: float = 0.5) str#
Return the dominant extraction pattern of raw parcel ids.
Offline helper used to (re)generate the per-admin-unit conversion table; not used on the ingest path. Picks the active pattern with the highest match ratio, breaking ties by lower complexity.
- openplaces.geo.ids.compute_parcel_id_local(series: pandas.Series, admin_unit_id=None, instruction: dict | None = None, kind: str = 'parcel', tolerance: float = 0.005) pandas.Series#
Compute the standardized
parcel_id_localkey for a parcel id column.Resolves the admin-unit-specific conversion (recipe
instructionthen the bundled default table, by sourcekind'parcel'or'tax'), and applies a hardened duplicate guard: if the conversion would collapse distinctparcel_id_assessorvalues beyond tolerance, it falls back tosimpleand then to the raw (uppercased, alphanumeric) id, never adding new duplicates over the source.