Skip to content

Templater property maps: where and how to enforce/set/change dtypes? #145

Description

@EllieKallmier

ATM - for generator/storage units, in the templater *_PROPERTY_MAP dicts (templater/mappings.py) you can only set numeric=True|False to give explicit information about the dtype of a property. Then all numeric properties become float by pd.to_numeric() and a float scaling factor; this means that 'year' values also becomes floats, e.g. '2035.0'.

Some numeric values e.g. years would be nice to keep as ints, and as @nick-gorman suggested in a comment in the review of #143 we could add a 'dtype=...' field to the property maps to allow other types like int/date/str etc to be explicitly set, e.g.:

_GENERATORS_EXISTING_PLANNED_PROPERTY_MAP = {
    ...
    "closure_year": dict(
        table="expected_closure_years",
        key_col="IASR ID",
        value_col="Expected Closure Year (Calendar year)",
        dtype="int"
    ),
    ....
}

My thoughts and sticking points:

  • NaNs. Converting some values in a column to int64/int32 etc doesn't work with NaN values in the column, which is allowed and expected in many cases.
  • Round-trip on CSV write/read after templating - if columns retain NaNs (which they can and likely will in some cases) an int-type will get lost again
  • Potential for lots of repeated dtype-handling if we add something similar in the validation step(s) enforcing column dtypes (and maybe makes more sense to do wild-card expansion/filling first as well? Not sure)
  • Currently the *_PROPERTY_MAP dicts tell you about what the input looks like and how to handle it - the 'numeric' field is True by default and only False where strings are explicitly expected, and the pd.to_numeric() use is an easy guard against typos/unexpected values otherwise. The maps don't necessarily tell you about what the values will look like after being templated (atm)
  • It's not very many columns that would really benefit from this (in generator/storage templating - 'closure_year') -> unless we also handle date-string formatting e.g. for 'commissioning_date' columns (but then I think we might need to add a 'format' field too, maybe?) or other similar string format style inputs (?)

Potential solutions:

  • Use pandas extension dtypes like Int64 that are nullable (good option for saving CSV results that look nicer, i.e. '2035' instead of '2035.0') -> and include something on the read/write CSV boundary if type enforcement there feels important?
  • Don't do anything in the templater mappings and just do type enforcing/setting at validation or translation boundaries (or elsewhere that's not at this templater step lol)
  • Other ideas?

***Despite the amount of text here I'm not that fussed either way if anyone else has a strong opinion!!

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    category: data-validationRelates to data validation practices across any module - e.g tables, schema or enforcementtype: questionFurther information is requested

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions