ATM - for generator/storage units, in the templater *_PROPERTY_MAP dicts (templater/mappings.py) you can only set numeric=True|False to give explicit information about the dtype of a property. Then all numeric properties become float by pd.to_numeric() and a float scaling factor; this means that 'year' values also becomes floats, e.g. '2035.0'.
Some numeric values e.g. years would be nice to keep as ints, and as @nick-gorman suggested in a comment in the review of #143 we could add a 'dtype=...' field to the property maps to allow other types like int/date/str etc to be explicitly set, e.g.:
_GENERATORS_EXISTING_PLANNED_PROPERTY_MAP = {
...
"closure_year": dict(
table="expected_closure_years",
key_col="IASR ID",
value_col="Expected Closure Year (Calendar year)",
dtype="int"
),
....
}
My thoughts and sticking points:
- NaNs. Converting some values in a column to
int64/int32 etc doesn't work with NaN values in the column, which is allowed and expected in many cases.
- Round-trip on CSV write/read after templating - if columns retain NaNs (which they can and likely will in some cases) an int-type will get lost again
- Potential for lots of repeated dtype-handling if we add something similar in the validation step(s) enforcing column dtypes (and maybe makes more sense to do wild-card expansion/filling first as well? Not sure)
- Currently the *_PROPERTY_MAP dicts tell you about what the input looks like and how to handle it - the 'numeric' field is True by default and only False where strings are explicitly expected, and the pd.to_numeric() use is an easy guard against typos/unexpected values otherwise. The maps don't necessarily tell you about what the values will look like after being templated (atm)
- It's not very many columns that would really benefit from this (in generator/storage templating - 'closure_year') -> unless we also handle date-string formatting e.g. for 'commissioning_date' columns (but then I think we might need to add a 'format' field too, maybe?) or other similar string format style inputs (?)
Potential solutions:
- Use pandas extension dtypes like
Int64 that are nullable (good option for saving CSV results that look nicer, i.e. '2035' instead of '2035.0') -> and include something on the read/write CSV boundary if type enforcement there feels important?
- Don't do anything in the templater mappings and just do type enforcing/setting at validation or translation boundaries (or elsewhere that's not at this templater step lol)
- Other ideas?
***Despite the amount of text here I'm not that fussed either way if anyone else has a strong opinion!!
ATM - for generator/storage units, in the templater *_PROPERTY_MAP dicts (
templater/mappings.py) you can only setnumeric=True|Falseto give explicit information about the dtype of a property. Then all numeric properties become float by pd.to_numeric() and a float scaling factor; this means that'year'values also becomes floats, e.g.'2035.0'.Some numeric values e.g. years would be nice to keep as ints, and as @nick-gorman suggested in a comment in the review of #143 we could add a
'dtype=...'field to the property maps to allow other types like int/date/str etc to be explicitly set, e.g.:My thoughts and sticking points:
int64/int32etc doesn't work with NaN values in the column, which is allowed and expected in many cases.Potential solutions:
Int64that are nullable (good option for saving CSV results that look nicer, i.e. '2035' instead of '2035.0') -> and include something on the read/write CSV boundary if type enforcement there feels important?***Despite the amount of text here I'm not that fussed either way if anyone else has a strong opinion!!