Skip to content

Commit 90949aa

Browse files
Merge pull request #408 from ludwiglierhammer/no_dtype_conv
Adjustments for GLAMOD processing scripts using parquet
2 parents b8e1911 + d74661c commit 90949aa

3 files changed

Lines changed: 29 additions & 8 deletions

File tree

‎.pre-commit-config.yaml‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,5 @@
11
default_language_version:
2-
python: python3.13
2+
python: python3.12
33

44
repos:
55
- repo: https://github.com/asottile/pyupgrade

‎CHANGES.rst‎

Lines changed: 15 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -3,9 +3,19 @@
33
Changelog
44
=========
55

6+
2.4.1 (2026-04-07)
7+
------------------
8+
Contributor to this version: Ludwig Lierhammer (:user:`ludwiglierhammer`)
9+
10+
Bug fixes
11+
^^^^^^^^^
12+
13+
* `duplicates`: do not change data types when updating quality flags and history description (:pull:`408`)
14+
15+
616
2.4.0 (2026-04-01)
717
------------------
8-
Contributors to this version: Ludwig Lierhammer (:user:`ludwiglierhammer`), Jan Marius Willruth (:user:`JanWillruth`)
18+
Contributors to this version: Ludwig Lierhammer (:user:`ludwiglierhammer`) and Jan Marius Willruth (:user:`JanWillruth`)
919

1020
New features and enhancements
1121
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
@@ -25,7 +35,7 @@ Breaking changes
2535

2636
* `mdf_reader`/`cdm_mapper`: use parquet as default instead of csv when reading and writing data from/to disk (:pul:`401`)
2737
* `cdm_mapper`: do not convert data types to strings while mapping to the CDM (:issue:`398`, :pull:`401`)
28-
* `cdm_mapper`: set default decimal_places from `0` to `1` for `location_accuracy`, `report_time_accuracy`, `station_speed` and ``station_course` (:pull:`401`)
38+
* `cdm_mapper`: set default decimal_places from `0` to `1` for `location_accuracy`, `report_time_accuracy`, `station_speed` and `station_course` (:pull:`401`)
2939

3040
Bug fixes
3141
^^^^^^^^^
@@ -66,7 +76,7 @@ Breaking changes
6676
* `cdm_mapper.read_tables`
6777
* `cdm_mapper.write_tables`
6878

69-
* set default for `extension` from ``csv` to specified `data_format` in `mdf_reader.write_data` (:pull:`363`)
79+
* set default for `extension` from `csv` to specified `data_format` in `mdf_reader.write_data` (:pull:`363`)
7080
* `mdf_reader.read_data`: save `dtypes` in return DataBundle as `pd.Series` not `dict` (:pull:`363`)
7181
* remove ``common.pandas_TextParser_hdlr`` (:issue:`8`, :pull:`348`)
7282
* ``cdm_reader_mapper`` now raises errors instead of logging them (:pull:`348`)
@@ -126,7 +136,7 @@ Internal changes
126136
^^^^^^^^^^^^^^^^
127137
* implement map_model test for Pub47 data (:issue:`310`, :pull:`327`)
128138
* rename test data class from test_data to TestData (:pull:`327`)
129-
* update .gitignore (:pull:``324`)
139+
* update .gitignore (:pull:`324`)
130140
* update and add docstrings for multiple functions (:pull:`324`)
131141
* ``cdm_reader_mapper.cdm_mapper``: update mapping functions for more readability (:pull:`324`)
132142
* ``cdm_reader_mapper.cdm_mapper``: introduce some helper functions (:pull:`324`)
@@ -345,7 +355,7 @@ Internal changes
345355
* ``cdm_mapper.codes.common``: convert range-key properties to list (:pull:`221`)
346356
* ``testing_suite``: new chunksize test with icoads_r300_d721 (:pull:`222`)
347357
* ``mdf_reader``, ``cdm_nmapper``: use model-depending encoding while writing data on disk (:pull:`222`)
348-
* code restructuring (:pull:``224`)
358+
* code restructuring (:pull:`224`)
349359
* remove unused functions and methods (:pull:`224`)
350360

351361

‎cdm_reader_mapper/duplicates/duplicates.py‎

Lines changed: 13 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -131,7 +131,10 @@ def _add_dups(row):
131131
df["duplicates"] = ""
132132

133133
report_ids = df["report_id"]
134-
return df.apply(lambda x: _add_dups(x), axis=1)
134+
135+
dtypes = df.dtypes
136+
result = df.apply(lambda x: _add_dups(x), axis=1)
137+
return result.astype(dtypes)
135138

136139

137140
def add_report_quality(df: pd.DataFrame, indexes_bad: Iterable[int]) -> pd.DataFrame:
@@ -356,6 +359,9 @@ def replace_keeps_and_drops(df, keep_):
356359

357360
self.get_duplicates(keep=keep, limit=limit, equal_musts=equal_musts)
358361
result = self.data.copy()
362+
363+
dtypes = result.dtypes
364+
359365
result["duplicate_status"] = 0
360366
if not hasattr(self, "matches"):
361367
self.get_matches(limit="default", equal_musts=equal_musts)
@@ -385,7 +391,9 @@ def replace_keeps_and_drops(df, keep_):
385391
result = add_report_quality(result, indexes_bad=indexes_bad)
386392
result = add_history(result, indexes)
387393
result = result.sort_index(ascending=True)
388-
self.result = add_duplicates(result, duplicates)
394+
result = add_duplicates(result, duplicates)
395+
396+
self.result = result.astype(dtypes)
389397
self.data = self.data.sort_index(ascending=True)
390398

391399
return self.result
@@ -678,6 +686,8 @@ def duplicate_check(
678686
if offsets:
679687
compare_kwargs = change_offsets(compare_kwargs, offsets)
680688

689+
dtypes = data.dtypes
690+
681691
Compared_ = Comparer(
682692
data=data,
683693
method=method,
@@ -720,4 +730,5 @@ def duplicate_check(
720730

721731
compared = pd.concat(compared)
722732
data.set_index(index, inplace=True)
733+
data = data.astype(dtypes)
723734
return DupDetect(data, compared, method, method_kwargs, compare_kwargs)

0 commit comments

Comments
 (0)