Skip to content

fix(dtypes): the report no longer calls a datetime parse "converted to object" - #508

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/fd2015-audit-record
Sep 28, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/fd2015-audit-record

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Bug

A column of datetime strings with different UTC offsets is common: a daylight-saving changeover, or exports from two regions merged together. fd.clean parses it correctly, cell by cell, into timestamps that each keep their own offset. No single datetime64 dtype can hold several offsets, so pandas returns them in an object column. The data is right. The audit record is wrong:

vals = ["2021-01-05 00:00:00+01:00", "2021-01-06 00:00:00+02:00",
        "2021-01-07 00:00:00+03:00", "2021-01-08 00:00:00+04:00"]
fd.clean(pd.DataFrame({"t": vals}))
# report action: "converted to object", risk low

Read plainly, "converted to object" says nothing was converted, when every cell was in fact parsed to a timestamp. fix_dtypes built the description from the resulting dtype alone and ignored the conversion target.

Fix

The description now says what happened:

parsed to timestamps; kept as object because the values carry different UTC offsets, so no single datetime64 dtype can hold them

A column mixing timezone-aware and naive values takes the same path; it gets its own reason: "… mix timezone-aware and timezone-naive timestamps …". The coercion warning for such a column now says "could not be parsed as timestamps" rather than "as object". The existing (N unparseable value(s) set to missing) suffix is unchanged.

Report text only. Cell values, dtypes, risk and confidence are unchanged, and were checked byte-for-byte against main over 12 cases on both interpreters. Every other column keeps its converted to <dtype> description, including single-offset and naive datetime columns.

The sibling "converted to … after semantic repair" description cannot receive such a column, because refine_numeric_after_semantic only produces real numeric dtypes. A test pins that.

Tests

  • tests/test_mixed_offset_audit_record.py, 20 tests. Controls check that single-offset, naive, numeric and bool descriptions are exactly as before, that the cell values are unchanged, and that the unparseable suffix still appears.
  • tests/test_dtypes_threshold_mutants.py: one test pinned the literal "converted to object". It now pins the target, dtype and casualty count instead.
  • 12 fail without the src/ change, on both py3.12 / pandas 2.3.3 and py3.9 / pandas 1.5.3. All pass with it (3 pre-existing pandas ≥ 2 skips on py3.9).
  • Full suite: py3.12 7559 passed, 0 failed. py3.9: 7526 passed, 1 failed. The failure (test_build_engine_cache_called_once_per_clean) is a pre-existing order-dependent test, reproduced on unmodified main: it fails on py3.9 whenever test_plugins.py::…test_import_freshdata_does_not_import_testing runs first. CI's fixed test order hides it. It is unrelated to this change.

Not changed here

fd.profile still previews such a column as would convert to object — the same misleading text on a different surface.

Compatibility impact

Report text only. Code that matched the literal converted to object for such a column needs to match the new description.

@coderabbitai

coderabbitai Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 0f5ca460-f8b2-4e60-9a21-6e857ca9ccfc


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

…o object"

A column of datetime strings carrying different UTC offsets -- the ordinary
result of a DST changeover or of merging exports from two regions -- is
parsed cell by cell into timestamps that each keep their own offset. No
single datetime64 dtype can hold several offsets, so pandas returns them in
an object column. The parse is faithful and lossless, but fix_dtypes built
its description from the dtype alone and recorded "converted to object", risk
low -- which reads as though nothing was converted.

The description now says what happened ("parsed to timestamps; kept as
object because the values carry different UTC offsets, so no single
datetime64 dtype can hold them", or "... mix timezone-aware and
timezone-naive timestamps"), and the coercion warning says "could not be
parsed as timestamps" instead of "as object". Report text only: cell values,
dtypes, risk and confidence are unchanged (checked byte-for-byte against
main over 12 cases), as is every other column's description.

The "after semantic repair" sibling description cannot receive such a column
(refine_numeric_after_semantic only produces real numeric dtypes); a test
pins that.

tests/test_mixed_offset_audit_record.py: 20 tests. With the updated
threshold test, 12 fail without this change on both py3.12/pandas 2.3.3 and
py3.9/pandas 1.5.3.
@kevincostner17
kevincostner17 merged commit 887e0ae into main Sep 28, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant