Skip to content

Make text_duplicates list the duplicates it counts - #4

Closed
tonycoder-hub wants to merge 1 commit into
mainfrom
cursor/fix-text-duplicates-dict-consistency-a3f1
Closed

Make text_duplicates list the duplicates it counts#4
tonycoder-hub wants to merge 1 commit into
mainfrom
cursor/fix-text-duplicates-dict-consistency-a3f1

Conversation

@tonycoder-hub

Copy link
Copy Markdown
Owner

Problem

text_duplicates can return two fields that contradict each other in a single call:

>>> from evaluate import load
>>> load("text_duplicates").compute(data=["hello sun", "hello sun ", "hello moon"], list_duplicates=True)
{'duplicate_fraction': 0.33333333333333337, 'duplicates_dict': {}}

duplicate_fraction is derived from get_hash, which strips surrounding whitespace, so "hello sun" and "hello sun " are counted as duplicates. duplicates_dict, however, was built with Counter(data) over the raw strings, so those same duplicates never show up in the listing. The caller is told that a third of the data is duplicated and simultaneously that there are no duplicates.

Fix

Count the duplicates over the same stripped strings that the fraction is based on:

c = Counter(d.strip() for d in data)

Inputs without stray whitespace are unaffected, so the documented examples and the module doctests keep their current output.

Testing

Added test_text_duplicates_lists_the_duplicates_it_counts in tests/test_metric_common.py, next to the existing per-module test for seqeval. It fails on main (assert {} == {'hello sun': 2}) and passes with the fix.

python -m pytest tests/test_metric_common.py -k text_duplicates
# 2 passed, 1 skipped (test_load_measurement_text_duplicates doctest + new test)

Wider offline run of tests/test_metric_common.py tests/test_metric.py tests/test_save.py tests/test_file_utils.py: 26 failed / 69 passed with the change versus 26 failed / 68 passed without it. The 26 failures are pre-existing ImportErrors for optional metric dependencies that are not installed in this environment (nltk, sacrebleu, jiwer, torch, ...) and are identical before and after.

Open in Web Open in Cursor 

duplicate_fraction is computed from hashes of stripped strings, so strings
that differ only in surrounding whitespace count as duplicates. duplicates_dict
was built from the raw strings, so a single call could report a non-zero
duplicate fraction together with an empty duplicates_dict.

Co-authored-by: Tony Coder <407243179@qq.com>
@tonycoder-hub

Copy link
Copy Markdown
Owner Author

Closing as stale — opened on or before 2026-08-17 and still unmerged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants