Make text_duplicates list the duplicates it counts - #4
Closed
tonycoder-hub wants to merge 1 commit into
Closed
Conversation
duplicate_fraction is computed from hashes of stripped strings, so strings that differ only in surrounding whitespace count as duplicates. duplicates_dict was built from the raw strings, so a single call could report a non-zero duplicate fraction together with an empty duplicates_dict. Co-authored-by: Tony Coder <407243179@qq.com>
Owner
Author
|
Closing as stale — opened on or before 2026-08-17 and still unmerged. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
text_duplicatescan return two fields that contradict each other in a single call:duplicate_fractionis derived fromget_hash, which strips surrounding whitespace, so"hello sun"and"hello sun "are counted as duplicates.duplicates_dict, however, was built withCounter(data)over the raw strings, so those same duplicates never show up in the listing. The caller is told that a third of the data is duplicated and simultaneously that there are no duplicates.Fix
Count the duplicates over the same stripped strings that the fraction is based on:
Inputs without stray whitespace are unaffected, so the documented examples and the module doctests keep their current output.
Testing
Added
test_text_duplicates_lists_the_duplicates_it_countsintests/test_metric_common.py, next to the existing per-module test forseqeval. It fails onmain(assert {} == {'hello sun': 2}) and passes with the fix.Wider offline run of
tests/test_metric_common.py tests/test_metric.py tests/test_save.py tests/test_file_utils.py: 26 failed / 69 passed with the change versus 26 failed / 68 passed without it. The 26 failures are pre-existingImportErrors for optional metric dependencies that are not installed in this environment (nltk, sacrebleu, jiwer, torch, ...) and are identical before and after.