Detects near-duplicate images across dataset splits using perceptual hashing, reports the resulting train/validation/test contamination, and produces a corrected, leak-free split assignment. Intended for machine learning researchers who need to verify that image classification splits do not share near-duplicate samples across partitions before reporting model metrics.
Detects near-duplicate images across train/validation/test splits using perceptual hashing (dHash), reports the resulting leakage, and produces a corrected, leak-free split assignment.
When building an image classification dataset, near-duplicate images (same photo re-saved, resized, lightly cropped, or recompressed) commonly end up split across train and test. This silently inflates test-set performance, since the model has effectively already seen a near-identical copy of the "unseen" example.
# development version
remotes::install_github("anakincodex/leakaudit")
# from CRAN, once published
install.packages("leakaudit")
library(leakaudit)
hashes <- compute_hashes(image_paths, split = split_labels)
grouped <- find_duplicate_groups(hashes, threshold = 5)
report <- dhash_audit(grouped)
print(report)
clean <- clean_splits(grouped, priority = c("train", "val", "test"))
A few CRAN packages deal with "leakage" in machine learning workflows, but none of them work at the image level:
dhash(), phash()) but has no concept of dataset splits or leakage
reporting.leakaudit is specifically about auditing and fixing near-duplicate
contamination across image dataset splits, and its dhash_audit() function
is unrelated to bioLeak::audit_leakage(), which audits fitted models on
tabular data via permutation testing.