Redundancy due to cut-paste operations in text creates bias in machine learning for NLP.
This module takes a directory and produces a subset of the files in that directory (in a list) with an upper bound on similarity between two files.
Features
- Identify copy paste redundancy in a document corpus
- Input: a folder with text documents and similarity threshold
- Output (a) a list of non-redundant documents (a non-redundant subset of the corpus)
- Output (b) list of document pairs found to be redundant with the amount of redundancy for the pair
- Python script (2.6) - tested on various Linux flavours + Windows XP/7
License
GNU General Public License version 3.0 (GPLv3)Follow Corpus redundancy manager
Other Useful Business Software
AestheticsPro Medical Spa Software
AestheticsPro is the most complete Aesthetics Software on the market today. HIPAA Cloud Compliant with electronic charting, integrated POS, targeted marketing and results driven reporting; AestheticsPro delivers the tools you need to manage your medical spa business. It is our mission To Provide an All-in-One Cutting Edge Software to the Aesthetics Industry.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of Corpus redundancy manager!