The TextProcessor Java package is a text processing toolkit, which provides some frequently used text processing functions such as stemming, removing stop-words, generating a term vocabulary, and calculating the term-doc frequency matrix. Basic topic mining models such as LDA and sparse NMF are also supported. The package can also generate feature files from a given text dataset with LDA and LIBSVM format for posterior procedures such as classification or clustering. The toolkit is also being extended for more advanced text analysis tasks based on natural language processing techniques.

Features

  • Well documented source code
  • Friendly API, very easy to use
  • Topic mining / modeling methods
  • Triplet extraction (NP VP NP triplet)

Project Activity

See All Activity >

Follow TextProcessor

TextProcessor Web Site

Other Useful Business Software
Entity Management Software Icon
Entity Management Software

Filejet’s entity management software organizes & files all of your entity reports: every year, in every jurisdiction – with total visibility.

Automate everything: entity compliance, registered agent services, annual report and BOI filings, org charts, and DBA/fictitious name and business registration renewals, so you can focus on higher-value work.
Learn More
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of TextProcessor!

Additional Project Details

Programming Language

Java

Related Categories

Java Artificial Intelligence Software, Java Natural Language Processing (NLP) Tool

Registered

2013-03-13