mlampros/textTinyR: Text Processing for Small or Big Data Files
Version 1.0.9

Processes big text data files in batches efficiently. For this purpose, it offers functions for splitting, parsing, tokenizing and creating a vocabulary. Moreover, it includes functions for building either a document-term matrix or a term-document matrix and extracting information from those (term-associations, most frequent terms). Lastly, it embodies functions for calculating token statistics (collocations, look-up tables, string dissimilarities) and functions to work with sparse matrices. The source code is based on 'C++11' and exported in R through the 'Rcpp', 'RcppArmadillo' and 'BH' packages.

Getting started

Package details

AuthorLampros Mouselimis <[email protected]>
MaintainerLampros Mouselimis <[email protected]>
Package repositoryView on GitHub
Installation Install the latest version of this package by entering the following in R:
mlampros/textTinyR documentation built on Jan. 21, 2018, 10:58 a.m.