| getvocab | R Documentation |
Extract words and phrases from a corpus of documents.
getvocab(
corpus,
mincount = 5,
minphrasecount = NULL,
ngram = 1,
lang = "en",
stopwords = lang,
excludewords = NULL,
removesinglechars = TRUE,
...
)
corpus |
The corpus of documents (a vector of characters). |
mincount |
Minimum word count to be considered as frequent. |
minphrasecount |
Minimum collocation of words count to be considered as frequent. |
ngram |
maximum size of n-grams. |
lang |
The language of the documents (NULL if no stemming). |
stopwords |
The language whose stop words are removed ( |
excludewords |
An optional custom vector of additional words to exclude from the vocabulary (e.g. corpus-specific stop words), on top of (or instead of) the language stopwords given through |
removesinglechars |
Whether single-character tokens are removed during cleanup. |
... |
Other parameters. |
The vocabulary used in the corpus of documents.
plotzipf, stopwords, create_vocabulary
data (capitals)
vocab1 = getvocab (capitals, mincount = 2) # With stemming
nrow (vocab1)
vocab2 = getvocab (capitals, mincount = 2, lang = NULL) # Without stemming
nrow (vocab2)
# Excluding additional, corpus-specific words
vocab3 = getvocab (capitals, mincount = 2, excludewords = c ("capital", "europe"))
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.