General Statistics for a Character Vector

Description

This function gives general statistics for a character vector, e.g. obtained by loading a text file with the readLines or stri_read_lines function, where each text line' is represented by a separate string.

Usage

1

Arguments

str

character vector to be aggregated

Details

Any of the strings must not contain \r or \n characters, otherwise you will get at error.

Below by 'white space' we mean the Unicode binary property WHITE_SPACE, see stringi-search-charclass.

Value

Returns an integer vector with the following named elements:

  1. Lines - number of lines (number of non-missing strings in the vector);

  2. LinesNEmpty - number of lines with at least one non-WHITE_SPACE character;

  3. Chars - total number of Unicode code points detected;

  4. CharsNWhite - number of Unicode code points that are not WHITE_SPACEs;

  5. ... (Other stuff that may appear in future releases of stringi).

See Also

Other stats: stri_stats_latex

Examples

1
2
3
4
5
s <- c("Lorem ipsum dolor sit amet, consectetur adipisicing elit.",
       "nibh augue, suscipit a, scelerisque sed, lacinia in, mi.",
       "Cras vel lorem. Etiam pellentesque aliquet tellus.",
       "")
stri_stats_general(s)

Want to suggest features or report bugs for rdrr.io? Use the GitHub issue tracker.