principal_component_analysis: Principal Component Analysis

View source: R/pca.R

principal_component_analysisR Documentation

Principal Component Analysis

Description

A user-friendly wrapper around PCA, made to work with ggfacto functions like interpret, ggfacto and hierarchical_clust. Variables are selected the way of the 'tidyverse', as in tabxplor::tab(). 'PCA2()' is its name in ggfacto 0.3.2, kept for former code.

Usage

principal_component_analysis(
  data,
  active_vars,
  wt,
  col.w = NULL,
  ind_name,
  scale.unit = TRUE,
  ind.sup = NULL,
  ncp = Inf,
  graph = FALSE,
  na = "mean",
  filter,
  ...
)

PCA2(
  data,
  active_vars,
  wt,
  col.w = NULL,
  ind_name,
  scale.unit = TRUE,
  ind.sup = NULL,
  ncp = Inf,
  graph = FALSE,
  na = "mean",
  filter,
  ...
)

Arguments

data

The data frame. To analyse a subset of the population, give the whole data frame and 'filter', or filter it inside the call with the native pipe, 'data |> dplyr::filter(...) |> principal_component_analysis(...)': the analysis then remembers which rows it used, so that ggpca, hierarchical_clust or is_in_analysis can be given the whole data frame afterwards.

active_vars

<tidy-select> The names of the active variables.

wt

<tidy-select> The weight variable, if any.

col.w

The weights of the columns, as a numeric vector of the same length than 'active_vars.'

ind_name

<tidy-select> Possibly, the variable holding the names of the individuals.

scale.unit

A boolean, if 'TRUE' (value set by default) then data are scaled to unit variance.

ind.sup

A vector indicating the indexes of the supplementary individuals, rows of 'data'.

ncp

Number of dimensions kept in the results. All of them by default: the eigenvalue table is how one chooses how many axes to interpret, and a truncated one cannot show the drop. To cluster on the first axes, give hierarchical_clust its own 'ncp'.

graph

A boolean, set to 'TRUE' to display the base graph.

na

How missing values of the active variables are treated. '"mean"', the default, places each one at its variable's weighted mean, where it adds nothing to the axes (as an excluded level does in a specific multiple correspondence analysis); the tooltips and interpret count them. '"drop"' leaves out the rows with a missing value.

filter

A condition on the rows of 'data', as in dplyr::filter(): only the rows where it is 'TRUE' are analysed ('filter = AGE >= 18'). Supplementary individuals are kept.

...

Additional arguments to pass to PCA.

Value

A 'PCA' object from FactoMineR, with one more element, 'source', which records for each row of 'data' its row in the analysis ('NA' if it was not analysed).

Examples

cars <- dplyr::mutate(mtcars, cyl = factor(cyl))
res.pca <- principal_component_analysis(cars, c(mpg, disp, hp, drat, wt, qsec))
interpret(res.pca)                           # the eigenvalues, then the axes

ggfacto(res.pca, cars, sup_vars = cyl)       # the individuals and the variables (biplot)
ggfacto(res.pca, profiles = FALSE)           # the circle of correlations alone
ggfacto(res.pca, cars, sup_vars = cyl, interactive = TRUE)  # hover: the means


ggfacto documentation built on Sept. 23, 2026, 1:08 a.m.