get_tokens: Get tokens

View source: R/get.R

get_tokensR Documentation

Get tokens

Description

Get tokens

Usage

get_tokens(
  collection = NULL,
  language = NULL,
  corpus = NULL,
  target_child = NULL,
  role = NULL,
  role_exclude = NULL,
  age = NULL,
  sex = NULL,
  token,
  stem = NULL,
  part_of_speech = NULL,
  replace = TRUE,
  connection = NULL,
  db_version = "current",
  db_args = NULL
)

Arguments

collection

A character vector of one or more names of collections

language

A character vector of one or more languages

corpus

A character vector of one or more names of corpora

target_child

A character vector of one or more names of children

role

A character vector of one or more roles to include

role_exclude

A character vector of one or more roles to exclude

age

A numeric vector of an single age value or a min age value and max age value (inclusive) in months. For a single age value, participants are returned for which that age is within their age range; for two ages, participants are returned for whose age overlaps with the interval between those two ages.

sex

A character vector of values "male" and/or "female"

token

A character vector of one or more token patterns ('%' matches any number of wildcard characters, '_' matches exactly one wildcard character)

stem

A character vector of one or more stems

part_of_speech

A character vector of one or more parts of speech

replace

A boolean indicating whether to replace "gloss" with "replacement" (i.e. phonologically assimilated form), when available (defaults to TRUE)

connection

Deprecated, ignored (childesr now reads from the childes-db dataset on Redivis)

db_version

String of the name of database version to use

db_args

Deprecated, ignored

Value

A 'tbl' of Token data, filtered down by supplied arguments

Identifiers

Numeric ids in childes-db ('transcript_id', 'utterance_id', token 'id', and so on) are internal to a database release: they are not stable across versions of childes-db and should never be used to link data across releases. The TalkBank persistent identifier (the 'pid' column returned by 'get_transcripts()') is the stable, externally-facing identifier for a transcript; use it to match transcripts across database versions or with other TalkBank tools. For reproducible analyses, pin the database version with the 'db_version' argument.

Examples

## Not run: 
get_tokens(token = "dog")

## End(Not run)

childesr documentation built on Sept. 16, 2026, 9:09 a.m.