knitr::opts_chunk$set( collapse = TRUE, comment = "#>", warning = FALSE ) data.table::setDTthreads(2)
library(OpenSpecy)
This example combines a few files bundled with OpenSpecy into a small reference library. Real library builds usually need larger lookup tables and more curation, but the same helper functions apply.
c_spec() and build_lib() default to the widest range represented by the
source spectra at resolution 6. Values outside each source's original range are
kept as NA, so useful spectral regions are not discarded.
This compact vignette example uses the shared overlapping range to keep the
rendered output small.
mini_files <- c( read_extdata("raman_hdpe.csv"), read_extdata("ftir_ldpe_soil.asp"), read_extdata("raman_atacamit.spc") ) mini_sources <- lapply(mini_files, read_any) mini_sources <- lapply(mini_sources, function(x) { x$metadata$intensity_units <- "absorbance" attr(x, "intensity_unit") <- "absorbance" x }) mini_raw <- c_spec(mini_sources, range = "common", res = 6) check_OpenSpecy(mini_raw) dim(mini_raw$spectra) mini_raw$metadata[, "file_name", with = FALSE]
Lookup templates help users see which metadata values need curation. When
path is not supplied, the template is returned as a data.table; users can
write it to CSV by supplying path. Lookup joins use exact values, so edit the
template until the key column matches the object metadata.
template <- make_lib_lookup_template( mini_raw, columns = "file_name", add = c("material", "material_type") ) template
For this small example, create the lookup table directly in R. The same shape could come from a CSV edited outside R.
lookup <- data.table::data.table( file_name = basename(mini_files), library_type = c("example", "example", "example"), spectrum_type = c("raman", "ftir", "raman"), material = c("hdpe", "ldpe in soil", "atacamite") ) hierarchy <- data.table::data.table( material = c("hdpe", "ldpe in soil", "atacamite"), material_class = c("polyethylene", "polyethylene", "copper mineral"), material_type = c("plastic", "plastic", "mineral") ) join_lib_metadata(mini_raw, lookup, by = "file_name", require_complete = TRUE)$metadata[ , c("file_name", "library_type", "spectrum_type", "material"), with = FALSE ]
build_lib() can run the ordinary lookup and material hierarchy joins before
applying named recipes. Recipe names become names in the returned list. Empty
recipes keep the merged spectra unchanged; other recipe lists are passed to
process_spec(). Missing values and processing attributes are handled
automatically. Metadata column names are also cleaned to lowercase underscore
names. Known aliases are coalesced using an editable lookup table. Variants
that differ only by underscores or one terminal plural s match
automatically.
Before sources are merged, build_lib() also converts declared reflectance and
transmittance spectra to absorbance by default. A nonempty
attr(x, "intensity_unit") is the primary truth for the whole object;
otherwise, metadata$intensity_units is evaluated spectrum by spectrum.
Unknown or missing units are left unchanged with a warning. Use
convert_intensity = FALSE when unit handling has already been completed
outside the builder.
The normal input is one or more file paths; one OpenSpecy or a list of
OpenSpecy objects supports sources already loaded in memory. An RDS path may
store either one object or a list of objects. Other formats are read with
read_any(). Large same-axis source lists are prepared in bulk, which is much
faster for legacy RDS files that store many one-spectrum objects. Named stages
and elapsed time are reported while the library is built; use
progress = FALSE when quiet output is preferred. Supplying
restrict_range_args triggers the existing restrict_range() operation before
deduplication and recipes; multiple retained ranges can exclude a known silent
region without custom workflow code.
name_lookup <- lib_metadata_name_lookup( project_code = c("campaign id", "study code"), regex = list(instrument_mode = "^method_[0-9]+$") ) name_lookup[ canonical_name %in% c("material_color", "number_of_accumulations") ] lib_clean_name(c("User Name", "Laser (%)", "Method...3"))
Named arguments add exact aliases to the defaults, while regex adds patterns
evaluated against cleaned names. Overlapping regex patterns produce an error
that identifies the source column and matching rules. Pass the result as
metadata_name_lookup. Set clean_metadata_values = TRUE in build_lib()
or clean_values = TRUE in lib_clean_metadata() to lowercase, trim, and
ASCII-normalize character metadata values before joins. Ordinary and
hierarchical joins run whenever their corresponding lookup input is non-NULL.
spectrum_identity receives one additional deterministic cleanup: recognizable
paths are reduced to their basename and trailing file extensions supported by
read_any() are removed. Numeric OPUS extensions include .10 and any other
terminal period followed only by digits. Exact lookup keys receive the same
cleanup. Each built library records changed values and counts in its
spectrum_identity_cleanup_report attribute. Keep flexible class patterns in a
separate table and apply them afterward with predict_class_reference().
Automatic ordinary lookups use the single shared column with overlapping values
and unique lookup keys. Lookups with no usable shared key are skipped with a
message; lookups with multiple usable shared keys are treated as ambiguous, so
wrap a lookup as list(lookup = table, by = "key") when an explicit key is
needed. Named by vectors map metadata names to different lookup names.
Canonical source metadata is standardized before these external joins. Every
spectrum receives library_name from a populated organization, otherwise
from user_name; this is the canonical source-library grouping used by build
assessments. A blank organization is still filled from the reviewed internal
user_name alias for the existing type lookup. The older lookup-level
fallback_by field is deprecated. Set fill_only = TRUE to fill blank values
such as library_type or spectrum_type without replacing populated source
metadata.
mini_libs <- build_lib( mini_files, recipes = list( raw = list(), derivative = list( conform_spec = FALSE, smooth_intens = TRUE, smooth_intens_args = list(window = 15, derivative = 1), make_rel = TRUE ), nobaseline = list( conform_spec = FALSE, smooth_intens = FALSE, subtr_baseline = TRUE, make_rel = TRUE ) ), metadata_lookups = lookup, material_hierarchy = hierarchy, clean_metadata_values = TRUE, convert_intensity = FALSE, assess = TRUE, dedupe = FALSE ) names(mini_libs) check_OpenSpecy(mini_libs$raw) check_OpenSpecy(mini_libs$derivative) attr(mini_libs$derivative, "derivative_order") attr(mini_libs$nobaseline, "baseline") mini_libs$raw$metadata[ , .(file_name, material, material_class, material_type, sn, assessment_flag, assessment_checks) ]
prune_lib() performs the spectrum-supported cleanup used before medoid or
model creation. With cross_class = TRUE, it first flags correlations strictly
above cross_class_threshold between different reviewed classes within each
source library. The spectrum with the most active wrong-class neighbors is
removed first and scores are recalculated; adjacent maximum-score ties remove
both endpoints. A second pass then uses the number of independent opposing
libraries, removing the higher-evidence endpoint or both on a tie. Generic and
unclassified labels are excluded. The audit records phase, view, component,
identities/classes/libraries, active degree or independent evidence, decision
round, threshold, and reason. Within the same
spectral-technique pool, other may then be
reassigned to the nearest established class, other plastic only to a
candidate whose material_type is plastic, and other material only to
organic matter or mineral. The matched material_type is copied with the
class. It then evaluates material classes from largest to smallest using
bounded correlation blocks. FTIR and NIR share a candidate pool; Raman is
separate; 2200--2420 cm^-1^ is excluded by default. After generic reassignment,
each spectrum_type and material-class group must contain at least min_n
spectra across the complete input database; min_n is not applied separately
to each source library. Smaller resolved groups are reassigned as a whole when an eligible
destination exists and removed only when no destination exists; groups exactly
at the threshold are retained whole and larger groups are never reduced below
it. Constant or unmatched spectra in eligible classes are retained, and
deterministic IDs resolve correlation ties. Use return = "report" for the
retained IDs, frozen class schedule, reassignment correlations, spectrum-level
removals, and excluded_classes, which reports observed support and the number
of additional or reassigned spectra needed. End-to-end builds combine these
class rows in assessments$cleanup$summary with the recipe name.
build_lib(prune = ...) maps recipe names to prune_lib() argument lists, so
each selected spectral representation is pruned independently. Resolved
class/type groups below min_n are reassigned as a whole to the
most-correlated established class in the same technique pool and material type;
they are dropped only when no eligible correlated destination exists. The
pruning assessment records the destination, mean class correlation, action,
and reason. In the official
workflow, an otherwise unresolved identity is temporarily labeled other.
The default remove_other = TRUE removes blank identities and the unresolved
literal other class before quality control. Reviewed broad other plastic
and other material categories stay in the reference libraries and remain
eligible for constrained reassignment during derivative/no-baseline pruning.
The builder keeps reviewed identifiers, source metadata, prior labels, reasons,
actions, and typed before/after counts in assessments$cleanup$summary. The
separate assessments$cleanup$dropped_spectrum_identities table contains only
the sorted distinct source identities removed anywhere in cleanup. Set
remove_other = FALSE to retain generic rows for the
constrained semisupervised prune_lib() reassignment described above.
pruned <- prune_lib( mini_libs$derivative, min_n = 1, return = "report", progress = FALSE ) pruned$summary pruned$schedule pruned$excluded_classes
The composable default leaves the new first pass off. Enable it explicitly for
a library carrying canonical library_name metadata; the official derivative
and no-baseline recipes do this by default:
prune_lib( reference_library, cross_class = TRUE, cross_class_threshold = 0.9 )
The version-controlled
workflows/OpenSpecy_reference_library.R
script retains the source and output path finders and passes those paths to one
build_lib() call. The complete visual flow,
including checkpoint reuse and old/new assessments, is maintained in
.specify/memory/build-lib-diagram.html.
Canonically named exact-class, regex-class, library-type, material-hierarchy,
known-bad-ID, material-form-regex, and common-use tables live under
workflows/data/. Raw-data corrections are completed externally before this
workflow runs. The source-library paths and output directory are explicit
inputs. When workflow_data is not supplied, build_lib() looks for the
helper tables under data/ beside the calling script, then under data/ or
workflows/data/ in the current working directory. Calling build_lib() with
no x now stops with an actionable error instead of guessing external paths.
The official build derives library_name from organization first and user name
second, fills a missing organization from user_name before one source-type
lookup, and asserts complete library/spectrum-type coverage. Per-recipe
retention rows in assessments$cleanup$summary report counts at preparation,
exclusion, deduplication, broad-category review, quality control, pruning,
post-transform, and final partition stages. A completely dropped source names
the first empty stage and its reason. The
exact class lookup runs first;
predict_class_reference() then evaluates the separate regex table only for
blank materials. Exact/regex overlaps are audited without overwriting exact
values, and conflicting regex predictions stop for review. material_class
describes chemistry rather than physical form. Every reviewed standard plastic
class begins with poly for discoverability (other plastic is the explicit
catch-all exception); polyethylene and polypropylene remain separate, and
polyhydroxy(meth)acrylates retains its chemically meaningful optional group
notation. Poly(styrene-butadiene) rubber and
poly(ethylene-propylene-diene) rubber have distinct classes. Paint binder labels
such as acrylic, alkyd, urethane, nitrocellulose, and vinyl combinations remain
in spectrum_identity, paint remains in material_form, and the hierarchy
maps each reviewed binder to its relevant chemical polymer family. Remaining
uncertain identities follow the selected generic-row review policy.
The official workflow also standardizes two enrichment fields. It normalizes
and concatenates all atomic metadata values once per spectrum, then evaluates
the reviewable material_form_regex.csv. A uniquely matching existing
material_form takes precedence; otherwise the full metadata row is searched.
Cross-category matches remain NA and are retained in the clash audit. The
controlled vocabulary currently covers paint, rubber, hard plastic, fiber,
pellet, film plastic, fragment, foam, and sphere/bead.
common_use_reference.csv is a material-class lookup for the proposed
predominant global ultimate end-market. Quantitative rows retain consumer,
industrial, and unresolved mass shares: consumer and industrial require
more than 50%, while mixed requires at least 25% on each side, at least 75%
classified, and neither side above 50%. Where comparable mass shares do not
exist, a qualitative proposal is allowed only with a cited application source,
retrieval date, and explicit review note; share cells remain blank. Consumer
endpoints include products used directly by individuals, such as passenger
tires, household goods, clothing, and patient-facing healthcare products.
Classes with substantial consumer and industrial endpoints use mixed.
Ambiguous catch-alls and poorly supported specialty groups remain NA.
Commodity shares use the OECD 2019 polymer-by-application mass
dataset;
PTFE uses a separate material-specific end-use source recorded in the CSV.
Coverage, provenance, matched evidence, and clashes are retained with the
upstream build assessments. Derivative and
nobaseline spectra pass three quality gates before pruning: FTIR CO2 is
flattened when the 2200--2420 maximum divided by the 2420--2550 silent-region
maximum is greater than two; high-tail detection ignores NA padding, trims each
spectrum on its finite support, and drops failed corrections; and finite
running SNR below two (or unavailable SNR) is removed. Same-library majority
closure then precedes independent-library evidence, generic reassignment, and
conventional top-match pruning. After rounding and type partitioning, the full
and model-range views are closed again. Raw remains unpruned. Every spectrum
excluded by cross-class closure is preserved in the review quarantine.
Each completed library, medoid, model, and assessment component is written with
a SHA-256 manifest under output_dir/checkpoints. reuse = TRUE loads it only when
the source/lookup signatures, arguments, package version, and builder contract
still match. Validated output is promoted into a versioned release directory,
which contains the seven legacy RDS names, explicit algorithm-named model RDS
files, one aggregate build RDS, and a release manifest. Existing versioned
payloads are immutable: a different hash stops promotion. The release also
contains quarantined_spectra.rds, a versioned bundle of valid OpenSpecy
objects by recipe/type, row-level removal metadata, the long conflict audit, and
build provenance. Its review copy and CSV mirrors are checkpointed under
output_dir/review before downstream model and comparison stages.
The returned object has four top-level elements: libraries, medoids,
models, and assessments. Libraries are recipe-first and then keyed by
ftir, raman, or nir: full Raman uses 200--4000, FTIR uses 400--4000,
and NIR uses 4000--12000 cm^-1^. FTIR/Raman medoids and models use 800--3200;
the NIR identification interval is derived from the longest region with at
least 90% finite coverage (falling back to available coverage) and recorded on
the object. All-blank metadata columns are dropped independently by type.
The enrichment columns are retained even when every value for one type is
missing; other all-blank metadata columns are dropped independently by type.
Each spectrum must observe at least 10% of its type-specific identification
axis to enter medoid selection or model fitting. PAM operates on a temporary
relative matrix where each spectrum's missing values are filled by that
spectrum's finite mean. Selected medoid IDs are then pulled from the original
range-restricted object, so published medoids retain their original NA
positions. Groups with more than 3,000 spectra use five deterministic
1,000-spectrum PAM samples; each candidate set is scored against the complete
group with the same correlation distance, avoiding a full oversized
dissimilarity matrix. models is algorithm-first (logistic_regression or
random_forest), then recipe and type. Logistic regression retains
derivative/nobaseline medoid training and fills restored gaps with the finite
mean at each wavenumber across training medoids. Random forests use the complete
eligible raw, derivative, and nobaseline type libraries, relative normalization,
wavenumber-mean filling, inverse-frequency balanced case sampling, 500
probability trees, permutation importance, and all available ranger threads.
Balanced sampling improves minority representation in each bootstrap without
also applying a class-vote correction. Because this resampling can make
out-of-bag accuracy optimistic, treat OOB results as fit diagnostics. The
release assessment separately applies each deployed model to its complete
corresponding source dataset.
train_spec_model() exposes both engines; build_model_lib() remains a
compatible wrapper. Each model stores its filler, support audit, training
diagnostics, and one tidy tests table. Logistic lambda selection remains
maximum out-of-fold macro class accuracy with its calibrated alpha,
no-intercept, grouped multinomial, and weighting settings unchanged.
Full assessment uses every candidate and downloaded legacy artifact. Candidate
and legacy artifacts independently use a seeded approximately ten-percent
class/type-stratified sample. Rows sharing a physical identifier or exact
spectral content are one group and cannot cross training and test. This
source-local policy tolerates taxonomy evolution
without fuzzy cross-version class matching; its results describe each
artifact on its own source population and are not paired causal deltas.
Held-out groups are removed from full reference libraries before their
identification assessment. Medoid artifacts are instead tested as deployed:
each medoid library identifies every spectrum in its complete corresponding
processed library (for example, medoid_derivative.rds searches
derivative.rds). Existing logistic and random-forest artifacts likewise
identify their complete corresponding source libraries once; assessment does
not select new medoids or call train_spec_model(). match_spec() uses stored
fillers for partial spectra. Macro class
accuracy is the primary metric, followed by evaluated coverage, overall
accuracy, and clear old/new
assess_spec(report = "all") shifts. The review surface contains at most ten
tables nested under cleanup, ref_lib, medoid, model, and
functionality. Old and new values occupy adjacent columns; accuracy,
misidentification counts, absolute diagnostic correlations, and warning/error
rate shifts are sorted from largest to smallest. Pass rows are omitted from
quality shifts. Machine-level split and row-test evidence is available through
attr(reference_library_build$assessments, "evidence"). Published accuracy
tables contain aggregate overall and macro metrics, not per-class accuracy
rows. A versioned release writes all global cleanup, quality, pruning, model,
and comparison evidence to assessments.rds; its library, medoid, and model
files contain only runtime data and prediction state. The accompanying
reference_library_build.rds is a lightweight index that can be supplied to
rebuild_lib_artifacts(). The smaller 1,000-spectrum run in the manual
benchmark is development-only and is not part of build_lib() or final
acceptance evidence.
The standard call therefore has this shape:
reference_library_build <- build_lib( files, output_dir = output_dir, previous_library_dir = "system", remove_other = TRUE )
After upstream libraries have already completed, rebuild only the downstream artifacts into a new explicit output location:
updated_reference_library_build <- rebuild_lib_artifacts( completed_build_dir, output_dir = downstream_output_dir, previous_library_dir = "system" )
completed_build_dir may instead be a reference_library_build.rds, a
libraries.rds checkpoint, or the corresponding in-memory build object.
The script is excluded from package builds but remains available in the GitHub repository so library releases can be reviewed and reproduced.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.