screaming_frog_internal: Import Screaming Frog Internal: All node facts

View source: R/screaming_frog_internal.R

screaming_frog_internalR Documentation

Import Screaming Frog Internal: All node facts

Description

Normalizes a Screaming Frog **Internal: All** CSV or data frame into node, redirect, canonical, and indexability tables. This is a node-side adapter: it does not reconstruct links from aggregate Inlinks/Outlinks counts.

Usage

screaming_frog_internal(x)

Arguments

x

A path to an Internal: All CSV file or an equivalent data frame.

Details

URLs are preserved as exported. Redirects are emitted only for valid 3xx rows with a non-blank destination. Canonicals are derived independently, including self-canonicals for audit.

Value

An object of class 'screaming_frog_internal' with components:

nodes

Normalized node facts in input row order. Columns are 'url', 'segments', 'content_type', 'http_status', 'status', 'indexability', 'indexability_status', 'canonical', 'redirect_to', 'redirect_type', 'crawl_allowed', 'indexing_allowed', robots fields, language/timestamps, and selected crawl metrics. Optional absent fields are typed 'NA' columns.

redirects

Raw 'from' / 'to' redirect pairs from valid 3xx rows.

canonicals

Raw 'from' / 'to' canonical pairs, including self-canonicals.

indexability

URL-level facts compatible with ‘pagerank()'’s indexability input.

diagnostics

Deterministic counts, missing optional and ignored columns, duplicate addresses, and row-level structural issues.

provenance

Input identity, retained input-row IDs, and the normalized-to-detected column manifest.

Examples

internal <- data.frame(
  Address = c("https://example.com/", "https://example.com/old"),
  `Status Code` = c("200", "301"),
  `Redirect URL` = c("", "https://example.com/new"),
  check.names = FALSE
)
imported <- screaming_frog_internal(internal)
imported$nodes
imported$redirects

pagerankr documentation built on Oct. 1, 2026, 5:09 p.m.