Hongxiang Li and Tsung Fei Khang
PBGoF is an R package for assessing whether a numeric sample is compatible with a univariate skew-normal distribution when the model parameters are estimated from the same data. It provides Kolmogorov-Smirnov (KS) and Cramer-von Mises (CvM) tests using either a parametric bootstrap or precomputed simulation quantiles, together with robust parameter estimation procedures.
Install PBGoF from a local source archive with:
install.packages("PBGoF_0.1.0.tar.gz", repos = NULL, type = "source")
Install PBGoF from GitHub with devtools:
if (!"devtools" %in% rownames(installed.packages())) {
install.packages("devtools")
}
devtools::install_github("Divo-Lee/PBGoF")
or with pak:
if (!"pak" %in% rownames(installed.packages())) {
install.packages("pak")
}
pak::pkg_install("Divo-Lee/PBGoF")
PBGoF depends on the R packages sn and methods.
Let $X_1,\ldots,X_n$ be an independent sample. Under the null hypothesis,
$$ H_0:\quad X_i \overset{\mathrm{iid}}{\sim} \mathrm{SN}(\xi,\omega,\alpha), \qquad \xi\in\mathbb{R},\quad \omega>0,\quad \alpha\in\mathbb{R}. $$
Writing $z=(x-\xi)/\omega$, the skew-normal density is
$$ f_{\mathrm{SN}}(x;\xi,\omega,\alpha)=\frac{2}{\omega} \phi(z) \Phi(\alpha z), $$
where $\phi$ and $\Phi$ are the standard normal density and distribution functions. Because $\boldsymbol\theta=(\xi,\omega,\alpha)^{\mathsf T}$ is estimated from the sample, this is a composite goodness-of-fit problem; the usual KS null distribution for a fully specified model is not applicable.
PBGoF estimates the model through sn.fit.robust(): MLE is attempted first, followed by MPLE and then MPLE with the matching-prior penalty. A fit is accepted only if the estimates and standard errors are finite and the scale is positive.
For order statistics $X_{(1)}\leq\cdots\leq X_{(n)}$, define
$$ U_{(i)}=F_{\mathrm{SN}} \left(X_{(i)};\widehat{\boldsymbol\theta}\right), \qquad i=1,\ldots,n. $$
The two-sided Kolmogorov-Smirnov discrepancy and PBGoF scaling are
$$ D_n=\max_{1\leq i\leq n}\left(\frac{i}{n}-U_{(i)},\;U_{(i)}-\frac{i-1}{n}\right), \qquad T_{\mathrm{KS}}=\sqrt{n_{\mathrm{eff}}}\,D_n. $$
The Cramer-von Mises statistic is
$$ W_n^2=\frac{1}{12n}+\sum_{i=1}^{n}\left[U_{(i)}-\frac{2i-1}{2n}\right]^2. $$
The parametric-bootstrap CvM test uses $W_n^2$, whereas the precomputed-quantile CvM test uses
$$ T_{\mathrm{CvM}}=\sqrt{n_{\mathrm{eff}}}\,W_n^2. $$
Within the table range, $n_{\mathrm{eff}}=n$.
The functions sn.para.bootstrap.ks.test() and sn.para.bootstrap.cvm.test() reproduce the full estimation procedure:
For $B_{\mathrm{valid}}$ finite bootstrap statistics, PBGoF uses
$$ \widehat p_{\mathrm{boot}}=\frac{1+\displaystyle\sum_{b=1}^{B_{\mathrm{valid}}}\mathbf{1} \left(T_b^*\geq T_{\mathrm{obs}}\right)}{B_{\mathrm{valid}}+1}. $$
Failed fits are excluded and reported. Re-estimation in every replicate calibrates the statistic for the composite null rather than incorrectly treating the fitted distribution as fixed.
PBGoF_ks_test() and PBGoF_cvm_test() use tables generated from 100,000 Monte Carlo replicates per available sample-size and centered-skewness combination. The data are fitted in two parameterizations:
The lookup value is
$$ \gamma_{1,\mathrm{used}}=\min \left(0.99,\max \left[0.01,\mathrm{round} \left(\left|\widehat\gamma_1\right|,2\right)\right]\right). $$
The absolute value follows reflection symmetry: if $X\sim\mathrm{SN}(\xi,\omega,\alpha)$, then $-X$ has shape $-\alpha$. The signs of $\alpha$ and $\gamma_1$ reverse, but the null distributions of the reflection-invariant EDF statistics do not. Therefore, estimated skewness values with the same absolute magnitude but opposite signs use the same reference-table row. PBGoF retains the signed gamma1_hat and returns the non-negative lookup value as gamma1_used.
The bundled tables cover sample sizes through 500. For $n>500$, all observations remain in the fit and EDF, but
$$ n_{\mathrm{eff}}=\min(n,500)=500 $$
is used for the external statistic multiplier and table lookup, following the finding that fitted-skew-normal EDF critical values above 500 are nearly identical to those at 500.
For stored quantiles $q_p$, define
$$ p^{*}=\min_{q_p\geq T_{\mathrm{obs}}}p,\qquad \widehat p_{\mathrm{table}}=1-p^{*}. $$
The probability grid is 0.01 to 0.99 in increments of 0.01, so the result is a conservative step-function approximation. PBGoF does not interpolate across sample size or skewness.
A small p-value is evidence against the fitted skew-normal model. A large p-value does not prove skew-normality; it means the test did not detect a departure at the available sample size and calibration resolution. Use sn.plot.check() alongside the formal tests.
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.