r-cidian 读取中文输入法词库,并把不同格式转换成统一的
cidian_dictionary 对象。目前支持以下格式:
| 格式 | 扩展名 | 来源 |
|-------|----------|---------|
| SCEL | .scel | 搜狗 |
| QCEL | .qcel | QQ 拼音 |
| QPYD | .qpyd | QQ 拼音 |
| BDICT | .bdict | 百度 |
| BCD | .bcd | 百度 |
底层解析由 Rust 库
cidian-rs
进行。解析结果保持词条在源文件中的顺序,不会自动规范化、排序或去重。
r-cidian 是 qinwf/cidian
的现代化替代方案。开发初衷是仿照已经删库且无法编译的原 cidian
包的功能,为 jiebaRS
提供方便的自定义词典导入功能。
你可以从 CRAN 安装 r-cidian 的发布版本:
install.packages("cidian")
如果您使用 Linux,您可以尝试从 P3M 为不同的 Linux 发行版安装预编译的二进制文件,避免从源代码编译。
R-universe 和 R-multiverse 上提供了预编译的二进制包:
install.packages("cidian", repos = "https://yousa-mirage.r-universe.dev")
install.packages("cidian", repos = "https://community.r-multiverse.org")
使用 pak 从源码安装:
pak::pak("Yousa-Mirage/r-cidian")
注意: 从源码构建需要 Rust 工具链来编译 Rust 后端。
library(cidian)
dictionary <- read_cidian("计算机科技.qcel")
summary(dictionary)
#> <summary.cidian_dictionary>
#> Format: QCEL
#> Name: 计算机名词
#> Category: 计算机科技
#> Entries: 9,646
#> Weighted entries: 9,646
entries <- cidian_entries(dictionary)
head(entries)
#> word code weight
#> 1 阿里通 a, li, tong 8
#> 2 阿姆达尔定律 a, mu, d.... 9646
#> 3 阿姆斯特朗公理 a, mu, s.... 9645
#> 4 阿帕网 a, pa, wang 9644
#> 5 埃尔布朗基 ai, er, .... 9643
#> 6 埃尔米特函数 ai, er, .... 9642
# cidian_metadata(dictionary) 可以获取该词典的元信息
其中,entries$code 是 list-column,每个元素都是一个字符向量。例如:
dictionary$entries$code[[1]]
#> [1] "a" "li" "tong"
write_words(dictionary, "words.txt")
可以把词典中所有词语每行一个写入文本文件,可用于为 jiebaRS
等分词工具提供自定义词典(jiebaRS 已内置了 import_cidian()
函数用于从输入法词库中导入自定义词语)。
MIT License
Any scripts or data that you put into this service are public.
Add the following code to your website.
For more information on customizing the embed code, read Embedding Snippets.