WikiNER-fr-gold:一个黄金标准的命名实体识别语料库
WikiNER-fr-gold: A Gold-Standard NER Corpus
October 29, 2024
作者: Danrun Cao, Nicolas Béchet, Pierre-François Marteau
cs.AI
摘要
本文讨论了WikiNER语料库的质量,这是一个多语言命名实体识别语料库,并提供了其整合版本。WikiNER的标注是以半监督的方式生成的,即没有进行事后手动验证。这种语料库被称为银标准。本文提出了WikiNER-fr-gold,这是WikiNER法语部分的修订版本。我们的语料库包括原始法语子语料库的随机抽样的20%(26,818个句子,70万个标记)。我们首先总结了每个类别中包含的实体类型,以制定标注准则,然后我们开始修订语料库。最后,我们对WikiNER-fr语料库中观察到的错误和不一致性进行了分析,并讨论了潜在的未来工作方向。
English
We address in this article the the quality of the WikiNER corpus, a
multilingual Named Entity Recognition corpus, and provide a consolidated
version of it. The annotation of WikiNER was produced in a semi-supervised
manner i.e. no manual verification has been carried out a posteriori. Such
corpus is called silver-standard. In this paper we propose WikiNER-fr-gold
which is a revised version of the French proportion of WikiNER. Our corpus
consists of randomly sampled 20% of the original French sub-corpus (26,818
sentences with 700k tokens). We start by summarizing the entity types included
in each category in order to define an annotation guideline, and then we
proceed to revise the corpus. Finally we present an analysis of errors and
inconsistency observed in the WikiNER-fr corpus, and we discuss potential
future work directions.Summary
AI-Generated Summary