Manually annotated corpora: Difference between revisions
No edit summary |
|||
Line 31: | Line 31: | ||
*[https://github.com/AylaRT/ACTER Github page] | *[https://github.com/AylaRT/ACTER Github page] | ||
* Rigouts Terryn, Ayla, 2019, ACTER (Annotated Corpora for Term Extraction Research) v1.3, Eurac Research CLARIN Centre | * Rigouts Terryn, Ayla, 2019, ACTER (Annotated Corpora for Term Extraction Research) v1.3, Eurac Research CLARIN Centre | ||
==Dutch Archaeology NER Training Dataset== | |||
A manually annotated NER dataset, consisting of Dutch archaeological excavation reports. The following entity types are labelled: Artefacts, Time periods, Materials, Places (geographical locations), Archaeological contexts and Species. | |||
The dataset is provided in the BIO format, with each token on 1 line and empty lines denoting sentence boundaries. On each line you can find the token, PoS tag, morphological segmentation and finally the label, separated by spaces. The PoS tag and morphological segmentation are assigned by Frog. | |||
* Version 1.0 (2019) | |||
* [https://zenodo.org/record/3544544#.YmEbqehBzZQ Download page] |
Revision as of 08:55, 21 April 2022
Manual corpora are collections of texts containing manually validated or manually assigned linguistic information, such as morphosyntactic tags, lemmas, syntactic parses, named entities etc. These corpora can be used to train new language annotation tools as well as to test the accuracy of existing annotation tools.
Corpus Gesproken Nederlands
This corpus is at manually transcribed and annotated for all its layers.
Eindhoven Corpus
The Eindhoven corpus (VU version) is a collection of Dutch written and transcribed spoken texts from the period from 1960 to 1976. The corpus contains approximately 768,000 tokens. It is the first collection of Dutch written and (transcribed) spoken texts.
- version 2.0.1 (2014)
- 3 MB
- Download link
Lassy Small
The Lassy Small Corpus is a corpus of approximately 1 million words with manually verified syntactical annotations. The lemmas and POS-tags were generated with Tadpole (now Frog) and the syntactical depency structures were generated with Alpino. The lemmas, POS-tags and syntactic tree structures were manually verified and corrected.
- version 6.0
- Download
- Online treebank search
SoNaR-1
Size: 1 million words. This is a manually annotated subset of the much larger (approx. 500 million) word) SoNaR corpus.
ACTER: Annotated Corpora for Term Extraction Research
The ACTER (Annotated Corpora for Term Extraction Research) is an annotated dataset for term extraction. Terms and Named Entities have been manually annotated in specialised comparable corpora covering 3 languages (English, French, and Dutch), and 4 domains (corruption, dressage, heart failure, and wind energy).
- CLARIN page
- Github page
- Rigouts Terryn, Ayla, 2019, ACTER (Annotated Corpora for Term Extraction Research) v1.3, Eurac Research CLARIN Centre
Dutch Archaeology NER Training Dataset
A manually annotated NER dataset, consisting of Dutch archaeological excavation reports. The following entity types are labelled: Artefacts, Time periods, Materials, Places (geographical locations), Archaeological contexts and Species. The dataset is provided in the BIO format, with each token on 1 line and empty lines denoting sentence boundaries. On each line you can find the token, PoS tag, morphological segmentation and finally the label, separated by spaces. The PoS tag and morphological segmentation are assigned by Frog.
- Version 1.0 (2019)
- Download page