Parallel Monolingual Corpora/nl: Difference between revisions

From Clarin K-Centre
Jump to navigation Jump to search
No edit summary
(Updating to match new version of source page)
 
(9 intermediate revisions by one other user not shown)
Line 4: Line 4:
==DAESO-corpus==
==DAESO-corpus==


Het DAESO Corpus is een parallelle eentalige treebank van Nederlandse teksten. Het corpus bevat ruim 2,1 miljoen woorden parallelle en vergelijkbare tekst. Ongeveer 678.000 woorden werden handmatig uitgelijnd en ongeveer 1,5 miljoen woorden werden automatisch uitgelijnd. Er is een semantische relatie toegevoegd aan de gealigneerde woorden/zinnen.
Het DAESO-corpus is een parallelle monolinguale treebank van Nederlandse teksten. Het corpus bevat ruim 2,1 miljoen woorden parallelle en vergelijkbare tekst. Ongeveer 678.000 woorden werden handmatig gealigneerd en ongeveer 1,5 miljoen woorden werden automatisch gealigneerd. Er is een semantische relatie toegevoegd aan de gealigneerde woorden/zinnen.


*92.5 MB
* 92.5 MB
*versie 1.0 (2010)
* versie 1.0 (2010)
*[http://hdl.handle.net/10032/tm-a2-h9 Download page]
* [http://hdl.handle.net/10032/tm-a2-h9 Downloadpagina]


<span id="Bible_Corpus"></span>
<span id="Bible_Corpus"></span>
Line 15: Line 15:
Een diachroon en synchroon parallel corpus van bijbelvertalingen in het Nederlands, Engels, Duits en Zweeds, met teksten van de 14e eeuw tot nu.  
Een diachroon en synchroon parallel corpus van bijbelvertalingen in het Nederlands, Engels, Duits en Zweeds, met teksten van de 14e eeuw tot nu.  


*[https://spraakbanken.gu.se/en/resources/openedges OpenEdges Download]
*[https://spraakbanken.gu.se/en/resources/openedges OpenEdges-download]


<span id="Simplification_Data"></span>
<span id="Simplification_Data"></span>
==[[Simplification_Data/nl|Simplificatiedata]]==
==[[Simplification_Data/nl|Simplificatiedata]]==
<div lang="en" dir="ltr" class="mw-content-ltr">
'''Manually created datasets'''
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
The Dutch municipal corpus is a parallel monolingual corpus for the evaluation of sentence-level simplification in the Dutch municipal domain created by Amsterdan Intelligence. It contains 1,311 translated parallel sentence pairs, automatically aligned from 50 documents from the Communications Department of the City of Amsterdam that were manually simplified to evaluate simplification for Dutch.
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
* 265 KB
* [https://github.com/Amsterdam-AI-Team/dutch-municipal-text-simplification/tree/master/complex-simple-sentences Download dataset (CSV file)]
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
'''Automatically created datasets'''
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
1) The first dataset is created by Bram Vanroy for Dutch text simplification tasks using text-to-text transfer transformers.It comprises Dutch source sentences along with their corresponding simplified sentences, generated with ChatGPT.
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
# Training = 1013 sentences (262 KB)
# Validation = 126 sentences (32.6 KB)
# Test = 128 sentences (33 KB)
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
* [https://huggingface.co/datasets/BramVanroy/chatgpt-dutch-simplification Download page (CSV files)]
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
2) The second dataset is created by UWV Nederland as part of the "Leesplank" project to ensure ethical and legal soundness. It comprises 2.87 million paragraphs and its simplified text as corresponding result. The paragraphs are based on the Dutch Wikipedia extract from [http://gigacorpus.nl/ Gigacorpus]. The text was filtered and cleaned by [https://learn.microsoft.com/en-us/azure/ai-services/openai/concepts/content-filter?tabs=warning%2Cpython-new using GPT-4 1106 preview].
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
* 3.02 MB
* [https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications Download page]
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
A more extended version of this dataset was made by Michiel Buisman and Bram Vanroy. This datasets contains a first, small set of variations of Wikipedia paragraphs in different styles (jargon, official, archaïsche_taal, technical, academic, and poetic).
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
* 3.02 MB
* [https://huggingface.co/datasets/UWV/veringewikkelderingen Download page]
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
3)  The third dataset is the comparable corpus created by Nick Vanackere. It contains a comparable corpus of 12,687 Wablieft articles between 2012-2017  from 206,466 De Standaard articles from 2013-2017. To ensure comparability, only articles from 08/01/2013 till 16/11/2017 were considered, resulting in 8,744 Wablieft articles and 202,284 De Standaard articles. The difference in the number of articles is due to the publication frequency, with Wablieft being weekly and De Standaard daily.
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
* 17.5 MB
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
* [https://github.com/nivack/comparable_corpus_Wablieft_deStandaard/blob/main/comparable_corpus_Wablieft_DeStandaard.txt Download page]
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
'''Translated datasets'''
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
1 The first translated dataset is created by Netherlands Forensic Institute using Meta's [https://ai.meta.com/research/no-language-left-behind/ No Language Left Behind model]. It comprises 167,000 aligned sentence pairs and serves as a Dutch translation of the SimpleWiki [https://cs.pomona.edu/~dkauchak/simplification/ dataset].
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
* 8.67 MB
* [https://huggingface.co/datasets/NetherlandsForensicInstitute/simplewiki-translated-nl Download dataset]
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
2 The second translated dataset is created by Theresa Seidl in the context of Controllable sentence simplification in Dutch. This is a synthetic dataset which is a combination of the first 10,000 rows of the parallel [https://github.com/XingxingZhang/dress WikiLarge dataset], and [https://github.com/facebookresearch ASSET (Abstractive Sentence Simplification Evaluation and Tuning) dataset]. By combining these two datasets, Theresa translated them to Dutch using [https://arxiv.org/pdf/1609.08144 Google Neural Machine Translation].
</div>
<div lang="en" dir="ltr" class="mw-content-ltr">
* [https://github.com/tsei902/simplify_dutch/tree/main/resources/datasets Download page]
</div>

Latest revision as of 11:22, 2 July 2024

Other languages:

DAESO-corpus

Het DAESO-corpus is een parallelle monolinguale treebank van Nederlandse teksten. Het corpus bevat ruim 2,1 miljoen woorden parallelle en vergelijkbare tekst. Ongeveer 678.000 woorden werden handmatig gealigneerd en ongeveer 1,5 miljoen woorden werden automatisch gealigneerd. Er is een semantische relatie toegevoegd aan de gealigneerde woorden/zinnen.

Bijbelcorpus

Een diachroon en synchroon parallel corpus van bijbelvertalingen in het Nederlands, Engels, Duits en Zweeds, met teksten van de 14e eeuw tot nu.

Simplificatiedata