Export translations

Settings

Group

Language

Format

Export for off-line translation

Export in native format

Export in CSV format

<div lang="en" dir="ltr" class="mw-content-ltr">
==Automatically translated datasets==
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
===ASSET Simplification Corpus===
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
The Abstractive Sentence Simplification Evaluation and Tuning (ASSET) Dataset (Alva-Manchego et al, 2020) was automatically translated to Dutch (Seidl et al., 2023), and is freely available.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
*[https://github.com/tsei902/simplify_dutch/tree/main/resources/datasets/asset Github download]
* Alva-Manchego, F., Martin, L., Bordes, A., Scarton, C., Sagot, B., & Specia, L. (2020). ASSET: A dataset for tuning and evaluation of sentence simplification models with multiple rewriting transformations. arXiv preprint arXiv:2005.00481.
* Seidl, T., Vandeghinste, V., & Van de Cruys, T. (2023). [https://kuleuven.limo.libis.be/discovery/fulldisplay?docid=alma9993527112601488&context=L&vid=32KUL_KUL:KULeuven&lang=en&search_scope=All_Content&adaptor=Local%20Search%20Engine&tab=all_content_tab&query=any,contains,seidl%20theresa&offset=0 Controllable Sentence Simplification in Dutch]. KU Leuven. Faculteit Ingenieurswetenschappen.
* Seidl, T., Vandeghinste, V. (2024). [https://clinjournal.org/clinj/article/view/171 Controllable Sentence Simplification in Dutch.] Computational Linguistics in the Netherlands Journal, 13, 31–61.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
===Wikilarge Dataset===
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
Automatic translation of the Wikilarge dataset, useful for automatic simplification (Seidl et al., 2023), freely available. Original dataset from Zhang & Lapata
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
*[https://github.com/tsei902/simplify_dutch/tree/main/resources/datasets/wikilarge Github download]
* Seidl, T., Vandeghinste, V., & Van de Cruys, T. (2023). [https://kuleuven.limo.libis.be/discovery/fulldisplay?docid=alma9993527112601488&context=L&vid=32KUL_KUL:KULeuven&lang=en&search_scope=All_Content&adaptor=Local%20Search%20Engine&tab=all_content_tab&query=any,contains,seidl%20theresa&offset=0 Controllable Sentence Simplification in Dutch]. KU Leuven. Faculteit Ingenieurswetenschappen.
*Seidl, T., Vandeghinste, V. (2024). [https://clinjournal.org/clinj/article/view/171 Controllable Sentence Simplification in Dutch.] Computational Linguistics in the Netherlands Journal, 13, 31–61.
* Zhang, X. & Lapata, M. (2017). Sentence Simplification with Deep Reinforcement Learning. In ''Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing'', pages 584–594, Copenhagen, Denmark. Association for Computational Linguistics.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
===NFI SimpleWiki dataset===
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
Translated dataset created by Netherlands Forensic Institute using Meta's [https://ai.meta.com/research/no-language-left-behind/ No Language Left Behind model]. It comprises 167,000 aligned sentence pairs and serves as a Dutch translation of the SimpleWiki [https://cs.pomona.edu/~dkauchak/simplification/ dataset].
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
* 8.67 MB
* [https://huggingface.co/datasets/NetherlandsForensicInstitute/simplewiki-translated-nl Download dataset]
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
==Comparable Corpus Wablieft De Standaard==
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
Corpus created by Nick Vanackere. It contains a comparable corpus of 12,687 Wablieft articles between 2012-2017 from 206,466 De Standaard articles from 2013-2017. To ensure comparability, only articles from 08/01/2013 till 16/11/2017 were considered, resulting in 8,744 Wablieft articles and 202,284 De Standaard articles. The difference in the number of articles is due to the publication frequency, with Wablieft being weekly and De Standaard daily.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
*[https://github.com/nivack/comparable_corpus_Wablieft_deStandaard Github]
*<small>[https://kuleuven.limo.libis.be/discovery/fulldisplay?docid=alma9993153812401488&context=L&vid=32KUL_KUL:KULeuven&lang=en&search_scope=All_Content&adaptor=Local%20Search%20Engine&tab=all_content_tab&query=any,contains,nick%20vanackere&offset=0 Vanackere, N., & Vandeghinste, V. (2022). Building a comparable corpus between easy-to-read Dutch Wablieft and De Standaard. KU Leuven. Faculteit Ingenieurswetenschappen.]</small>
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
==Synthetic datasets==
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
===SONAR WRPEI Simplification Dataset===
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
The Synthetic Simplification Dataset was compiled within the Duidelijke Taal project and is based on the WR-P-E-I component (websites) of the SoNaR corpus. The dataset consists of three parts: 6,986 sentences from the SoNaR corpus, a synthetic simplification of the SoNaR sentences created by GPT-4 and sentence pairs consisting of one SoNaR sentence and its simplified version each.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
*[https://hdl.handle.net/10032/tm-a2-y7 Weblink]
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
====Human evaluation of automated text simplification: crowdsourcing results====
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
The language material ‘Human evaluation of automated text simplification: crowdsourcing results’ was compiled as part of the Duidelijke Taal project. The dataset consists of sentences from the SoNaR corpus, a version simplified by GPT-4 and the human evaluations of those simplifications with respect to simplicity, accuracy and fluency. The sentence pairs are a subset of the SONAR WRPEI Simplification Dataset
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
*[https://hdl.handle.net/10032/tm-a2-y8 Weblink]
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
===UWV Leesplank NL wikipedia===
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
The set contains 2,391,206 pragraphs of prompt/result combinations, where the prompt is a paragraph from Dutch Wikipedia and the result is a simplified text, which could include more than one paragraph. This dataset was created by UWV, as a part of project "Leesplank", an effort to generate datasets that are ethically and legally sound.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
* [https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications/blob/main/README.md HuggingFace ReadMe file]
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
*[https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications Dataset]
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
A more extended version of this dataset was made by Michiel Buisman and Bram Vanroy. This datasets contains a first, small set of variations of Wikipedia paragraphs in different styles (jargon, official, archaïc language, technical, academic, and poetic).
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
* 3.02 MB
* [https://huggingface.co/datasets/UWV/veringewikkelderingen Download page]
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
===ChatGPT generated dataset by Van de Velde===
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
Created in light of a master thesis by Charlotte Van de Velde. The dataset contains Dutch source sentences and aligned simplified sentences, generated with ChatGPT. All splits combined, the dataset consists of 1267 entries.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
# Training = 1013 sentences (262 KB)
# Validation = 126 sentences (32.6 KB)
# Test = 128 sentences (33 KB)
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
* [https://huggingface.co/datasets/BramVanroy/chatgpt-dutch-simplification Download page (CSV files)]
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
==Manually simplified==
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
===Dutch municipal data===
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
The Dutch municipal corpus is a parallel monolingual corpus for the evaluation of sentence-level simplification in the Dutch municipal domain. The corpus was created by Amsterdam Intelligence. It contains 1,311 translated parallel sentence pairs that were automatically aligned. The sentence pairs originate from 50 documents from the Communications Department of the City of Amsterdam that were manually simplified to evaluate simplification for Dutch.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
*265 KB
*[https://github.com/Amsterdam-AI-Team/dutch-municipal-text-simplification Github]
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
===Dutch Contextualized Lexical Simplification Evaluation Dataset===
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
As part of her internship with us last year, Eliza Hobo developed the first contextual lexical simplification model for Dutch. Due to the lack of Dutch evaluation data for lexical simplification, we developed a pilot benchmark dataset for the task using authentic municipal data. We select sentences from a collection of 48 municipal documents based on the presence of a complex word from a list curated by domain experts and based on their word count (less than 20 words).
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
The sentences were simplified by 23 native speakers of Dutch who pursued or obtained an academic degree. They were shown a sentence with the highlighted complex word and five simplification options that LSBertje generated. The annotators could select from these options and propose additional simplifications.
</div>

<div lang="en" dir="ltr" class="mw-content-ltr">
*[https://amsterdamintelligence.com/resources/geen-makkie-data Website]
*[https://aclanthology.org/2023.bea-1.42/ Paper]
</div>