User contributions for Griet

16:3916:39, 11 June 2024 diff hist −1‎ m K-Dutch ‎No edit summary
15:2215:22, 11 June 2024 diff hist +7‎ Parallel Monolingual Corpora/nl ‎ Created page with "* 3.02 MB * [https://huggingface.co/datasets/UWV/veringewikkelderingen Downloadpagina]"
15:2215:22, 11 June 2024 diff hist +2‎ Translations:Parallel Monolingual Corpora/17/nl ‎No edit summary current
15:2015:20, 11 June 2024 diff hist +5‎ Translations:Parallel Monolingual Corpora/14/nl ‎No edit summary current Tag: Manual revert
15:1915:19, 11 June 2024 diff hist 0‎ Translations:Parallel Monolingual Corpora/2/nl ‎No edit summary current
15:1515:15, 11 June 2024 diff hist −163‎ Parallel Monolingual Corpora/nl ‎ Created page with "* 8.67 MB * [https://huggingface.co/datasets/NetherlandsForensicInstitute/simplewiki-translated-nl Download dataset]"
15:1515:15, 11 June 2024 diff hist +35‎ Parallel Monolingual Corpora ‎No edit summary
15:1515:15, 11 June 2024 diff hist +89‎ N Translations:Parallel Monolingual Corpora/26/nl ‎ Created page with "* [https://github.com/tsei902/simplify_dutch/tree/main/resources/datasets Downloadpagina]" current
15:1515:15, 11 June 2024 diff hist +572‎ N Translations:Parallel Monolingual Corpora/25/nl ‎ Created page with "2 De tweede vertaalde dataset is gemaakt door Theresa Seidl in het kader van Controllable sentence simplification in het Nederlands. Dit is een synthetische dataset die een combinatie is van de eerste 10.000 rijen van de parallelle [https://github.com/XingxingZhang/dress WikiLarge-dataset], en [https://github.com/facebookresearch ASSET-dataset (Abstractive Sentence Simplification Evaluation and Tuning)]. Door deze twee datasets te combineren, vertaalde Theresa ze naar he..." current
15:1315:13, 11 June 2024 diff hist +116‎ N Translations:Parallel Monolingual Corpora/24/nl ‎ Created page with "* 8.67 MB * [https://huggingface.co/datasets/NetherlandsForensicInstitute/simplewiki-translated-nl Download dataset]" current
15:1215:12, 11 June 2024 diff hist +354‎ N Translations:Parallel Monolingual Corpora/23/nl ‎ Created page with "1 De eerste vertaalde dataset is gemaakt door het Nederlands Forensisch Instituut met behulp van Meta's [https://ai.meta.com/research/no-language-left-behind/ No Language Left Behind-model]. De dataset bestaat uit 167.000 uitgelijnde zinsparen en dient als Nederlandse vertaling van de [https://cs.pomona.edu/~dkauchak/simplification/ SimpleWiki-dataset]" current
15:1215:12, 11 June 2024 diff hist −202‎ Parallel Monolingual Corpora/nl ‎ Created page with "* 17.5 MB"
15:1015:10, 11 June 2024 diff hist +24‎ N Translations:Parallel Monolingual Corpora/22/nl ‎ Created page with "'''Vertaalde datasets'''" current
15:1015:10, 11 June 2024 diff hist +136‎ N Translations:Parallel Monolingual Corpora/21/nl ‎ Created page with "* [https://github.com/nivack/comparable_corpus_Wablieft_deStandaard/blob/main/comparable_corpus_Wablieft_DeStandaard.txt Downloadpagina]" current
15:0915:09, 11 June 2024 diff hist +9‎ N Translations:Parallel Monolingual Corpora/20/nl ‎ Created page with "* 17.5 MB" current
15:0915:09, 11 June 2024 diff hist +514‎ N Translations:Parallel Monolingual Corpora/19/nl ‎ Created page with "3) De derde dataset is het vergelijkbare corpus gemaakt door Nick Vanackere. Het bevat 12.687 Wablieft-artikelen uit de periode 2012-2017 en 206.466 De Standaard-artikelen uit de periode 2013-2017. Om de vergelijkbaarheid te garanderen, werden alleen artikels van 08/01/2013 tot 16/11/2017 bekeken, wat resulteerde in 8.744 Wablieft-artikels en 202.284 De Standaard-artikels. Het verschil in het aantal artikelen is te wijten aan de verschijningsfrequentie: Wablieft verschij..." current
15:0215:02, 11 June 2024 diff hist −138‎ Parallel Monolingual Corpora/nl ‎ Created page with "2) De tweede dataset is gemaakt door UWV Nederland als onderdeel van het “Leesplank”-project, een poging om datasets te genereren die ethisch en juridisch verantwoord zijn. De dataset bestaat uit 2,87 miljoen alinea's en de bijbehorende vereenvoudigde tekst. De paragrafen zijn gebaseerd op het Nederlandse Wikipedia-extract uit [http://gigacorpus.nl/ Gigacorpus]. De tekst is gefilterd en opgeschoond door [https://learn.microsoft.com/en-us/azure/ai-services/openai/conc..."
15:0015:00, 11 June 2024 diff hist +86‎ N Translations:Parallel Monolingual Corpora/18/nl ‎ Created page with "* 3.02 MB * [https://huggingface.co/datasets/UWV/veringewikkelderingen Downloadpagina]" current
15:0015:00, 11 June 2024 diff hist +265‎ N Translations:Parallel Monolingual Corpora/17/nl ‎ Created page with "Een uitgebreidere versie van deze dataset is gemaakt door Michiel Buisman en Bram Vanroy. Deze dataset bevat een eerste, kleine set variaties van Wikipediaparagrafen in verschillende stijlen (jargon, officieel, archaïsche taal, technisch, academisch en poëtisch)."
14:5914:59, 11 June 2024 diff hist +103‎ N Translations:Parallel Monolingual Corpora/16/nl ‎ Created page with "* 3.02 MB * [https://huggingface.co/datasets/UWV/Leesplank_NL_wikipedia_simplifications Downloadpagina]" current
14:5814:58, 11 June 2024 diff hist +554‎ N Translations:Parallel Monolingual Corpora/15/nl ‎ Created page with "2) De tweede dataset is gemaakt door UWV Nederland als onderdeel van het “Leesplank”-project, een poging om datasets te genereren die ethisch en juridisch verantwoord zijn. De dataset bestaat uit 2,87 miljoen alinea's en de bijbehorende vereenvoudigde tekst. De paragrafen zijn gebaseerd op het Nederlandse Wikipedia-extract uit [http://gigacorpus.nl/ Gigacorpus]. De tekst is gefilterd en opgeschoond door [https://learn.microsoft.com/en-us/azure/ai-services/openai/conc..." current
14:5514:55, 11 June 2024 diff hist −119‎ Parallel Monolingual Corpora/nl ‎ Created page with "* [https://huggingface.co/datasets/BramVanroy/chatgpt-dutch-simplification Downloadpagina (CSV-bestanden)]"
14:5314:53, 11 June 2024 diff hist +106‎ N Translations:Parallel Monolingual Corpora/14/nl ‎ Created page with "* [https://huggingface.co/datasets/BramVanroy/chatgpt-dutch-simplification Downloadpagina (CSV-bestanden)]"
14:5314:53, 11 June 2024 diff hist +96‎ N Translations:Parallel Monolingual Corpora/13/nl ‎ Created page with "# Training = 1013 zinnen (262 KB) # Validatie = 126 zinnen (32.6 KB) # Test = 128 zinnen (33 KB)" current
14:5214:52, 11 June 2024 diff hist −34‎ Parallel Monolingual Corpora/nl ‎ Created page with "Het Nederlandse gemeentelijke corpus is een parallel monolinguaal corpus voor de evaluatie van zinsvereenvoudiging in het Nederlandse gemeentelijke domein. Het corpus is gemaakt door Amsterdam Intelligence. Het bevat 1.311 vertaalde parallelle zinsparen die automatisch gealigneerd werden. De zinsparen zijn afkomstig uit 50 documenten van de communicatieafdeling van de gemeente Amsterdam die handmatig werden vereenvoudigd om de vereenvoudiging voor het Nederlands te evalu..."
14:5214:52, 11 June 2024 diff hist +258‎ N Translations:Parallel Monolingual Corpora/12/nl ‎ Created page with "1) De eerste dataset is door Bram Vanroy gemaakt voor Nederlandsetekstvereenvoudigingstaken met behulp van text-to-text transfer transformers en bestaat uit Nederlandse bronzinnen samen met hun corresponderende vereenvoudigde zinnen, gegenereerd met ChatGPT."
14:4814:48, 11 June 2024 diff hist +55‎ Parallel Monolingual Corpora ‎No edit summary
14:4814:48, 11 June 2024 diff hist −81‎ Parallel Monolingual Corpora/nl ‎ Created page with "* 265 KB * [https://github.com/Amsterdam-AI-Team/dutch-municipal-text-simplification/tree/master/complex-simple-sentences Download dataset (CSV-bestand)]"
14:4814:48, 11 June 2024 diff hist +38‎ N Translations:Parallel Monolingual Corpora/11/nl ‎ Created page with "'''Automatisch gecreëerde datasets'''" current
14:4814:48, 11 June 2024 diff hist +153‎ N Translations:Parallel Monolingual Corpora/10/nl ‎ Created page with "* 265 KB * [https://github.com/Amsterdam-AI-Team/dutch-municipal-text-simplification/tree/master/complex-simple-sentences Download dataset (CSV-bestand)]" current
14:4814:48, 11 June 2024 diff hist +480‎ N Translations:Parallel Monolingual Corpora/9/nl ‎ Created page with "Het Nederlandse gemeentelijke corpus is een parallel monolinguaal corpus voor de evaluatie van zinsvereenvoudiging in het Nederlandse gemeentelijke domein. Het corpus is gemaakt door Amsterdam Intelligence. Het bevat 1.311 vertaalde parallelle zinsparen die automatisch gealigneerd werden. De zinsparen zijn afkomstig uit 50 documenten van de communicatieafdeling van de gemeente Amsterdam die handmatig werden vereenvoudigd om de vereenvoudiging voor het Nederlands te evalu..." current
14:4414:44, 11 June 2024 diff hist 0‎ Parallel Monolingual Corpora ‎No edit summary
14:4314:43, 11 June 2024 diff hist −53‎ Parallel Monolingual Corpora/nl ‎ Created page with "'''Manueel gecreëerde datasets'''"
14:4214:42, 11 June 2024 diff hist +34‎ N Translations:Parallel Monolingual Corpora/8/nl ‎ Created page with "'''Manueel gecreëerde datasets'''" current
14:4114:41, 11 June 2024 diff hist +9‎ Parallel Monolingual Corpora/nl ‎No edit summary
14:4114:41, 11 June 2024 diff hist 0‎ Translations:Parallel Monolingual Corpora/6/nl ‎No edit summary current
14:4114:41, 11 June 2024 diff hist +4‎ Translations:Parallel Monolingual Corpora/3/nl ‎No edit summary current
14:4114:41, 11 June 2024 diff hist +5‎ Translations:Parallel Monolingual Corpora/2/nl ‎No edit summary
14:3914:39, 11 June 2024 diff hist +32‎ Parallel Monolingual Corpora/nl ‎No edit summary
14:3814:38, 11 June 2024 diff hist 0‎ Translations:Parallel Monolingual Corpora/1/nl ‎No edit summary current
14:3814:38, 11 June 2024 diff hist +3‎ Translations:Parallel Monolingual Corpora/Page display title/nl ‎No edit summary current
13:4413:44, 11 June 2024 diff hist −167‎ Parallel Multilingual Corpora/nl ‎ Created page with "==Dutch Government Website Corpus=="
13:4413:44, 11 June 2024 diff hist −1‎ Parallel Multilingual Corpora ‎No edit summary
13:4313:43, 11 June 2024 diff hist +95‎ N Translations:Parallel Multilingual Corpora/50/nl ‎ Created page with "* [https://live.european-language-grid.eu/catalogue/corpus/2877/ European Language Grid-pagina]" current
13:4313:43, 11 June 2024 diff hist +50‎ N Translations:Parallel Multilingual Corpora/49/nl ‎ Created page with "Parallel (EN-NL) corpus van 6.532 vertaaleenheden." current
13:4313:43, 11 June 2024 diff hist +35‎ N Translations:Parallel Multilingual Corpora/48/nl ‎ Created page with "==Dutch Government Website Corpus==" current
13:4313:43, 11 June 2024 diff hist −213‎ Parallel Multilingual Corpora/nl ‎ Created page with "==The Open Parallel Corpus (OPUS)=="
13:4313:43, 11 June 2024 diff hist +214‎ N Translations:Parallel Multilingual Corpora/47/nl ‎ Created page with "* [https://lt3.ugent.be/resources/multiling-en-nl/ Projectinformatie en downloadinstructies] * [https://sites.google.com/site/centretranslationinnovation/tpr-db/public-studies#h.p_iVVuCQOHJx2O MultiLing-informatie]" current
13:4213:42, 11 June 2024 diff hist +554‎ N Translations:Parallel Multilingual Corpora/46/nl ‎ Created page with "De multiLing-dataset is gebaseerd op zes Engelse bronteksten die in verschillende talen zijn vertaald. Vier daarvan (teksten 1-4) zijn nieuwsartikelen en de andere twee (teksten 5-6) zijn sociologische teksten uit een encyclopedie. De Nederlandse data bestaat uit twee delen. ENDU20: tien Nederlandse vertalingen van de multiLing-set door tien vertalers die recent hun mastersdiploma gehaald hebben en die Nederlands als moedertaal hebben. En ENDU20-MT: twee Nederlandse mach..." current
13:4013:40, 11 June 2024 diff hist +19‎ N Translations:Parallel Multilingual Corpora/45/nl ‎ Created page with "==MultiLing EN-NL==" current

User contributions for Griet

11 June 2024

Navigation menu

Search