== Is there a text only version of the Corpus Spoken Dutch? ==
== Is there a text only version of the Corpus Spoken Dutch? ==
The download files of the Corpus Spoken Dutch (CGN) do not contain the text only. The <code>ort</code> files contain ortographic transcriptions and timestamps and the <code>plk</code> files contain part-of-speech and lemma information. The following perl script takes a list of plk files as input and prints the text. If you run this script from the command line in your terminal, then you can create text files.
Latest revision as of 14:19, 5 July 2024
Is there a text only version of the Corpus Spoken Dutch?