Jump to content

Translations:Q&A/94/en: Difference between revisions

From Clarin K-Centre
FuzzyBot (talk | contribs)
Importing a new version from external source
 
(No difference)

Latest revision as of 14:19, 5 July 2024

Information about message (contribute)
This message has no documentation. If you know where or how this message is used, you can help other translators by adding documentation to this message.
Message definition (Q&A)
The download files of the Corpus Spoken Dutch (CGN) do not contain the text only. The <code>ort</code> files contain ortographic transcriptions and timestamps and the <code>plk</code> files contain part-of-speech and lemma information.  The following perl script takes a list of plk files as input and prints the text. If you run this script from the command line in your terminal, then you can create text files.

The download files of the Corpus Spoken Dutch (CGN) do not contain the text only. The ort files contain ortographic transcriptions and timestamps and the plk files contain part-of-speech and lemma information. The following perl script takes a list of plk files as input and prints the text. If you run this script from the command line in your terminal, then you can create text files.