Learning Outcomes
The students will be able to:
• familiarize with basic notions of corpus compilation and analysis.
• comprehend the various ways of usage of corpora for linguistic research.
• combine the theory and praxis of compilation and annotation of a small language corpus that they will create and also the usage of corpus software packages.
Course Content (Syllabus)
The goal of the course is to familiarize students with the basic theoretical and technical tools of compiling and analysing natural language corpora. Throughout course issues on the usefulness of corpora in linguistic research will be discussed and basic definitions and criteria as to how to compile and use natural language corpora. Furthermore, an important part of the course will be devoted to presenting the notions of "wordlists", "concordances" in various usage contexts. Moreover, existing corpora of Greek language will be presented and the current trends of corpus (stand-off) annotation with XML, a markup language that is currently used in many language corpora projects.
Last but not least, participants will learn the most common and well-known methods of statistical analysis of language data.
Keywords
Corpora, (stand-off) annotation, markup languages, wordlists, concordances, corpus software packages
Course Bibliography (Eudoxus)
Τάντος, Α., Μαρκαντωνάτου, Σ., Αναστασιάδη-Συμεωνίδη, Ά., Κυριακοπούλου, Π., 2015. Υπολογιστική γλωσσολογία. [ηλεκτρ. βιβλ.] Αθήνα:Σύνδεσμος Ελληνικών Ακαδημαϊκών Βιβλιοθηκών. Διαθέσιμο στο: http://hdl.handle.net/11419/2205