Please use this identifier to cite or link to this item:

http://hdl.handle.net/10609/70667
Title: Predicció de l'ús del català mitjançant la classificació supervisada
Author: Grimaldo Moreno, Francisco
López Iñesta, Emilia
Perucho Pla, Manel
Querol Puig, Ernest
Keywords: linguistic use
prediction
artificial intelligence
machine learning
supervised classification
Issue Date: Jan-2016
Publisher: Treballs de Sociolingüística Catalana
Citation: Grimaldo, F., López-Iñesta, E., Perucho, M. & Querol Puig, E. (2016). "Predicció de l'ús del català mitjançant la classificació supervisada". Treballs de Sociolingüística Catalana, (26), pp. 181-197. ISSN 0211-0784. doi: 10.2436/20.2504.01.115
Abstract: One of the main challenges that the sociology of language has faced is the determination of the variables that govern the use of a language. Inspired by the field of artificial intelligence, in this study we make use of machine learning as a suitable approach to implement computational methods that permit the induction of linguistic use models derived from the available data. We aim to improve the level of prediction for the degree of use of the Catalan language achieved up to now. To this end, we have used three supervised classification techniques: Naive Bayes, decision trees, and support vector machines. We needed an empirical corpus that would allow us to test the prediction level of a theoretical model as well as its validity within different sociolinguistic situations. To the best of our knowledge, the work by Querol is the one providing the highest prediction success in all the Catalan-speaking territories. Thus, the research presented in this paper uses that data to conclude that supervised classification can be used to successfully determine prediction models for the degree of use of Catalan that outperform previous attempts and that allow us to identify the most relevant variables of the problem. Moreover, it also helps us to solve the methodological problem of the division of linguistic groups and shows that the use of a language is a continuous system rather than a discrete one.
Language: Catalan
URI: http://hdl.handle.net/10609/70667
ISSN: 0211-0784
Appears in Collections:Articles

Share:
Export:
Files in This Item:
File SizeFormat 
Grimaldo_TSC16_Predicció.pdf951.23 kBAdobe PDFView/Open

Items in repository are protected by copyright, with all rights reserved, unless otherwise indicated.