Bereiche | Tage | Auswahl | Suche | Aktualisierungen | Downloads | Hilfe
SOE: Fachverband Physik sozio-ökonomischer Systeme
SOE 8: Focus Session: Complex Systems Approaches to Language and Communication
SOE 8.1: Vortrag
Dienstag, 1. April 2014, 10:15–10:45, GÖR 226
A Comparative Study of Language Complexity in Wikipedia — •János Kertész1,2, Taha Yasseri2,3, and András Kornai2,4 — 1Central European University — 2Budapest University of Technology and Economics — 3University of Oxford — 4Computer and Automation Research Institute of the Hungarian Academy of Sciences.
We present statistical analysis of English texts from Wikipedia [1]. We try to address the issue of language complexity empirically by comparing the Simple English Wikipedia (Simple) to comparable samples of the main English Wikipedia (Main). Simple is supposed to use a more simplified language with a limited vocabulary, and editors are explicitly requested to follow this guideline, yet in practice the vocabulary richness of both samples are at the same level. Detailed analysis of longer units (n-grams of words and part of speech tags) shows that the language of Simple is less complex than that of Main primarily due to the use of shorter sentences, as opposed to drastically simplified syntax or vocabulary. Comparing the two language varieties by the Gunning readability index supports this conclusion. We also report on the topical dependence of language complexity, that is, that the language is more advanced in conceptual articles compared to person-based (biographical) and object-based articles. Finally, we investigate the relation between conflict and language complexity by analyzing the content of the talk pages associated to controversial and peacefully developing articles, concluding that controversy has the effect of reducing language complexity.
[1] Yasseri T, Kornai A, Kertész J (2012) PLoS ONE 7(11): e48386.