Measuring Linguistic Diversity During COVID-19

Dunn, J.; Coupe, T.; & Adams, B. (2020). "Measuring Linguistic Diversity During COVID-19." Proceedings of the 4th Workshop on NLP and Computational Social Science. Association for Computational Linguistics. 1-10. Abstract. Computational measures of linguistic diversity help us understand the linguistic landscape using digital language data. The contribution of this paper is to calibrate measures of

Mapping Languages: The Corpus of Global Language Use

Dunn, J. (2020). "Mapping Languages: The Corpus of Global Language Use." Language Resources and Evaluation. 54: 999-1018. Abstract. This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages

Geographically-Balanced Gigaword Corpora for 50 Language Varieties

Dunn, J. & Adams, B. (2020). "Geographically-Balanced Gigaword Corpora for 50 Language Varieties." In Proceedings of the Language Resources and Evaluation Conference. European Language Resources Association. 2528-2536. Abstract. While text corpora have been steadily increasing in overall size, even very large corpora are not designed to represent global population demographics. For example, recent work has

Mapping Languages and Demographics

Dunn, J. and Adams, B. (2019). "Mapping Languages and Demographics with Georeferenced Corpora." In Proceedings of Geocomputation 2019. Abstract. This paper evaluates large georeferenced corpora, taken from both web-crawled and social media sources, against ground-truth population and language-census datasets. The goal is to determine (i) which dataset best represents population demographics; (ii) in what parts