BNCLab

Exploring individual and social variation

BNCLab was created as part of the ESRC-funded project The British National Corpus (BNC) as a sociolinguistic dataset: Exploring individual and social variation, carried out by a team including Vaclav Brezina, Dana Gablasova, Tony McEnery, Miriam Meyerhoff and Susan Reichelt.

The project focuses on a variety of sociolinguistic variables and aims to provide new insights into theories of language change on community, generation, and individual levels. The data used in this platform were extracted from the demographic part of the British National Corpus 1994 and the Spoken British National Corpus 2014.

 http://corpora.lancs.ac.uk/bnclab/

Screenshot of the Wmatrix tag wizard

Wmatrix

Corpus tagging, analysis, comparison and exploration
Screenshot of web page for the First Workshop on Language Models for Low-Resource Languages

Workshop on Language Models for Low-Resource Languages

20 January 2025