Publishing Partner: Cambridge University Press CUP Extra Publisher Login
amazon logo
More Info


New from Oxford University Press!

ad

Raciolinguistics

Edited by H. Samy Alim, John R. Rickford, and Arnetha F. Ball

Raciolinguistics "Brings together a critical mass of scholars to form a new field dedicated to theorizing and analyzing language and race together."


New from Cambridge University Press!

ad

Sociolinguistics from the Periphery

By Sari Pietikäinen, FinlandAlexandra Jaffe, Long BeachHelen Kelly-Holmes, and Nikolas Coupland

Sociolinguistics from the Periphery "presents a fascinating book about change: shifting political, economic and cultural conditions; ephemeral, sometimes even seasonal, multilingualism; and altered imaginaries for minority and indigenous languages and their users."


Academic Paper


Title: Evaluating authorship distance methods using the positive Silhouette coefficient
Author: Robert Layton
Institution: University of Sheffield
Author: Paul Watters
Homepage: http://www.comp.mq.edu.au/~pwatters
Institution: University of Sheffield
Author: Richard Dazeley
Institution: The University of Ballarat
Linguistic Field: Computational Linguistics
Abstract: Unsupervised Authorship Analysis (UAA) aims to cluster documents by authorship without knowing the authorship of any documents. An important factor in UAA is the method for calculating the distance between documents. This choice of the authorship distance method is considered more critical to the end result than the choice of cluster analysis algorithm. One method for measuring the correlation between a distance metric and a labelling (such as class values or clusters) is the Silhouette Coefficient (SC). The SC can be leveraged by measuring the correlation between the authorship distance method and the true authorship, evaluating the quality of the distance method. However, we show that the SC can be severely affected by outliers. To address this issue, we introduce the Positive Silhouette Coefficient, given as the proportion of instances with a positive SC value. This metric is not easily altered by outliers and produces a more robust metric. A large number of authorship distance methods are then compared using the PSC, and the findings are presented. This research provides an insight into the efficacy of methods for UAA and presents a framework for testing authorship distance methods.

CUP AT LINGUIST

This article appears IN Natural Language Engineering Vol. 19, Issue 4.

Return to TOC.

Add a new paper
Return to Academic Papers main page
Return to Directory of Linguists main page