Predicting twitter user demographics using distant supervision from website traffic data

Aron Culotta, Nirmal Kumar Ravi, Jennifer Cutler

Research output: Contribution to journalArticlepeer-review

36 Scopus citations


Understanding the demographics of users of online social networks has important applications for health, marketing, and public messaging. Whereas most prior approaches rely on a supervised learning approach, in which individual users are labeled with demo-graphics for training, we instead create a distantly labeled dataset by collecting audience measurement data for 1,500 websites (e.g., 50% of visitors to are estimated to have a bachelor's degree). We then fit a regression model to predict these demographics from information about the followers of each website on Twitter. Using patterns derived both from textual content and the social network of each user, our final model produces an average held-out correlation of .77 across seven difierent variables (age, gender, education, ethnicity, income, parental status, and political preference). We then apply this model to classify individual Twitter users by ethnicity, gender, and political preference, finding performance that is surprisingly competitive with a fully supervised approach.

Original languageEnglish (US)
Pages (from-to)389-408
Number of pages20
JournalJournal of Artificial Intelligence Research
StatePublished - Feb 2016

ASJC Scopus subject areas

  • Artificial Intelligence


Dive into the research topics of 'Predicting twitter user demographics using distant supervision from website traffic data'. Together they form a unique fingerprint.

Cite this