Using machine learning to build public policy agenda from social media conversations
This work addresses the time- and labor-intensive process of issue identification for public policy makers, but the results are preliminary and the approach is incremental.
The authors propose a human-augmented ML approach to identify public policy issues from Twitter, using LDA, Top2Vec, and GPT-2. They achieved 'very good' and 'good' inter-rater agreement on narrative readability and coherence, and above-average cosine similarity on three of five narrative themes.
Issue identification and agenda setting represents an important stage in the public policy making process. Traditional approaches for carrying out activities under this stage are time- and labor-intensive on data collection and analysis, in addition to being costly to scale over large geographic areas. In this work we propose a human-augmented machine learning (ML) approach for identifying matters of public interest from social media conversations. The approach consists of five stages namely, input data cleaning and preprocessing, keywords extraction and issue identification, narrative creation, narrative validation, and agenda validation. We implemented experiments to validate the output of our method based on a Twitter dataset and using Latent Dirichlet Allocation (LDA) and Top2Vec for topic modeling. Natural Language Generation (NLG) was achieved using GPT-2 while narrative and agenda validation were based on similarity analysis and human evaluation. We achieved "very good" and "good" inter-rater agreement (IRA) on readability and coherence of agenda narrative generation by our GPT-2 model. On the other hand, IRA was "good" for generated agenda items. We also achieved above average cosine similarity score on at least three out of five reference text (narrative) themes. These results demonstrate that the ML approach represents a promising methodology for identifying issues of public interest from social media conversations.