Leveraging GPT for the Generation of Multi-Platform Social Media Datasets for Research
Social media datasets are essential for research on disinformation, influence operations, social sensing, hate speech detection, cyberbullying, and other significant topics. However, access to these datasets is often restricted due to costs and platform regulations. As such, acquiring datasets that span multiple platforms which are crucial for a comprehensive understanding of the digital ecosystem is particularly challenging. This paper explores the potential of large language models to create lexically and semantically relevant social media datasets across multiple platforms, aiming to match the quality of real datasets. We employ ChatGPT to generate synthetic data from a real dataset consisting of posts from three different social media platforms. We assess the lexical and semantic properties of the synthetic data and compare them with those of the real data. Our empirical findings suggest that using large language models to generate synthetic multi-platform social media data is promising. However, further enhancements are necessary to improve the fidelity of the outputs.
doi
10.1145/3648188.3675153
isbn
979-8-4007-0595-3
name
Leveraging GPT for the Generation of Multi-Platform Social Media Datasets for Research
source
publisher_html_pdf
acm_url
https://dl.acm.org/doi/10.1145/3648188.3675153
authors
Henry Tari, M. Danial Khan, Justus Rutten, Darian Othman, Rishabh Kaushal, Thales Bertaglia, Adriana Iamnitchi
doi_url
https://doi.org/10.1145/3648188.3675153
license
© Copyright held by the owner/author(s). Publication rights licensed to ACM.
summary
Social media datasets are essential for research on disinformation, influence operations, social sensing, hate speech detection, cyberbullying, and other significant topics. However, access to these datasets is often restricted due to costs and platform regulations. As such, acquiring datasets that span multiple platforms which are crucial for a comprehensive understanding of the digital ecosystem is particularly challenging. This paper explores the potential of large language models to create lexically and semantically relevant social media datasets across multiple platforms, aiming to match the quality of real datasets. We employ ChatGPT to generate synthetic data from a real dataset consisting of posts from three different social media platforms. We assess the lexical and semantic properties of the synthetic data and compare them with those of the real data. Our empirical findings suggest that using large language models to generate synthetic multi-platform social media data is promising. However, further enhancements are necessary to improve the fidelity of the outputs.
arxiv_url
https://arxiv.org/abs/2407.08323
published
2024-09-10
conference
HT '24: 35th ACM Conference on Hypertext and Social Media, Poznan, Poland, September 2024
open_access
true
acm_html_url
https://dl.acm.org/doi/full/10.1145/3648188.3675153
displayAuthor
Henry Tari , Maastricht University, Netherlands, h.tari@student.maastrichtuniversity.nl; M. Danial Khan , Maastricht University, Netherlands, md.khan@student.maastrichtuniversity.nl; Justus Rutten , Maastricht University, Netherlands, justus.rutten@student.maastrichtuniversity.nl; Darian Othman , Maastricht University, Netherlands, do.othman@student.maastrichtuniversity.nl; Thales Bertaglia , Utrecht University, Netherlands, t.f.costabertaglia@uu.nl; Rishabh Kaushal , Maastricht University, Netherlands, rishabh.kaushal@maastrichtuniversity.nl; Adriana Iamnitchi , Maastricht University, Netherlands, a.iamnitchi@maastrichtuniversity.nl
displayPublishTime
2024-09-10