Identifying Social Media Accounts That Explain Fandom Hashtags
While many hashtags on social media platforms are self-explanatory, the meanings of proliferating community-specific custom hashtags are often opaque for outsiders.

Identifying Social Media Accounts That Explain Fandom Hashtags

Liuyun Ling, Keishi Tajima

Published in: HT ’26 · DOI: 10.1145/3800935.3830960 · License: CC BY 4.0

Keywords: social network, entity linking, polysemy, disambiguation

Abstract

While many hashtags on social media platforms are self-explanatory, the meanings of proliferating community-specific custom hashtags are often opaque for outsiders. Fandom hashtags created by fan communities of celebrities or influencers are a notable subset of these community-specific hashtags. This paper proposes a method to identify the key accounts that best explain the meaning of these fandom hashtags. Our baseline approach collects fan users employing a target hashtag, retrieves their followees, and ranks the followees by their follower count within the collected fan users. This method encounters two primary challenges. First, in tightly-knit communities, the most-followed users are not necessarily the key figure in the community. To identify the key figure in the community, we combine an influence estimation model with the ranking by follower count. Second, some hashtags are polysemous, i.e., used with different meanings across disparate communities. To distinguish polysemous hashtags from single-meaning ones, we cluster the followees of the collected fan users to isolate distinct communities that use the same hashtag with different meanings. To detect less prominent usages of polysemous hashtags, we repeat this process for multiple distinct time periods. Experimental results using data from X demonstrate the effectiveness of our method.

CCS Concepts: • Information systems → Social networking sites ;

Full text

1 Introduction

Hashtags play a crucial role on contemporary social media platforms, such as X (formerly Twitter). Users rely on hashtags to navigate vast amounts of information, follow ongoing events, and engage with communities around shared interests. Thus, identifying the meanings of emerging hashtags is an important task.

Many widely used hashtags consist of ordinary words or phrases that directly convey their intended meanings. Examples include #Christmas, #ChampionsLeague, and #MondayMotivation. These tags are largely self-explanatory and accessible to a broad audience.

In addition to such general-purpose markers, many hashtags originate within specific communities. These community-originated hashtags often carry specialized meanings that are not immediately apparent from the words themselves. For instance, #TBT refers to the practice of posting nostalgic content on Thursdays (Throwback Thursday). While now widely recognized, its meaning was initially clear only to members of a particular community.

A prominent subcategory of community-originated hashtags is fandom hashtags, which are associated with fan communities built around some key figures. Examples include #SwiftOutfit, used for posts about American singer Taylor Swift's fashion, and #ArmySelcaDay, popular within the fan community of K-pop group BTS. As fandom culture continues to grow on social platforms, understanding newly emerging fandom hashtags is increasingly important for casual users and trend analysis systems.

Although the meanings of these hashtags are rarely explicitly documented, these fandom hashtags are usually centered around one or more social media accounts corresponding to the key figures. In many cases, identifying and highlighting these accounts to users can provide immediate insight into the meaning of the hashtags.

This paper propose a simple yet effective method to identify social media accounts that best explain the meanings of fandom hashtags. First, we collect fan users who have posted with the target hashtag. Next, we retrieve the accounts they follow, and rank these followees based on their follower counts within this fan user set. This localized ranking serves as a baseline proxy for identifying the accounts representing the key figures associated with the hashtag.

While effective in many cases, this approach faces two primary challenges. First, in tightly-knit fan communities, the most-followed accounts are not always the most influential one, nor do they necessarily represent the key figures associated with the fandom of the communities. To address this discrepancy, we assess the influence of candidate accounts using an influence estimation model, and select the most influential one.

Second, some hashtags are polysemous, for which we must identify the most representative accounts for each distinct meaning. Therefore, we need to distinguish between accounts associated with the same meaning and accounts associated with different meanings. To achieve this, we cluster the candidate accounts selected by the process above based on the similarity of their follower bases, identifying distinct subgroups. Additionally, we repeat this process at several distinct time intervals to increases our chances of detecting multiple meanings including less prominent ones.

We evaluate our method using 17 fandom hashtags collected from X. The results demonstrate that our approach effectively identifies the accounts that best explain the meanings of fandom hashtags for both single-meaning hashtags and polysemous ones.

2 Related Work

This study is most closely related to two research areas: word sense disambiguation and entity linking. In this section, we survey these fields and also discuss the related task of named-entity recognition.

2.1 Word Sense Disambiguation

Polysemy has long been a fundamental challenge in natural language processing [2, 5, 29]. Classical word sense disambiguation (WSD) methods have relied on curated lexical resources such as WordNet [4, 22] and supervised models trained on manually annotated corpora [26]. These approaches typically use contextual features, collocation statistics, or semantic similarity to determine the most appropriate sense of a word. Recent advances in language models such as BERT [9] and RoBERTa [19] have led to significant improvements in WSD [14, 21].

In traditional WSD, the task is a classification problem: given a fixed inventory of senses for a polysemous word, the system selects the correct one for each occurrence of the word based on its context [12]. In contrast, our task is to distinguish polysemous hashtags from single-meaning hashtags, and create the inventories for them.

Polysemy is also a critical issue in internet applications such as web search [1, 37]. In these domains, disambiguation is complicated by the informal, dynamic, and community-specific nature of language on the web. This is especially true for social media posts, which are often short and rely on context from previous posts.

Hashtags on social media present a particularly challenging form of polysemy, as their meanings shift over time and vary across communities. It has prompted the development of specialized clustering methods for hashtags. For instance, [35, 36] proposed clustering hashtags based on the temporal similarity of their occurrences. Alternatively, [39] clustered hashtags by detecting community structures within their co-occurrence graphs, while [44] integrated both temporal proximity and co-occurrence relationships.

We cluster hashtags based on user community structure. A similar idea was first proposed for disambiguating tags in folksonomies [3, 43]. These studies construct a graph of users, tags, and tagged items, and detect user or item communities to disambiguate polysemous tags. Because a similar correlation between hashtag adoption and community structure has been observed on Twitter [33], this approach is also expected to be effective for hashtags on Twitter.

Our work draws inspiration from these community-aware approaches but shifts the focus from clustering polysemous hashtags to identifying social media accounts that best explain the meanings of previously unseen hashtags. This account-centric perspective is particularly suitable for fandom hashtags, where key meanings are often embodied in central accounts in the community.

In summary, while prior research on polysemous words or polysemous hashtags has primarily focused on clustering hashtag usages or assigning them to predefined senses, the emphasis of our method is on distinguishing polysemous hashtags and on linking each detected meaning to an explanatory account—a novel approach particularly effective for understanding fandom hashtags.

2.2 Entity Linking

Entity linking (EL) is the task of connecting a textual mention of an entity (such as a person, organization, or event) to a corresponding entry in a structured knowledge base (e.g., Wikipedia). It is a critical step in information extraction and knowledge graph construction. Traditional EL approaches typically rank candidates based on contextual similarity, popularity, and coherence with nearby entities in the text [7, 23, 30]. Recent advances have leveraged deep learning and pretrained language models to improve EL performance. Models such as BLINK [41] and GENRE [6] use dense embeddings and sequence generation to directly link mentions to entities with high accuracy. Some methods also incorporate global coherence by modeling entity relations across the entire document [13, 16].

While most EL research focuses on linking to encyclopedic knowledge bases, several recent studies explore linking to domain-specific or user-generated content, such as product catalogs, biomedical databases, or social media profiles [20, 27].

Our task shares commonalities with EL, as we also link a surface-level expression (a hashtag) to an informative external resource (a social media account). In our task, however, the linking is not based on direct textual matching. Because our targets are fandom hashtags, we detect the key figures in the communities through the estimation of their popularity and influence in the community.

2.3 Named Entity Recognition

Named Entity Recognition (NER) is another foundational task in natural language processing that involves identifying and classifying entities such as people, organizations, locations, and creative works in text [18, 25]. Classical NER approaches relied on rule-based systems and statistical models such as hidden Markov models (HMM) and conditional random fields (CRF), which require extensive annotated data and handcrafted features [10, 28].

More recently, neural architectures have also significantly improved the performance of NER, especially in low-resource and noisy environments [38, 40]. These models leverage contextual embeddings to better capture the syntactic and semantic nuances required for accurate entity classification.

Several studies have proposed adaptations of NER for social media platforms like Twitter, where text is often informal, abbreviated, or creatively structured [8, 24, 32]. These methods incorporate additional social or contextual signals such as user mentions, hashtags, or metadata to improve robustness.

NER is conceptually related to our task in that both involve linking linguistic expressions to real-world referents. However, NER focuses more on the detection of entity mentions within the text, whereas in our task, the target hashtags are given. Furthermore, we aim to identify specific social media accounts that serve as semantic references for interpreting the meaning of fandom hashtags. In many cases, the accounts we identify may themselves be named entities (e.g., the official account of a celebrity), but their names are not always explicitly mentioned in the text surrounding the hashtags and must be inferred through social context.

3 Proposed Method

Our method consists of the following steps:

    Collect users who have posted with the target hashtag (assumed to be "fan users"). We then retrieve the followee lists of these users and rank the followees by their follower count within this fan user set. The top-ranked accounts form our initial candidate list.\

    Assess the influence of each candidate account and rerank the list to avoid prioritizing accounts that have many followers but lack significant influence within the fan community.\

    Determine if the target hashtag is polysemous by extracting subgroup structure from the follow graph consisting of the candidate accounts and the fan users.\

    Repeat the process above using data from different time periods to detect less prominent usages that may be prominent only in specific time periods.\

    If the hashtag is determined to have a single meaning, select the most influential account. If we detect multiple meanings, select the most influential account from the candidate subset corresponding to each meaning.\

In this section, we explain the details of each step.

3.1 Identifying Candidate Accounts

We first collect posts containing the target hashtag and identify the users who posted them. We then obtain the followee lists of these users and aggregate them to construct a follow graph. From this graph, we compile a ranked list of the accounts most commonly followed by the users in our collected set.

This simple approach could potentially rank accounts that are popular across the entire platform, not just within the target community. However, this phenomenon did not occur in our experiments. If necessary, one could normalize an account's follower count within the fan community by its total follower count on the entire platform. In our experiments, however, this approach over-penalized important community accounts while sometimes elevating niche fan accounts that have followers only within the community. Therefore, we adopted the simpler ranking method.

To mitigate the risk of selecting less influential accounts, we use this initial ranking only to create a candidate set, which we then rerank based on influence. In our experiments, we select the top 10 accounts from this initial ranking to form the candidate set.

3.2 Reranking by Influence

To rerank the candidates, we assess their influence using a modification of the Linear Influence Model [42]. This model evaluates a user's influence based on the number of likes and reposts their posts receives. We consider only original posts (authored by the user) and quoted reposts (reposts with added commentary), as reactions to these posts are direct indicators of the user's influence. In contrast, we exclude direct reposts (without comments) and simple replies, as the engagement they receive is heavily influenced by the original poster.

To formalize this, let $p{u,1}, ldots, p{u,nu}$ denote user $u$’s relevant posts over a time period, and let $t{u,1}, ldots, t{u,nu}$ be their timestamps. We include all the original posts and the quoted posts by $u$ no matter it includes the target hashtag or not. Let $E{u,i}$ be the total number of likes and reposts for post $p{u,i}$. We define $F(u)$, the total influence of user $u$ over the period as:


This formulation, rather than a simple sum of $E{u,i}$, prevents users from achieving high scores simply by posting frequently. Our goal is to prioritize users who consistently post content that receives high engagement. While there is room for improvement in this formula, it performed better in our experiments than other variations.

Table 1: The 17 Hashtags (10 Single-meaning and 7 Polysemous) in our Experiment and their Meanings

Hashtag

Meaning(s)


10 single-meaning hashtags



#iroastream

clips from live streaming by YouTuber @Laureniroas


#KuzuArt

fan art collection for YouTuber @VampKuzu


#KanatArt (in Japanese)

fan art collection for YouTuber @amanekanatach


#SakunaArt (in Japanese)

fan art collection for YouTuber @yuukisakuna


#SaEgusa (in Japanese)

fan art collection for YouTuber @333akina


#IroEs (in Japanese)

fan art collection for YouTuber @Laureniros


#SaraSeizu (in Japanese)

fan art collection for YouTuber @SaraHoshikawa


#Esuko-to (in Japanese)

fan art collection for YouTuber @FuwaMinato


#KongoRikiyaZou (in Japanese)

fan art collection for YouTuber @reiToyarei


#HiguchiKaedeLIVEBREAKING (in Japanese)

live streaming by live streamer @HiguchiKaede


7 polysemous hashtags



#lisa

a member of Korean girl group Blackpink

a Japanese singer LiSA

#Ras

a game live streamer Ras

a character in a game BangDream

#Haru (in Japanese)

a member of a Japanese pop group Bullet Train

a member of a K-pop group Nexz

#peach (in Japanese)

an airline company

a kind of fruit

#Luka (in Japanese)

a lesser panda in Ueno Zoo

a virtual vocaloid

#AGF

Animate Girls Festival

a food/beverage company AGF

#Eve

Christmas Eve

a singer Eve

3.3 Clustering Candidates

While ordinary polysemous words can take on multiple meanings within a single community depending on context, a single hashtag is rarely used with multiple meanings within the same group, as hashtags function as identifiers of specific topics. Therefore, polysemous hashtags typically emerge only when two disjoint communities happen to adopt the same hashtag to refer to different topics.

Furthermore, users employing a community-specific hashtag with the same intended meaning inherently form a unified community structure. This is particularly evident with fandom hashtags, where users sharing a hashtag's meaning naturally form a network centered around the key figure associated with that hashtag.

In summary, we assume that users employing the same hashtag with different meanings belong to distinct communities, whereas users sharing a hashtag's meaning belong to a single community. Based on this premise, we distinguish polysemous hashtags from single-meaning hashtags by detecting whether multiple, separate communities have adopted the same hashtag.

To detect these community structures within the network of candidate followees and their followers, we cluster the candidate followees based on the structural similarity of their follower bases. Specifically, we measure the similarity between any pair of candidate followees using the Jaccard similarity coefficient [15] applied to the aggregated followee sets of their respective followers. The set of accounts followed by a user's followers serves as an effective proxy for the collective interests of that user's audience.

Formally, let C be the set of candidate accounts. Let I(u) be the set of followers (in-neighbors) of an account u, and let O(u) be the set of followees (out-neighbors) of an account u. For any pair of accounts u, v ∈ C, we calculate the similarity J(u, v) as:


(1)

An alternative would be to compute the Jaccard similarity directly between the follower sets of u and v (I(u) and I(v)), a technique that has proven valuable in other social media analyses [11, 17, 34]. However, the strength of our approach is its ability to utilize both direct and indirect relationships in case of sparse graph structure.

Given these pairwise similarities, we cluster the accounts using a bottom-up agglomerative clustering algorithm. We start by assigning each account to its own cluster and repeatedly merge the most similar pair of clusters until the similarity between the closest pair falls below a threshold. In our experiments, we used a threshold of 0.3; however, the results remained unchanged for any threshold value between 0.1 and 0.4. In our experiments, the similarity between two users belonging to different communities that use the same hashtag with different meanings was always near zero, which made it very easy to distinguish them from pairs belonging to the same community. It is probably because two communities that partially overlap rarely use the same hashtag with different meanings. Thus, our method is not sensitive to this threshold value.

3.4 Multiple Time Periods

We repeat the entire process at different time periods, generating distinct rankings consisting of different sets of candidate followees. Different sets of candidates are obtained not because the follow graph structure changes over time but because different sets of posts containing the target hashtag are obtained in the first step.

This increases our chances of detecting multiple hashtag meanings. When a hashtag is polysemous, one meaning is often dominant, making it difficult to detect other meanings from a single, long-term data sample. However, less popular meanings sometimes experience short bursts of activity, temporarily surpassing the dominant one. By analyzing shorter time windows, we can capture these bursts. If we aggregated all data over a long period, these signals would be lost.

For example, our experiments on the hashtag #lisa revealed two meanings: a Korean singer and a Japanese singer. Posts about the Korean singer are typically dominant. However, during a week when the Japanese singer held a live performance, related posts surged, causing accounts associated with her to enter our top-10 candidate list. Clustering the candidates from that specific week revealed two distinct clusters, allowing us to successfully identify both meanings of the hashtag.

3.5 Candidate Selection

For hashtags determined to have a single meaning, we select the top-ranked account from our final influence-based ranking. For hashtags with multiple meanings, we identify the highest-ranked account within each detected meaning cluster to serve as the explanatory account for that sense.

4 Experiments

We evaluated our approach using data collected from X. Since no prior research has addressed the identification of social media accounts that best explain fandom hashtags, we establish the validity of our approach by evaluating its overall effectiveness alongside an analysis of its core components.

4.1 Dataset

We manually selected 17 fandom hashtags, detailed in Table 1, consisting of 10 single-meaning and 7 polysemous hashtags. While we used well-known examples like #SwiftOutfit in Section 1, such hashtags are not ideal for our experiment because their meanings are easily discoverable via a simple web search. To simulate a more realistic scenario, we selected lesser-known fandom hashtags whose meanings are more difficult to ascertain.

The communities surrounding online streamers (e.g., YouTubers) serve as a rich source for such examples. Because many creators in this category are known only to niche audiences, it is difficult for outsiders to infer the context of their associated fandom hashtags.

Additionally, we avoided hashtags that are identical to the key figure's account name, as they present an artificially trivial task. Instead, we focused on hashtags representing specific subtopics related to the key figure, such as designated tags for fan art.

For users outside the corresponding fan communities, the meanings of the hashtags listed in Table 1 are difficult to deduce from the text alone. However, because each hashtag ultimately relates to a specific entity, we were able to establish their ground-truth meanings by manually identifying the representative accounts associated with those entities.

Our data collection encompasses both English and Japanese, as the United States and Japan are the countries with the largest (111.3M) and the second largest (75.8M) X user base by a big margin over the third largest India (27.3M) [31].

For each hashtag, we collected up to 500 recent posts over a 7-day period, repeating this process four times. As explained in Section 3, analyzing multiple datasets from shorter, distinct time periods—rather than a single dataset spanning a prolonged interval—increases the likelihood of detecting multiple meanings. We selected a 7-day window because posting activity on X exhibits a weekly cycle. While shorter intervals could increase the likelihood of detecting multiple meanings, they would also elevate the risk of noise.

We then retrieved the followees of each post's author to establish a pool of candidate accounts. Finally, to evaluate the influence of these candidates, we collected their posts published during the corresponding 7-day window.

4.2 Overall Effectiveness

Table 2: Summary of the Results


single-meaning

polysemous

Number of Cases

10

7

The best answer in top-10

10

7

The best answer at top-1

7

3

Second meaning identified

N/A

4

The results for the 17 hashtags are summarized in Table 2. For all 17 cases, the correct explanatory account was present within the top-10 candidate list. After reranking by influence, our method ranked the most explanatory account at the top for 7 of the 10 single-meaning hashtags and for 3 of the 7 polysemous hashtags. For the polysemous cases, our method successfully identified the secondary meaning for 4 of the 7 hashtags.

4.3 Analysis of Core Components

We next present several case studies to analyze the effectiveness of each component of our method.

Figure 1

In the graph at the top, the bar for VampKuzu is located at the top of the chart and is the second longest. The longest bar belongs to reiToyarei, which is located in the fifth position from the top. In the graph at the bottom, the bar for VampKuzu is located in the fifth position from the top and is the second longest. The longest bar belongs to FuwaMinato, which is located in the sixth position from the top.

Figure 1: The top 10 answers for #KuzuArt in the order of their initial ranking from top to bottom. X-axis represents the degree of influence. Top: based on the data in Aug.15-22, 2024. Bottom: based on the data in Aug. 28-Sep. 04, 2024.

Figure 1 shows the results for a single-meaning hashtag #KuzuArt, for which our method failed to rank the best answer, @VampKuzu, at the top. The graph at the top shows the result based on the data collected in Aug. 15-22, 2024 and the graph at the bottom shows the result for Aug. 29-Sep. 4, 2024. In both, the accounts are ordered vertically by their initial follower-based rank, while the horizontal bar length represents their influence score, F(u). The final ranking is determined by this influence score.

In the top graph, the correct account @VampKuzu is ranked first initially but is reranked to second place by influence, behind another YouTuber, @reiToyarei, with whom @VampKuzu collaborated that week. In the bottom graph, @VampKuzu starts at 5th and is reranked to 2nd, surpassed by @FuwaMinato, a YouTuber with a significant fan overlap with @VampKuzu. This example illustrates the challenge of correctly selecting the single most explanatory account when activities across communities (such as collaborations) or dense interconnections between communities exist.

The change in the initial top-10 ranking between the two weeks was not due to changes in follow relationships but resulted from sampling a different set of 500 posts, and thus a different set of fan users. Aggregating the influence scores (F(u)) over both weeks would have correctly ranked @VampKuzu first. However, extending the observation period reduces the chance of detecting secondary meanings for polysemous hashtags, as transient signals get diluted. This reveals a trade-off. A potential solution, left for future work, is a hybrid approach that combines long-term influence data with short-term polysemy detection, perhaps with dynamic observation windows.

The graphs in Figure 1 also validate our reranking step by showing that a high follower count does not guarantee high influence. For example, @asapclub, ranked 7th initially in the bottom graph, has a negligible influence score.

Table 3: The Jaccard Similarity matrix for 10 initial candidates for a single-meaning hashtag #KuzuArt for Dec. 22, 2024. The top-ranked account (1st row and 1st column) has high similarity with all the other accounts. Many single-meaning hashtags have similar similarity matrices.


V

n

K

F

r

c

L

k

h

V

VampKuzu

–

.79

.81

.69

.56

.58

.62

.60

.53

.59

nijisanjiapp

.79

–

.75

.63

.52

.61

.65

.64

.55

.57

Kanae2434

.81

.75

–

.66

.56

.66

.73

.63

.65

.64

FuwaMinato

.69

.63

.66

–

.60

.73

.52

.67

.46

.51

reiToyarei

.56

.52

.56

.60

–

.57

.44

.54

.35

.38

Laureniroas

.58

.61

.66

.73

.57

–

.55

.82

.49

.49

chronoirinfo

.62

.65

.73

.52

.44

.55

–

.52

.83

.75

honmonoibrahim

.60

.64

.63

.67

.54

.82

.52

–

.45

.49

kanaeiroiro

.53

.55

.65

.46

.35

.49

.83

.45

–

.71

Vllexceed

.59

.57

.64

.51

.38

.49

.75

.49

.71

–

Table 4: The Jaccard Similarity matrix for 10 initial candidates for a multi-meaning hashtag #lisa for the week of Nov. 17, 2024. This hashtag is used to mean a Korean singer Lisa and also to mean a Japanese singer LiSA. The candidates clearly partition into two clusters: those related to a Korean singer Lisa (those ranked at ranks 2nd and 4th to 9th) and those related to a Japanese singer LiSA (those ranked at 1st, 3rd, and 10th). The intra-cluster similarities are high (≥ 0.49), while inter-cluster similarities are near zero.


L

w

L

L

B

L

L

L

T

i

LiSAOLive

–

.01

.76

.00

.01

.00

.00

.00

.00

.64

wearelloud

.01

–

.01

.66

.55

.53

.54

.53

.56

.00

LiSASTAFF

.76

.01

–

.00

.01

.00

.00

.00

.00

.74

LISANATIONS

.00

.66

.00

–

.56

.72

.71

.78

.75

.00

BLACKPINK

.01

.55

.01

.56

–

.52

.49

.53

.51

.00

LaliceUpdates

.00

.53

.00

.72

.52

–

.70

.77

.73

.00

LiliesHome

.00

.54

.00

.71

.49

.70

–

.78

.86

.00

Lsglobal

.00

.53

.00

.78

.53

.77

.78

–

.88

.00

TeamLisaPH

.00

.56

.00

.75

.51

.73

.86

.88

–

.00

itadakimasu47

.64

.00

.73

.00

.00

.00

.00

.00

.00

–

Table 5: Accounts that appear in at least one of four candidate lists for the hashtag #lisa produced based on the data collected in four different weeks. Accounts related to the Japanese singer LiSA (LiSAOLiVE, LiSASTAFF, itadakimasu47) appear in the top 10 only for Nov. 17, 2024. If we create a single ranking by using all the data in these four weeks, they would not be included in the top 10 candidate list.


2024.11.17


2024.11.29


2024.12.05


2024.12.12










userid


I(u)


rank


I(u)


rank


I(u)


rank


I(u)


rank

wearelloud

112

2

39

1

126

1

102

1









LiSAOLiVE

115

1















LISANATIONS

87

4

33

4

112

2

92

2









BLACKPINK

82

5

39

1













LaliceUpdates

80

6



93

3

87

3









LiSASTAFF

93

3















ygofficialblink



38

3













LILITEAMTH327





88

4

79

4









Lsglobal

66

8



85

7

79

4









lsloops





88

4

76

7









jennierubyiane



31

5













LiliesHome

72

7



86

6

79

4









numberoneHQ



27

6













officialBLISSOO



27

7













TeamLisaPH

63

9





74

8









LisaRadio





77

8

68

10









oddatelier



27

8













BBUBLACKPINK



27

8













USNation0327





75

9











LSMFRANCE







69

9









itadakimasu47

63

9















blackpinkbabo



26

10













lalaluvlalisa





75

9











Table 3 displays the Jaccard similarity matrix for the candidates for #KuzuArt. The consistently high similarity scores across all pairs (most > 0.5) strongly indicate a single, coherent community, making it easy to classify this as a single-meaning hashtag.

In contrast, Table 4 shows the matrix for the polysemous hashtag #lisa. The candidates clearly partition into two clusters: those related to a Korean singer Lisa (those ranked at ranks 2nd and 4th to 9th) and those related to a Japanese singer LiSA (those ranked at 1st, 3rd, and 10th). Note that hashtags on X are case-insensitive. The intra-cluster similarities are high (≥ 0.49), while inter-cluster similarities are near zero. This is an expected result. Polysemy typically arises when a term is adopted by two largely disjoint communities; if communities overlap significantly, a shared understanding prevents divergent meanings from forming. The clustering is therefore robust and not sensitive to the specific threshold chosen. Our method's failures in identifying secondary meanings were primarily due to the associated accounts not appearing in the initial top-10 candidate list.

Table 5 shows the initial candidate lists for the #lisa hashtag from four different weeks. The three accounts related to the Japanese singer LiSA (LiSAOLiVE, LiSASTAFF, and itadakimasu47) appear in the top 10 only during the week of Nov 17. Without analyzing that specific period, this secondary meaning would have been missed. Aggregating all posts from the four weeks into a single analysis would have suppressed this signal, preventing these accounts from making the candidate list. This confirms that separate rankings for different time periods is crucial for detecting multiple meanings.

5 Conclusion

In this paper, we proposed and evaluated a method for identifying the social media accounts that best explain the meanings of fandom hashtags. Our approach first identifies candidate accounts based on their follower counts within the extracted fan community, then rerank them by their influence. To address polysemy, we identify community structure within the network consisting of the users who use the target hashtag, and their followees. We repeat the whole process at multiple time periods to increase the chance to detect less popular usages of polysemous hashtags.

We evaluated our approach for 17 fandom hashtags collected from X. The experimental results demonstrate that our method is effective. Furthermore, the findings confirm that clustering candidate accounts across multiple time windows successfully detects polysemous hashtags. While our ranking method proved effective, it leaves room for future improvement.

By illuminating the latent meanings of emerging community-oriented hashtags, this work provides a valuable tool for trend analysis systems, content recommendation engines, and also for researchers studying semantic evolution of digital vernacular.

Acknowledgments

This work was supported by JSPS KAKENHI Grant Number 26K02918 and Kayamori Foundation of Informational Science Advancement Grant Number K37XXX668.

References

[1] Vibhanshu Abhishek and Kartik Hosanagar. 2007. Keyword generation for search engine advertising using semantic similarity between terms. In Proc. of the 9th Intl. Conf. on Electronic Commerce (Minneapolis, MN, USA) (ICEC ’07). ACM, New York, NY, USA, 89–94. https://doi.org/10.1145/1282100.1282119

[2] Eneko Agirre and Philip Edmonds (Eds.). 2007. Word sense disambiguation: Algorithms and applications. Text, Speech and Language Technology, Vol. 33. Springer. https://doi.org/10.1007/978-1-4020-4809-8

[3] Ching-man Au Yeung, Nicholas Gibbins, and Nigel Shadbolt. 2009. Contextualising tags in collaborative tagging systems. In Proc. of the 20th ACM Conf. on Hypertext and Hypermedia (Torino, Italy) (HT ’09). ACM, New York, NY, USA, 251–260. https://doi.org/10.1145/1557914.1557958

[4] Satanjeev Banerjee and Ted Pedersen. 2002. An adapted Lesk algorithm for word sense disambiguation using WordNet. In Intl. Conf. on Intelligent Text Processing and Computational Linguistics (Mexico City, Mexico) (CICLing ’02). Springer, 136–145. https://doi.org/10.1007/3-540-45715-111

[5] Michele Bevilacqua, Tommaso Pasini, Alessandro Raganato, Roberto Navigli, 2021. Recent trends in word sense disambiguation: A survey. In Proc. of the 13th Intl. Joint Conf. on Artificial Intelligence (Montreal, Canada) (IJCAI ’21). 4330–4338. https://doi.org/10.24963/ijcai.2021/593

[6] Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. 2021. Autoregressive entity retrieval. In Proc. of the 9th Intl. Conf. on Learning Representations(ICLR 2021). OpenReview.net, 20 pages.

[7] Silviu Cucerzan. 2007. Large-scale named entity disambiguation based on Wikipedia data. In Proc. of the Joint Conf. on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (Prague, Czech Republic) (EMNLP-CoNLL 2007). 708–716.

[8] Leon Derczynski, Diana Maynard, Giuseppe Rizzo, Marieke van Erp, Genevieve Gorrell, Raphael Troncy, Johann Petrak, and Kalina Bontcheva. 2015. Analysis of named entity recognition and linking for tweets. Information Processing & Management 51, 2 (2015), 32–49. https://doi.org/10.1016/j.ipm.2014.10.006

[9] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proc. of the 2019 Conf. of the North American Chapter of the ACL: Human Language Technologies (Minneapolis, MN, USA) (NAACL-HLT 2019). 4171–4186. https://doi.org/10.18653/v1/n19-1423

[10] Jenny Rose Finkel, Trond Grenager, and Christopher Manning. 2005. Incorporating non-local information into information extraction systems by Gibbs sampling. In Proc. of the 43rd Annual Meeting of the ACL (Ann Arbor, MI, USA) (ACL 205). 363–370. https://doi.org/10.3115/1219840.1219885

[11] Michael Fire, Lena Tenenboim, Ofrit Lesser, Rami Puzis, Lior Rokach, and Yuval Elovici. 2011. Link prediction in social networks using computationally efficient topological features. In 2011 IEEE 3rd Intl. Conf. on Social Computing (Boston, MA, USA) (SocialCom 2011). 73–80. https://doi.org/10.1109/PASSAT/SocialCom.2011.20

[12] Stephani Foraker and Gregory L. Murphy. 2012. Polysemy in sentence comprehension: Effects of meaning dominance. Journal of Memory and Language 67, 4 (2012), 407–425. https://doi.org/10.1016/j.jml.2012.07.010

[13] Octavian-Eugen Ganea and Thomas Hofmann. 2017. Deep joint entity disambiguation with local neural attention. In Proc. of the 2017 Conf. on Empirical Methods in Natural Language Processing (EMNLP) (Copenhagen, Denmark). 2619–2629. https://doi.org/10.18653/v1/D17-1277

[14] Luyao Huang, Chi Sun, Xipeng Qiu, and Xuanjing Huang. 2019. GlossBERT: BERT for word sense disambiguation with gloss knowledge. In Proc. of the 2019 Conf. on Empirical Methods in Natural Language Processing and the 9th Intl. Joint Conf. on Natural Language Processing (Hong Kong, China) (EMNLP-IJCNLP ’19). ACL, 3509–3514. https://doi.org/10.18653/v1/D19-1355

[15] Paul Jaccard. 1901. Distribution de la flore alpine dans le bassin des Dranses et dans quelques régions voisines. Bulletin de la Société Vaudoise des Sciences Naturelles 37 (1901), 241–272.

[16] Hangfeng Le and Ivan Titov. 2018. Improving entity linking by modeling latent relations between mentions. In Proc. of the 56th Annual Meeting of the ACL (Melbourne, Australia). 1595–1604. https://doi.org/10.18653/v1/P18-1148

[17] David Liben-Nowell and Jon Kleinberg. 2007. The link-prediction problem for social networks. Journal of the American Society for Information Science and Technology 58, 7 (2007), 1019–1031. https://doi.org/10.1002/asi.20591

[18] Xiao Ling and Daniel S Weld. 2012. Fine-grained entity recognition. In Proc. of the 26th AAAI Conf. on Artificial Intelligence (Toronto, Canada) (AAAI 2012, Vol. 12). 94–100. https://doi.org/10.1609/aaai.v26i1.8122

[19] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A robustly optimized BERT pretraining approach. CoRR abs/1907.11692 (2019), 13 pages. arXiv:1907.11692 https://doi.org/10.48550/arXiv.1907.11692

[20] Lajanugen Logeswaran, Kenton Lee, Kristina Toutanova, Jacob Devlin, Saleema Amershi, and Ming-Wei Chang. 2019. Zero-shot entity linking by reading entity descriptions. In Proc. of the 57th Conf. of the ACL (Florence, Italy) (ACL 2019). 3449–3460. https://doi.org/10.18653/v1/p19-1335

[21] Daniel Loureiro and Alípio Jorge. 2021. Analysis and evaluation of language models for word sense disambiguation. Computational Linguistics 47, 2 (2021), 387–443. https://doi.org/10.1162/colia00405

[22] George A. Miller. 1995. WordNet: a lexical database for English. Commun. ACM 38, 11 (Nov. 1995), 39–41. https://doi.org/10.1145/219717.219748

[23] David Milne and Ian H. Witten. 2008. Learning to link with Wikipedia. In Proc. of the 17th ACM Conf. on Information and Knowledge Management (Napa Valley, CA, USA). ACM, New York, NY, USA, 509–518. https://doi.org/10.1145/1458082.1458150

[24] Seungwhan Moon, Leonardo Neves, and Vitor Carvalho. 2018. Multimodal named entity recognition for short social media posts. In Proc. of the 2018 Conf. of the North American Chapter of the ACL: Human Language Technologies (New Orleans, LA, USA) (NAACL-HLT 2018). 852–860. https://doi.org/10.18653/v1/n18-1078

[25] David Nadeau and Satoshi Sekine. 2007. A survey of named entity recognition and classification. Lingvisticae Investigationes 30, 1 (2007), 3–26. https://doi.org/10.1075/li.30.1.03nad

[26] Roberto Navigli. 2009. Word sense disambiguation: A survey. ACM Comput. Surv. 41, 2, Article 10 (Feb. 2009), 69 pages. https://doi.org/10.1145/1459352.1459355

[27] Yaroslav Nechaev, Francesco Corcoglioniti, and Claudio Giuliano. 2017. Linking knowledge bases to social media profiles. In Proc. of the Symp. on Applied Computing (Marrakech, Morocco) (SAC ’17). ACM, New York, NY, USA, 145–150. https://doi.org/10.1145/3019612.3019645

[28] Nita Patil, Ajay Patil, and B.V. Pawar. 2020. Named entity recognition using conditional random fields. Procedia Computer Science 167 (2020), 1181–1188. https://doi.org/10.1016/j.procs.2020.03.431

[29] Alessandro Raganato, Jose Camacho-Collados, and Roberto Navigli. 2017. Word sense disambiguation: A unified evaluation framework and empirical comparison. In Proc. of the 15th Conf. of the European Chapter of the ACL (Valencia, Spain) (EACL ’17). 99–110. https://doi.org/10.18653/v1/e17-1010

[30] Lev Ratinov, Dan Roth, Doug Downey, and Mike Anderson. 2011. Local and global algorithms for disambiguation to Wikipedia. In Proc. of the 49th Annual Meeting of the ACL (Portland, OR, USA) (ACL 2011). 1375–1384.

[31] World Population Review. 2025. Twitter/X Users by Country 2025. https://worldpopulationreview.com/country-rankings/twitter-users-by-country. Retrieved on October 17, 2025.

[32] Alan Ritter, Sam Clark, Mausam, and Oren Etzioni. 2011. Named entity recognition in tweets: an experimental study. In Proc. of the 2011 Conf. on Empirical Methods in Natural Language Processing (Edinburgh, UK) (EMNLP 2011). 1524–1534.

[33] Daniel Romero, Chenhao Tan, and Johan Ugander. 2013. On the interplay between social and topical structure. In Proc. of the 7th Intl. Conf. on Weblogs and Social Media (Cambridge, MA, USA). 516–525. https://doi.org/10.1609/icwsm.v7i1.14411

[34] Satu Elisa Schaeffer. 2007. Graph clustering. Computer Science Review 1, 1 (2007), 27–64. https://doi.org/10.1016/j.cosrev.2007.05.001

[35] Giovanni Stilo and Paola Velardi. 2014. Temporal semantics: Time-varying hashtag sense clustering. In Proc. of 19th Intl. Conf. on Knowledge Engineering and Knowledge Management (Linköping, Sweden) (EKAW ’14). Springer, 563–578. https://doi.org/10.1007/978-3-319-13704-942

[36] Giovanni Stilo and Paola Velardi. 2017. Hashtag sense clustering based on temporal similarity. Computational Linguistics 43, 1 (2017), 181–200. https://doi.org/10.1162/COLIa00277

[37] Christopher Stokoe, Michael P. Oakes, and John Tait. 2003. Word sense disambiguation in information retrieval revisited. In Proc. of the 26th Annual Intl. ACM SIGIR Conf. on Research and Development in Informaion Retrieval (Toronto, Canada) (SIGIR ’03). ACM, New York, NY, USA, 159–166. https://doi.org/10.1145/860435.860466

[38] Kaveh Taghipour and Hwee Tou Ng. 2015. Semi-supervised word sense disambiguation using word embeddings in general and specific domains. In Proc. of the 2015 Conf. of the North American Chapter of the ACL: Human Language Technologies (Denver, CO, USA) (NAACL HLT 2015). 314–323. https://doi.org/10.3115/v1/n15-1035

[39] Mengmeng Wang and Mizuho Iwaihara. 2015. Hashtag sense induction based on co-occurrence graphs. In Proc. of 17th Asia-PacificWeb Conf. (Guangzhou, China) (APWeb ’15). Springer, 154–165. https://doi.org/10.1007/978-3-319-25255-113

[40] Zhenghui Wang, Yanru Qu, Liheng Chen, Jian Shen, Weinan Zhang, Shaodian Zhang, Yimei Gao, Gen Gu, Ken Chen, and Yong Yu. 2018. Label-aware double transfer learning for cross-specialty medical named entity recognition. In Proc. of the 2018 Conf. of the North American Chapter of the ACL: Human Language Technologies (New Orleans, LA, USA) (NAACL-HLT 2018). 1–15. https://doi.org/10.18653/v1/n18-1001

[41] Ledell Wu, Thibaut Fevry, Edouard Grave, Allen Nie, Yann Dauphin, and Jason Weston. 2020. Scalable zero-shot entity linking with dense entity retrieval. In Proc. of the 2020 Conf. on Empirical Methods in Natural Language Processing(EMNLP 2020). 6397–6407. https://doi.org/10.18653/v1/2020.emnlp-main.519

[42] Jaewon Yang and Jure Leskovec. 2010. Modeling information diffusion in implicit networks. In Proc. of the 10th IEEE Intl. Conf. on Data Mining (Sydney, Australia) (ICDM 2010). 599–608. https://doi.org/10.1109/ICDM.2010.22

[43] Ching-man Au Yeung, Nicholas Gibbins, and Nigel Shadbolt. 2007. Mutual contextualization in tripartite graphs of folksonomies. In Proc. of the 6th Intl. Semantic Web Conf. (Busan, Korea) (ISWC). Springer, 966–970.

[44] Minxin Zou and Mizuho Iwaihara. 2016. Hashtag sense disambiguation based on content and temporal proximities. In DEIM Forum (Fukuoka, Japan). 8 pages.

Source


    Imported from ACM’s structured HTML source. ACM Reference Format: Liuyun Ling and Keishi Tajima. 2026. Identifying Social Media Accounts That Explain Fandom Hashtags. In 37th ACM Conference on Hypertext (HT '26), September 14--18, 2026, London, United Kingdom. ACM, New York, NY, USA 7 Pages. https://doi.org/10.1145/3800935.3830960

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime