Pith. sign in

REVIEW 1 cited by

Entropy-based Classification of 'Retweeting' Activity on Twitter

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1106.0346 v1 pith:7WBRWIRY submitted 2011-06-02 cs.SI cs.CY

classification cs.SIcs.CY
keywords activitytwitterclassificationretweetinguseractivitiescontentautomatic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Twitter is used for a variety of reasons, including information dissemination, marketing, political organizing and to spread propaganda, spamming, promotion, conversations, and so on. Characterizing these activities and categorizing associated user generated content is a challenging task. We present a information-theoretic approach to classification of user activity on Twitter. We focus on tweets that contain embedded URLs and study their collective `retweeting' dynamics. We identify two features, time-interval and user entropy, which we use to classify retweeting activity. We achieve good separation of different activities using just these two features and are able to categorize content based on the collective user response it generates. We have identified five distinct categories of retweeting activity on Twitter: automatic/robotic activity, newsworthy information dissemination, advertising and promotion, campaigns, and parasitic advertisement. In the course of our investigations, we have shown how Twitter can be exploited for promotional and spam-like activities. The content-independent, entropy-based activity classification method is computationally efficient, scalable and robust to sampling and missing data. It has many applications, including automatic spam-detection, trend identification, trust management, user-modeling, social search and content classification on online social media.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Entropy and type-token ratio in gigaword corpora

    cs.CL 2024-11 conditional novelty 5.0 of 10

    Word entropy and type-token ratio in billion-token corpora are linked by an asymptotic formula built from Zipf and Heaps laws, confirmed across English, Spanish and Turkish texts.

Pith tools