{"id":390211,"date":"2024-06-29T09:05:53","date_gmt":"2024-06-29T09:05:53","guid":{"rendered":"http:\/\/savepearlharbor.com\/?p=390211"},"modified":"-0001-11-30T00:00:00","modified_gmt":"-0001-11-29T21:00:00","slug":"","status":"publish","type":"post","link":"https:\/\/savepearlharbor.com\/?p=390211","title":{"rendered":"<span>Detecting attempts of mass influencing via social networks using NLP. Part 2<\/span>"},"content":{"rendered":"<div><!--[--><!--]--><\/div>\n<div id=\"post-content-body\">\n<div>\n<div class=\"article-formatted-body article-formatted-body article-formatted-body_version-2\">\n<div xmlns=\"http:\/\/www.w3.org\/1999\/xhtml\">\n<p>In<a href=\"https:\/\/habr.com\/ru\/post\/674126\/\" rel=\"noopener noreferrer nofollow\"> Part 1<\/a> of this article, I built and compared two classifiers to detect trolls on Twitter. You can check it out.<\/p>\n<p>Now, time has come to look more deeply into the datasets to find some patterns using exploratory data analysis and topic modelling.<\/p>\n<p><strong>EDA<\/strong><\/p>\n<p>To do just that, I first created a word cloud of the most common words, which you can see below.<\/p>\n<figure class=\"full-width\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/habrastorage.org\/r\/w1560\/getpro\/habr\/upload_files\/a9b\/42a\/3b8\/a9b42a3b84c362e12df39644ee20a198.png\" width=\"904\" height=\"868\" data-src=\"https:\/\/habrastorage.org\/getpro\/habr\/upload_files\/a9b\/42a\/3b8\/a9b42a3b84c362e12df39644ee20a198.png\"\/><figcaption><\/figcaption><\/figure>\n<p>So, \u201cbutthis\u201d and \u201cbestUSAtoday\u201d are the most frequently used words, which seemingly makes no sense. However, <a href=\"http:\/\/butthis.com\" rel=\"noopener noreferrer nofollow\">butthis.com<\/a> is actually one of the most famous websites within the Twitter account bestUSAtoday considered to be related to the IRA agency. The account itself was suspended by Twitter for violation of the Twitter Rules. Moreover, the <a href=\"http:\/\/butthis.com\" rel=\"noopener noreferrer nofollow\">butthis.com<\/a> domain no longer exists as well, so I could only look at the screenshots of the site and, from what I saw, all its content was political. This word was mentioned 350 times in 2288 tweets, which makes this external resource popular among trolls.\u00a0<\/p>\n<p>Other most common words were \u201disis\u201d (ISIS, i.e. Islamic State of Iraq and the Levant), \u201dpolice\u201d, \u201dpotus\u201d (President of the United States), \u201dnews\u201d and words related to the main competitors of the 2016 US Presidential Elections Hillary Clinton and Donald Trump.<\/p>\n<p><strong>Topic modelling<\/strong><\/p>\n<p>The next step, topic modelling, showed me the most common topics in the tweets under study. For this purpose, I used LDA, which requires a bag-of-words representation of the tweets as its input.\u00a0<\/p>\n<p>As part of this research, I compared Gensim LDA with Scikit Learn LDA, and it turned out that Scikit Learn does not provide a convenient coherence calculation model, which could allow me to quickly obtain the coherence measure given a certain number of topics. That\u2019s why I chose Gensim LDA as the model with a broader range of the required features.\u00a0<\/p>\n<p>Speaking of the coherence measure, it shows the level of correlation between words in a topic. The higher the coherence is the more sense in a topic we will observe. In my experiments, I was looking for the so-called coherence k-value, which represents the peak of the rapid growth of coherence, and found out that the optimal number of the topics would be 7.<\/p>\n<figure class=\"full-width\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/habrastorage.org\/r\/w1560\/getpro\/habr\/upload_files\/497\/415\/2d5\/4974152d53e7f5a2e105f421aa2d07d5.png\" width=\"857\" height=\"742\" data-src=\"https:\/\/habrastorage.org\/getpro\/habr\/upload_files\/497\/415\/2d5\/4974152d53e7f5a2e105f421aa2d07d5.png\"\/><figcaption><\/figcaption><\/figure>\n<p>As for the results of the topic modelling:\u00a0<\/p>\n<figure class=\"full-width\"><img loading=\"lazy\" decoding=\"async\" src=\"https:\/\/habrastorage.org\/r\/w1560\/getpro\/habr\/upload_files\/88b\/196\/4c9\/88b1964c9e3961e91b0f189aa13eaa5e.png\" width=\"827\" height=\"494\" data-src=\"https:\/\/habrastorage.org\/getpro\/habr\/upload_files\/88b\/196\/4c9\/88b1964c9e3961e91b0f189aa13eaa5e.png\"\/><figcaption><\/figcaption><\/figure>\n<p>Topic: 0 is about the above-mentioned troll resources \u201dbestUSAtoday\u201d, <a href=\"http:\/\/butthis.com\" rel=\"noopener noreferrer nofollow\">butthis.com<\/a> and other online resources, such as <a href=\"http:\/\/Newsmax.com\" rel=\"noopener noreferrer nofollow\">Newsmax.com<\/a> or <a href=\"http:\/\/TheHuffingtonPost.com\" rel=\"noopener noreferrer nofollow\">TheHuffingtonPost.com<\/a>, which are private online news and opinion websites in the US. Topic: 5 concerns Donald Trump with words like \u201dIran\u201d, \u201ddeal\u201d, \u201dirandeal\u201d,\u201dpotus\u201d, \u201dirannucleardeal\u201d or \u201dtrump2016\u201d. This topic covers the year of the nuclear deal between Iran and a group of six countries led by POTUS Donald Trump. Another interesting Topic: 4 is dedicated to refugees and migration and includes such words as \u201cattacks\u201d, \u201ccould\u201d or \u201ckill\u201d. In addition, there is Topic: 2, which is about the immigration caused by the Syrian war in the Middle East and about how candidates for the position of the President were going to deal with it. Topic: 6 is fully dedicated to the 2016 elections and the sources of propaganda.<\/p>\n<p><strong>Conclusions and further work<\/strong><\/p>\n<p>So, two classifiers with very high overall accuracy, precision, recall and F-1 score have been built and tested on several features. The experimental results have shown that the tweet text as a feature gives better accuracy than hashtags. The models were trained on large amounts of data and, thus, can be used as a correct solution for detecting attempts at mass influencing. The exploratory analysis and topic modelling allowed me to delve deeper into the actual goals of the trolling. All this makes it possible to conclude that the ability to affect internal affairs of another state exists but, probably, depends on the amount of the troll army used and how they are used.<\/p>\n<p>This study in two parts is initial and can be furthered with deep learning and analysis of visual components (such as images and videos attached to the tweets). To continue with the research, I\u2019m going to build a classifier that will take into account such aspects as the user name, the user picture, kind of account activities (number of retweets and likes), number of tweets per unit of time and some deeper parameters like the type of followers or the sentiment of the tweets.<\/p>\n<p>The only concern related to more advanced analysis and deep learning is that troll accounts will soon be automated and converted into bot accounts. This means that we will soon need to extract features of bots instead of trolls. This might be quite a complicated thing because bots usually have many similar features, but not all of them represent trolls.\u00a0<\/p>\n<\/div>\n<\/div>\n<\/div>\n<p><!----><!----><\/div>\n<p><!----><\/p>\n<div class=\"tm-article-poll-container\"><!--[--><\/p>\n<div class=\"tm-article-poll tm-article-poll_variant-bordered\">\n<div class=\"tm-notice tm-notice_positive tm-article-poll__notice\"><!----><\/p>\n<div class=\"tm-notice__inner\"><!----><\/p>\n<div class=\"tm-notice__content\" data-test-id=\"notice-content\"><!--[--><span>\u0422\u043e\u043b\u044c\u043a\u043e \u0437\u0430\u0440\u0435\u0433\u0438\u0441\u0442\u0440\u0438\u0440\u043e\u0432\u0430\u043d\u043d\u044b\u0435 \u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u0435\u043b\u0438 \u043c\u043e\u0433\u0443\u0442 \u0443\u0447\u0430\u0441\u0442\u0432\u043e\u0432\u0430\u0442\u044c \u0432 \u043e\u043f\u0440\u043e\u0441\u0435. <a rel=\"nofollow\" href=\"\/kek\/v1\/auth\/habrahabr\/?back=\/ru\/articles\/674130\/&#038;hl=ru\">\u0412\u043e\u0439\u0434\u0438\u0442\u0435<\/a>, \u043f\u043e\u0436\u0430\u043b\u0443\u0439\u0441\u0442\u0430.<\/span><!--]--><\/div>\n<\/div>\n<\/div>\n<p><!--[--><\/p>\n<div class=\"tm-article-poll__header\">Do you think it&#8217;s possible to influence a significant number of citizens via social networks?<\/div>\n<div class=\"tm-article-poll__answers\"><!--[--><\/p>\n<div class=\"tm-article-poll__answer\">\n<div class=\"tm-article-poll__answer-data\"><span class=\"tm-article-poll__answer-percent tm-article-poll__answer-percent_winning\">100% <\/span><span class=\"tm-article-poll__answer-label\">Yes<\/span><span class=\"tm-article-poll__answer-votes\">2<\/span><\/div>\n<div class=\"tm-article-poll__answer-bar\">\n<div class=\"tm-article-poll__answer-progress tm-article-poll__answer-progress_winning\" style=\"width: 100%\"><\/div>\n<\/div>\n<\/div>\n<div class=\"tm-article-poll__answer\">\n<div class=\"tm-article-poll__answer-data\"><span class=\"tm-article-poll__answer-percent\">0% <\/span><span class=\"tm-article-poll__answer-label\">No<\/span><span class=\"tm-article-poll__answer-votes\">0<\/span><\/div>\n<div class=\"tm-article-poll__answer-bar\">\n<div class=\"tm-article-poll__answer-progress\" style=\"width: 0%\"><\/div>\n<\/div>\n<\/div>\n<p><!--]--><\/div>\n<div class=\"tm-article-poll__stats\"> \u041f\u0440\u043e\u0433\u043e\u043b\u043e\u0441\u043e\u0432\u0430\u043b\u0438 2 \u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u0435\u043b\u044f.    \u0412\u043e\u0437\u0434\u0435\u0440\u0436\u0430\u0432\u0448\u0438\u0445\u0441\u044f \u043d\u0435\u0442. <\/div>\n<p><!--]--><\/div>\n<p><!--]--><\/div>\n<p> \u0441\u0441\u044b\u043b\u043a\u0430 \u043d\u0430 \u043e\u0440\u0438\u0433\u0438\u043d\u0430\u043b \u0441\u0442\u0430\u0442\u044c\u0438 <a href=\"https:\/\/habr.com\/ru\/articles\/674130\/\"> https:\/\/habr.com\/ru\/articles\/674130\/<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<div><!--[--><!--]--><\/div>\n<div id=\"post-content-body\">\n<div>\n<div class=\"article-formatted-body article-formatted-body article-formatted-body_version-2\">\n<div xmlns=\"http:\/\/www.w3.org\/1999\/xhtml\">\n<p>In<a href=\"https:\/\/habr.com\/ru\/post\/674126\/\" rel=\"noopener noreferrer nofollow\"> Part 1<\/a> of this article, I built and compared two classifiers to detect trolls on Twitter. You can check it out.<\/p>\n<p>Now, time has come to look more deeply into the datasets to find some patterns using exploratory data analysis and topic modelling.<\/p>\n<p><strong>EDA<\/strong><\/p>\n<p>To do just that, I first created a word cloud of the most common words, which you can see below.<\/p>\n<figure class=\"full-width\"><figcaption><\/figcaption><\/figure>\n<p>So, \u201cbutthis\u201d and \u201cbestUSAtoday\u201d are the most frequently used words, which seemingly makes no sense. However, <a href=\"http:\/\/butthis.com\" rel=\"noopener noreferrer nofollow\">butthis.com<\/a> is actually one of the most famous websites within the Twitter account bestUSAtoday considered to be related to the IRA agency. The account itself was suspended by Twitter for violation of the Twitter Rules. Moreover, the <a href=\"http:\/\/butthis.com\" rel=\"noopener noreferrer nofollow\">butthis.com<\/a> domain no longer exists as well, so I could only look at the screenshots of the site and, from what I saw, all its content was political. This word was mentioned 350 times in 2288 tweets, which makes this external resource popular among trolls.\u00a0<\/p>\n<p>Other most common words were \u201disis\u201d (ISIS, i.e. Islamic State of Iraq and the Levant), \u201dpolice\u201d, \u201dpotus\u201d (President of the United States), \u201dnews\u201d and words related to the main competitors of the 2016 US Presidential Elections Hillary Clinton and Donald Trump.<\/p>\n<p><strong>Topic modelling<\/strong><\/p>\n<p>The next step, topic modelling, showed me the most common topics in the tweets under study. For this purpose, I used LDA, which requires a bag-of-words representation of the tweets as its input.\u00a0<\/p>\n<p>As part of this research, I compared Gensim LDA with Scikit Learn LDA, and it turned out that Scikit Learn does not provide a convenient coherence calculation model, which could allow me to quickly obtain the coherence measure given a certain number of topics. That\u2019s why I chose Gensim LDA as the model with a broader range of the required features.\u00a0<\/p>\n<p>Speaking of the coherence measure, it shows the level of correlation between words in a topic. The higher the coherence is the more sense in a topic we will observe. In my experiments, I was looking for the so-called coherence k-value, which represents the peak of the rapid growth of coherence, and found out that the optimal number of the topics would be 7.<\/p>\n<figure class=\"full-width\"><figcaption><\/figcaption><\/figure>\n<p>As for the results of the topic modelling:\u00a0<\/p>\n<figure class=\"full-width\"><figcaption><\/figcaption><\/figure>\n<p>Topic: 0 is about the above-mentioned troll resources \u201dbestUSAtoday\u201d, <a href=\"http:\/\/butthis.com\" rel=\"noopener noreferrer nofollow\">butthis.com<\/a> and other online resources, such as <a href=\"http:\/\/Newsmax.com\" rel=\"noopener noreferrer nofollow\">Newsmax.com<\/a> or <a href=\"http:\/\/TheHuffingtonPost.com\" rel=\"noopener noreferrer nofollow\">TheHuffingtonPost.com<\/a>, which are private online news and opinion websites in the US. Topic: 5 concerns Donald Trump with words like \u201dIran\u201d, \u201ddeal\u201d, \u201dirandeal\u201d,\u201dpotus\u201d, \u201dirannucleardeal\u201d or \u201dtrump2016\u201d. This topic covers the year of the nuclear deal between Iran and a group of six countries led by POTUS Donald Trump. Another interesting Topic: 4 is dedicated to refugees and migration and includes such words as \u201cattacks\u201d, \u201ccould\u201d or \u201ckill\u201d. In addition, there is Topic: 2, which is about the immigration caused by the Syrian war in the Middle East and about how candidates for the position of the President were going to deal with it. Topic: 6 is fully dedicated to the 2016 elections and the sources of propaganda.<\/p>\n<p><strong>Conclusions and further work<\/strong><\/p>\n<p>So, two classifiers with very high overall accuracy, precision, recall and F-1 score have been built and tested on several features. The experimental results have shown that the tweet text as a feature gives better accuracy than hashtags. The models were trained on large amounts of data and, thus, can be used as a correct solution for detecting attempts at mass influencing. The exploratory analysis and topic modelling allowed me to delve deeper into the actual goals of the trolling. All this makes it possible to conclude that the ability to affect internal affairs of another state exists but, probably, depends on the amount of the troll army used and how they are used.<\/p>\n<p>This study in two parts is initial and can be furthered with deep learning and analysis of visual components (such as images and videos attached to the tweets). To continue with the research, I\u2019m going to build a classifier that will take into account such aspects as the user name, the user picture, kind of account activities (number of retweets and likes), number of tweets per unit of time and some deeper parameters like the type of followers or the sentiment of the tweets.<\/p>\n<p>The only concern related to more advanced analysis and deep learning is that troll accounts will soon be automated and converted into bot accounts. This means that we will soon need to extract features of bots instead of trolls. This might be quite a complicated thing because bots usually have many similar features, but not all of them represent trolls.\u00a0<\/p>\n<\/div>\n<\/div>\n<\/div>\n<p><!----><!----><\/div>\n<p><!----><\/p>\n<div class=\"tm-article-poll-container\"><!--[--><\/p>\n<div class=\"tm-article-poll tm-article-poll_variant-bordered\">\n<div class=\"tm-notice tm-notice_positive tm-article-poll__notice\"><!----><\/p>\n<div class=\"tm-notice__inner\"><!----><\/p>\n<div class=\"tm-notice__content\" data-test-id=\"notice-content\"><!--[--><span>\u0422\u043e\u043b\u044c\u043a\u043e \u0437\u0430\u0440\u0435\u0433\u0438\u0441\u0442\u0440\u0438\u0440\u043e\u0432\u0430\u043d\u043d\u044b\u0435 \u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u0435\u043b\u0438 \u043c\u043e\u0433\u0443\u0442 \u0443\u0447\u0430\u0441\u0442\u0432\u043e\u0432\u0430\u0442\u044c \u0432 \u043e\u043f\u0440\u043e\u0441\u0435. <a rel=\"nofollow\" href=\"\/kek\/v1\/auth\/habrahabr\/?back=\/ru\/articles\/674130\/&#038;hl=ru\">\u0412\u043e\u0439\u0434\u0438\u0442\u0435<\/a>, \u043f\u043e\u0436\u0430\u043b\u0443\u0439\u0441\u0442\u0430.<\/span><!--]--><\/div>\n<\/div>\n<\/div>\n<p><!--[--><\/p>\n<div class=\"tm-article-poll__header\">Do you think it&#8217;s possible to influence a significant number of citizens via social networks?<\/div>\n<div class=\"tm-article-poll__answers\"><!--[--><\/p>\n<div class=\"tm-article-poll__answer\">\n<div class=\"tm-article-poll__answer-data\"><span class=\"tm-article-poll__answer-percent tm-article-poll__answer-percent_winning\">100% <\/span><span class=\"tm-article-poll__answer-label\">Yes<\/span><span class=\"tm-article-poll__answer-votes\">2<\/span><\/div>\n<div class=\"tm-article-poll__answer-bar\">\n<div class=\"tm-article-poll__answer-progress tm-article-poll__answer-progress_winning\" style=\"width: 100%\"><\/div>\n<\/div>\n<\/div>\n<div class=\"tm-article-poll__answer\">\n<div class=\"tm-article-poll__answer-data\"><span class=\"tm-article-poll__answer-percent\">0% <\/span><span class=\"tm-article-poll__answer-label\">No<\/span><span class=\"tm-article-poll__answer-votes\">0<\/span><\/div>\n<div class=\"tm-article-poll__answer-bar\">\n<div class=\"tm-article-poll__answer-progress\" style=\"width: 0%\"><\/div>\n<\/div>\n<\/div>\n<p><!--]--><\/div>\n<div class=\"tm-article-poll__stats\"> \u041f\u0440\u043e\u0433\u043e\u043b\u043e\u0441\u043e\u0432\u0430\u043b\u0438 2 \u043f\u043e\u043b\u044c\u0437\u043e\u0432\u0430\u0442\u0435\u043b\u044f.    \u0412\u043e\u0437\u0434\u0435\u0440\u0436\u0430\u0432\u0448\u0438\u0445\u0441\u044f \u043d\u0435\u0442. <\/div>\n<p><!--]--><\/div>\n<p><!--]--><\/div>\n<p> \u0441\u0441\u044b\u043b\u043a\u0430 \u043d\u0430 \u043e\u0440\u0438\u0433\u0438\u043d\u0430\u043b \u0441\u0442\u0430\u0442\u044c\u0438 <a href=\"https:\/\/habr.com\/ru\/articles\/674130\/\"> https:\/\/habr.com\/ru\/articles\/674130\/<\/a><br \/><\/br><\/br><\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[],"tags":[],"class_list":["post-390211","post","type-post","status-publish","format-standard","hentry"],"_links":{"self":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts\/390211","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=390211"}],"version-history":[{"count":0,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=\/wp\/v2\/posts\/390211\/revisions"}],"wp:attachment":[{"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=390211"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=390211"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/savepearlharbor.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=390211"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}