German lecturer with Ph.D, ..., I started using Python and R, .and stumbled about emojis ...

Sunday, February 5, 2023

Heat Stains on Twitter (first observations)

In order to find spots where twitter gets hot, with people writing down emphatic statements or opinions, we could start with emojis, as the creation of hashtags is not predictable, while the number of emojis is relatively limited. 

Yet, by the most common packages for Python and R we still get between 1400 and 5000 emojis. These still are far too much, for me, to loop through these lists and make a tweet search request for each of them. I decided to begin with my own list of emojis which will be enriched.

The emojis used in a very lively Italian gossip group ("#jerù"), where often anger about the stars is expressed, will be the first ones, already 34. For the beginning, I will keep within the realm of Italian tweets. As the case of Polish 😆 teaches, the use of emojis largely depends on cultures defined by languages. 

My first Italian list is : 👀 🐻 😅 🌜 👑 🌈 ☕ 🤍 😍 🥰 ♥️ 🤦 ❤️ 🌲 😂 🦋 📸 🤣 ⛰️ 💜 💚 ♀️ 😁 🔥 💖 💗 🙃 😋 🕛 😎 😭 😜 🌺 ✨.

Looking for signs of excitement, I will keep record of tweets searched with my emoji list only if on 100 tweets I get more than 15 exclamation marks. Angry and happy people love doubling and tripling these signs. 

In a tweet search done on November 12, 2022, the most prolific emoji was 🤣

In 100 tweets with this emoji, it appeared 231 times. Exclamation marks: 27 on 9234 characters, with eight double "!" and three "!!!!". 

The most frequent hashtags were 

"#taleequaleshow" "#merito" "#konpetenza" "#ottoemezzo" "#calenda" "#novax", 

i.e. two about political tv shows, two are hashtags used by a journalist (book title: “Damned Pacifists”), two about politics, with the famous hashtag "novax". Maybe this could indicate the right way to find angry people. Should we move on following the hashtags?


A hashtag "novax" search (07/ 01/ 23, n=100) results in only 15 emojis,

💯 🔝 ⚧️ 👋 👿 💉 👇 🐧 💳 💩 🤣 🤡 🏳️ 🇪🇺 🪳

with only four exclamation marks. Where has the excitation gone? We see ten times 💩, following 35 🤣.


The same for "#calenda" (an Italian politician). A search results in 18 exclamation marks, with three times double "!" and one "!!!". But, according to R and the relative emoji package, only five out of 100 tweets contain emojis, 44 in all (26 unique). Five ➡️ , five 🇮🇹, five 🇪🇺. Is it that in politics, Italians use only few emojis? This could be due to the higher age of people interested in politics here.


Let us try with other "controversial topics". A "#Salvini" search (right wing politician) results in 11 on 100 tweets with emojis. 22 🤡 , 18 😂, 13 👏 and five 🇮🇹. Other "hot" topics like #bce (European Central Bank) or "nosbarchi" (no acceptance of refugees) give similar results. People get angry about politics, but do not use emojis in these fields.


Still, by searching with emojis we can find angry people. The clown 🤡 is used when people find laughable something or somebody. Again, in Italy, we get top hashtags about soccer and about the Reality Show Big Brother VIP.

We see

feature frequency     rank 1 🤡 432      1 2 😂 72      2 3 🤣 38      3 4 🤮 26      4 5 😅 16      5 6 👇🏻 12 6 7 💩 12 6

People are angry. Searching for anger with emojis around could be helpful.

The number of unique emojis here is only 23. 
🤬 🤣 ✅ 🤦 😏 😱 😡 ⚫ 🖕 😵 💸 🤡 😂 🤢 💩 🤮 💫 ⏩ 😁 📹.
Maybe this should be the basic list of emojis, on the search of angry people, i.e. heat stains on 
twitter. 

With the middle finger, for example, tweets supposedly are quite aggressive. With an apisearch (n=10) on November 12, 2022, we receive antisemitic content, but only two exclamation marks. This number is again growing with 🤮, becoming 7 in ten tweets, and 2 double "!". 😡 gives eight exclamation marks (two double), tweets about animals rights. The 💣 brings tweets about conspiracy theories, without any "!"  
😡 search gives 42 emojis in 10 tweets, namely
🤬 🤣 💯 👇 😠 💪 🤦 ➡️ 🙏 😢 👌 ♀️ 💥 😤 👿 😱 😡 🤔 ❣️ 🇮🇹 🔻 🖕 ♂️ 🔴 👎 😈 🤞 😂 🤢 💩 🤮 😅 🥲 🤪 🥺 ‼️ 😖 ⏩ 🤨 😁 📹 ☕
The exclamation marks are here, not among the 
characters. We should getting to know emojis by the company they keep. 

Emoji diversity


For comparison, a ☕ search results in 62 unique emojis and 421 total emojis, lexical emoji diversity  .147. Lexical diversity among emojis is low in all cases, maximums are .23 with 🔥, .21 with 😎, .22 with the ❤️ and .2 with 😜. This does not depend on the simple number of emojis. (for example 254 correspond to .10, 601 to .15.). The variety of accompanying emojis rather seems to be a characteristic of the emoji itself.




technically

(ok, it is not elegant, but I had learned programming with Algol W, in 1975)

import pandas as pd

import tweepy

import csv

import emoji

import emojis

import regex

from collections import Counter

#https://stackoverflow.com/questions/49113909/split-and-count-emojis-and-words-in-agiven-string-in-python


<authentication stuff>


api = tweepy.API(auth)

with open('emojilistit22.csv','r') as mine:

    leser = csv.reader(mine, delimiter=',')

    leserl = list(*leser)

    for kw in leserl:

        print (kw)

        container = []

        tweetCount = 100

        results = api.search_tweets(kw, count=tweetCount, lang="it")

        for tweet in results:

            container.append(tweet.text)

        row_count = len(container)

        print("number of tweets ", row_count)

        filename = "exclamation_basis_" + kw + "2022-11" + ".csv"

       

        f = open(filename, 'w')

        writer = csv.writer(f)

        writer.writerow(container)

        f.close()

        zeichendf = pd.read_csv(filename)

        zeichenkette = zeichendf.to_string()

        lang = len(zeichenkette)

        print("Number of characters: ", lang)

        x = zeichenkette.count(".")

        print("Number of single points (full stops? .)", x)

prozent = x/lang*100

        print(prozent, "%")

        y = zeichenkette.count("...")

        print("Number of three points (ellipsis ...)", y)

        prozent = y/lang*100

        print(prozent, "%")

        if y>0:

            relation = x/y

            print("Relation: ", relation)

        x = zeichenkette.count("!")

        print("Number of exclamation marks", x)

        prozent = x/lang

        print(prozent, "%")

        if x>15:

            anteil = lang/x

            print (kw, "exclamation marks:", x, "on", lang, "characters", anteil, "%")

        y = zeichenkette.count("!!")

        print("Number of two exclamation marks", y)

        prozent = y/lang*100

        print(prozent, "%")

        if y>0:

            relation = x/y

            print("Relation: ", relation)

        x = zeichenkette.count("!!!")

        print("Number of three exclamation marks", x)

        y = zeichenkette.count("!!!!")

        print("Number of four exclamation marks", y)

# just having a look at the emojis, will construct an archive later

        material = str(zeichenkette)

        emoji_hier = emojis.get(material)

        print (*emoji_hier)

        zahlein = emojis.count(material, unique=True)

        zahlall = emojis.count(material)

        print (zahlein, zahlall)

 


 

Monday, January 9, 2023

the meaning of ⚫

A black hole, a blackout, "be careful"? Even Emojipedia does not assign a specific meaning to the ⚫ emoji. I have encountered a frequent use during the burial of Gianluca Vialli, a very famous and beloved soccer champion who passed away at age 58. 

Among the hearts ❤️ 💙 and prayer 🙏🏻 and soccerballs ⚽ contained in the 307 (of 710, search n=1000) tweets, ⚫ appeared 33 times and kept the tenth position. 

Searching only tweets with ⚫, in 08/01/23 with n=100, the following 53 emojis are used:

🚨 7️⃣ 🤬 ✅ 🔝 🌸 💮 💛 🌷 ❤️ 📋 🐎 🖤 🧡 💻 📰 ⚪ 🎂 🎶 ♠️ 🎙️ 🇦🇷 ▪️ 💜 ⚫ 💙 🤝 🕷️ ◼️ 🔴 ♥️ ⏳ 📄 💚 🔲 🌹 😂 🏟️ ⭐ 🤎 🤮 🔳 🔥 🏴 ⛔ 📝 🔵 ♣️ 🥺 ✍️ 🚫 ◾ ☕

We can try to group these emojis
💛❤️🖤🧡💜🤎💙♥️💚 hearts
🌸💮🌷🌹         flowers
😂🎶🎂⭐☕🤝🔝✅     positive feelings
📄📋 📰 ✍️📝      writing/media
🤬⛔🤮🕷️🚫🚨      negative feelings
🥺🔥
♠️♣️
The others concern time (⏳), soccer (🏟️, 🇦🇷), a horse (🐎), or are probably color symbols for soccer teams, used in 
sentences like "#MonzaInter ⚫🔵\nForza ragazzi!". 
⚪⚫🔴🔵 🎙️▪️ ◼️  💻 🔲 🔳  🏴◾  7️⃣ .

The emoji contexts are

⚫ 🤬

⚫ 🤮 🤮 🤮 🤮 🤮

🔴⚫

⚪⭐⭐⭐⚫

🗞📰 📎💻

⚪🔴⚪🔴⚫🔴

🔴⚫,

⚪⚫

🔵⚪🔴⚫⚪🔵

⚫⛔

💪🔴⚫

🤣🤣⚪⚫

💪⚫🔵

🖤💙🐍☕⚫🔵

⚫🔵🐍

🚨🤝🏻🔥⏳🔴⚫

What the black circle means? We can see it from the combination. Things like ⚫🔵 are color games. We know, they

are referred to soccer teams, but it could be anything else, like marmalade brands or political parties. And here? ⚫ 🤮 🤮 🤮 🤮 🤮. Just what it says.









technically 

with R & emoji package, 

ntoken(kette)

 text1 

916269 

> emoji_count(kette)

[1] 8815

> textstat_frequency(alle_emojis, n=30)

   feature frequency rank docfreq group

1       ⚫      2379    1       1   all

2       ⚪       925    2       1   all

3       🔵       922    3       1   all

4       🔴       738    4       1   all

5       💪       316    5       1   all

6       ⚽       140    6       1   all

7       🖤       130    7       1   all

8       🔥       115    8       1   all

9        ❤️       106    9       1   all

10      🇮🇹        87   10       1   all

11      😂        80   11       1   all

12      💙        76   12       1   all

13      🏆        69   13       1   all

14      😍        69   13       1   all

15      👇        64   15       1   all

16    💪🏻        54   16       1   all

17       🎙️        53   17       1   all

18       1️⃣        49   18       1   all

19      🇾🇪        44   19       1   all

20      🤍        43   20       1   all

21      💦        38   21       1   all

22      👈        38   21       1   all

23       ⚠️        38   21       1   all

24      🟢        34   24       1   all

25      ⭐        34   24       1   all

26      ✅        33   26       1   all

27      ☕        32   27       1   all

28      😘        32   27       1   all

29      🐍        31   29       1   all

30       ♥️        30   30       1   all


  emoji_tweets total_tweets

         <int>        <int>

1         1000         1000


[1] "number of emojis"

 8815 

> print("unique emojis")

[1] "unique emojis"

> zeichenarten

text1 

  350 

[1] "emoji diversity"

0.03970505 


head(toptag)

[1] "#internapoli"      "#forzainter"      

[3] "#milan"            "#inter"           

[5] "#salernitanamilan" "#seriea"  

(soccer teams and First League)

Saturday, January 7, 2023

Emoji Translation? 👀 🐻 😅 🌜 👑 🌈 ☕ ?

Whether you are a clueless boomer or a vibrant data scientist, the question: "How to know the meaning ....?" of 👀 🐻 😅 🌜 seemingly has an easy answer: Look it up!

Emojipedia, for example, explains a heart means "romance and love", while vomit stands for "disease or disgust". This could be the bridge even to convincing sentiment analyses. Instead of the heart, we will count the occurrences of the word "love".

Now, take emojis from tweets like "❤️🌼🍀", and write "love" "love, appreciation and happiness" and "good luck".

See "❤️✈️" and understand "love" "overseas vacation or airplane mode" or "love travel". Looking good?

What about "🤣✈️✈️✈️✈️✈️✈️❤️❤️❤️❤️❤️"? "Rolling ... laughing" "travel" "travel" "travel" "travel" "travel" "travel" "love" "love" "love" "love" "love"! OK, the repetition can be understood, and coded, as intensification. But how to get an intensification of "travel"? Maybe ✈️ is meaning something else, one time or the other.

If we have a look at the company ✈️ keeps, in Italian (100 tweets, 07/01/23) we get the following list of 41 emojis

🎶 📀 🇹🇭 🤩 💪 😂 ✈️ 🍾 🎥 📷 ☠️ 🌻 💙 🤣 👨 📽️ 🇳🇴 🇺🇸 ✨ 🥂 🐿️ 🌲 🔥 💀 🎵 ➡️ 🐮 ♥️ 🛬 😃 🇧🇾 ⏰ ❣️ 💛 🇦🇷 💜 🇰🇷 🇷🇺 🇺🇾 ⚰️ ❤️

In German, we find 39
⤵️ 🇹🇭 🤩 ✈️ 🎥 😉 💫 🆑 🧑 ♨️ 🚚 👨 🤣 🔊 ‼️ 😍 🌫️ 💢 🚽 😵 🖤 🇩🇪 😶 🐈 🚀 👀 🇦🇹 🚛 🙇 ♂️ 🚒 💵 📖 🐥 🐷 👩 😏 🛫 🇷🇺

In Italian, six different hearts and maybe the fire refer to love/positive emotion. In German, one loving face and a black heart give a different impression. Two red exclamation marks and a pig may indicate an aggressive tone. 

Doing the same tweet search with R, in 100 Italian tweets 
the program counts 683 emojis, with 61 unique ones. 21 red hearts and 17 coffins look rather emotional. Looking at the top hashtags, we 
understand why. 
head(toptag)
[1] "#donnalisi"      "#gfvip"         
[3] "#denzzzers"      "#incorvassi"    
[5] "#danieledalmoro" "#delizon"
All of them are related to a Reality Show, the Italian "Big BrotherVIP". 
The people tweeters are talking about are closed in an 
apartment in Rome. Nobody is going to fly. Indeed, the airplane here
is often used as "you make me fly" (mi fai volare) in the sense
of "with you I am living intense sensations". Which explains the 
✈️✈️✈️✈️✈️✈️❤️❤️❤️❤️❤️. 

In German, in 75 tweets collected we get 476 emojis, with 95 unique ones, which means the emoji diversity
is quite high (.2). 
Top hashtags are concerning the topic "flying": 
[1] "#boeing"    "#ryanair"   "#scotradar"
[4] "#lufthansa" "#3c4a04"    "#ireland" 
Main topics are linked to soccer (Bayern München ...). 
Studying single tweets, we also find "Wer randaliert fliegt" (whoever riots, gets kicked out)and 
"Randalierer abschieben" (deport rioters), which might account for aggressiveness
in some posts. 



technically

The emoji combinations have been fished with a Python request, n = 100, lang = it/de, on November, 22, 2022 and on January, 7, 2023. Use of emojis and emoji packages was essential.


Sunday, November 13, 2022

signs of excitement!!! Excitation marks in tweets 🤪.

Years ago, at Berlin there were demonstrations on the streets, with pensioners asking loudly to shoot these jerks 👾, while some elderly people 🧓 at Milan had heated discussions about politics on piazza Duomo, and groups of youngsters were strolling around everywhere, shouting or laughing 🥳. Today, on the streets and on the busses, we are watching our telephones, in silence. 

The lack of noise does not mean people today were less excited. New and old wars, brand new clothes, children, latest iPhones, euthanasia, foreigners in town: Everybody is excited, I read on Facebook and in the chats. 

Things change, and so does our way of expressing ourselves. Even what seemed rock hard, and menacing, when we were kids at school, like punctuation, is changing. 

Once we would have written: "I am very, more, most, terribly ... angry". today we write things like "It's a shame!!!!"

What is happening here to our good old exclamation mark? Repetition of signs is something we know from Social Media and chats. People repeat emojis. 

Fishing 100 red heart tweets in German with R quanteda, we find 145 heart symbols with 48 pairs of hearts, 41 triplets (for details, see below "technical notes"). 100 Italian vomit tweets give 238 🤮 symbols, with 125 repetitions and 70 triplets. This tendency can, weaker or stronger, be found any day with any emoji of a certain kind (not with personifications, for example 🐻. Elsewhere we will try to distinguish types of emojis). 

We could follow that people who close, like a friend of mine last Saturday, the first sentence of a mail with two exclamation marks and the last one with four of them, are using our punctuation mark as an emoji. 


Nice hypothesis, but how to be controlled? A keyword tweet search with punctation signs does not work, not in R, not in Python. For a "!" you get your 100 documents, but they contain only few exclamation marks (lang="de", 2.11.2022: 14).  Do not ask me why. An exclamation mark is not counted as an emoji, for now.


How will we find the signs of exitement, then?  Searching with emojis as keywords. 


To make a beginning, I choose something nice and friendly. 

We make a tweet_search (with R), or an api.search_tweets (Python) with keyword = "😊" and n=100.  

Number of documents 100

Italian (29/10/2022): Number of characters 6297. Exclamation marks 6. Two exclamation marks 0

German: Number of characters 8942. Number of exclamation marks 22
Two exclamation marks 1

The difference between Italian (6) and German (22) usage of "!" looks interesting, although a probability of .25, compared to .2 we find in a novel from about 1900 ,. is not really exciting. 

But maybe collecting only nice tweets we are on the wrong way. 

Trying out "😔", results change only slightly.

Italian
Number of characters:  9551
Number of exclamation marks 5
Number of two exclamation marks 0

German
Number of characters:  8890
Number of exclamation marks 9
Number of two exclamation marks 0

Still, we did not enter the world of strong feelings. 
Trying with the nauseated face "🤢" emoji:

German
Number of characters:  8892
Number of exclamation marks 20
Number of two exclamation marks 4

Italian
Number of characters:  8036
Number of exclamation marks 10
Number of three exclamation marks 1


and with vomit "🤮", 

German 
Number of characters:  11170
Number of exclamation marks 20
Number of two exclamation marks 2

Italian
Number of characters:  9889
Number of exclamation marks 26
Number of two exclamation marks 4

Here they are, few of them. Four double exclamation marks in Italian, two in German. 

Would trying with controversial topics give stronger results?  Something like #deutschland (in German), #italia (in Italian)

German
Number of characters:  13868
Number of exclamation marks 8
Number of two exclamation marks 0

Italian
Number of exclamation marks 14
Number of two exclamation marks 0


No. No repetition of "!". The trend is confirmed with conspiracy theory hashtags like "#plandemie" in German. 
Number of characters:  12479
Number of exclamation marks 5
Number of two exclamation marks 1

In Italian, even "#novax" produces few exclamation marks (2.11.: 5)

Things change again when we take up again searching (in Italian) with special emojis we had found elsewhere, keywords like "🦋".
Italian
Number of characters:  7190
Number of exclamation marks 15
Number of two exclamation marks 2

In Italian, the number of exclamation marks is higher here. Six of them were found in smiley tweets, 15 here. How can this growth be 
explained? The butterfly emoji does not look really exciting, it is not 🤮. With R, we can look up correlated hashtags.
> head(toptag)
[1] "#art"       "#artlovers" "#catlovers"
[4] "#a"         "#amemici20" "#jeru"   

These hashtags are partly linked to a conspiracy theory born during a Big Brother VIP show in Italian television. it is all about a couple ... 
Trying these hashtags, for example #jeru, finally we can hear excited 
tweeters:

Number of characters:  10298
Number of exclamation marks 37
0.35929306661487664 %
Number of two exclamation marks 5

37 simple, five double exlamation marks is quite a result. Useless to say these tweets are full of emojis, too. 

In Germany, Reality Shows do not provoke similar emotional discourses on twitter. With #temptationisland, only 18 tweets are found, with 
three exclamation marks. #badgirlsclub has zero tweets (always on October, 29). When and where German twitters get loud? Try "#impfgegner" 8opponents of vaccination): 
Number of characters:  14160
Number of exclamation marks 28
Number of three exclamation marks 1

"#putintrolle":
Number of characters:  12883
Number of exclamation marks 27
0.2095785143211985 %
Number of three exclamation marks 1

While in Italian we get an increase in (double) exclamation marks when communicating about gossip, in Germany only few, very specific keywords create a similar, though minor, effect. 

In short, yes, repetitions of exclamation marks,  signs of excitement, can be found on Twitter, but only in certain areas that, from one culture to the other, 
are different. Probably the differences in usage are correlated in some mysterious way to the chosen emojis.  

When starting our research of excitement with emojis, we do not need to know anything about the world, like those unforeseeable hashtags.





Technically

A simple Python search

    api = tweepy.API(auth) 

    container = []

    with open('zeichenbasis.csv', 'w') as tweet_csv_file: 

        pass 

keywords = "😊"

tweetCount = 1000

results = api.search_tweets(keywords, count=tweetCount, lang="it")

for tweet in results:

    container.append(tweet.text)

f = open('zeichenbasis_smile.csv', 'w')

   writer = csv.writer(f)

   writer.writerow(container)

   f.close()

zeichendf = pd.read_csv('zeichenbasis_smile.csv')

zeichenkette = zeichendf.to_string()

print(zeichenkette) 

x = zeichenkette.count(".")

print("Number of single points (full stops? .)", x)

y = zeichenkette.count("...")

print("Number of three points (ellipsis ...)", y)

if y>0: 

    relation = x/y

    print("Relation: ", relation)




x = zeichenkette.count("!")

print("Number of exclamation marks", x)


y = zeichenkette.count("!!!")

print("Number of three exclamation marks", y)

if y>0: 

    relation = x/y

    print("Relation: ", relation)


R search with rtweet and quanteda

toks_ngram <- tokens_ngrams(tokens(corpse), n = 2:4)

> ngr_matrix <- dfm(toks_ngram)

> textstat_frequency(ngr_matrix, n=20)

feature frequency

1 . . 45

2 😊😊 28


Searching for heart tweets (❤️)  in German, we obtain 145 hearts and

> textstat_frequency(ngr_matrix, n=20)

feature frequency rank 

1   ❤️ 

❤️ 

     48 1

2   ❤️ 

 

❤️ 

 

❤️

 41 2

3  ❤️  ❤️  ❤️  ❤️  37 3


Collecting tweets with the vomit emoji in Italian gives: 


Total frequency 238

> textstat_frequency(ngr_matrix, n=20)

feature frequency rank docfreq

1 🤮_🤮 124 1 45

2 🤮_🤮_🤮 79 2 39

3 🤮_🤮_🤮_🤮 40 3 19

Image flow, image row. Understanding emojis? with Vilém Flusser

Little pictures in between We used to send letters to each other. We used to write down, letter after letter, word by word,  what we hoped w...