A two-part analysis of Aleksandar Vučić’s address in Mrkonjić Grad, 4 August 2026 (Part 1)

Reading Time: 7 minutes

Introduction: What can text analysis tell us about a political speech?

Political speeches are designed not merely to transmit information. They construct stories, identify groups, assign responsibility, evoke emotions, recall selected episodes from the past and suggest ways in which an audience should understand the present and future. Some of this can be detected through close reading. Computational text analysis adds another perspective: it allows us to count systematically what is repeated, identify recurring combinations of words, examine shifts in emotional vocabulary and locate rhetorical patterns that may be difficult to notice while simply listening to a speech.

This two-part analysis uses one speech as a worked example. The purpose is descriptive and exploratory rather than causal. It does not attempt to determine whether particular historical or political assertions made in the speech are correct. Nor does it claim that frequency automatically equals importance. Instead, the question is methodological: what can relatively simple text-analysis techniques reveal about the structure of a political message, and where do those techniques begin to fail?

That distinction is particularly important for political discourse. A computer can tell us that Serbia occurs 54 times. It cannot, from that number alone, tell us what “Serbia” means in each passage, whether the speaker invokes the state, a political community, a historical subject or an emotional object of identification. Quantitative indicators therefore become most informative when they are combined with qualitative reading.

The political setting

The speech analysed here was delivered by Serbian President Aleksandar Vučić in Mrkonjić Grad on 4 August 2026 during the joint Serbia–Republika Srpska commemoration of Serbs killed and displaced during and after Croatia’s 1995 military operation Storm. The event had been announced as a joint commemoration organised by Serbia and Republika Srpska, with the Serbian Orthodox patriarch and the leadership of Republika Srpska also expected to attend. Mrkonjić Grad was deliberately selected as a symbolically charged location because the town itself experienced warfare and displacement during the final months of the Bosnian war in 1995.

Operation Storm began in August 1995 and restored most territory then controlled by the self-proclaimed Republic of Serbian Krajina to Croatian government control. It also produced a mass exodus of Croatian Serbs. UNHCR has estimated that approximately 250,000 Croatian Serbs left Croatia during this period, while also placing that displacement within the broader wartime context in which large numbers of Croats and other non-Serbs had previously been expelled from Serb-controlled areas.

This background matters because the speech is fundamentally a commemorative political text. Memory, victimhood, responsibility, national identity and promises of protection are therefore not peripheral topics: they constitute the communicative setting in which the speech was delivered.

Why analyse an English translation?

The original transcript is in Serbian, while the computational analysis of this exercise was conducted on an English translation. There are practical reasons for doing so. English has a particularly rich ecosystem of text-analysis resources, including the Bing sentiment dictionary and the NRC Emotion Lexicon used here. Translation also facilitates comparison with English-language political corpora and makes some semantic analyses easier to reproduce for an international audience.

There is nothing inherently illegitimate about this strategy. Research on cross-lingual sentiment analysis has shown that analysing translated text can sometimes produce useful results, particularly where high-quality analytical resources are much better developed for the target language. At the same time, translation can alter the sentiment associated with words and passages, which is why translated sentiment should be treated as an approximation rather than as a direct measurement of the source-language text.

For political speeches the problem is wider than sentiment. Translation involves choices about lexical equivalence, syntax, cultural references and pragmatic force. Studies of political-speech translation show that modulation, reduction, addition and other translation strategies may affect how a political message is represented in the target language.

There are several particularly important problems in this case.

Serbian grammatical structure is different from English. Cases, grammatical gender, verbal aspect, clitic placement and relatively flexible word order can carry information that disappears or changes in translation. Serbian is also a pro-drop language: speakers can omit subject pronouns because the verb often identifies the grammatical person. English normally requires an explicit subject. An English translation can therefore insert I, we, you or they where no corresponding pronoun occurred in Serbian. That makes pronoun-frequency analysis especially sensitive to translation.

Agency can also change. Passive and impersonal Serbian constructions may become active English sentences, potentially making responsibility appear more explicit. The referent of Serbian mi may also remain strategically ambiguous, government, political movement, nation, audience or Serbian people, whereas a translator may implicitly resolve the ambiguity.

The English transcript presents an additional problem: in places it is visibly literal and syntactically awkward. That is important because machine-readable text may become easier to count while becoming less faithful as political language. Idioms, irony, rhetorical questions, culturally marked terminology and emotionally charged expressions can shift or disappear.

Finally, translation is not the only source of uncertainty. Transcription punctuation determines what the computer sees as a “sentence”. Quoted speech is currently analysed together with Vučić’s own assertions. When the speaker quotes or paraphrases an opponent, a sentiment dictionary does not know that the words may represent a position he is rejecting.

The strongest design is consequently bilingual: analyse Serbian when investigating linguistic form and rhetoric, and use the English version primarily for broad semantic patterns, exploratory sentiment measures and comparability.

Preparing the transcript

The R workflow correctly performs one important separation before lexical analysis: audience annotations enclosed in brackets, such as [APPLAUSE], are extracted and then removed from the version used for word and sentiment calculations. This prevents the audience’s behaviour from being counted as the president’s vocabulary.

The resulting cleaned corpus contains 4,936 word tokens and 215 detected sentences.

Table 1. Basic descriptive statistics of the translated speech

statisticvalue
words4936
sentences215
unique_words1081
type_token_ratio0.219
mean_words_per_sentence23.884
median_words_per_sentence18
mean_characters_per_word4.44

How long and lexically diverse is the speech?

The translated corpus contains 4,936 words and 1,081 different word forms. The resulting type-token ratio is approximately 0.219, meaning that unique forms account for about 21.9% of all tokens.

This gives a rough indication of lexical diversity, but it should not be interpreted as a quality score. Type-token ratio decreases mechanically as documents become longer, and English morphology differs substantially from Serbian. It is therefore most useful for comparing documents of similar length processed in exactly the same way.

The mean sentence contains 23.9 words, while the median is only 18. The difference tells us something useful: the speech contains a number of relatively long sentences that pull the average upward. This fits the nature of oral political rhetoric, in which clauses are frequently chained together through repetition, contrast and accumulation.

The mean word length is approximately 4.44 characters. The analysis therefore does not suggest unusually complicated vocabulary at the individual-word level. Much of the rhetorical complexity lies instead in sentence construction, historical references and shifts of speaker, addressee and temporal perspective.

No formal Flesch or comparable readability index was calculated, and it would be inadvisable to infer one from these figures alone. More importantly, applying an English readability formula to a literal translation of spoken Serbian would create an impressive-looking number whose substantive meaning would be doubtful.

Which words dominate?

The frequency results provide the clearest first picture of the speech.

The most frequent content word is Serbia, appearing 54 times. It is followed by people (42), Serbian (34), Serbs (33), never (33), always (27), Srpska (26) and Republika (23). Other prominent words include thank, president, Montenegro, love, Belgrade, forget, church and care.

Figure 1. Most frequent content words in the translated transcript
Figure 2. Word cloud of frequently occurring content words

A word cloud should never substitute for a frequency table, but visually it makes one feature particularly obvious: the speech is organised around collective identity. Serbia, Serbian, Serbs, people, Republika and Srpska dominate the lexical landscape.

Two less obvious words are equally revealing: never and always. These are not policy terms. They are words of temporal absolutism. They help transform particular historical events into a narrative of continuity: what “always” happened, what must “never” happen again, who “always” supported whom and what should “never” be forgotten. In commemorative rhetoric this creates a bridge between past memory and future commitment.

Words such as thank and love point in a different direction. They become especially prominent toward the end of the address and anticipate a pattern that becomes much clearer in Part 2: the speech moves from suffering and accusation toward solidarity, gratitude, protection and collective strength.

From individual words to phrases

Single-word frequencies lose context. Bigrams and trigrams begin to restore some of it.

The most frequent bigram is Republika Srpska, appearing 23 times. Serbian people occurs nine times, Banja Luka six, and world war six. Other recurring combinations include take care, dear friends, four times, last time, polycentric Serbia, Serbian church and us Serbs.

Table 2. Most frequent bigrams

bigramn
republika srpska23
serbian people9
banja luka6
world war6
take care5
dear friends4
four times4
last time4
many many4
polycentric serbia4
serbian church4
us serbs4

Among the repeated trigrams are second world war, Serbs practically disappear, call us Serbs, eyes Siniše Dobrić, last time like president and many many Serbs.

Table 3. Most frequent trigrams

trigramn
second world war3
serbs practically disappear3
call us serbs2
dear Nevenka dear2
eyes Siniše Dobrića2
last time like2
Leibniz sufficient reason2
many many many2
many many serbs2
Nevenka dear Nevenka2
time like president2
world war ii2

What Part 1 tells us

Even this relatively simple set of measures produces a coherent preliminary picture. The speech is centred lexically on Serbia, Serbs and Republika Srpska. It repeatedly uses categorical temporal markers such as never and always, while common phrases connect collective identity with historical memory, unity and political obligation.

At the same time, the analysis shows why text analysis requires methodological restraint. Word counts reveal emphasis, not meaning. Dictionaries reveal what their categories are capable of detecting. Translation can alter grammar, pronouns and emotion.

In Part 2, the analysis moves from what the speech talks about to how its political message is constructed: emotional trajectory, pronouns, rhetorical markers, audience applause, collocations and the larger frames connecting victims, history, strength, unity and political leadership.

Appendix

Here you can download a zip file containing the transcript of the speech in Serbian and English, the r code used in the text analysis and the output folder with all the results.

Komentariši

Vaša email adresa neće biti objavljivana. Neophodna polja su označena sa *