Part 2. From Emotion to Political Rhetoric: How the Speech Builds Its Message
Is the speech positive or negative?
A first look at the Bing sentiment dictionary produces an apparently surprising result. Of the words recognised by the lexicon, 197 are classified as positive and 144 as negative. That corresponds to 57.8% positive and 42.2% negative among matched sentiment words.

Table 4. Bing sentiment classification of matched words
| sentiment | n | proportion |
| positive | 197 | 0.58 |
| negative | 144 | 0.42 |
This does not mean that the speech is “57.8% positive”. Only 341 sentiment-bearing token matches are involved, a small subset of the 4,936-word transcript. More importantly, a simple lexicon evaluates words rather than propositions.
The transcript contains an excellent demonstration. One early sentence describing the indifference of others to Serbian victims is assigned a positive score because the translated sentence contains words corresponding to care, celebrated and victory. Semantically, however, the speaker is condemning that celebration. Similarly, a sarcastic reference to “democratic, European values” can be read by a dictionary as positive even though its rhetorical function is criticism.
Negation creates the same problem. “Never” may reverse or intensify the meaning of a neighbouring expression, but the bag-of-words method simply adds positive and negative lexical scores. Reported speech creates another difficulty: an accusation quoted in order to reject it is still counted as vocabulary inside the speech.
The sentiment totals should consequently be read as lexical orientation, not as an automatic judgement of the speaker’s attitude.
The emotional trajectory is more informative
The sentence-level analysis offers a more interesting pattern.
The 215 sentences were divided into ten equal sequential groups. These should be called speech deciles rather than substantive “sections”: they are produced automatically by sentence order and do not correspond to hand-coded rhetorical divisions.
Mean sentiment is negative in the first two deciles, at approximately –0.45 and –0.41. It approaches neutrality in the third, turns strongly positive in the fourth, remains mostly positive thereafter and reaches its highest average in the final decile, approximately +0.86.

Table 5. Mean and total sentiment by sequential speech decile
| section | mean_sentiment | total_sentiment |
| 1 | -0.45 | -10 |
| 2 | -0.41 | -9 |
| 3 | -0.09 | -2 |
| 4 | 0.73 | 16 |
| 5 | 0.27 | 6 |
| 6 | 0.48 | 10 |
| 7 | 0.10 | 2 |
| 8 | 0.33 | 7 |
| 9 | 0.71 | 15 |
| 10 | 0.86 | 18 |
Here the quantitative pattern corresponds reasonably well with a close reading. The speech begins with murdered children, grieving mothers, expulsion, Jasenovac, war and accusations of historical silence. The opening is consequently dominated by vocabulary likely to be classified negatively.
Later, the rhetorical centre of gravity changes. The language increasingly concerns peace, strength, protection, unity, gratitude, love, economic development and the relationship between Serbia and Republika Srpska. The final passages include repeated thanks and “long live” formulations. A movement from grievance and mourning toward reassurance and collective affirmation is therefore visible both qualitatively and in the sentiment trajectory.
This is one of the more defensible findings because it concerns a broad change in semantic orientation rather than the precise emotional value of an individual Serbian expression.
The black LOESS curve in the figure should nevertheless be understood as a descriptive smoother, not as a statistical proof that the speech “becomes positive”. Sentence scores are dependent, translation-sensitive and derived from a dictionary rather than human coding.
Which emotions appear?
The NRC Emotion Lexicon adds eight emotional categories. Trust is the largest with 132 matches, followed by fear with 82, anticipation with 80, joy with 79, sadness with 53, anger with 38, surprise with 32 and disgust with 28.

Table 6. NRC emotion-category matches
| sentiment | n | proportion |
| trust | 132 | 0.25 |
| fear | 82 | 0.16 |
| anticipation | 80 | 0.15 |
| joy | 79 | 0.15 |
| sadness | 53 | 0.10 |
| anger | 38 | 0.07 |
| surprise | 32 | 0.06 |
| disgust | 28 | 0.05 |
Again, these are lexicon assignments, not measurements of what percentage of the audience “felt trust” or “felt fear”. An individual word can also be associated with more than one NRC category.
The mixture itself is nevertheless revealing. Fear, sadness and anger fit the commemorative account of violence and victimhood. Anticipation corresponds to repeated statements about what must happen in the future. Joy and trust are consistent with the later emphasis on unity, support, love and confidence in collective strength.
Rather than describing the speech as either fearful or hopeful, the results point toward an emotional combination: remembered vulnerability is paired with promised security. Past suffering supplies the problem; collective strength supplies the proposed answer.
Who are “we”, “you” and “they”?
The rhetorical dictionary counts 147 occurrences of you, 111 of we, 71 of I and 65 of they. Negation words not and never appear 105 times. Words associated with obligation, must, should and need, appear 11 times.
Table 7. Pronouns, negation, modality and selected rhetorical markerst
| category | frequency | frequency_per_1000 |
| second_person | 147 | 29.8 |
| first_person_plural | 111 | 22.5 |
| negation | 105 | 21.3 |
| first_person_singular | 71 | 14.4 |
| third_person_plural | 65 | 13.2 |
| nation | 45 | 9.1 |
| obligation | 11 | 2.2 |
| future | 5 | 1.0 |
| economy | 3 | 0.6 |
| democracy | 1 | 0.2 |
The prominence of we is unsurprising in collective political rhetoric, but the interaction among we, you and they is more interesting.
The speech repeatedly constructs a collective we: Serbs, Serbia, Serbia together with Republika Srpska, or at times speaker and audience. The second person shifts between intimate address to the audience and confrontational address to external actors. They is similarly flexible: political opponents, Croatian political or media actors, international actors, former Serbian leaders or other groups can occupy that grammatical position.
This creates a rhetorical geometry rather than a simple pronoun count: we identifies community, you can produce intimacy or confrontation, and they identifies an out-group against which collective identity becomes sharper.
But this is precisely where translation matters most. Serbian can omit pronouns that English requires. The 111 English wes therefore cannot safely be interpreted as 111 deliberate uses of Serbian mi. A rigorous linguistic study should repeat this analysis on the original Serbian and distinguish explicit pronouns from subjects inferred through verb morphology.
The high number of negations is more robust conceptually. The speech is filled with boundaries: what will not be forgotten, what Serbia will not allow, what others did not do, and what political actors should not accept. Negation is therefore not merely grammatical; it contributes to a rhetoric of prohibition and historical determination.
Collocations: which words repeatedly travel together?
Collocation analysis asks a slightly different question from simple bigram counting. It identifies words that occur together more strongly than would be expected from their individual frequencies.
Among the strongest recurring combinations are world war, take care, Serbian people, last time, dear friends, much love, Republika Srpska, Banja Luka, common state, Serbian church and polycentric Serbia.
Table 8. Selected statistically prominent collocations
| collocation | count | count_nested | length | lambda | z |
|---|---|---|---|---|---|
| world war | 6 | 6 | 2 | 7.18 | 8.15 |
| take care | 6 | 6 | 2 | 7.18 | 8.15 |
| serbian people | 9 | 9 | 2 | 3.25 | 7.73 |
| last time | 4 | 4 | 2 | 6.02 | 7.61 |
| dear friends | 4 | 4 | 2 | 6.11 | 7.51 |
| much love | 4 | 4 | 2 | 5.08 | 7.37 |
| leadership republika | 5 | 5 | 2 | 5.26 | 7.28 |
| like president | 4 | 4 | 2 | 4.45 | 6.97 |
| long live | 3 | 3 | 2 | 7.15 | 6.86 |
| second world | 3 | 3 | 2 | 5.46 | 6.82 |
| dear nevenka | 3 | 3 | 2 | 5.71 | 6.80 |
| republika srpska | 23 | 23 | 2 | 10.32 | 6.76 |
| belgrade banja | 3 | 3 | 2 | 5.59 | 6.73 |
| many many | 4 | 4 | 2 | 3.74 | 6.31 |
| common state | 3 | 3 | 2 | 6.44 | 6.22 |
| serbian church | 4 | 4 | 2 | 3.97 | 6.19 |
| many times | 3 | 3 | 2 | 4.30 | 6.08 |
| never allow | 3 | 3 | 2 | 4.64 | 5.47 |
| banja luka | 6 | 6 | 2 | 10.99 | 5.39 |
| four times | 4 | 4 | 2 | 8.22 | 5.30 |
| one said | 3 | 3 | 2 | 3.36 | 5.27 |
| bosnia herzegovina | 4 | 4 | 2 | 10.62 | 5.17 |
| interests specific | 3 | 3 | 2 | 8.17 | 5.16 |
| thank republika | 3 | 3 | 2 | 3.28 | 5.16 |
| polycentric serbia | 4 | 4 | 2 | 4.89 | 5.12 |
| practically disappear | 3 | 3 | 2 | 10.37 | 5.01 |
| eyes sinise | 3 | 3 | 2 | 7.80 | 5.00 |
| serbia none | 3 | 3 | 2 | 4.10 | 4.89 |
| representative state | 3 | 3 | 2 | 7.54 | 4.86 |
| many serbs | 3 | 3 | 2 | 2.74 | 4.45 |
| serbia love | 3 | 3 | 2 | 2.88 | 4.42 |
| serbia republika | 4 | 4 | 2 | 2.31 | 4.27 |
| us serbs | 4 | 4 | 2 | 2.20 | 4.13 |
| serbs practically | 3 | 3 | 2 | 6.25 | 4.10 |
| srpska serbia | 3 | 3 | 2 | 1.85 | 3.14 |
| people serbia | 4 | 4 | 2 | 1.62 | 3.14 |
| serbs practically disappear | 3 | 0 | 3 | -2.36 | -0.73 |
| second world war | 3 | 0 | 3 | -2.42 | -1.04 |
| serbia republika srpska | 4 | 0 | 3 | -3.30 | -1.29 |
| thank republika srpska | 3 | 0 | 3 | -4.86 | -1.88 |
| belgrade banja luka | 3 | 0 | 3 | -5.59 | -1.90 |
| leadership republika srpska | 5 | 0 | 3 | -5.74 | -2.20 |
How to read a table. count indicates how many times a term appears, while lambda (λ) measures the strength of association of its words: larger positive values indicate that the words appear together more often than would be expected based on their individual frequencies. The z statistic shows the strength of that connection in relation to its statistical uncertainty; larger positive values provide stronger evidence that the collocation is not the result of random co-occurrence. Neither λ nor z are percentages or probabilities.
Several rhetorical domains emerge. Historical memory is represented by world war. Collective identity appears through Serbian people, Republika Srpska and Serbian church. Interpersonal proximity appears through dear friends, much love and take care. The phrase last time reflects the speaker’s repeated presentation of the speech as an important concluding presidential message.
One especially interesting collocation is polycentric Serbia. It occurs only four times, but repetition turns a relatively technical phrase into a political label. In context, “polycentrism” is constructed negatively, as fragmentation rather than pluralism. This is a useful example of why frequency alone is not enough: a phrase need not occur dozens of times to become rhetorically salient if it is concentrated in a highly emphatic passage.
Where does the audience react?
The transcript records 30 audience-reaction markers. Twenty-seven are ordinary [APPLAUSE], one combines applause with a raised speaker tone, one records the chant “Aco Srbine”, and one is marked as unclear. In other words, 28 of the 30 recorded events are applause-like reactions.
Table 9. Frequency and sequence of recorded audience reactions
| reaction_type | frequency |
| APPLAUSE | 27 |
| CHEERING : ACO SRBINE, ACO SRBINE, ACO SRBINE, ACO SRBINE | 1 |
| UNCLEAR | 1 |
| APPLAUSE – RAISED SPEAKER'S TONE | 1 |
Their distribution is not uniform. Approximately two reactions occur in the first quarter of the transcript, 13 in the second quarter, five in the third and ten in the final quarter. The first substantial burst therefore comes only after the opening memorial narrative has been established.
Applause clusters around several kinds of statements: rejection of external narratives; apologies to Krajina Serbs; the promise that there will be no new “Storm” or pogrom against Serbs; assertions that Serbia is now strong enough to protect its people; insistence on a single Serbian collective identity; rejection of “polycentric” Serbian identity; support for Republika Srpska; and statements of gratitude and mutual love between Serbia and Republika Srpska.
Near the end, the transcript records the chant “Aco Srbine” immediately before the speaker returns to Nevenka Dobrić and the image of her murdered son. That moment is especially revealing because personal leadership, collective identity and memorial emotion converge in the same part of the speech.
Audience reactions cannot tell us what the entire audience believed, and transcription may omit quieter or overlapping responses. But applause locations can identify passages that achieved immediate public resonance inside that particular gathering.
The major frames
Quantitative results become most useful when brought back to the actual speech. Close reading suggests at least seven interconnected frames.
The first is a memory and victimhood frame. The recurring image of Siniša Dobrić’s eyes personalises collective suffering. Historical events are repeatedly narrated through individual victims, mothers, children and families.
The second is a continuity-of-history frame. The First World War, the Second World War, Jasenovac, the wars of the 1990s, Operation Storm, Kosovo and more recent regional controversies are placed within a broad historical sequence. The speech therefore does not represent 1995 as an isolated event but as part of a recurring pattern.
The third is a unity frame. Serbia and Republika Srpska are repeatedly represented as politically distinct but belonging to a single Serbian people. Church, language, family ties and shared memory reinforce that collective boundary.
The fourth is a vulnerability-to-strength frame. Earlier Serbia is represented as weak, silent or unable to protect Krajina Serbs. Present and future Serbia are represented as economically, politically and militarily stronger. The repeated “never again” promise converts retrospective memory into a prospective security commitment.
The fifth is an external double-standard frame. Zagreb, Sarajevo, Podgorica, Europe and other international actors appear at different points as institutions or audiences that allegedly misunderstand, ignore or apply unequal standards to Serbian experiences.
The sixth is an internal-adversary frame. The speech distinguishes not only Serbs from external others but also politically acceptable and unacceptable positions within the Serbian community. “Salon nationalists”, those allegedly acting against Serbian interests, and advocates of “polycentric Serbia” become internal counter-figures.
Finally, there is a leadership and legacy frame. The repeated phrase that this is the speaker’s “last time” addressing the audience as president, the apologies for mistakes, claims of experience and promises of protection personalise the national narrative. The political community is linked not only to institutions but to the speaker’s own record and responsibility.
These frames overlap. Memory justifies unity; historical vulnerability justifies strength; strength justifies leadership; and the audience’s applause helps turn several of those propositions into moments of collective affirmation.
Conclusion
This case study shows both the attraction and the limits of computational political-text analysis.
The algorithms identify genuine structure. Collective identity dominates the vocabulary. The opening is lexically darker than the conclusion. Historical memory, national unity and promises of protection repeatedly intersect. Audience applause is concentrated around some of the strongest identity and security statements. Repetition and collocation expose phrases that organise the speech around Serbia, Republika Srpska, the Serbian people and the promise that earlier forms of suffering must not recur.
But the analysis also produces warnings. A positive word can occur in a negative accusation. Sarcasm can look positive to a sentiment dictionary. Translation can create pronouns that were implicit in Serbian.
That is perhaps the most useful lesson. Text analysis works best not as a substitute for reading, but as a disciplined way of deciding where to read more closely.
Methodology appendix
| Column | English explanation | Objašnjenje na srpskom |
| collocation | The multi-word expression being evaluated, e.g. world war or Serbian people. | Višerečni izraz koji se analizira, npr. world war ili Serbian people. |
| count | Number of times the complete expression occurs in the transcript. | Broj pojavljivanja kompletnog izraza u transkriptu. |
| count_nested | Number of occurrences that are also contained within a longer detected collocation. It is mainly relevant when collocations of different lengths are analysed simultaneously. | Broj pojavljivanja izraza koja su istovremeno deo duže identifikovane kolokacije. Posebno je relevantan kada se istovremeno analiziraju kolokacije različitih dužina. |
| length | Number of words in the collocation: 2 = bigram, 3 = trigram, etc. | Broj reči u kolokaciji: 2 = bigram, 3 = trigram itd. |
| lambda (λ) | Strength of association between the words. A larger positive λ means that the words occur together much more strongly than would be expected from their individual frequencies. It is an association score, not a percentage or probability. | Jačina povezanosti između reči. Veća pozitivna vrednost λ znači da se reči pojavljuju zajedno mnogo češće nego što bi se očekivalo na osnovu njihovih pojedinačnih frekvencija. To je mera asocijacije, a ne procenat niti verovatnoća. |
| z | Wald z-statistic for λ. It indicates how large the estimated association is relative to its statistical uncertainty. Larger positive z-values provide stronger evidence that the collocation represents a genuine association rather than random co-occurrence. | Waldova z-statistika za λ. Pokazuje koliko je procenjena povezanost velika u odnosu na njenu statističku neizvesnost. Veća pozitivna z-vrednost pruža snažniji dokaz da je reč o stvarnoj kolokaciji, a ne slučajnom zajedničkom pojavljivanju reči. |
How to read lambda and z
The simplest is:
lambda = how strong the bond is.
z = how statistically confident we are that the relationship is not random.
The R package quanteda calculates λ as an interaction parameter from a log-linear model, while z is the ratio of the estimated λ to its standard error. Therefore, a high lambda does not automatically mean the highest z: an expression may have a very strong estimated relationship, but if it occurs infrequently, the estimate may be less precise.
For example, the first row of the table:
world war — count = 6, λ = 7.18, z = 8.15
can be read like this:
“World war” occurs six times and shows a very strong association between its two component words (λ = 7.18), with a very high z-score (8.15), indicating that their co-occurrence is highly unlikely to be accidental.
As a very rough statistical rule, |z| > 1.96 corresponds approximately to a significance level of 5%, |z| > 2.58 to a level of about 1%, and values above 3 represent a very strong signal. However, for this blog I would avoid formal “hypothesis testing” and use z primarily to rank how convincingly related the collocations are, since you are analyzing a single speech exploratively, not a sample of speech from some population.
Director of Wellington based My Statistical Consultant Ltd company. Retired Associate Professor in Statistics.
Has a PhD in Statistics and over 45 years experience as a university professor, consultant, international researcher and government advisor.