LT3, Ghent University

SanreMello Corpus

Seven decades of Italy's Sanremo and Sweden's Melodifestivalen, as data.

Two national song contests with different histories of English use. Sanremo has stayed mostly Italian and picks up the odd English word. Melodifestivalen sang in Swedish until 2002 and then mostly moved to English. The corpus lets you look at both, in the lyrics and in the sound.

Italy · 1951 to present
Sanremo

Lyrics have had to be mostly Italian across its history. English turns up as individual borrowed words rather than as a change of language.

Sweden · 1958 to present
Melodifestivalen

Swedish was required until 2002. After that English became common, and many entries are now sung entirely in English.

3,369
entries
2
contests, 7 decades
97.4%
lyric coverage
2,168
tracks scored for emotion
Corpus

Two contests, seven decades

The corpus has 3,369 competing entries, 2,220 from Sanremo and 1,149 from Melodifestivalen, spanning the 1950s to 2026.

Number of competing entries per decade for each contest
Entries per decade. Counts for each contest, by decade.
Language

Use of English

Each entry is tagged for its language, then measured for English borrowing and for switching between languages within a song.

Share of Melodifestivalen entries sung in English, by year
Song language. Melodifestivalen entries were almost all Swedish until 2002, when the English share rises quickly. Sanremo has no comparable change.
Live anglicism density per decade for both contests
Borrowed words. English loanwords per 1000 tokens inside the majority language, over time, for both contests.
Mean code-mixing index before and after 2002
Code-switching. Average within-song code-mixing. Melodifestivalen's rises after 2002; Sanremo's stays about the same.
Language of each lyric line for four example songs
Line by line. Each cell is one lyric line, colored by its language. Melodifestivalen songs mix languages across lines; Sanremo songs stay in one language and borrow single words.
Sound

How the two contests sound

Valence (how positive) and arousal (how energetic) are estimated from the audio, along with mood tags.

Valence and arousal distributions and valence over time
Valence and arousal. Melodifestivalen tends to be more positive and more energetic than Sanremo, across the whole period.
Rise and fall of individual mood tags over time
Moods over time. The share of songs each year carrying each mood tag.
Emotional quadrant composition by decade for each contest
Quadrant mix. The share of joyful, tense, tender, and sad songs per decade.
How far each winner sits from its own year's average
Winners. How far each winning song sits from the average sound of its own year.
Method

How it was built

Corpus

Entries, dates, artists, songwriter credits, and results come from Wikipedia, Wikidata, and contest archives.

Lyrics

Lyrics were found for 97.4 percent of entries from public sources. They are used only to compute the language features and are not shared.

Language features

Automatic language identification, anglicism counts against reviewed lexicons (the Italian lexicon was checked by a native speaker), and a code-mixing index.

Emotion features

Valence, arousal, and mood tags predicted from the audio with Music2Emo and linked to 1,995 songs. This part starts in 1972, where audio is available.

Use it

Data, code, and an explorer

The features and metadata are released without the lyric text, under CC BY-NC-SA 4.0. The code is MIT.

SanreMello Corpus
Thomas Moerman, Pranaydeep Singh, Alessandra Teresa Cignarella, and Joni Kruijsbergen. LT3, Ghent University.