Seven decades of Italy's Sanremo and Sweden's Melodifestivalen, as data.
Two national song contests with different histories of English use. Sanremo has stayed mostly Italian and picks up the odd English word. Melodifestivalen sang in Swedish until 2002 and then mostly moved to English. The corpus lets you look at both, in the lyrics and in the sound.
Lyrics have had to be mostly Italian across its history. English turns up as individual borrowed words rather than as a change of language.
Swedish was required until 2002. After that English became common, and many entries are now sung entirely in English.
The corpus has 3,369 competing entries, 2,220 from Sanremo and 1,149 from Melodifestivalen, spanning the 1950s to 2026.
Each entry is tagged for its language, then measured for English borrowing and for switching between languages within a song.
Valence (how positive) and arousal (how energetic) are estimated from the audio, along with mood tags.
Entries, dates, artists, songwriter credits, and results come from Wikipedia, Wikidata, and contest archives.
Lyrics were found for 97.4 percent of entries from public sources. They are used only to compute the language features and are not shared.
Automatic language identification, anglicism counts against reviewed lexicons (the Italian lexicon was checked by a native speaker), and a code-mixing index.
Valence, arousal, and mood tags predicted from the audio with Music2Emo and linked to 1,995 songs. This part starts in 1972, where audio is available.
The features and metadata are released without the lyric text, under CC BY-NC-SA 4.0. The code is MIT.