TL;DR
This project presents a comprehensive computational analysis of 1,891 Egyptian songs spanning the period from 1920 to 2022. By leveraging Python-based automation and Natural Language Processing (NLP), the study quantified shifts in song popularity, sentiment polarity, and the prevalence of offensive language. The findings reveal a significant transition toward “neutral” and “negative” content in the 21st century, correlating specifically with the rise of Mahraganat and Egyptian Rap.
Problem
While Western musicology has extensively documented the evolution of lyrical content, there is a critical lack of empirical data regarding the Egyptian musical landscape. Specifically, the cultural shifts reflected in lyrics over the last century remain largely undocumented. This research gap leaves the impact of emerging genres, such as Mahraganat and Egyptian Rap, on public sentiment and social norms unquantified. The objective of this project was to develop a data-driven framework to identify these trends, providing the Egyptian Musician’s Syndicate with evidence-based insights to better understand and regulate modern lyrical content.
Design Iterations
The project architecture was divided into two primary phases: Data Acquisition and Analytical Processing.
Phase 1: Multi-Source Data Collection
To ensure a statistically robust sample size, a multi-source scraping strategy was implemented:
-
Historical Data (1920–2000): Utilized
Discogs-clientand Selenium to scrape Wikipedia. This approach was necessary because older tracks are frequently absent from modern streaming databases. -
Modern Data (2000–2022): Leveraged the
Spotipylibrary to interface with the Spotify API, specifically targeting the “Egyptian” genre to capture contemporary trends such as Mahraganat. -
Lyric Extraction: A hybrid methodology was employed where
LyricsGeniusserved as the primary API for modern tracks, while a custom Selenium-based scraper targeted Google results to retrieve lyrics for older tracks where automated tools failed.
Phase 2: Data Refinement & Analysis
The raw dataset underwent a rigorous multi-stage preprocessing pipeline:
- Normalization: Removal of diacritics (Tashkeel), punctuation, and non-standard characters using Python’s
remodule to ensure consistency across the corpus. - Sentiment Analysis: Implemented the CAMeL model to categorize lyrics into positive, negative, or neutral sentiments—selected for its high accuracy in Arabic sentiment detection.
- Offensiveness Detection: A custom script cross-referenced a curated list of 245 culturally inappropriate terms to quantify the “offensiveness” of the dataset over time.
Technical Details
The technology stack was selected to maximize automation and ensure data integrity:
-
Automation: Selenium was employed for complex web scraping where dynamic content or multi-step navigation was required (e.g., navigating Wikipedia’s varying page structures).
-
API Integration:
Discogs-clientandSpotipywere utilized to handle high-frequency requests efficiently, significantly reducing the time and computational resources required compared to manual scraping. -
NLP Pipeline: The analysis utilized CAMeL for sentiment classification and a custom Python script to perform keyword-based filtering for offensive content.
-
Statistical Analysis: Pearson’s correlation coefficient () was calculated to determine the strength of the relationship between time (decades/years) and variables such as popularity and sentiment.
Results & Key Metrics
The analysis yielded several critical data points regarding the evolution of Egyptian music:
- Popularity Trends: A strong positive correlation () was identified between time and song popularity, with scores increasing from an average of 2.40 in the 1920s to 40.77 in the 2020s.
- Sentiment Shift: A statistically significant shift toward “neutral” and “negative” sentiments was observed starting in the 2000s. Notably, positive sentiment showed almost no correlation with time ().
- Offensiveness Growth: The percentage of songs containing offensive terms exhibited a sharp upward trajectory after 1950, peaking at 30% in the 2020s.
- Table of Key Metrics: Results from the sentiment and offensiveness analysis are summarized in the following table per decade:
| Decade | Number of Songs | Positive | Negative | Neutral | Song Popularity Score | Songs with Offensive Words | Songs Without Offensive Words | Percentage of Songs Containing Offensive Words |
|---|---|---|---|---|---|---|---|---|
| 1970 | 24 | 16 | 2 | 6 | 22.17 | 1 | 23 | 4% |
| 1980 | 14 | 6 | 0 | 8 | 17.00 | 2 | 12 | 17% |
| 1990 | 36 | 15 | 3 | 18 | 28.67 | 2 | 34 | 6% |
| 2000 | 196 | 108 | 15 | 73 | 29.74 | 3 | 193 | 2% |
| 2010 | 222 | 93 | 39 | 90 | 34.89 | 14 | 208 | 7% |
| 2020 | 69 | 20 | 15 | 34 | 45.78 | 9 | 60 | 15% |
Academic Poster
Publication Link
The full research paper is available for download at the following link: Download PDF