How It Works
An 8-month journey from a Big Data experiment to a production-ready music recommendation system.
The Project
The Music Recommendation Engine was built by a team of 3 classmates over 8 months (on and off). We used it across multiple courses: Big Data, Machine Learning, and Deep Learning.
What started as a Big Data + ML project evolved into a full recommendation system similar to what you'd see in a streaming app's "next song" logic.
3 Members
8 Months
1,100+
How It Evolved
In the beginning, as someone who just got into how system design works for databases—while I was also working on making my own database from scratch side-by-side (Find-DB )—my simple goal was to put a ton of data and spread it into different tables and databases by some logic, just to get our hands dirty.
Data We Extracted:
- Audio features from WAV files
- Metadata about music (album, artist, track info)
- Compiled spectrograms of music files
- Extracted lyrics and sentiment analysis
- Wave features and acoustic properties
Then it slowly evolved—we started putting ML into our pipeline for recommendations using different embeddings from WAV files, lyrics, and even vision transformers to predict the next sequence of songs.
My Role & Responsibilities
I was mainly responsible for leading the team of 3 for the data ingestion cycle and making strategic decisions by conducting a literature review study. The goal was to ensure we didn't fall into the same loop as other music recommendation projects using the same old techniques.
We evaluated different deep learning models and have sent 2 research papers for this project. We're expecting a response in February 2026.
Scope Philosophy
"There was no predefined spec. I decided what to build, how far to take it, and when to stop."
Key Strategic Decisions
Based on our literature review of papers less than 20 years old, we made three critical decisions:
Treat This as a Data Engineering Problem First, ML Second
We invested early in ETL and schema design to build a clean and stable data warehouse that could evolve with progressive needs and new findings.
The Trade-off:
We sacrificed the first few months for slower visible progress because recommendation quality is bounded by data quality. This made debugging in later stages much easier when tables failed to give expected results.
Use Content-Based Similarity Over Collaborative Filtering
One of the most regrettable parts: we could not get real data for user interactions with the music.
In our efforts to falsify fake data for user interactions, we realized that compiling the scale of fake users to mimic these interactions would consume the vast majority of our time. So we had to drop user interaction data.
Testing Approach:
For all testing purposes, each of the 3 group members generated 10 playlists each from the 1,100 songs in our database.
Hybrid Framework Beyond Spectrogram-Based CNNs
We built a continuous voting ensembler where each feature set had different weights contributing to similarity from the last n songs played.
With the diverse dataset we acquired for each song, this approach helped connect every dot for the next potential song recommendation.
Future Plans (Phase 3)
We envision adding a real-time interface to capture whether music is liked or not via implicit + explicit features like:
- Skipping the song
- Skipping parts of it
- Replay behavior
- Volume adjustments
This would cater to taste first and labels later. Unfortunately, due to time and effort constraints, we couldn't achieve this.
The Challenge:
Each music data point containing such a diverse feature set means the periodic computation required for a song would be very computationally expensive.
From our research, a periodic time queue for songs would be carried out in a horizontally scaled way to update the social dynamics of the song. To keep things light on the user end, the updated model would be the new federated model for all until its next cycle of update.
These update cycles should be based on factors like streaming frequency, jumps of hype, popularity, etc.
What We Shipped
We shipped a working internal recommendation system that could generate next-song predictions from any playlist in under 4 seconds.
While it wasn't production-ready for consumers, it was stable enough that all three of us used it regularly to evaluate recommendation quality and debug failures.
The biggest win? How quickly we could iterate on features because of the data foundation we built early.