Automatic Classical Music and Pop Song Mashup Systems

The Hong Kong University of Science and Technology
Department of Computer Science and Engineering


MPhil Thesis Defence


Title: "Automatic Classical Music and Pop Song Mashup Systems"

By

Mr. Yu Foon Darin CHAU


Abstract:

Music mashups combine material from multiple songs into a coherent novel
composition, requiring the alignment of harmony, rhythm, phrase structure,
texture, and other factors. While human mashup artists rely on musical
intuition and production skills, automatic mashup systems must formalize
compatibility in computationally measurable and musically meaningful terms,
exploring transparent, explainable machine creativity and controlled
generation. This thesis presents two complementary automatic mashup systems
for pop and classical music that operationalize musical compatibility and
investigate automated mashups in both domains as principled, interpretable
compositions. We first generalize the notion of musical compatibility into a
unified, time-local, distance-based framework, subsuming prior harmonic and
spectral heuristics and admitting learned, structure-aware representations.
Building on this framework, we develop a pop-song pipeline operating in the
audio domain. We extract and analyze harmonic and rhythmic features using
transformer-based models over large musical corpora, retrieve compatible
segments using a beat-aligned distance, and generate mashups via source
separation, alignment, and mastering. We additionally propose a latent-feature
compatibility metric derived from the internal activations of a chord-
recognition model. Subjective evaluation shows that our system is broadly
comparable to the commercial RaveDJ service, while the latent-feature metric
outperforms both a discrete-chord baseline and AutoMashupper. We also
investigate automated mashups for symbolic music over the Western classical
repertoire. We first design and verify symbolic musical compatibility based
on traditional music-theoretic practices. This notion of compatibility
motivates a parallel pipeline in which the system standardizes scores to
MusicXML, extracts voice streams, motivic fingerprints, and cadential
annotations, and calculates musical compatibility among candidates using a
counterpoint-based voice-leading penalty. Across both systems, results show
that compatibility is necessary but not sufficient for an enjoyable mashup,
and that richer, structure-aware, and rule-aware representations meaningfully
improve perceived quality, pointing toward explainable, controllable
generative music.


Date:                   Thursday, 23 July 2026

Time:                   10:00am - 12:00noon

Venue:                  Room 5504
                        Lifts 25/26

Chairman:               Prof. Gary CHAN

Committee Members:      Prof. Andrew HORNER (Supervisor)
                        Dr. Xiaojuan MA