Introduction to Music Information Retrieval (MIR)

  • Extracts meaningful features from music for indexing and retrieval
  • Enables content-based search and recommendation systems
  • Incorporates user behavior and preferences into models
  • Recognizes the multimodal nature of music

Evolution and Growth of MIR

  • Emerged in the last two decades
  • Driven by audio compression, computing power, and mobile music players
  • Expanded with music streaming services
  • Shifted from symbolic representations (MIDI) to direct audio signal processing

Key Applications of MIR

  • Music retrieval through similarity-based search
  • Computational music theory and comparative analysis
  • Music creation via “audio mosaicing”
  • Enhancing expert knowledge with large datasets

Challenges and Future Directions

  • Understanding user preferences and behavior
  • Developing high-level music descriptors
  • Expanding research beyond Western music traditions
  • Building scalable, robust MIR systems

The Future of MIR

  • Personalized music recommendation systems
  • Integration of multiple sensory modalities
  • Considering technological, social, and cultural perspectives
  • Potential for AI-driven musical companionship

Audio Content Analysis

Music Feature Extraction in MIR

  • Extracts meaningful characteristics (descriptors) from music
  • Features classified by abstraction level, temporal scope, and musical facet
  • Low-level features describe signal properties like timbre and energy
  • High-level features represent musical elements like melody and harmony

Low-Level Music Features

  • Derived directly from the audio signal, mainly in the frequency domain
  • Represent timbre, loudness, and spectral properties
  • Examples: Zero Crossing Rate, MFCCs, spectral moments
  • Used for timbre representation and audio fingerprinting

See: Freesound API, Freesound Timbral Search

High-Level Music Features

  • Semantically meaningful descriptors inferred from low-level features
  • Melody features derived from pitch content and periodicity
  • Harmony features include chroma representation of tonal structure
  • Rhythm features analyze timing, tempo, and meter patterns

ex: Real-time pitch extraction

Find Similar Sounds with Freesound

similar sounds

Signal Processing in MIR

  • Fundamental for analyzing audio signals and extracting features
  • Two main approaches: time-domain and frequency-domain analysis
  • Time-domain captures amplitude changes over time
  • Frequency-domain reveals spectral content and pitch information

Time-Domain Analysis

  • Represents audio signals as amplitude over time
  • Extracts low-level features like Zero Crossing Rate and RMS energy
  • Identifies transients and note onsets through waveform analysis
  • Useful for loudness, attack time, and rhythmic structure analysis

ex: Essentia.js

ex: Mel-spectrogram

Frequency-Domain Analysis

  • Uses Fourier Transform to represent signals in the frequency domain
  • Short-Time Fourier Transform (STFT) provides time-varying spectral analysis
  • Extracts features like spectral moments, MFCCs, and chroma features
  • Essential for timbre description, pitch estimation, and tonality analysis

ex: Chroma

Machine Learning in MIR

  • Infers high-level musical information from extracted features
  • Classification and auto-tagging predict genres, instruments, and moods
  • Deep learning models (CNNs) enhance music recognition tasks
  • HMMs, GMMs, and SVMs assist in segmentation and pattern recognition

Musical Structure Analysis

Segmentation

  • Divides an audio stream into homogeneous sections
  • Used in speech/music separation, note identification, and structure analysis
  • Methods include model-free (feature change detection) and model-based (trained probabilistic models)
  • Related to onset detection, novelty detection, and music structure analysis

ex: Audio Slicer

Pattern Recognition

  • Identifies motifs, chords, and melodies in music
  • Uses self-similarity analysis and novelty detection for motif detection
  • Chord recognition relies on chroma features, templates, and probabilistic models
  • Melody extraction applies f₀ estimation and source separation techniques

See: Music21 or Sonic Visualiser or Musipedia

Music Metadata

Descriptive Metadata

  • Categorizes music using title, artist, album, genre, and style
  • Genre classification models music based on content and context
  • Auto-tagging assigns semantic labels from user-generated data
  • Used in music retrieval, recommendation, and indexing systems

See: MusicBrainz or Discogs

Semantic Metadata

  • Describes music’s meaning and expressive content (lyrics, mood, instrumentation)
  • Lyrics provide textual context and influence perception
  • Mood and emotion recognition classify expressive intent
  • Instrumentation helps define timbre, genre, and music content

See: Musixmatch

Music Retrieval Systems

Query-by-Example

  • Uses an audio sample to retrieve similar music
  • Similarity measured locally (within excerpts) or globally (between pieces)
  • Humming or singing can serve as a query input
  • Commercial systems (e.g., SoundHound) match queries to a database

ex: SoundHound Music free ex: Real-time music autotagging

Query-by-Text

  • Uses keywords, semantic tags, or lyrics to retrieve music
  • Semantic search engines fuse metadata with audio features
  • Tag-based retrieval assigns descriptive labels (e.g., mood, genre)
  • Lyric search models the contextual and multimodal aspects of music

ex: Spotify, Apple Music, etc.

Music Recommendation Systems

Collaborative Filtering

  • Uses user preferences or item similarity to generate recommendations
  • User-based filtering recommends music based on similar users’ tastes
  • Item-based filtering recommends based on similarity between songs
  • Explicit (likes, ratings) and implicit (skips, listening history) feedback shape recommendations

ex: Spotify

Content-Based Filtering & Hybrid Systems

  • Content-based filtering analyzes music features for recommendations
  • Uses low-level (timbre, rhythm) and high-level (genre, mood) features
  • Hybrid systems combine collaborative and content-based methods
  • Context-aware systems factor in location, time, and user state

ex: Pandora

Music Classification

  • Genre classification uses machine learning to assign genre labels
  • Features like timbre, rhythm, and harmony aid classification
  • Mood classification analyzes emotional content using semantic labels
  • Machine learning models (e.g., SVM, kNN) improve classification accuracy

ex Genre Classifier ex: Mood Classifier

Music Clustering

  • Groups music based on similarity in audio features or user interactions
  • Similarity-based clustering analyzes feature vectors from audio content
  • User-generated clustering uses collaborative tags and playlist data
  • Supports visualization tools and personalized recommendation systems

Music Creation Tools

Music browsing/playlists

More Projects at Freesound Labs