How we identify songs from audio, humming, and fragments.
Play the song, sing or hum it, type a misheard lyric, or describe what you remember. Here's the five-stage recognition and discovery pipeline that handles those different clues.
1. Audio and humming recognition
The mobile microphone records up to 10 seconds. ACRCloud compares the clip with its catalog of recorded music and melodies to identify songs that are playing nearby, sung, or hummed. Search That Song does not retain the clip.
2. Embedding match
Your query becomes a vector via OpenAI embeddings, then we run a nearest-neighbour search across millions of indexed lyric chunks. Paraphrases and misheard lines still find the right song.
3. Web reasoning
For vibe queries like "90s song with whistling, sounds melancholy", we run a real-time search and feed the results to an LLM that pulls out candidate matches.
4. Listening verification
Top candidates link to YouTube so you can confirm by ear in one click. The first 30 seconds is usually enough to know.
5. Neighbour discovery
Once you've confirmed a match, we surface five embedding-nearest songs (same era, same mood, same energy) for the rabbit hole.
Why this is harder than it looks
Music search has four properties that break standard search engines:
- The clue may be sound rather than words. A user may only have a noisy recording, a melody they can hum, or a few notes they can sing.
- Memory is fuzzy. People rarely remember the exact words. They remember the cadence, a similar word, or the wrong word entirely.
- Lyrics are repetitive. A four-word fragment can match thousands of songs. Disambiguation needs a model that understands which match is most likely your match.
- The query language is everything. Sometimes you have a phrase, sometimes a feeling, sometimes a description of the music video. A single retrieval method can't handle all three.
Recorded audio recognition versus humming recognition
Recorded-audio recognition looks for the fingerprint of an existing recording, even when a phone microphone adds background noise. Humming recognition focuses on the melody, allowing a sung or hummed version to lead back to the original song. The microphone feature supports both approaches so users do not need to know which kind of clue they have.
What makes our index different
Most lyric search tools index whole songs and rank by exact-phrase match. We index lyric chunks (overlapping windows of a few lines each) as embeddings, which means a fragment like "I'm walking on broken glass and dreaming" can match a song whose actual lyric is "walking through the broken parts, dreaming wide awake". Same meaning, different words.
What gets fed back into the system
Every confirmed match (when a user clicks through and confirms the song) gets re-embedded and added to the index, alongside the query that found it. Over time the system learns the natural mapping between how people describe songs and the songs themselves. The more it's used, the better it gets at the long tail.
What we don't do
We show only short lyric excerpts, just enough to confirm a match. We don't claim ownership of artwork or audio. We don't sell anonymized search data. The site is built to find songs, not to repackage them.