What is a stem?

In this workflow, a stem is an audio track containing a particular part, such as vocals, drums or bass. Combining the estimated stems recreates an approximation of the original mix, with each part available for individual control.

Why separation is difficult

A finished song combines many sounds that overlap in time and frequency. A singer and guitar may share similar frequencies. Reverb and effects can spread one source across the mix. The original parts cannot simply be removed like layers in an image file.

What the AI contributes

Source-separation models learn patterns associated with voices and instruments. They use those patterns to estimate components of mixed audio. Different models and split options can produce different results; the app’s internal model architecture is not specified here.

Why artefacts appear

When a model cannot clearly assign a sound, you may hear traces of another instrument or a watery quality in the stem. Compression, heavy effects, live crowd noise and dense arrangements can make the task harder.

How to evaluate a result

Compare a chorus, a quiet passage and a section with percussion. Use headphones and listen at a reasonable volume. Evaluate whether the separation serves your purpose: a practice backing track has different needs from a production-ready isolated vocal.

Keep exploring

Remove vocals: the step-by-step guide ↗Isolate the vocal part ↗Read the app FAQ ↗View the current App Store listing ↗