Audio Source Separator
Two official fp16 ONNX conversions of Deezer Spleeter 2-stem estimate vocal and accompaniment frequency masks. The selected segment is decoded to 44.1 kHz stereo, passed through a 4096/1024 STFT and the outputs can be compared with the original mix.
Key features
- Run SHA-256-checked MIT-licensed Spleeter 2-stem converted models in local WASM
- Use the model’s 44.1 kHz stereo, periodic Hann 4096/1024 STFT and 512-frame input contract
- Estimate actual vocal and accompaniment tracks using squared-normalized outputs from two models
- Listen to original, vocal, accompaniment and recombined audio and inspect RMS before and after residual consistency
- Select a 3–10 second segment from a ≤90 second file, attest rights, cancel the memory worker and save Float32 WAV/JSON
How to use
- Choose an audio file you have rights to process and confirm your permission.
- Set a 3–10 second segment within the file.
- Run separation and wait for model download, STFT, inference and synthesis.
- Listen to the original, both estimated stems and recombination; inspect leakage and artifacts before saving WAV and JSON.
Use cases
- Draft a new arrangement from your own guide vocal and backing music
- Listen to a reduced-vocal practice mix from a song you may process
- Compare estimated stems with the recombined mix to hear artifacts
- Keep model hashes and reconstruction measurements with an authorized short clip
Frequently asked questions
Will it restore the original studio tracks exactly?
No. These are learned estimates. Instruments may leak into vocals and voice into accompaniment. Low reconstruction error does not guarantee stem quality.
Does near-zero recombination error mean the separation is perfect?
No. The tool records raw independent inverse-STFT RMS error, then assigns the remaining residual to accompaniment to keep the final sum close to the original. Neither number measures stem correctness.
Why only 10 seconds at a time?
The model uses 512-frame input chunks and browser memory and processing time are limited. You can select any 3–10 second segment within a file up to 90 seconds.
Can the tool verify copyright automatically?
No. You must confirm that you own or have permission to process the file. You remain responsible for the original recording’s use and sharing conditions.
Is my audio uploaded?
No. Audio bytes stay in browser memory. Only fixed model and WASM code files are downloaded from this site, and cancellation terminates the inference worker.
Privacy
Your audio is decoded, inferred and exported in the browser and is not uploaded. First use downloads two fixed models (about 38 MiB) and a WASM runtime (about 13 MiB) from this site. Rights confirmation is self-attested, not automatically verified.
Comments & questions