
Mvsep
Audio AI Tools
MVSEP separates music into vocal and instrumental stems with AI and supports audio-to-text transcription.

What does Mvsep do?
MVSEP is an AI music separation tool that performs audio splitting into voice and music parts. Use drag-and-drop or upload a file (including remote and batch upload) and choose a separation type and output encoding.
You can export results in multiple formats such as MP3 (320 kbps), and registered users can access higher-quality lossless/uncompressed options (e.g., WAV 16-bit/FLAC 16-bit). Premium options include higher bit-depth exports like WAV 32-bit and FLAC 24-bit. Registering also provides higher priority in the processing queue and enables lossless exports.
Beyond separation, MVSEP includes text extraction from audio. It also offers a growing collection of specialized models for different instruments and tasks, along with demos and experimental voice/speech generation options.
What does MVSEP do with my audio?
It separates audio into voice and music parts using AI. It also supports extracting text from audio.
How do I upload files to run separation?
You can drag and drop to upload, browse for a file, use remote upload, or run batch uploads.
What separation types are available?
One available option is “BS Roformer SW,” which can separate into vocals, bass, drums, guitar, piano, and other categories.
What output formats can I export?
MP3 (320 kbps) is available. Lossless and uncompressed WAV/FLAC options are available for registered users, and higher bit-depth WAV 32-bit / FLAC 24-bit options are available with premium.
Does registering change anything?
Registration gives higher priority in the processing queue and unlocks lossless exports.
Are instrument-focused models available?
Yes. MVSEP includes additional instrument-specific models (with demos) such as guitar variants, plucked strings, percussion, keys, brass, woodwind, choir, and more in its model lineup.