Algorithm info: FourSepNet SqrtMC v1.0.
Single model, not an ensemble.
Checkpoint: best epoch 107.
Input: 44.1 kHz stereo audio.
Output stems: vocals, drums, bass, other and instrumental.
Instrumental is calculated as drums + bass + other.
STFT: n_fft=4096, hop_length=1024.
Inference framework: PyTorch 2.0.1, CUDA 11.8.
Inference device: NVIDIA GeForce RTX 4090 D.
Metrics:
Metric sdr for instrum: 8.8098
Metric si_sdr for instrum: 8.2388
Metric l1_freq for instrum: 31.9042
Metric log_wmse for instrum: 11.0848
Metric aura_stft for instrum: 9.6447
Metric aura_mrstft for instrum: 10.5098
Metric bleedless for instrum: 19.2745
Metric fullness for instrum: 26.8477
Metric sdr for vocals: 9.1060
Metric si_sdr for vocals: 8.3118
Metric l1_freq for vocals: 31.1030
Metric log_wmse for vocals: 11.0848
Metric aura_stft for vocals: 5.6228
Metric aura_mrstft for vocals: 6.5622
Metric bleedless for vocals: 15.0397
Metric fullness for vocals: 17.2687
Date added: 2026-08-28 |