bsmamba — live speaker extraction (causal, 44.1 kHz)
Give it three seconds of one person's voice. It then pulls that speaker out of whatever the microphone hears — other talkers, room reverb, noise — as it streams, with 46 ms of algorithmic latency. Trained on rooms, not studios.
STEP 1 — WHO TO FOLLOW
no speaker enrolled
STEP 2 — SEPARATE
idle — enroll a speaker first (headphones recommended)
this speakerblendeveryone else
this speaker
everyone else
speaker match
Live mode uses the causal model (hears nothing ahead of now). File mode uses the bidirectional model on the whole file for the highest quality. The "speaker match" bar is the server's absence gate: when the enrolled voice isn't there, the extracted output is muted and everything goes to "everyone else". Sessions are capped at 10 minutes and a few listeners; files at 2 minutes. Nothing is stored.