Keep your voice. Build the best performance.
Sunofriend is exploring lyric-aware, melody-aware vocal comping: record your own voice, compare several takes one musical phrase at a time, make the choices yourself, and assemble a natural dry vocal before considering any gentle correction.
Now planned as audio-native: no MIDI file required.
Target MIDI is no longer the canonical or required representation for vocal comping. The Semantic Musical State programme defines the forward plan: a shared, time-aligned musical state supports vocal comping and identity-preserving remix together, and a reviewed phrase can be recorded, compared and chosen with no MIDI input. The pilot below is the implemented v1 record and remains reproducible as historical evidence.
A working phrase pilot—not automatic comping.
A private browser prototype can play a reviewed phrase, record aligned vocal-only attempts, compare them with a reviewed melody and collect an explicit listening decision. Local analysis can rank evidence, but it does not select a take, tune a note, join phrases or render a replacement vocal.
The pilot has already exposed important design rules: supplied lyrics must remain canonical; speech transcription is only a rough phonetic clue; consonants and guttural closures are not failed pitch; and a singer must be able to request another relaxed pickup instead of accepting the highest score.
Three ways to cover a whole song.
Phrase-by-phrase recording
Hear one reviewed phrase, record several relaxed attempts, choose a benchmark and repeat only where needed.
Lowest stress and clearest feedback; slower for an entire song.
Complete takes, then repair gaps
Start from several full-song performances so breaths, tone and emotion carry naturally across lines.
Fast when a complete take is already close; harder when range or confidence varies sharply.
One base pass plus guided pickups
Keep the continuity of the best broad performance, then use the browser to replace only phrases that need another attempt.
Balances natural flow, manageable recording and a realistic path to finishing a whole song.
The proposed default is the hybrid route. It avoids a brittle word-by-word “Franken-vocal,” retains a broad human performance as the continuity anchor, and turns the browser into a focused pickup coach wherever that base is not good enough.
The proposed whole-song workspace.
One screen should answer three questions at all times: where am I in the song, what am I hearing now, and what decision is still needed?
“The lyric for this complete musical phrase”
Concept only. These controls illustrate the intended interaction; the public website does not record, upload or process audio.
From first phrase to export.
Map
Confirm lyrics, phrase boundaries and the intended melody before scoring any take.
Record
Loop a phrase with backing, melody or AI-reference cues; save every dry attempt at song zero.
Compare
Listen first, then reveal pitch, timing and signal evidence as supporting information.
Choose
Lock a human base for the phrase—or explicitly request another pickup or retain the AI voice.
Join
Preview transitions with breaths and handles preserved; reject any audible seam.
Polish
Optionally audition gentle correction on chosen regions only, then export a reviewed dry vocal stem.
What the design must protect.
- Your audio stays local. The public site has no upload or hosted vocal-processing endpoint.
- Listening outranks scoring. Automated evidence stays collapsed until the singer has heard the alternatives.
- The unit is usually a phrase. Word or syllable substitutions are rescue tools only after safe boundaries are reviewed.
- One broad base preserves identity. Switch penalties, breath handles, timbre and expression continuity matter as much as pitch.
- No acceptable take is a valid result. The interface asks for another pickup instead of hiding the gap with correction.
- Correction is optional and downstream. It may be auditioned gently on chosen audio only; originals and uncorrected renders remain intact.
- AI remains visibly separate. An authorised AI reference or duet region is a labelled fallback, never silently presented as the singer.
Go audio-native, then earn assembly.
The next useful step is the programme's Cycle 1: a minimum no-MIDI phrase decision with an exact source map, one or more fresh browser pickups and a playable replacement phrase against the backing—then the musical and workflow feedback, before adding analysis. In parallel, the first label snapshot is frozen and tiny local training experiments prove the pipeline; every trained output stays a research challenger that can never select a take.
Whole-song assembly, reviewed joins and gentle correction follow only after that audio-native loop is comfortable across complete songs. See the canonical programme plan ↗ and the remix research plan.