You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
I would like to propose adding a “Custom Vocal Integration” feature to ACE-Step 1.5, with two complementary modes:
Voice Performance Mode
Allow users to upload a recorded vocal performance (e.g., singing or humming), and use it as the main vocal track of the final song.
The system would:
Detect tempo, key, and structure from the input audio
Generate full instrumentation around the user’s voice (drums, bass, harmonies, etc.)
Preserve the original vocal recording without replacing it
This would be a natural extension of the existing Extract/Lego/Complete workflow.
Voice Cloning Mode
Allow users to provide a short reference of their voice so the system can generate entirely new songs sung in a similar voice.
The system would:
Learn vocal characteristics from a short audio sample
Generate new melodies and lyrics performed in the user’s voice
Optionally allow style control (emotion, intensity, genre adaptation)
Why this is useful
Makes music generation more personal and expressive
Bridges the gap between AI composition and human performance
Expands ACE-Step from a “music generator” into a true collaborative tool between user and model
I believe even implementing the first mode would already be a major step forward, while the second would unlock a much deeper level of creative control.
reacted with thumbs up emoji reacted with thumbs down emoji reacted with laugh emoji reacted with hooray emoji reacted with confused emoji reacted with heart emoji reacted with rocket emoji reacted with eyes emoji
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Custom Vocal Integration (Voice Input + Voice Cloning)
I would like to propose adding a “Custom Vocal Integration” feature to ACE-Step 1.5, with two complementary modes:
Allow users to upload a recorded vocal performance (e.g., singing or humming), and use it as the main vocal track of the final song.
The system would:
Detect tempo, key, and structure from the input audio
Generate full instrumentation around the user’s voice (drums, bass, harmonies, etc.)
Preserve the original vocal recording without replacing it
This would be a natural extension of the existing Extract/Lego/Complete workflow.
Allow users to provide a short reference of their voice so the system can generate entirely new songs sung in a similar voice.
The system would:
Learn vocal characteristics from a short audio sample
Generate new melodies and lyrics performed in the user’s voice
Optionally allow style control (emotion, intensity, genre adaptation)
Why this is useful
Makes music generation more personal and expressive
Bridges the gap between AI composition and human performance
Expands ACE-Step from a “music generator” into a true collaborative tool between user and model
I believe even implementing the first mode would already be a major step forward, while the second would unlock a much deeper level of creative control.
All reactions