In short
Voice data used for recreation or cloning needs explicit consent naming that use, a stated retention period, a withdrawal route and records that can be produced during an audit.

A few minutes of clean speech is now enough to recreate a convincing voice. Consent language written for transcription projects does not cover that use.
What the release has to say
- That the recording may be used to recreate a voice
- Which categories of output are permitted and which are excluded
- How long the recording and any derived output are kept
- How the contributor withdraws, and what withdrawal removes
Keep provenance attached
Every file should carry its consent reference. If a dataset cannot say which release covers which recording, it cannot honour a withdrawal request, and that failure surfaces at the worst possible moment.
Pay properly
A voice that will be reused indefinitely is worth more than an hour of recording time. Rates that reflect the actual use are both fairer and better for retention.
Practical checklist
- Define the acceptance rule before any volume starts
- Review a small pilot before committing the full budget
- Track errors by category, language and reviewer
- Keep consent, source notes and version history with the files
Before you ask for a quote
A clear brief saves days. Share a sample file, target language or region, expected volume, deadline, quality threshold and any privacy restrictions. A supplier can then price the work on real effort rather than assumptions.
- Which languages, markets or user groups must be represented?
- What format does the final file need to arrive in?
- Who will approve ambiguous cases during the pilot?
- voice
- consent
- ethics
Need this done rather than read about it?
We run collection, annotation, transcription and localization projects for teams who would rather spend their time on the product.
Start a conversation