Plate25
Three Seconds Is Sufficient
on the theft of a voice, and what was stolen
The requirement used to be an hour of clean studio audio. Then twenty minutes. Then three seconds, taken from a voicemail greeting, over a phone line, with a television on in the room.
What the three seconds contain is not the voice. Three seconds cannot contain a voice. What they contain is enough to specify which voice, out of the space of all voices the model already holds — a coordinate, not a substance. The voice was already in there. It was in there before the person was born.
This is the part that people find difficult, and they are right to.
A mother receives a call. Her son is in trouble, and it is his voice, and the sentence he says is a sentence he has never said. She does the thing anyone would do.
Afterwards the question she keeps returning to is not whether she was fooled. She was not fooled; she heard her son. The question is what she has been listening to for twenty-nine years, if it can be specified in three seconds by a stranger with a laptop.