How Voice Assistants Recognize Who is Talking
This patent describes a system that uses a recorded voice sample to identify if new speech comes from the same person, allowing voice assistants to only respond to specific users.
Patent Number
US 11514901
Status
Active
Filing Date
June 11, 2019
Grant Date
November 29, 2022
Expiration
June 11, 2039
Claims
22
Assignee
Amazon Technologies
Inventors
Roland Maas, Sree Hari Krishnan Parthasarathi, Bjorn Hoffmeister, Brian King
Citations
1 forward · 22 backward
What it covers
The system first captures a sample of a user's voice, called "first audio data," during an initial setup or when a wake word is spoken (Claim 1). From this, it creates a unique "reference feature vector" that represents that user's voice (Claim 1). Later, when new "second audio data" comes in, the system compares it against this stored reference using a "trained model," like a neural network (Claim 1). It then figures out which parts of the new audio match the original speaker and which do not. For example, if a child says "Alexa, play music" and then the parent says "Alexa, stop," the system would identify the parent's voice as matching the reference and execute the "stop" command, while ignoring the child's voice if it doesn't match the reference. It only processes commands from the recognized speaker, "excluding" other voices (Claim 1).
What it doesn't cover
- —Does not cover systems that identify any speaker, only those that compare against a pre-determined reference speaker.
- —Does not cover systems that authenticate a speaker using a password or passphrase, as it focuses on voice characteristics.
- —Does not cover ignoring speech based on content (e.g., profanity filters), only based on speaker identity.
- —Does not cover systems that perform speech recognition on all incoming audio regardless of speaker, as it specifically "excludes" non-matching speech from command processing.
The clever bit
The novelty lies in using a "reference feature vector" derived from a specific speaker's voice to filter incoming audio before full speech recognition, ensuring that only desired speech from the intended speaker is processed for commands. This improves efficiency and accuracy by ignoring irrelevant voices.
Why it matters
This technology is crucial for personalizing voice assistant experiences and enhancing privacy. By ensuring that a voice assistant like Alexa only responds to its designated user, it prevents unintended interactions from other household members or background conversations. This allows for features like personalized music profiles or access to sensitive information only by the account holder.
Real-world examples
- 1.Amazon Echo devices (Alexa)
- 2.Google Assistant's Voice Match
- 3.Apple HomePod (Siri)
- 4.Smart speakers with personalized voice recognition
Generated by PatentBrief · Not legal advice · patentbrief.org
US 11514901 · 2026