# How Voice Assistants Recognize Who is Talking

> This patent describes a system that uses a recorded voice sample to identify if new speech comes from the same person, allowing voice assistants to only respond to specific users.

- **Patent:** US 11514901
- **Original title:** Anchored speech detection and speech recognition
- **Owner:** Amazon Technologies
- **Granted:** 2022
- **Status:** Active
- **Times cited:** 1
- **Field:** consumer_electronics, software, ai_ml, telecommunications

## What it does

The system first captures a sample of a user's voice, called "first audio data," during an initial setup or when a wake word is spoken (Claim 1). From this, it creates a unique "reference feature vector" that represents that user's voice (Claim 1). Later, when new "second audio data" comes in, the system compares it against this stored reference using a "trained model," like a neural network (Claim 1). It then figures out which parts of the new audio match the original speaker and which do not. For example, if a child says "Alexa, play music" and then the parent says "Alexa, stop," the system would identify the parent's voice as matching the reference and execute the "stop" command, while ignoring the child's voice if it doesn't match the reference. It only processes commands from the recognized speaker, "excluding" other voices (Claim 1).

## What it does NOT cover

- Does not cover systems that identify any speaker, only those that compare against a pre-determined reference speaker.
- Does not cover systems that authenticate a speaker using a password or passphrase, as it focuses on voice characteristics.
- Does not cover ignoring speech based on content (e.g., profanity filters), only based on speaker identity.
- Does not cover systems that perform speech recognition on all incoming audio regardless of speaker, as it specifically "excludes" non-matching speech from command processing.

## The clever bit

The novelty lies in using a "reference feature vector" derived from a specific speaker's voice to filter incoming audio before full speech recognition, ensuring that only desired speech from the intended speaker is processed for commands. This improves efficiency and accuracy by ignoring irrelevant voices.

## Real-world examples

1. Amazon Echo devices (Alexa)
2. Google Assistant's Voice Match
3. Apple HomePod (Siri)
4. Smart speakers with personalized voice recognition

## Why it matters

This technology is crucial for personalizing voice assistant experiences and enhancing privacy. By ensuring that a voice assistant like Alexa only responds to its designated user, it prevents unintended interactions from other household members or background conversations. This allows for features like personalized music profiles or access to sensitive information only by the account holder.

## Frequently asked questions

### What does How Voice Assistants Recognize Who is Talking cover?

This patent describes a system that uses a recorded voice sample to identify if new speech comes from the same person, allowing voice assistants to only respond to specific users.

### Who owns patent US 11514901?

Amazon Technologies owns this patent, granted in 2022.

### When does this patent expire?

This patent is expected to expire on June 11, 2039, when the invention enters the public domain.

### What is patent US 11514901 cited by?

This patent has been cited by 1 later patents that build on its ideas.

### What problem does this patent solve?

This technology is crucial for personalizing voice assistant experiences and enhancing privacy. By ensuring that a voice assistant like Alexa only responds to its designated user, it prevents unintended interactions from other household members or background conversations. This allows for features like personalized music profiles or access to sensitive information only by the account holder.

### What does this patent NOT cover?

Does not cover systems that identify any speaker, only those that compare against a pre-determined reference speaker.

**Full plain-English explainer:** https://patentbrief.org/patent/us/11514901/anchored-speech-detection-and-speech-recognition

**Original patent:** https://patents.google.com/patent/US11514901

---

_Source: PatentBrief — https://patentbrief.org. Patent facts are from public records; the plain-English explanation is PatentBrief's._


## Related patents

Semantically similar inventions in the PatentBrief corpus:

- [How Smart Speakers Know You're Talking to Them After a Command](https://patentbrief.org/patent/us/11361763/detecting-system-directed-speech) — This patent describes how a smart speaker system can tell if follow-up speech is meant for it, even without a "wake word," by analyzing voice activity and partial speech recognition results using an AI model.
- [How Wearable Devices Use Location and Biometrics to Know Who You Are](https://patentbrief.org/patent/us/20220189458/speech-based-user-recognition) — This patent describes how a computer system can figure out who is using a wearable device by combining its location data with biometric information like a face scan or voice recording.
- [How Voice Assistants Change Their Speech Based on How You Talk](https://patentbrief.org/patent/us/10276149/dynamic-text-to-speech-output) — This patent describes a system where a voice-controlled device adjusts its text-to-speech output characteristics, like speed or tone, based on the user's speaking habits or current situation, making responses feel more natural and personalized.
- [How Sonos Speakers Use Personalized Wake Words to Recognize Different Users](https://patentbrief.org/patent/us/9965247/icloud-drive) — A system that lets multiple people control a shared speaker by using unique voice-trigger words to link their specific music accounts and preferences.
- [Making Computer Voices Sound More Expressive from Your Speech](https://patentbrief.org/patent/us/11062694/text-to-speech-processing-with-emphasized-output-audio) — This patent describes how a computer listens to your spoken words, figures out which parts you emphasized, and then makes its own computer-generated voice emphasize those same parts when it speaks back.
