Sharing AI Knowledge Between Private Datasets Using Synthetic Data
This patent describes a method for two separate AI systems to share learned information by generating artificial data based on common features, without directly exchanging their private training data.
Patent Number
US 12039012
Status
Active
Filing Date
October 23, 2021
Grant Date
July 16, 2024
Expiration
October 23, 2041
Claims
23
Assignee
Sharecare AI
Inventors
Gabriel Gabra ZACCAK, Srivatsa Akshay SHARMA, Salvatore Giuliano VIVONA, Marina TITOVA
Citations
1 forward · 71 backward
What it covers
This technology allows different AI systems, called "federated endpoints," to learn from each other even when their training data must remain private. It works by first identifying "shared sample features" that are common between the first and second training datasets (Claim 1). For example, two hospitals might have patient data, but only the patient's age and gender are shared features. A "generator" AI model is trained on the first endpoint using these shared features and other specific data, creating a "one-hot encoding" of class labels (Claim 1). This trained generator then produces "synthetic samples" (artificial data) that mimic the original data's characteristics. Finally, this generator is used by the second federated endpoint to help it perform its own tasks, like making predictions or classifications (Claim 1). This means the second endpoint gets the benefit of the first endpoint's learning without ever seeing its sensitive raw data.
What it doesn't cover
- —Does not cover federated learning methods that only exchange model weights or gradients without generating synthetic data.
- —Does not cover systems where the separate training datasets have no common or 'shared sample features' at all.
- —Does not cover transferring raw, un-processed training data directly between the federated endpoints.
- —Does not cover methods that use encoding schemes other than 'one-hot encoding' for the class labels of the shared features.
- —Does not cover scenarios where a 'generator' is not explicitly trained to produce 'synthetic samples' for inference.
The clever bit
The clever part is enabling knowledge transfer between distinct, private datasets by training a generator on one dataset using shared features, and then using that *trained generator* to create synthetic data for another dataset's tasks. This avoids directly sharing sensitive raw data while still allowing the AI models to benefit from each other's learning.
Why it matters
This patent addresses a critical challenge in AI: how to leverage large, distributed datasets for training without compromising data privacy or security. It enables collaboration in sensitive fields like healthcare, where data cannot be centrally pooled due to regulations or competitive concerns. By allowing AI models to learn from each other's insights through synthetic data, it can lead to more robust and accurate AI systems across various organizations.
Real-world examples
- 1.Collaborative medical research where hospitals train AI models on patient data without sharing individual records.
- 2.Financial fraud detection systems that learn from different banks' transaction patterns while keeping customer data private.
- 3.Autonomous vehicle systems sharing driving experiences from different car manufacturers without exchanging raw sensor data.
- 4.Personalized health recommendations where AI models learn across user groups while maintaining individual privacy.
Generated by PatentBrief · Not legal advice · patentbrief.org
US 12039012 · 2026