# How AI Transfers Visual Styles Between Images While Keeping Structure

> This patent describes an AI method to generate a new image that combines the structure of one input image with the visual style of another, semantically related image, using a pre-trained Vision Transformer model.

- **Patent:** US 12282696
- **Original title:** Method and system for semantic appearance transfer using splicing ViT features
- **Owner:** Yeda Research and Development Co
- **Granted:** 2025
- **Status:** Active
- **Times cited:** 2
- **Field:** ai_ml, software, consumer_electronics, telecommunications, gaming

## What it does

This method generates a new image by taking two inputs: a 'source structure image' and a 'target appearance image'. The goal is to create a 'third image' that keeps the original structure from the first image but adopts the visual look from the second. For example, if you provide a picture of a person (source structure) and a picture of a famous painting (target appearance), the system aims to redraw the person in the style of the painting. It achieves this by training a 'generator' (an Artificial Neural Network, or ANN) to minimize differences between the appearance of the generated image and the target appearance, and optionally, to minimize differences between the structure of the generated image and the source structure (Claim 1, Claim 3). A key component is using a pre-trained, fixed Vision Transformer (ViT) model as a 'semantic prior' to understand and correctly match objects between the two input images (Abstract, Claim 8). This allows the generator to be trained effectively with just a single pair of input images.

## What it does NOT cover

- Does not cover style transfer methods that do not use a Vision Transformer (ViT) model as a semantic prior for training the generator.
- Does not cover image generation where the structure is also created or significantly altered, as it focuses on preserving the 'first structure' (Claim 1).
- Does not cover methods that require large datasets for training the generator itself, as this patent emphasizes training with only a 'single structure/appearance image pair' (Abstract).
- Does not cover transferring appearance without semantic awareness, meaning it aims to 'paint' semantically related objects, not just apply a general filter.
- Does not cover methods that rely solely on adversarial training, as the patent specifies training 'without adversarial training' (Abstract).

## The clever bit

The truly clever part is using a pre-trained and fixed Vision Transformer (ViT) model as an 'external semantic prior'. This allows a separate generator to be trained efficiently with only a single pair of input images, without needing extensive datasets or complex adversarial training, while still achieving high-quality, semantically accurate appearance transfer.

## Real-world examples

1. Virtual try-on applications for clothing or accessories
2. Digital fashion design tools
3. Content creation for advertising and media
4. Personalized avatar generation
5. Artistic style transfer applications
6. Special effects in film and video production

## Why it matters

This technology is important because it enables high-quality, semantically aware image manipulation with less training data and computational effort. By leveraging pre-trained AI models, it simplifies the process of transferring complex visual styles onto new structures. This can accelerate content creation workflows and open new possibilities for personalized digital experiences across various industries.

## Frequently asked questions

### What does How AI Transfers Visual Styles Between Images While Keeping Structure cover?

This patent describes an AI method to generate a new image that combines the structure of one input image with the visual style of another, semantically related image, using a pre-trained Vision Transformer model.

### Who owns patent US 12282696?

Yeda Research and Development Co owns this patent, granted in 2025.

### When does this patent expire?

This patent is expected to expire on December 18, 2042, when the invention enters the public domain.

### What is patent US 12282696 cited by?

This patent has been cited by 2 later patents that build on its ideas.

### What problem does this patent solve?

This technology is important because it enables high-quality, semantically aware image manipulation with less training data and computational effort. By leveraging pre-trained AI models, it simplifies the process of transferring complex visual styles onto new structures. This can accelerate content creation workflows and open new possibilities for personalized digital experiences across various industries.

### What does this patent NOT cover?

Does not cover style transfer methods that do not use a Vision Transformer (ViT) model as a semantic prior for training the generator.

**Full plain-English explainer:** https://patentbrief.org/patent/us/12282696/method-and-system-for-semantic-appearance-transfer-using-splicing-vit-features

**Original patent:** https://patents.google.com/patent/US12282696

---

_Source: PatentBrief — https://patentbrief.org. Patent facts are from public records; the plain-English explanation is PatentBrief's._


## Related patents

Semantically similar inventions in the PatentBrief corpus:

- [How AI Generates Images Based on Style and Content Cues](https://patentbrief.org/patent/us/12141700/generative-adversarial-network-for-generating-images) — This patent describes an artificial intelligence system that creates new images by understanding both what the image should show (like a cat) and how it should look (like a high-quality photo), using a special internal layer to blend these instructions.
- [How to Build 3D Models Using Pictures from Different Angles](https://patentbrief.org/patent/us/12737973/3d-model-generation-using-multiple-textures) — This patent describes a system that creates a three-dimensional digital object by combining multiple flat images taken from different viewpoints, then lets a user fine-tune how those images wrap around the object.
- [Upscaling Images with AI and Sub-Pixel Information](https://patentbrief.org/patent/us/12737843/upsampling-an-image-using-one-or-more-neural-networks) — This patent describes how artificial intelligence, specifically neural networks, can make low-resolution images look sharper by cleverly guessing new pixel colors based on tiny details between existing pixels.
- [Creating Your Own Image in Virtual and Augmented Reality](https://patentbrief.org/patent/us/12737996/providing-a-secured-self-representation-by-writing-to-a-portion-if-a-frame) — This patent describes how an artificial reality system can show a user their own image, called a self-representation, within a virtual world by using machine learning to identify and display parts of their live image.
- [Automated AI for Adapting to New Data Without Retraining](https://patentbrief.org/patent/us/20230162023/system-and-method-for-automated-transfer-learning-with-domain-disentanglement) — This patent describes an automated system that builds artificial intelligence models capable of adapting to new, different data without needing full retraining, by learning to ignore irrelevant changes.
