How AI Generates Images Based on Style and Content Cues
This patent describes an artificial intelligence system that creates new images by understanding both what the image should show (like a cat) and how it should look (like a high-quality photo), using a special internal layer to blend these instructions.
Patent Number
US 12141700
Status
Active
Filing Date
September 17, 2020
Grant Date
November 12, 2024
Expiration
September 17, 2040
Claims
41
Assignee
Naver
Inventors
Naila Murray
Citations
3 forward · 5 backward
What it covers
This patent details a generative adversarial network (GAN) that creates synthetic images. It uses two main parts: a generator neural network and a discriminator neural network. The generator takes a random 'noise vector' and a pair of 'conditioning variables' as input (Claim 1). One variable describes 'semantic information,' like what objects are in the image, and the other describes 'aesthetic information,' such as its quality or style (Claim 1). A key part of the generator is a 'mixed-conditional batch normalization layer' which normalizes the data flow within the network (Abstract, Claim 1). This layer uses the semantic and aesthetic conditioning variables to adjust its internal parameters, allowing it to precisely control how the generated image reflects both content and style (Abstract, Claim 1). For example, you could ask it to generate a 'dog' (semantic) that looks 'beautiful' (aesthetic).
What it doesn't cover
- —Does not cover image generation systems that only use semantic information without also incorporating aesthetic information as a distinct conditioning variable.
- —Does not cover GANs that generate images without using a 'mixed-conditional batch normalization layer' that specifically adjusts its parameters based on both semantic and aesthetic inputs.
- —Does not cover image generation where the conditioning variables are not applied via an 'affine transformation' to compute the normalization layer's parameters.
- —Does not cover image generation where the aesthetic information is not provided as 'continuous information in the form of histogram score distributions' (Claim 1).
The clever bit
The clever part is how the patent uses a 'mixed-conditional batch normalization layer' within the generator. This layer dynamically adjusts its internal scaling and shifting parameters based on both semantic (what it is) and aesthetic (how it looks) instructions, allowing the AI to blend these two types of guidance seamlessly when creating an image.
Why it matters
This technology is important because it allows for more precise control over AI-generated images. Instead of just asking an AI to create 'a dog,' this system enables users to specify 'a beautiful dog' or 'a low-quality dog,' giving creators fine-grained control over both content and style. This level of control is crucial for applications where specific visual characteristics are desired, moving beyond simple object generation to nuanced artistic or functional outputs.
Real-world examples
- 1.AI art generators that allow users to specify style and content
- 2.Tools for generating synthetic datasets with controlled visual properties for training other AIs
- 3.Image editing software offering style transfer or content-aware generation features
- 4.Virtual try-on applications where clothing style and fit are controlled
Generated by PatentBrief · Not legal advice · patentbrief.org
US 12141700 · 2026