A Computer Chip for Training Many AI Agents at Once
This patent describes a specialized computer processor designed to quickly train multiple artificial intelligence agents simultaneously using a unique set of instructions, speeding up how AI learns.
Patent Number
US 9754221
Status
Active
Filing Date
March 9, 2017
Grant Date
September 5, 2017
Expiration
March 9, 2037
Claims
22
Assignee
Alphaics
Inventors
Nagendra Nagaraja
Citations
53 forward · 2 backward
What it covers
The system uses a "first processor" to set up AI "reinforcement learning agents" and their "environments," assigning each a unique ID (Claim 1). A "first memory module" stores special instructions, part of an "application-domain specific instruction set (ASI)," which include these agent or environment IDs as "operands" (Claim 1). A "complex instruction fetch and decode (CISFD) unit" then decodes these instructions, creating "threads" that also carry the IDs (Claim 1). A "second processor" with multiple "cores" executes these threads in parallel, applying instructions to the correct agents or environments (Claim 1). This second processor determines actions, "state-value functions," "Q-values," and "reward values" for the agents. These values are stored in a "second memory module" and then used to train a "neural network" via a "neural network data path" to improve the AI's learning (Claim 1). For example, a system could train hundreds of robotic agents to navigate different simulated environments simultaneously using these specialized instructions.
What it doesn't cover
- —General-purpose computer processors running reinforcement learning algorithms without specialized instruction sets.
- —Reinforcement learning systems that do not use agent or environment IDs as operands within their instructions.
- —Hardware architectures that do not separate the creation of agents/environments from the execution of learning operations into distinct processors.
- —Reinforcement learning implementations that do not communicate with a neural network via dedicated data paths for approximating reward or state-value functions.
- —Systems where multiple agents are processed sequentially rather than in parallel using "Single Instruction Multiple Agents (SIMA)" type instructions.
The clever bit
The core innovation is the "Single Instruction Multiple Agents (SIMA)" concept, where a single instruction can be applied simultaneously to many reinforcement learning agents interacting with their environments. This, combined with a specialized processor architecture and instruction set, allows for highly parallel and efficient training of AI.
Why it matters
As artificial intelligence becomes more complex, training AI models, especially through reinforcement learning, requires immense computing power. This patent addresses this challenge by proposing dedicated hardware and instruction sets. This approach can significantly speed up the development and deployment of AI systems by making the learning process more efficient.
Real-world examples
- 1.Google's DeepMind AlphaGo and AlphaStar training systems
- 2.NVIDIA's specialized AI accelerators for robotics simulation
- 3.Robotics training platforms
- 4.Autonomous vehicle simulation environments
Generated by PatentBrief · Not legal advice · patentbrief.org
US 9754221 · 2026