How AI Learns to Run Faster on Specific Computer Chips
This patent describes how a smart AI system, called a reinforcement learning agent, trains other AI models to run more efficiently on specific computer hardware by cleverly reducing their size without losing too much accuracy.
Patent Number
US 20250371349
Status
Active
Filing Date
August 21, 2025
Grant Date
—
Expiration
August 21, 2045
Claims
23
Assignee
Intel
Inventors
Xiaofan Xu, Zsolt Biro, Cormac Brick
Citations
2 forward · 0 backward
What it covers
The patent details a method for optimizing neural networks to run efficiently on a specific hardware device. It uses a "reinforcement learning agent" (Claim 1) that receives an "embedding state" (Claim 1) representing characteristics of a neural network layer, such as its "kernel size" or "number of weights" (Claim 4). Based on this information, the agent generates "actions" (Claim 1) to reduce the "computational cycles" (Claim 1) needed to execute the neural network. A common action is "pruning weights" (Claim 2) within the layer, determined by a "sparsity ratio" (Claim 2). After the modified neural network runs, the system determines a "reward" (Claim 1) for the agent by checking if the output's "accuracy" (Claim 1) meets a set threshold and if a "target cycle reduction" (Claim 6) was achieved. This reward then helps the agent update its "policy" (Claim 1) to make better optimization decisions in the future. For example, an AI model designed to detect objects in a security camera feed could be optimized this way to run quickly on the camera's embedded processor without missing important events.
What it doesn't cover
- —Optimizing neural networks using methods that do not involve a reinforcement learning agent to generate actions.
- —Training neural networks without specifically considering the reduction of computational cycles on a target hardware device.
- —Pruning weights in a neural network without evaluating the impact on both accuracy and computational cycles as part of a reward system.
- —Optimization techniques that focus solely on improving neural network accuracy without also aiming to reduce execution time.
- —Manual optimization of neural network layers by a human engineer without an automated learning agent.
The clever bit
The innovation lies in using a reinforcement learning agent to intelligently *learn* how to optimize a neural network for a specific hardware's performance, rather than relying on fixed rules or manual tuning. This agent dynamically balances the trade-off between the model's accuracy and its computational cost on the target device.
Why it matters
As AI models become more complex, they demand significant computing power. This patent addresses the critical challenge of deploying these models on devices with limited resources, like smartphones, drones, or IoT sensors. By automatically making AI models smaller and faster for specific hardware, it enables more widespread and efficient use of artificial intelligence outside of large data centers. This is crucial for real-time applications and reducing energy consumption.
Real-world examples
- 1.AI models running on smartphone processors
- 2.Machine learning inference on embedded systems in smart home devices
- 3.Neural networks deployed on edge AI accelerators
- 4.Computer vision tasks on autonomous vehicle hardware
- 5.Optimized AI for industrial IoT sensors
Generated by PatentBrief · Not legal advice · patentbrief.org
US 20250371349 · 2026