How Multiple Graphics Cards Keep Their Memory Caches Updated
This patent describes methods for keeping the specialized memory caches of multiple graphics processing units (GPUs) consistent and up-to-date when they work together in a computer system.
Patent Number
US 12737317
Status
Active
Filing Date
November 21, 2023
Grant Date
September 15, 2026
Expiration
~November 2043 (estimated)
Claims
0
Assignee
—
Inventors
—
Citations
0 forward · 0 backward
What it covers
This patent details a system for managing memory in computers with multiple graphics processing units (GPUs), often called a 'multi-tile architecture.' Each GPU has its own main memory, a fast 'memory side cache,' and a 'memory management unit' (MMU). These GPUs communicate through a 'communication fabric.' The core mechanism is that one GPU's MMU (e.g., the first MMU) is not only responsible for its own memory and cache updates but also determines if the 'memory side cache' of another GPU (e.g., the second memory side cache) needs to be updated. For example, if a high-performance computer uses two GPUs to render a complex scene, and one GPU updates a texture, its MMU checks if the other GPU's cache holds an outdated version of that texture and, if so, instructs it to refresh.
What it doesn't cover
- —Does not cover systems with only a single graphics processing unit (GPU).
- —Does not cover systems where a central processing unit (CPU) or other non-GPU component solely manages cross-GPU cache coherence.
- —Does not cover multi-GPU systems that do not employ 'memory side caches' as described.
- —Does not cover systems where each GPU's cache is updated entirely independently without coordination from another GPU's MMU.
- —Does not cover multi-GPU setups that lack a 'communication fabric' for direct GPU-to-GPU data exchange.
The clever bit
The clever part is having one GPU's Memory Management Unit (MMU) intelligently decide if another GPU's memory side cache needs updating. This allows for efficient data consistency across multiple graphics cards without constant, broad invalidations, improving overall system performance.
Why it matters
This technology is important for high-performance computing, especially in fields like artificial intelligence training and professional graphics. When multiple GPUs work together on large datasets, ensuring all their memory caches are consistent and up-to-date is critical. This approach helps prevent data errors and performance bottlenecks, allowing complex tasks to run more efficiently and reliably across powerful computing systems.
Real-world examples
- 1.AI training servers
- 2.Professional graphics workstations
- 3.High-performance computing clusters
- 4.Data centers with GPU accelerators
Generated by PatentBrief · Not legal advice · patentbrief.org
US 12737317 · 2026