Researchers are developing a new memory technology for GPUs and AI accelerators inspired by storage devices such as solid-state drives [1].

This development addresses a critical bottleneck in artificial intelligence hardware. While current accelerators provide immense speed, they often lack the memory capacity required to handle increasingly massive datasets without slowing down.

Most high-end GPUs currently rely on high bandwidth memory, known as HBM [1]. This technology allows hardware to shuffle data at speeds of multiple terabytes a second [1]. However, the capacity of HBM is generally limited to the gigabyte range [1].

By drawing inspiration from storage-class memory and SSD architecture, the new approach aims to bridge the gap between high-speed volatile memory and high-capacity permanent storage [1]. The goal is to allow GPUs to reach memory capacities in the multiple terabyte range while maintaining the performance levels required for AI workloads [1].

"Virtually every high-end GPU and AI accelerator relies on high bandwidth memory (HBM), which can shuffle data around at multiple terabytes a second but can only reach into the gigabytes," a reporter for The Register said [1].

The technology is currently in development. Once implemented, it could change how AI models are loaded and processed by reducing the need to constantly move data between a slow system drive and the fast GPU memory [1]. This shift would potentially enable the execution of larger models on fewer chips, a change that could lower the cost and energy requirements of AI scaling [1].

GPUs could explode to multiple TB with new storage-inspired memory tech

If successful, this technology would decouple AI performance from the physical limits of HBM. By integrating storage-like capacities directly into the GPU's memory architecture, the industry could move away from the 'memory wall,' where processing power outpaces the ability to feed data to the chip. This would likely accelerate the development of larger, more complex large language models that currently require massive clusters of hardware to function.