Front Lines of Technology Development in the AI Era

From Inference to the Physical AI That Lies Ahead
KIOXIA Faces the Challenge of Increased Demand for Processing Capacity Driven by Rapidly Evolving AI

September 15, 2026

The use of generative AI has become common in many workplaces. With the recent emergence of interactive forms of generative AI, as well as AI agents that are capable of handling work, an increasing number of companies are considering the adoption of such technologies. As the application of generative AI expands, the computational resources used for processing will also naturally increase, and GPU performance is increasing to address those needs. However, advanced GPU performance is not the only requirement. The memory and storage which pass data to the GPUs also require increased performance and expanded capacity. What types of initiatives are currently underway to address such requirements? To learn more, we spoke with an industry expert at Kioxia Corporation (Kioxia), one of the world's largest specialty manufacturers of flash memory and SSDs.

Confronting the limits of memory due to the rapid increase in "inference"

Generative AI has broadly penetrated both our personal and professional lives and businesses. And now the way we use AI is undergoing a significant transformation. Until recently, AI was typically used by "asking a question and receiving an answer." However, the field is rapidly evolving toward AI agents that "let AI do the work."

According to Makoto Hamada of Kioxia, "This year can truly be called year one of the AI agent." As AI adoption advances, it will undoubtedly embrace agentic AI, where AI agents use other AI agents to complete tasks.

Makoto Hamada, General Manager, SSD Strategic Product Marketing Managing Department, SSD Division, Kioxia

Hamada explains, "If this type of AI usage spreads, it will eliminate the time required for manual entry, and solve problems much faster than ever before."

Should this trend advance, the requirements placed on systems that run generative AI will significantly change. To date, massive resources have been devoted to "training" to create generative AI models. Going forward, resources for conducting "inference" at sites where generative AI is used will become critical.

"Once the use of agentic AI begins, AI agents will converse with each other without waiting for manual input from humans, and the number of systems operated autonomously by AI will increase. The volume of data this generates will grow exponentially, as will the volume of data to be processed, which will likely require even higher levels of performance from GPUs," said Hamada.

On the horizon is the likely expansion of AI applications beyond agentic AI to emerging world of physical AI. When the era of physical AI arrives, the demands on processing speeds will become even more acute. The reason is that in order for AI to operate in physical space, it must import data from three-dimensional space and time and feed the inference results, which incorporate those details, back into the operational scope in real time.

In anticipation of such changes, the processing capabilities of GPUs are already being rapidly increased. But Hamada points out that the limitation of the memory that supplies data to the GPU is a significant issue that also must be addressed.

"To provide data to the GPU, a type of DRAM (main memory) called 'High Bandwidth Memory (HBM)' is currently installed next to the GPU. It has a three-dimensional, laminated structure, and the limits of HBM have started to become apparent," (said Hamada.)

The solution is direct access from the GPU to the SSD

What limits the expansion of HBM? The limitation lies in the difficulty of increasing the capacity. In addition to the problem of increasing the number of lamination layers, the cost increase due to miniaturization is substantial, and potentially expensive. In other words, there are limits to higher capacity from a purely technical standpoint, but even if those limits could be exceeded, implementation would be economically challenging as we move to the stage of social implementation widespread adoption of AI continues.

Hamada adds, "If HBM hits capacity limits, the 'brute force' method of simply increasing GPU performance will no longer work." So how can this problem be solved? The answer lies in the approach of using SSDs rather than more DRAM.

Hamada explains, "NVIDIA is redefining the role of SSDs in AI-driven workloads. The goal is to use NVIDIA GPUDirect Storage to enable a direct data path between storage and GPU memory. In short, SSDs are no longer viewed merely as data storage, but as a high-performance source of data for GPUs."

To do so, major breakthroughs are needed in SSDs. In addition to significantly reducing data access latency, I/O performance must be dramatically improved. To meet this requirement, Kioxia is currently developing an SSD product line called the KIOXIA GP Series (photo).

Mockup of the KIOXIA GP Series
The company plans to begin providing evaluation samples of the KIOXIA GP Series to select customers by the end of 2026 (product images may differ from the actual product).

Hamada explains, "We already have XL-FLASH™ high-speed storage class memory, which is based on our BiCS FLASH™ 3D NAND technology. The KIOXIA GP Series incorporates XL-FLASH™ with a focus on dramatically increase random read performance. Moreover, it can reduce power consumption per I/O operation compared to conventional SSDs."

So why does the KIOXIA GP Series SSD focus on random read performance? The reason lies in the anticipated use case.

"The KIOXIA GP Series is specifically designed for use in vector database search. To perform inference using proprietary company information or real-time information, a method called RAG (Retrieval-Augmented Generation) is used to feed external information to generative AI without retraining. This is where vector databases are used. When performing that search, it is necessary to read finely dispersed data rather than a large set of data all at once. For that reason, the random read performance becomes important, says Hamada.

Kioxia initiatives advancing in four directions

The manufacturing cost per unit of capacity for SSDs is typically lower than that of DRAM, making it more cost-effective to expand SSD capacity. If data can move more directly between SSD storage and GPU memory, systems can help reduce storage I/O bottlenecks and support larger inference workloads. Furthermore, the KIOXIA GP Series SSD provides extremely high performance per unit of power, which contributes to the reduction of data center power consumption.

In addition to the KIOXIA GP Series, Kioxia is deploying multifaceted SSD solutions (see Figure).

Four directions of Kioxia SSD initiatives
In addition to the newly announced KIOXIA GP Series, many products including the KIOXIA CM Series and the KIOXIA LC Series are specifically designed to be used for AI.

Hamada says, "There are four major directions for our SSD initiatives." The first is the ultra-high-performance KIOXIA GP Series introduced above. The second direction is the KIOXIA CM Series, which balances high performance with high capacity.

"This series supports KV cache, which is equally important as RAG to AI inference. The KV cache is a location that temporarily stores previously generated inference results to enable the AI to answer without recalculation when the same question is asked. Since the data is read in a reasonably organized format from the KV cache, sequential read performance becomes more important than random read performance, and high reliability is also required due to the large number of rewrites, says Hamada.

The third direction is the KIOXIA LC Series, which focuses on high density and high capacity. Intended to be used for AI training and as a data pool for inference results, this series features extremely high-capacity QLC (high-density flash memory capable of storing 4 bits of data per cell). Finally, the fourth direction is Nearline SSDs aimed at TCO optimization by replacing HDDs.

In addition to these solutions, Kioxia is developing software to utilize SSDs more effectively. One result of that effort is KIOXIA AiSAQ™ vector database search software for generative AI. KIOXIA AiSAQ technology deploys the vector database, including the data and indexes for accelerating searches, onto the SSD, which enables search speeds comparable to previous methods that deploy data into memory. This is being provided as open-source software and has been integarted into the Milvus open-source vector database. It is expected to become an important solution for addressing future increases in RAG data volume.

Achieving storage innovation with the evolution of AI

Hamada explains, "While we are pursuing multiple initiatives to expand the utilization of AI, these development efforts are by no means easy." The world of generative AI continues to change at a bewildering pace, and new challenges to be solved are emerging one after another, which makes it difficult to finalize specifications. Sharing his outlook for the future, Hamada states, "For that reason, we will continue development while engaging in ongoing dialogue with NVIDIA and our users. To continue improving AI, we plan to stay at the forefront of storage innovation."

  • NVIDIA and GPUDirect are trademarks and/or registered trademarks of NVIDIA Corporation in the U.S. and other countries.
  • Other company names, product names, and service names may be trademarks of third-party companies.

Reprinted from "Nikkei xTECH" (Published June 9, 2026) with permission from Nikkei BP.
Department names and titles are as of the time of the interview.