Near-memory Dequantization Architecture In Custom HBM for LLM inference (SK hynix)


Researchers from SK hynix published a technical paper titled “StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration.” The paper proposes StreamDQ for “a lightweight architectural enhancement that enables on-the-fly dequantization in the memory subsystem for high-throughput, large-batch LLM inference,” and reports “up to 7.08× speedup and ... » read more

The Evolution of HBM


High-bandwidth memory originally was conceived as a way to increase capacity in memory attached to a 2.5D package. It has since become a staple for all high-performance computing, in some cases replacing SRAM for L3 cache. Archana Cheruliyil, senior product marketing manager at Alphawave Semi, talks about how and where HBM is used today, how it will be used in the future, why it is essential fo... » read more