flâneur — a map of the web's best reading

ATT: A Fault-Tolerant ReRAM Accelerator for Attention-based Neural Networks | IEEE Conference Publication | IEEE Xplore

ieeexplore.ieee.org · saved by 1 readers

Processing-in-Memory (PIM) platforms are more and more popular in accelerating neural network applications due to fewer data movements compared to FPGA and ASIC implementations [1]–​[3]. Essentially, computations in both convolutional layers and fully-connected layers can be transformed to matrix-matrix and matrix-vector multiplications. These linear algebra operations can be mapped to crossbars to achieve excellent performance [4], [5]. Weight pruning on PIM is seldom studied because mapping irregular computation patterns to crossbars is challenging. Attention-based Neural Networks (AttNNs) have been proven to significantly outperform convolutional neural networks (CNNs) and recurrent neural networks (RNNs) in the wide variety of Natural Language Processing (NLP) tasks [6]–​[9]. In general, it is impractical to deploy a deep learning network to an accelerator designed for a different network. Specifically for AttNNs, the gaussian error linear unit (gelu) activation function [10] is un

Explore this link on the map →

saved by