ReTransformer: ReRAM-based Processing-in-Memory Architecture for Transformer Acceleration | IEEE Conference Publication | IEEE Xplore
Natural Language Processing (NLP) is an important sector of Artificial Intelligence (AI) that enables computers to process human languages. Today, the resounding success of deep learning has advanced NLP research by introducing different architectures of deep neural network (DNN) models, such as LSTM [13], RNN [21], GRU [5]. Among the DNN models for sequence-based NLP tasks, attention mechanism is a popular and powerful tool for its ability to handle various lengths and focus on the most relevant parts in the sequence. Very recently, an outstanding self-attention based model- Transformer [28] significantly reduce the path length between long-range dependencies and thus achieve state-of-the-art performance in many transduction tasks. Nevertheless, Transformer models that have already optimized for simplicity still require considerable computational resources owing to the complicated nature of the NLP sequences. For instance, the original Transformer model proposed in [28], which is the
Explore this link on the map →