Deeploy: Enabling Energy-Efficient Deployment of Small Language Models on Heterogeneous Microcontrollers | IEEE Journals & Magazine | IEEE Xplore
The latest evolutions in mainstream artificial intelligence (AI) have been driven by Transformers, which have taken over from recurrent neural networks (RNNs) and convolutional neural networks (CNNs) as the leading edge models for language processing and multimodal applications [1], [2]. The success of Transformers can be primarily attributed to the emergence of the foundation model (FM) paradigm: large Transformer models extensively pretrained on the datasets spanning trillions of tokens and then fine tuned with a much lower volume of labeled data to solve the domain-specific problems. Following the success of FMs in natural language processing (NLP) [1], [3], an increasing number of fields are starting to formulate and adapt FMs for high dimensional sensor data that has traditionally been challenging to process, like decoding the neural data [4], [5], or training embodied AI agents [6], [7], which may incorporate the multimodal sensor inputs. Operating directly on the sensory data an
Download PDF Download References Request Permissions Save to Alerts Abstract: With the rise of embodied foundation models (EFMs), most notably small language models (SLMs), adapting Transformers for the edge applications has become a very active fi...Show More Metadata Abstract: With the rise of embodied foundation models (EFMs), most notably small language models (SLMs), adapting Transformers for the edge applications has become a very active field of research. However, achieving the end-to-end deployment of SLMs on the microcontroller (MCU)-class chips without high-bandwidth…
saved by
related reading
- How To Scale Your Modeljax-ml.github.io
- FTRANS: Energy-Efficient Acceleration of Transformers using FPGAarxiv.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- [2401.02721] A Cost-Efficient FPGA Implementation of Tiny Transformer Model using Neural ODEarxiv.org
- On the Tradeoffs of SSMs and Transformers | Goomba Labgoombalab.github.io
- BETA: Binarized Energy-Efficient Transformer Accelerator at the Edgearxiv.org
- Can LLMs Be Computers?percepta.ai
- Full Stack Optimization of Transformer Inference: a Surveyarxiv.org
- AI Revolution - Transformers and Large Language Models (LLMs)blog.eladgil.com
- Large Language Models Reading List | Sebastian Raschka, PhDsebastianraschka.com
- Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server | NVIDIA Technical Blogdeveloper.nvidia.com
- Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformerieeexplore.ieee.org