✳flâneur — a map of the web's best reading
Pavel
0 followers · 141 views
Open this reading profile →
on the atlas — 31
- [2310.08586] PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm1 savers
- google maps - Google Search1 savers
- ReTransformer: ReRAM-based Processing-in-Memory Architecture for Transformer Acceleration | IEEE Conference Publication | IEEE Xplore1 savers
- ATT: A Fault-Tolerant ReRAM Accelerator for Attention-based Neural Networks | IEEE Conference Publication | IEEE Xplore1 savers
- Energon: Toward Efficient Acceleration of Transformers Using Dynamic Sparse Attention | IEEE Journals & Magazine | IEEE Xplore1 savers
- Sanger: A Co-Design Framework for Enabling Sparse Attention using Reconfigurable Architecture | MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture1 savers
- [2012.09852] SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning1 savers
- ELSA: Hardware-Software Co-design for Efficient, Lightweight Self-Attention Mechanism in Neural Networks | IEEE Conference Publication | IEEE Xplore1 savers
- [2002.10941] A$^3$: Accelerating Attention Mechanisms in Neural Networks with Approximation1 savers
- Accelerating Transformer Networks through Recomposing Softmax Layers | IEEE Conference Publication | IEEE Xplore1 savers
- [2209.10797] DFX: A Low-latency Multi-FPGA Appliance for Accelerating Transformer-based Text Generation1 savers
- Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformer | IEEE Conference Publication | IEEE Xplore1 savers
- [2007.08563] FTRANS: Energy-Efficient Acceleration of Transformers using FPGA1 savers
- MnnFast: A Fast and Scalable System Architecture for Memory-Augmented Neural Networks | IEEE Conference Publication | IEEE Xplore1 savers
- ls - Google Search1 savers
- TensorRT-LLM/examples/gemma/README.md at main · NVIDIA/TensorRT-LLM1 savers
- [2312.15159] Understanding the Potential of FPGA-Based Spatial Acceleration for Large Language Model Inference1 savers
- [2310.09949] Chameleon: a heterogeneous and disaggregated accelerator system for retrieval-augmented language models1 savers
- 4. Nsight Compute CLI — NsightCompute 12.4 documentation1 savers
- [2403.00579] NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing1 savers
- [2403.06664] Smart-Infinity: Fast Large Language Model Training using Near-Storage Processing on a Real System1 savers
- bloomberg - Google Search1 savers
- pica task llm - Google Search1 savers
- pip how to check version of package - Google Search1 savers
- [2304.07493] OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair Quantization1 savers
- Quantization and Hardware Architecture Co-Design for Matrix-Vector Multiplications of Large Language Models | IEEE Journals & Magazine | IEEE Xplore1 savers
- [2201.06618] VAQF: Fully Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer1 savers
- [2011.14203] EdgeBERT: Sentence-Level Energy Optimizations for Latency-Aware Multi-Task NLP Inference1 savers
- A Fast and Flexible FPGA-based Accelerator for Natural Language Processing Neural Networks | ACM Transactions on Architecture and Code Optimization1 savers
- [2401.11459] AttentionLego: An Open-Source Building Block For Spatially-Scalable Large Language Model Accelerator With Processing-In-Memory Technology1 savers
- [2401.11851] BETA: Binarized Energy-Efficient Transformer Accelerator at the Edge1 savers