flâneur

Star Attention: Efficient LLM Inference over Long Sequences

arxiv.org · 4,505 words · saved by 1 readers

N/A

Star Attention: Efficient LLM Inference over Long Sequences Shantanu Acharya 1 Fei Jia 1 Boris Ginsburg 1 Abstract these methods improve training efficiency, autoregressive decoding during inference still requires the model to attend Inference with…

related reading