[2311.06668] In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
Abstract:Large language models (LLMs) demonstrate emergent in-context learning capabilities, where they adapt to new tasks based on example demonstrations. However, in-context learning has seen limited effectiveness in many settings, is difficult to quantitatively control and takes up context window space. To overcome these limitations, we propose an alternative approach that recasts in-context learning as in-context vectors (ICV). Using ICV has two steps. We first use a forward pass on demonstration examples to create the in-context vector from the latent embedding of the LLM. This vector captures essential information about the intended task. On a new query, instead of adding demonstrations to the prompt, we shift the latent states of the LLM using the ICV. The ICV approach has several benefits: 1) it enables the LLM to more effectively follow the demonstration examples; 2) it's easy to control by adjusting the magnitude of the ICV; 3) it reduces the length of the prompt by removing the in-context demonstrations; 4) ICV is computationally much more efficient than fine-tuning. We demonstrate that ICV achieves better performance compared to standard in-context learning and fine-tuning on diverse tasks including safety, style transfer, role-playing and formatting. Moreover, we show that we can flexibly teach LLM to simultaneously follow different types of instructions by simple vector arithmetics on the corresponding ICVs.
[2311.06668] In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Machine Learning arXiv:2311.06668 (cs) [Submitted on 11 Nov 2023 ( v1 ), last revised 13 Feb 2024 (this version, v3)] Title: In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering Authors: Sheng Liu , Haotian Ye , Lei Xing , James Zou View a PDF of the paper tit
Explore this link on the map →related reading
- How does in-context learning work? A framework for understanding the differences from traditional supervised learning | SAIL Blogai.stanford.edu
- In-context Learning and Induction Headstransformer-circuits.pub
- [2301.00234] A Survey on In-context Learningarxiv.org
- Are Emergent Abilities in Large Language Models just In-Context Learning?arxiv.org
- [2506.06266] Cartridges: Lightweight and general-purpose long context representations via self-studyarxiv.org
- [2309.01809] Are Emergent Abilities in Large Language Models just In-Context Learning?arxiv.org
- Why We Need Continual Learning | Andreessen Horowitza16z.com
- GenAI Handbookgenai-handbook.github.io
- Continual Learning in Token Space | Lettaletta.com
- Effective context engineering for AI agents \ Anthropicanthropic.com
- Recursive Language Models | Alex L. Zhangalexzhang13.github.io
- Graph Enabled Llama Index - siwei.iosiwei.io