Optimizing LLMs For Real World Applications - Lightspeed Venture Partners
We use cookies to enhance your browsing experience, serve personalized ads or content, and analyze our traffic. By clicking "Accept All", you consent to our use of cookies. No cookies to display. 11/28/2023 ENTERPRISE At our recent Generative London event, Gemma Garriga, Technical Director, Office of the CTO, Google Cloud and Meryem Arik, Co-Founder of TitanML spoke with Antoine Moyroud, Partner at Lightspeed, to decode the intricacies of optimizing large language model (LLM) inference. As the buzz around deploying LLMs for real-world applications intensifies, the discussion shared valuable thoughts on quantization, fine-tuning, hardware choices, benchmarking challenges, and the ever-evolving landscape of model architectures. Here are four areas of AI technology the panel covered in their discussion: The panel dived into quantization techniques, exploring methods to slim down models without compromising accuracy. Meryem highlighted activation weight quantization (AWQ) as a standout, ca
Optimizing LLMs For Real World Applications - Lightspeed Venture Partners 11/28/2023 Enterprise Share on twitter Share on facebook Share on linkedin Copy link Optimizing LLMs For Real World Applications Meryem Arik, Gemma Garriga, and Antoine Moyroud At our recent Generative London event, Gemma Garriga , Technical Director, Office of the CTO, Google Cloud and Meryem Arik , Co-Founder of TitanML spoke with Antoine Moyroud, Partner at Lightspeed, to decode the intricacies of optimizing large language model (LLM) inference. As the buzz around deploying LLMs for real-world applications intensifies
Explore this link on the map →related reading
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- GenAI Handbookgenai-handbook.github.io
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Optimizing inference · Hugging Facehuggingface.co
- A Guide to Quantization in LLMs | Symbl.aisymbl.ai
- Things we learned about LLMs in 2024simonwillison.net
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities (Version 1.0)arxiv.org
- Fine-Tuning Llama-2: Tailoring Models to Unique Applicationsanyscale.com
- 2025: The year in LLMssimonwillison.net
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Best practices to accelerate inference for large-scale production workloadstogether.ai