flâneur — a map of the web's best reading

Optimizing LLMs For Real World Applications - Lightspeed Venture Partners

lsvp.com · 913 words · saved by 1 readers

We use cookies to enhance your browsing experience, serve personalized ads or content, and analyze our traffic. By clicking "Accept All", you consent to our use of cookies. No cookies to display. 11/28/2023 ENTERPRISE At our recent Generative London event, Gemma Garriga, Technical Director, Office of the CTO, Google Cloud and Meryem Arik, Co-Founder of TitanML spoke with Antoine Moyroud, Partner at Lightspeed, to decode the intricacies of optimizing large language model (LLM) inference. As the buzz around deploying LLMs for real-world applications intensifies, the discussion shared valuable thoughts on quantization, fine-tuning, hardware choices, benchmarking challenges, and the ever-evolving landscape of model architectures. Here are four areas of AI technology the panel covered in their discussion: The panel dived into quantization techniques, exploring methods to slim down models without compromising accuracy. Meryem highlighted activation weight quantization (AWQ) as a standout, ca

Optimizing LLMs For Real World Applications - Lightspeed Venture Partners 11/28/2023 Enterprise Share on twitter Share on facebook Share on linkedin Copy link Optimizing LLMs For Real World Applications Meryem Arik, Gemma Garriga, and Antoine Moyroud At our recent Generative London event, Gemma Garriga , Technical Director, Office of the CTO, Google Cloud and Meryem Arik , Co-Founder of TitanML spoke with Antoine Moyroud, Partner at Lightspeed, to decode the intricacies of optimizing large language model (LLM) inference. As the buzz around deploying LLMs for real-world applications intensifies

Explore this link on the map →

related reading