✳flâneur — a map of the web's best reading
Llama2
huggingface.co · 5,421 words · saved by 1 readers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Llama 2 · Hugging Face Transformers documentation Llama 2 Transformers 🏡 View all docs AWS Trainium & Inferentia Accelerate Argilla AutoTrain Bitsandbytes CLI Chat UI Dataset viewer Datasets Deploying on AWS Diffusers Distilabel Evaluate Google Cloud Google TPUs Gradio Hub Hub Python Library Huggingface.js Inference Endpoints (dedicated) Inference Providers Kernels LeRobot Leaderboards Lighteval Microsoft Azure OpenEnv Optimum PEFT Reachy Mini Safetensors Sentence Transformers TRL Tasks Text Embeddings Inference Text Generation Inference Tokenizers Trackio Transformers Transformers.js Xet smo
Explore this link on the map →saved by
related reading
- Transformer Explainer: LLM Transformer Model Visually Explainedpoloclub.github.io
- API Reference — TensorRT LLMnvidia.github.io
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- I blame the tokenizer | David Quareldavidquarel.github.io
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Neuronpedianeuronpedia.org
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decodingarxiv.org
- How LLMs Actually Work | 0xkato0xkato.xyz
- Introducing Llama2-70B-Chat with MosaicML Inference | Databricks Blogmosaicml.com