✳flâneur — a map of the web's best reading
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- Accelerating Diffusion LLMs via Adaptive Parallel Decodingarxiv.org
- LLM Visualizationbbycroft.net
- CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Creditsarxiv.org
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- GitHub - karpathy/nanochat: The best ChatGPT that $100 can buy. · GitHubgithub.com
- GitHub - linkedin/Liger-Kernel: Efficient Triton Kernels for LLM Training · GitHubgithub.com
- Llama 2 · Hugging Facehuggingface.co
- LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelismarxiv.org
- GitHub - affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. · GitHubgithub.com
- AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Accelerationarxiv.org