✳flâneur — a map of the web's best reading
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
arxiv.org · saved by 1 readers
N/A
Explore this link on the map →related reading
- GitHub - brexhq/prompt-engineering: Tips and tricks for working with Large Language Models like OpenAI's GPT-4. · GitHubgithub.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- GitHub - openai/parameter-golf: Train the smallest LM you can that fits in 16MB. Best model wins! · GitHubgithub.com
- Branches · HazyResearch/intelligence-per-watt · GitHubgithub.com
- Jacobian Lens – Qwen3.6-27B | Neuronpedianeuronpedia.org
- Neuronpedianeuronpedia.org
- GitHub - linkedin/Liger-Kernel: Efficient Triton Kernels for LLM Training · GitHubgithub.com
- MatX: High-throughput chips for LLMsmatx.com
- GitHub - affaan-m/ECC: The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond. · GitHubgithub.com
- GitHub - x1xhlol/system-prompts-and-models-of-ai-tools: FULL Augment Code, Claude Code, Cluely, CodeBuddy, Comet, Cursor, Devin AI, Junie, Kiro, Leap.new, Lovable, Manus, NotionAI, Orchids.app, Perplexity, Poke, Qoder, Replit, Same.dev, Tragithub.com
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decodingarxiv.org
- GitHub - PaulPauls/llama3_interpretability_sae: A complete end-to-end pipeline for LLM interpretability with sparse autoencoders (SAEs) using Llama 3.2, written in pure PyTorch and fully reproducible. · GitHubgithub.com