flâneur — a map of the web's best reading

(6 封私信 / 3 条消息) Speculative Speculative Decoding (SSD) - 知乎

zhuanlan.zhihu.com · saved by 1 readers

Tri Dao实验室在发布FA4之余,发布了一篇关于投机采样的文章,展示了一种LLMSYS级别的投机采样优化方案。利用一块额外的GPU,掩盖了draft model的推理,并通过fan-out cache预测target model的接受结果。文章还提到了和现在热门的eagle方案的兼容,取得了不错的效果。或许未来DT分离(Draft/Target分离,not David Tao)会成为新的LLM推理范式。下面是对这篇文章的简单总结(almost generated by gemini)。 附文章链接:https://arxiv.org/pdf/2603.03251 github地址(不要用cu130,会有兼容性错误):https://github.com/tanishqkumar/ssd?tab=readme-ov-file ARM SoC的启动过程: RomBoot --> SPL --> u-boot --> Linux kernel --> file system --> start application (RomBoot是固化在SoC内部的。) spl的产生: 因为芯片厂商固化的… https://jcnr8jfbo9wd.feishu.cn/docx/QMMfd6jDsosKAgxJxO3czxOznbf?from=from_copylink更好的阅读体验请查看飞书。 设计哲学要是用 tensor core 那么最基础的方式就是 sass 指令 HMMA,可…

Explore this link on the map →

saved by