llm-algo-leetcode 推理优化笔记

299 字
1 分钟
llm-algo-leetcode 推理优化笔记

Datawhale
Datawhale

大家好,我是芯缘,是 Datawhale 社区发起的 2026 年 8 月“llm-algo-leetcode 推理优化方向”组队学习活动的运营助教。本页汇总了我在本次活动中的笔记。

目录#

任务学习内容学习笔记
Task 0:推理结构地基Task 0: Attention MHA GQA- 推理结构地基
Task 1:GPU 硬件与 KV CacheTask 1-1: GPU Architecture and Memory
Task 1-2: KV Cache and Memory Growth
- GPU 架构与内存
- KV Cache 与优化技术
Task 2:Attention 访存瓶颈与 FlashAttentionTask 2-1: FlashAttention Sim
Task 2-2: FlashAttention Memory Model
- FlashAttention 1~4 全解汇总
Task 3:解码算法Task 3-1: Decoding Strategies
Task 3-2: Speculative Decoding
Task 3-3: Multi-Token Decoding
Task 3-4: Decode Scheduling
- 从 Logits 到 Token:解码采样策略
- Speculative Decoding
- 调度策略
Task 4:KV Cache 与推理服务内存管理Task 4-1: KV Cache and Memory Growth
Task 4-2: vLLM PagedAttention
Task 4-3: SGLang RadixAttention
Task 4-4: Prefix Caching and Chunked Prefill
Task 4-5: KV Cache Scheduling
- KV Cache 与推理服务内存管理
- 调度策略
Task 5:量化推理与部署Task 5-1: Quantization Theory and INT4 INT8
Task 5-2: Quantization W8A16
Task 5-3: GPTQ and AWQ Weight Quantization
Task 5-4: FP8 and KV Cache Quantization
- 模型量化策略
Task 6:综合项目Task 6-1: Inference Performance Comparison
Task 6-2: Quantized Inference and Deployment
Task 6-3: Speculative Decoding Benchmark
Task 6-4: Prefix Caching Benchmark
- 推理优化 Benchmark

2026/8/30 分享会内容合集#

文章分享

如果这篇文章对你有帮助,欢迎分享给更多人!

llm-algo-leetcode 推理优化笔记
https://pinghaoyang.com.cn/posts/content/
作者
平昊阳
发布于
2026-08-19
许可协议
CC BY-NC-SA 4.0

评论区

Profile Image of the Author
平昊阳
乘长风,破巨浪, 展鸿图于未央!
--
总访问量
--
访客数
公告
欢迎来到我的个人博客!欢迎关注交流吖!
更多相关公告,见
社交-留言」。
音乐
封面

音乐

暂未播放

0:000:00
暂无歌词
站点统计
文章
118
分类
20
标签
164
总字数
1,050,769
运行时长
0
最后活动
0 天前

文章目录