Skip to content

Commit 19e2fb5

Browse files
committed
Enable AI Infra Wiki synchronization
1 parent befd4af commit 19e2fb5

13 files changed

Lines changed: 1796 additions & 12 deletions

README.md

Lines changed: 7 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -52,19 +52,21 @@ node scripts/feishu/register-app.mjs
5252
API Scope 之外,应用还必须拥有目标知识空间的资源权限。进入知识空间的
5353
「设置 → 权限设置/成员设置」把应用加入可查看成员;如果界面不能直接选择
5454
应用,可以先把应用作为机器人加入一个飞书群,再把该群加入知识空间成员。
55-
否则整库遍历会返回 `131006 wiki space permission denied`
55+
否则列举知识空间顶层节点会返回 `131006 wiki space permission denied`
56+
只把单篇页面共享给应用时,应用可能可以读取该页面及其后代,但这不等于拥有
57+
整个知识空间的遍历权限。
5658

5759
### 内容映射
5860

59-
编辑 `content/feishu/sessions.json``wiki` 配置用于递归同步一个完整知识库
60-
目录
61+
编辑 `content/feishu/sessions.json``wiki` 配置用于递归同步一个根页面及其
62+
全部后代
6163

6264
```json
6365
{
6466
"wiki": {
65-
"rootNodeToken": "Bzqww95pBiox3Tkszkncphw6nBh",
67+
"rootNodeToken": "SURAw8B4riKkhRkUKPYc43k7nHf",
6668
"sourceBaseUrl": "https://lcpu-club.feishu.cn/wiki",
67-
"title": "AI Infra 小组"
69+
"title": "AI Infra Wiki"
6870
}
6971
}
7072
```

content/feishu/sessions.json

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,11 @@
77
"end": "2026-10-01"
88
},
99
"publishMeetingUrl": false,
10+
"wiki": {
11+
"rootNodeToken": "SURAw8B4riKkhRkUKPYc43k7nHf",
12+
"sourceBaseUrl": "https://lcpu-club.feishu.cn/wiki",
13+
"title": "AI Infra Wiki"
14+
},
1015
"sessions": [
1116
{
1217
"id": "01",

docs/.vitepress/data/generated/feishu.json

Lines changed: 76 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,5 +10,81 @@
1010
"revisionId": 27
1111
}
1212
}
13+
},
14+
"wiki": {
15+
"title": "AI Infra Wiki",
16+
"rootNodeToken": "SURAw8B4riKkhRkUKPYc43k7nHf",
17+
"sourceUrl": "https://lcpu-club.feishu.cn/wiki/SURAw8B4riKkhRkUKPYc43k7nHf",
18+
"documentId": "CG7cdsC45ooyVTxorPGcI2Y8nld",
19+
"objectType": "docx",
20+
"revisionId": 1119,
21+
"pages": [
22+
{
23+
"wikiNodeToken": "E3ltwZcMEiBJy2kusiCcKlk5nzb",
24+
"parentWikiNodeToken": null,
25+
"title": "Topic 1 - Kernel and Compilers",
26+
"route": "/wiki/E3ltwZcMEiBJy2kusiCcKlk5nzb",
27+
"depth": 0,
28+
"order": 0,
29+
"documentId": "LopWdSBvxoi1yJxq5ZYckSwlnJg",
30+
"objectType": "docx",
31+
"revisionId": 14
32+
},
33+
{
34+
"wikiNodeToken": "VIJPw24sqiPOhXkfEwXcEvVInhg",
35+
"parentWikiNodeToken": "E3ltwZcMEiBJy2kusiCcKlk5nzb",
36+
"title": "Session 1 7.23",
37+
"route": "/wiki/VIJPw24sqiPOhXkfEwXcEvVInhg",
38+
"depth": 1,
39+
"order": 1,
40+
"documentId": "QCpddjfi3oR3LxxRattcszYrnvg",
41+
"objectType": "docx",
42+
"revisionId": 27
43+
},
44+
{
45+
"wikiNodeToken": "I6BFwTffUiAKGUk8TgncV6DIn1b",
46+
"parentWikiNodeToken": "VIJPw24sqiPOhXkfEwXcEvVInhg",
47+
"title": "1.1 & 1.2",
48+
"route": "/wiki/I6BFwTffUiAKGUk8TgncV6DIn1b",
49+
"depth": 2,
50+
"order": 2,
51+
"documentId": "GaJGdUDlRowCZSxpOMScacMSncd",
52+
"objectType": "docx",
53+
"revisionId": 148
54+
},
55+
{
56+
"wikiNodeToken": "QmwywJ9UjiysBakT3vychu0PnJe",
57+
"parentWikiNodeToken": "VIJPw24sqiPOhXkfEwXcEvVInhg",
58+
"title": "推送",
59+
"route": "/wiki/QmwywJ9UjiysBakT3vychu0PnJe",
60+
"depth": 2,
61+
"order": 3,
62+
"documentId": "Vr27dsf11ooFf2xhtYzcvFaZnkw",
63+
"objectType": "docx",
64+
"revisionId": 342
65+
},
66+
{
67+
"wikiNodeToken": "UuLEwoEJRiSaBak8gdnc7UfPn9c",
68+
"parentWikiNodeToken": null,
69+
"title": "Weiming HPC Training Camp x LCPU AI Infra Seminars 暑期活动方案",
70+
"route": "/wiki/UuLEwoEJRiSaBak8gdnc7UfPn9c",
71+
"depth": 0,
72+
"order": 4,
73+
"documentId": "Hjbud6tOJotp6YxipiIcP23PnM5",
74+
"objectType": "docx",
75+
"revisionId": 1125
76+
},
77+
{
78+
"wikiNodeToken": "BtqywAZHNiA38KkEwb0cNvTUnNf",
79+
"parentWikiNodeToken": null,
80+
"title": "Weiming HPC Training Camp x LCPU AI Infra Seminars (筹)",
81+
"route": "/wiki/BtqywAZHNiA38KkEwb0cNvTUnNf",
82+
"depth": 0,
83+
"order": 5,
84+
"documentId": "D02XdMhnkoBo6GxdXzacH16KnHb",
85+
"objectType": "docx",
86+
"revisionId": 51
87+
}
88+
]
1389
}
1490
}

docs/wiki/.gitkeep

Lines changed: 0 additions & 1 deletion
This file was deleted.
Lines changed: 214 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,214 @@
1+
---
2+
title: "Weiming HPC Training Camp x LCPU AI Infra Seminars (筹)"
3+
description: "Weiming HPC Training Camp x LCPU AI Infra Seminars (筹) · AI Infra Wiki"
4+
outline: deep
5+
lastUpdated: false
6+
---
7+
8+
<!-- 此文件由 npm run sync:feishu 生成,请在飞书 Wiki 中编辑正文。 -->
9+
10+
[AI Infra Wiki](/wiki/)
11+
12+
# Weiming HPC Training Camp x LCPU AI Infra Seminars (筹)
13+
14+
[在飞书中查看原文 ↗](https://lcpu-club.feishu.cn/wiki/BtqywAZHNiA38KkEwb0cNvTUnNf)
15+
16+
> 暂时考虑GPU Programming和集群、环境、通信、Linux & Tools等部分也许可以合并 (好像没想到不能合并的部分、似乎即使CPU上的优化也可以在AI Infra的scope里
17+
>
18+
> 暂时考虑的形式是讨论会形式,每周一到两次活动,每次活动一组同学讲同一个Topic相关的一些sessions, sessions可以动态补充,大体要有条理顺序,但可以按照兴趣发散。要求所有报名参加的同学至少讲一次、pkusc的同学至少讲两次;讲完之后大家讨论、提意见。 除了大伙讲的之外, 可以多几次Guest Lectures, (对接赞助?)以及超算队学长们、等等。
19+
>
20+
> 每个session:20-40min,最好有配套的代码和handout、exercises,需要有质量要求,肯定是比较精致比较好,别全部都是vibe出来的就好。
21+
>
22+
> 每次的session发布到bilibili、youtube, handout / execrcises发布到社团公众号、官网、hpcwiki、(也许)社团知乎。允许同学们转载
23+
>
24+
> 考虑找赞助商给Lectures & 课题 & 算力,给大家一些GPU开发节点用,每次配套的小练习也许可以上hpcgame评测,这些内容对社团infra开发压力较大
25+
>
26+
> 考虑给Lecturers发劳务? 300-600 可以收集评价分档次发放?给handout、exercises也发?
27+
28+
**飞书群:** AI Infra学习小组
29+
30+
# Topics & Sessions (暂)
31+
32+
### GPU Programming:怎么写Kernel
33+
34+
- 01 - GPU体系结构与GPU编程模型 (上)
35+
- 02 - GPU体系结构与GPU编程模型(下)
36+
- 03 - CUDA初识:把代码跑到GPU上去
37+
- 04 - CUDA生态:CU{x} CUB与并行算法
38+
- 05 - GPU的特点与更高效的CUDA代码(上):Warp divergence 、 bank conflict 与 I/O coalesce
39+
- 06 - GPU的特点与更高效的CUDA代码(中):Memory hierarchy、TMA 与 Memory Bound
40+
- 07 - GPU的特点与更高效的CUDA代码(下):专用指令、SIMD、Tensor Core 与 Compute Bound
41+
- 08 - 写写Kernel(上):Speed of Light 、Optimization Goal 与 Profiling
42+
- 09 - 写写Kernel(中):Fusing, Tiling and Prefetch
43+
- 10 - 写写Kernel(下):Soft pipeline, Warp Specialization and Persistent Kernel
44+
- 11 - 初识DSL(I):Triton primitives and Kernels in Triton
45+
- 12 - 初识DSL(II):Disadvantage of Triton and TileLang primitives
46+
- 13 - 初识DSL(III):Kernels in TileLang and CuTeDSL Primitives
47+
- 14 - 初始DSL(IV):Kernels in CuTeDSL and Trading off in DSL
48+
- 15 - 硬件侧的演进:Ampere, Hopper and Blackwell
49+
- 16 - Kernel进阶:SoL GEMM的优化之旅
50+
- 17 - Kernel进阶:SoL Attention的优化之旅:MHA & MLA
51+
- 18 - Kernel进阶:SoL MoE的优化之旅
52+
53+
### **ML Compilation & Parallelization**
54+
55+
- 01 - 数据并行 DP:PyTorch DDP & 梯度累加 - 双卡训练初尝试
56+
- 02 - 数据并行 DP:ZeRo Stage1-3 & FSDP - 显存优化
57+
- 03 - 张量并行 TP:MLP列切分/行切分、Attention头切分
58+
- 04 - 流水线并行 PP: GPipe、1F1B
59+
- 05 - 流水线并行 PP:Interleaved-1F1B
60+
- 05 - 流水线并行 PP:DualPipe
61+
- 07 - 序列并行 SP:序列并行与上下文并行原理
62+
- 08 - 序列并行 SP:FlashAttention与RingAttention
63+
- 09 - 专家并行 EP:MoE原理
64+
- 10 - 专家并行 EP:负载均衡、MoE训练&推理特性
65+
- 11 - 专家并行 EP:DeepEP07 - (Advanced ML Parallelization)Automated ML Parallelization:FlexFlow & Alpa
66+
67+
> sifer: 我感觉Parallelization没必要和ML Compilie在一块讲?
68+
>
69+
> richwei:感觉可以拆开,当时一起想到的,就写一起了
70+
71+
### ML Compiler: Automatic! Automatic!
72+
73+
- 从CUDA到DSL:天下人苦CUDA久矣!
74+
- 编译原理速览:DSL → cubin
75+
- … 待zyx补充
76+
77+
### LLM Training Framework:我要训练!
78+
79+
- 01 - 梯度下降、反向传播与模型训练简介
80+
- 02 - 我不止一张卡(上):怎么训得更快 — Data Parallelism
81+
- 03 - 我不止一张卡(中):模型好像有点大 — Model Parallelism(TP, EP, SP, CP, PP)
82+
- 04 - 我不止一张卡(下):
83+
- 05 - (
84+
-
85+
86+
### LLM Inference and Serving Framework:再推快一点!
87+
88+
- 01 - 从KV Cache说起:Paged Attention, Radix Cache & Prefix Cache
89+
- 02 - 推理侧也有并行:推理框架下的TP, EP, SP, CP, PP
90+
- 03 - 专门的人干专门的事:PD分离、KV Cache Transfer、AF分离
91+
- 04 - 公平的推理:Continuous Batching & Chunked Prefill
92+
- 05 - 想占点便宜:Speculative Decoding
93+
- 06 -
94+
95+
### RL Framework:既要推理又要训练怎么办
96+
97+
- 01 -
98+
99+
字节开源框架:[VeRL](https://github.com/verl-project/verl)
100+
101+
阿里开源框架:[ROLL](https://github.com/alibaba/ROLL)
102+
103+
### 藏在背后的一环:集群通信
104+
105+
- 从TCP/IP到RDMA
106+
- IB、ROCE、iWarp简介
107+
- RDMA详解(I)
108+
- RDMA详解(II)
109+
- RDMA详解(III)
110+
- RDMA详解(IV)
111+
- 从RDMA到NCCL、UCCL:通信库干了什么
112+
- torch.distributed与NCCL Primitives
113+
- xxx
114+
- xxx
115+
- 高性能带通信Kernels:我要自己写通信!
116+
-
117+
118+
### Generative AI(LLM 算法基础)
119+
120+
Generative models of text:RNN LMs / Autodiff & Transformer LMs & Learning LLMs / Decoding & Pre-training, fine-tuning / Modern Transformers
121+
122+
Generative models of images:CNN/Encoder-only Transformers/ViT & GAN/PGM & Diffusion models
123+
124+
Applying and adapting foundation models:VAE & In-Context Learning / Prompt Engineering / Instruction Fine-tuning / Reinforcement learning with human feedback (RLHF)
125+
126+
Multimodal foundation models:Direct Preference Optimization (DPO) / Text-to-image generation / Latent diffusion model & VLM & Cross-Attention / Diffusion Transformer / Prompt-to-Prompt
127+
128+
Scaling Up:Querying Transformer / Scaling Laws & Mixture of Experts & Distributed training & Flash Attention / Efficient decoding strategies
129+
130+
Advanced Topics:Long Context in LLM & Reasoning Models & State Space Models / Hybrid Models & Code Generation / Autonomous Agents & Audio understanding and synthesis & Generative Models for Videos & Interactive World Models + Science of Alignment
131+
132+
# 时间线(暂)
133+
134+
频率:每周1-2次
135+
136+
## 重要时间节点
137+
138+
**北大暑假**
139+
140+
校本部:6月29日起
141+
142+
医学部:7月6日起
143+
144+
**北大秋季学期开学**
145+
146+
校本部:9 月 7 日
147+
148+
医学部:
149+
150+
- 在校本科生:8 月 31 日
151+
- 本科新生、研究生:9 月 7 日
152+
153+
# TODO List
154+
155+
## Timeline确定
156+
157+
频率:每周1-2次
158+
159+
07.08-07.15 前期筹备
160+
161+
宣发预热(07.15)
162+
163+
活动正式开始(07.17/07.20)(?)
164+
165+
正式活动周期:07.17-09.07
166+
167+
每次活动一个topic的一个子主题?暑假以2-3个模块为目标(?)
168+
169+
## 组织形式
170+
171+
线上:腾讯会议
172+
173+
线下:提前预约活动教室(maybe社团活动教室?)
174+
175+
## 课程设计
176+
177+
待完善的topics:
178+
179+
ML Compiler: Automatic! Automatic 部分
180+
181+
LLM Training Framework:我要训练!部分
182+
183+
LLM Inference and Serving Framework:再推快一点!部分
184+
185+
RL Framework:既要推理又要训练怎么办
186+
187+
藏在背后的一环:集群通信 部分
188+
189+
## 赞助
190+
191+
**赞助形式?**
192+
193+
暂时考虑的赞助形式是:
194+
195+
- 给卡,容器 / VM / Baremetal
196+
- 给钱
197+
198+
权益:
199+
200+
- Guest Lectures
201+
- 联系方式
202+
- 内推 / hiring
203+
204+
### (待补充)
205+
206+
## 参考资料
207+
208+
Generative AI:[https://www.cs.cmu.edu/~mgormley/courses/10423/schedule.html](https://www.cs.cmu.edu/~mgormley/courses/10423/schedule.html)
209+
210+
(15-779)Advanced Topics in Machine Learning Systems:[https://www.cs.cmu.edu/~zhihaoj2/15-779/schedule.html](https://www.cs.cmu.edu/~zhihaoj2/15-779/schedule.html)
211+
212+
## 关联会议:
213+
214+
<table><colgroup><col /><col /><col /></colgroup><tbody><tr><td>日期</td><td>会议主题</td><td>记录</td></tr><tr><td>2026.07.10</td><td>Weiming HPC Training Camp x LCPU AI Infra Seminars 筹备启动会</td><td>《Weiming HPC Training Camp x LCPU AI Infra Seminars (筹)——0710会议议程》</td></tr></tbody></table>
Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,16 @@
1+
---
2+
title: "Topic 1 - Kernel and Compilers"
3+
description: "Topic 1 - Kernel and Compilers · AI Infra Wiki"
4+
outline: deep
5+
lastUpdated: false
6+
---
7+
8+
<!-- 此文件由 npm run sync:feishu 生成,请在飞书 Wiki 中编辑正文。 -->
9+
10+
[AI Infra Wiki](/wiki/)
11+
12+
# Topic 1 - Kernel and Compilers
13+
14+
[在飞书中查看原文 ↗](https://lcpu-club.feishu.cn/wiki/E3ltwZcMEiBJy2kusiCcKlk5nzb)
15+
16+

0 commit comments

Comments
 (0)