Study Group on

ML Systems

in Production

Ak,Cerulean Labs
01

Pretraining

Learn from data.

+

Post-training

Refine behavior.

02

Post-training

A crowded frontier.

Frontier
Pretraining

OpenAIOpenAI
Google DeepMindGoogle DeepMind
AnthropicAnthropic
MetaMeta
DeepSeekDeepSeek
Alibaba / QwenAlibaba / Qwen
KimiKimi
GrokGrok
Midjourney
TII
StepFun
AWS
Essential AI
IBM
Arcee AI
Poolside
Upstage
Hugging Face
Replit
ElevenLabs
01.AI
SenseNova
Cohere
Tencent
Inception
NVIDIA
AssemblyAI
Voyage AI
Luma
iFLYTEK
Google
AI21 Labs
Magic
Anyscale
Inflection
Mistral
Baidu
Jina AI
ByteDance
Together AI
SambaNova
Cursor
Deep Cogito
Nous Research
Zhipu AI
Huawei
Snowflake
Moonshot AI
Fireworks AI
Adobe
Ideogram
BAAI
MiniMax
Perplexity
Liquid AI
Pika
Microsoft
xAI
Ai2
Runway
Stability AI
Apple
Cerebras
Aleph Alpha
Groq
Baichuan
03

Frontier

Pretraining

OpenAI
Google DeepMind
Anthropic
Meta
DeepSeek
Alibaba / Qwen
Kimi
Grok
04

DeepSeek-V3 pretraining

US$5.328M

Estimated compute cost

05

We overthrow

pretraining.

06

More room to explore.

Architecture choices.

Learning efficiency.

07

Beyond CUDA.

More hardware.

More people doing science.

08

Open the whole process.

Code.

Experiments.

Results.

09

The study group

A place to become
a researcher.

10

Built around three things.

01

Hands-on work

Labs-centered

Automated scoring

02

People to think with

Deep-dives into papers

Office hours to chat with

03

Compute

11

16,000

H200 GPU-hours

Secured for the project.

12

People to think with.

Team image 1
Team image 2
Team image 3
Team image 4
Team image 5
Team image 6
Team image 7
Team image 8
13

3

Papers

First Author, Paper in AAAI 2024 Student Abstracts

First Author, Paper in IJCAI 2023 Symposium on LLMs

Authors, Paper in TAICHI 2026

TPC Member, FLLM 2023 - 2025

Dept. of CS Course Lecturer

Honorable Mention, HiPAC 2025

Participant, ISC 2026

Honorable Mention, HiPAC 2026

Creator of Superfile

Founding Product Manager, Clustron

Senior Backend Engineer, Clustron

14

3

HPC Competitions

First Author, Paper in AAAI 2024 Student Abstracts

First Author, Paper in IJCAI 2023 Symposium on LLMs

Authors, Paper in TAICHI 2026

TPC Member, FLLM 2023 - 2025

Dept. of CS Course Lecturer

Honorable Mention, HiPAC 2025

Participant, ISC 2026

Honorable Mention, HiPAC 2026

Creator of Superfile

Founding Product Manager, Clustron

Senior Backend Engineer, Clustron

15

23k+

GitHub Stars

First Author, Paper in AAAI 2024 Student Abstracts

First Author, Paper in IJCAI 2023 Symposium on LLMs

Authors, Paper in TAICHI 2026

TPC Member, FLLM 2023 - 2025

Dept. of CS Course Lecturer

Honorable Mention, HiPAC 2025

Participant, ISC 2026

Honorable Mention, HiPAC 2026

Creator of Superfile

Founding Product Manager, Clustron

Senior Backend Engineer, Clustron

16

The learning arc

01

Foundations

GPT-2 architecture and training

Optimizers and experiment design

02

Research sprints

Embeddings, data, and architecture

Build a deep learning framework

03

Your research

One week to plan

Three weeks to investigate

17

6–9

credits of workload

Every Tuesday

20:00 – 22:30

In person

ML and software engineering experience expected.

18

Join us.

Applications close

September 25, 23:59

Results by September 28

Apply through the SDC registration form.

19