Build & Train a GLM-5.3-Flash Model From Scratch with Python

From the creator

Learn how to design, build, and train an advanced multimodal language model from scratch using modern techniques like mixture-of-experts, sparse attention, and reinforcement learning. This hands-on guide walks you through the entire lifecycle from tokenization and pre-training to post-training optimization and evaluation. Video from @vukrosic GitHub: https://github.com/vukrosic/glm-5.3-flash-from-scratch Become AI researcher in 90 days: https://www.skool.com/become-ai-researcher-2669/about Slides: https://github.com/vukrosic/glm-5.3-flash-from-scratch/blob/main/slides/slides.html Blog: https://z.ai/blog/glm-5.3-flash ❤️ Support for this channel comes from our friends at Scrimba – the coding platform that's reinvented interactive learning: https://scrimba.com/freecodecamp ⭐️ Chapters ⭐️ - 00:00 Introduction & What We're Building - 01:02 The Modern AI Researcher Role & Asking Research Questions - 04:50 Tokenization & Byte-Level Vocabulary (Why Small Vocab Matters) - 07:01 Embeddings & Transformer Forward Pass Overview - 08:53 GLM-5.3 Architecture Overview & Model Specifications - 10:25 Code Walkthrough: Embeddings & Token Representation - 11:56 Manifold Constrained Hyperconnections (DeepSeek Residuals) - 13:27 Output Projection & Weight Tying - 14:27 RMSNorm & Normalization Layers - 15:10 Positional Encodings (RoPE vs. NoPE) & Sparse Attention Indexer - 18:30 Linear Attention (State-Space Memory) vs. Sparse Attention - 20:46 Mixture of Experts (MoE) & Shared Experts - 23:06 Adding Vision: Patch Embeddings & 2D RoPE - 26:14 Pre-Training Pipeline, Loss & Optimization (AdamW) - 28:05 Pre-Training Experiments: Data Diversity, Interleaving & Curriculums - 30:27 Post-Training & Reinforcement Learning (RL) Setup - 34:44 Designing Reward Functions & Group Relative Policy Optimization (GRPO) - 37:37 Parameter-Efficient RL Updates & Freezing Layers - 39:35 Evaluating RL Results: Task Gains & Regression Risks - 40:47 RL Hyperparameter Experiments: Group Size, Temperature & Seeds - 43:03 Summary & Advice for Aspiring AI Researchers 🎉 Thanks to our Champion and Sponsor supporters: 👾 @omerhattapoglu1158 👾 @goddardtan 👾 @akihayashi6629 👾 @kikilogsin 👾 @anthonycampbell2148 👾 @tobymiller7790 👾 @rajibdassharma497 👾 @CloudVirtualizationEnthusiast 👾 @adilsoncarlosvianacarlos 👾 @martinmacchia1564 👾 @ulisesmoralez4160 👾 @_Oscar_ 👾 @jedi-or-sith2728 👾 @justinhual1290 -- Learn to code for free and get a developer job: https://www.freecodecamp.org Read hundreds of articles on programming: https://freecodecamp.org/news

Choose to Build with AI
Matched to Advanced

AI Maker Residence at KOKO

The third AI workshop taught by our legendary teacher, Nick Sarafa. In one of the last events we did an asset manager raised an additional £25M on their fund within a space of 9 months. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence. One person did.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence at KOKO
Live event
AI Maker Residence at KOKO
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.