2026 (7)
July (5)
- How LLMs go from base models to assistants July 23, 2026
- Scaling up a DiT: replicating Peebles & Xie, then optimizing the training loop July 23, 2026
- A toy assistant: SFT and DPO on my from-scratch GPT-2 July 16, 2026
- Faster Qwen-Image-Edit 2511 by caching KV July 9, 2026
- Diffusion on a 2D spiral July 2, 2026
June (2)
2025 (1)
July (1)
- Transformers Don't Need LayerNorm at Inference Time July 23, 2025
2024 (1)
October (1)
- ARENA Capstone: Hyperparameter tuning for MELBO October 5, 2024