Posts
- Writing the Code Is Only the Beginning
- A Backend Can Be Correct and Still Be Impossible to Operate
- The Database Update Succeeded. The Event Didn't. Now What?
- The Hardest Distributed Systems Problem: Not Knowing What Happened
- What Happens After You Put Something on a Queue?
- My Code Was Correct Until Two Requests Ran at Once
- Retries Can Make a Failure Worse: How One Slow Service Takes Down Another
- The Fastest Backend Operation Is the One You Don't Do
- Why Databases Get Slow: A Query Is Work, Not a Lookup
- The Journey of a Backend Request: Every Arrow Was Hiding Work
- Day 7: Deployment Is Not The End
- Day 6: Why Do Models Fail Even When Training Accuracy Looks Great?
- Day 5: When Does a Training Script Become a Pipeline?
- Day 4: Training One Model Is Easy. Finding The Best One Is Hard.
- Day 3: If Git Versions Code, What Versions Data?
- Day 2: From Notebook to Reproducible Training
- Day 1: Why DevOps Is Not Enough for Machine Learning
- Day 7: What Does a Complete AI Platform Actually Look Like?
- Day 6: Why vLLM Exists
- Day 5: Why Ray Exists
- Day 4: Why AI Needs Different Scheduling
- Day 3: Why GPU Sharing Is Hard
- Day 2: How Kubernetes Sees GPUs
- Day 1: Why AI Workloads Are Different from Traditional Applications
- LFX Mentorship: My Experience as Mentee with CNCF Thanos