Blog
Views are my own
Notes on building 0->1 products, better metrics, trustworthy AI, and personalization, and the messy middle between research and shipped product. Occassionally you might find notes about women empowerment, education, environment, policy, books I have been reading, and everthing in between!
-
Hops, Skips, and Spirals: How Reasoning Models Fail and Why It Compounds in Agent Loops
August 01, 2026
Reasoning models don't fail randomly, they fail in patterns: missed hops, coverage gaps, misread questions, and overthinking spirals. A single wrong answer is forgiving; an agent loop that plans, a...
-
When the Test Stops Telling the Truth: Evaluations, Benchmark Saturation, and the Metrics We Trust
June 21, 2026
Frontier models now sit within a point or two of each other at the top of benchmarks that defined 'hard' only a few years ago. This post walks through why benchmarks saturate, takes a deep dive int...
-
Reading the residual stream: a mechanistic study of misgendering in Gemma and Aya
June 06, 2026
Two years ago I made a deliberate bet to focus on mechanistic interpretability of language models. The thesis was simple: the shortest path to a better product is understanding why the model behave...
-
What were my odds? Estimating FIFA 2026 ticket lottery chances probabilistically
May 09, 2026
I applied for 2 tickets in the FIFA Random Selection Draw and won 1. This post walks through a transparent probability model — and an interactive page where you can play with the assumptions yourself.
-
Welcome — what I'll be writing about
May 04, 2026
Welcome — what I’ll be writing about