Transformer internals: attention without the mysticism
A mechanical look at scaled dot-product attention, KV caching memory geometry, and rotary embeddings without the hand-waving metaphors.
Available for interesting problems
I'm a software engineer working on distributed systems and developer tooling. This is where I keep my notes on technology, the places I've been, and the things I've built. No newsletter pop-ups, no cookie banner — just writing.
A mechanical look at scaled dot-product attention, KV caching memory geometry, and rotary embeddings without the hand-waving metaphors.
Four days walking from Mestia to Ushguli through Georgia's high Caucasus. Medieval stone towers, glacial mud, and guesthouse hospitality.
Public leaderboards suffer from severe data contamination and metric hacking. How we design calibrated, private eval suites for production systems.
A multi-package Raft consensus engine and key-value store in Go with snapshot compaction and a partition-injection test harness.
A fast structured log streaming CLI in Rust with zero-allocation predicate filtering, autodetection, and colorized formatting.
I'm happy to talk about backend architecture, developer experience, or where to eat in Lisbon. The inbox is open.
Say hello →