Jev from typesafe.ai fixed 30% of my problem. Measuring it fixed the rest.
I tried typesafe.ai's decision model to make an LLM scoring step cheaper. It would have, by a third. The harness I built to test it found the other two-thirds.
Engineer, builder, tinkerer
Software, agents, music, and hardware. Notes on what I'm building and what I've learned.
I tried typesafe.ai's decision model to make an LLM scoring step cheaper. It would have, by a third. The harness I built to test it found the other two-thirds.
A first-pass CI had crept to eight minutes. I assumed slow tests. An agent and one afternoon of measuring found a two-core runner and a Postgres per test file.
How Google-only sign-in gets a server-verifiable identity on Next.js and Firebase: a session cookie, a static layout, and a proof that never clicks the popup.
Two npm modules from 2017–18: a 1966 chatbot behind a Promise, and an intent parser with 34 tests. I ran them again in 2026 next to a local model.