Hey there,

Do you have an AI that works well in demo, but is flaky in production? And you don't quite know why? If so, I can help.

I'm Chong Han, a builder and designer based in SF with 20 years of experience working on complex software workflows. My experience includes time at Microsoft, Honeycomb, and CodeSee. I've also been building AI for a while, previously computer vision, and recently the AI coding agent at Empromptu.

I built Axle, the open-source AI runtime that powers Sunnyday. I've spent a lot of time working with AI APIs, thinking about control loops, and building evals to ensure that Sunnyday is reliable and consistent. Let me apply that to your project too.

We can take 30 minutes so you can tell me about your project and what you want to achieve. I'm happy to give you my opinions on how to make it better – yours to keep for free – and if it's a good fit, we can work on it together.

Where I can help

Make your AI in production reliable

We'll define what success actually means for your system, then build evals that measure it. Instead of finding out from users when something breaks, you'll see exactly where and why and ship changes with confidence.

Start with this

Automate a process with AI, reliably

We'll decompose your process together: code where code works, models where they're needed. And we'll build the evals alongside it, so it's testable from version one — not bolted on after things go wrong.

Start with this

How we work together

  1. 01

    Introductory Call

    We take 30 minutes and you walk me through your AI project: what it does, where it breaks, and what you want it to do. You leave with my honest read on what it needs, whether or not we go further.

  2. 02

    Scope the project

    If we both think it's a fit, I'll write up a fixed-scope proposal. We will define together what needs to get built, how I work with you and your team, how we measure success. A typical engagement runs weeks, not months.

  3. 03

    Build and Ship

    We will build this out together, decomposing the process, writing the code and building the evals alongside it. You and your team will leave with a thorough understanding of how it works and why.

  4. 04

    Watch it run

    Shipping isn't the finish line. Once it's live, the evals run against real traffic and we watch what they catch. I stick around for the first few weeks of production to tune until it holds.