AI that survives 70 million users.
Taking AI features from demo to production at scale: evaluation before release, humans accountable for the output, and a growth curve the business could rely on.
- Industry
- Creator platform
- Focus
- AI product delivery · Evaluation and guardrails · Engineering practice
- Role
- Head of Grow Engineering
The problem
Everyone can build an AI demo. Very few can put one in front of millions of users and stand behind what it says.
The gap is not the model. It is everything around it: knowing whether a change made the feature better or worse, catching the bad answer before a customer sees it, and being able to ship weekly without holding your breath each time.
What we did
- Made AI features testable before releaseAn evaluation framework, so every AI feature is measured against real cases before it ships rather than judged by whoever tried it last. The team can answer “is this version better?” with evidence instead of opinion.
- Scaled an agentic product to half a million usersTwo teams, led through their managers, took agentic AI insights and peer-to-peer messaging from twenty thousand to five hundred thousand monthly active users in three months, and kept it reliable enough to leave running.
- Put AI inside how the team works, not just the productImplementation moved to AI tooling with engineers owning problem definition, review and verification. Delivery time on medium-sized work dropped by roughly a quarter. The judgement stayed with people; the typing did not.

What changed
Two companies became one platform and customers did not notice. An entire user base migrated with zero downtime, with unified billing and login behind it. The AI features that drove growth went out on a weekly cadence with evaluation in front of them.
“The Plann and Linktree teams have been extremely fortunate to have such a smart and genuinely caring leader at the helm. I’ll miss our open and candid weekly chats.”
More of where we’ve worked
Not sure where your version of this starts?
Thirty minutes to work out what’s worth doing, and what isn’t.


