AI has changed the physics of product development.
Not the philosophy.
The philosophy has always been simple: get close to users, understand what they’re trying to accomplish, build something, watch what happens, learn, and iterate.
Qualitative feedback gives you depth. Quantitative measurement gives you breath. That loop hasn’t changed.
What has changed is how fast we can run it. And when building gets dramatically faster, the bottleneck moves.
The cost of building has collapsed
An MVP that once took weeks or months can now be assembled in days.
A designer can generate several interfaces in an afternoon. An engineer can wire up a workflow before the end of the day. A PM can change an interaction and try another version straight away.
But faster building doesn’t automatically mean better products. It means you can make more attempts.
The advantage goes to the team that learns from those attempts fastest.
That is the real AI-era advantage.
Qualitative discovery has to move at AI speed
The answer isn’t to go back to months of formal user research.
It’s to make qualitative discovery run at the same speed as development.
You don’t need 10,000 beta users. You need a handful of people who deeply understand the problem.
This is how we do it at Tupai.ai. We work with two kinds of super user.
The Teacher
The first is the person who forces you to go through the product step by step.
Don’t ask them, “Do you like this?”
Watch them try to accomplish the task.
Where do they hesitate? What do they misunderstand? What do they expect to happen next?
The Teacher is valuable because they are willing to go through the product practice by practice. They aren’t evaluating the demo. They’re evaluating whether the product actually helps students to understand.
They expose the tiny things a product team goes blind to after staring at the same screen for too long.
And those tiny things matter.
A confusing label. A missing step. The wrong default.
These are almost impossible to catch with an eval, unless you already know to test for them.
The PM-Mom
The second is what I think of as the PM-Mom.
The ruthless pragmatist.
Someone who doesn’t care how clever the underlying AI is. They care whether the product is useful for their child.
They ask:
“Why is this here?”
“Who actually needs this?”
This feedback can be uncomfortable. That’s precisely why it’s useful.
The PM-Mom cuts through all the sophistication of the technology and gets back to the only question that really matters:
Does this solve a real problem?
You need both.
The Teacher shows you how the product fails in practice. The PM-Mom tells you whether the product deserves to exist in the first place.
Neither is statistically representative. That’s not the point.
The value is information density.
One great user can expose a failure mode that thousands of synthetic test cases would never reveal.
Behavioral exhaust is feedback too
Increasingly, the richest qualitative signal isn’t what users tell you.
It’s what they do.
A user rewriting a prompt three times is feedback.
A user repeatedly correcting the same field is feedback.
A user ignoring a feature you spent three months building is feedback.
They don’t need to tell you, “Your product doesn’t understand my intent.” Their behavior already did.
AI makes it possible to analyze this behavioral exhaust at scale: support tickets, failed queries, prompt histories, corrections, abandoned workflows, sales conversations, community discussions.
But I don’t think the goal is to have an LLM tell you what users want.
The goal is to use AI to surface patterns that humans can investigate much faster.
AI can accelerate discovery. It shouldn’t replace it.
This is where evals become incredibly powerful
Evals are one of the most powerful tools an AI product team has.
They let you test hundreds or thousands of cases quickly. Compare model versions. Iterate on prompts. Catch regressions. Protect critical workflows.
Most importantly, they turn a discovered failure into institutional memory.
You find a problem. You fix it. You add it to the eval suite.
Now the team doesn’t have to rediscover it every time the model, prompt, or workflow changes.
This is where eval-driven development earns its reputation.
But there is an important boundary.
An eval can only test what you’ve decided to measure.
You have to choose the cases. You have to define what good looks like. You have to identify the failure modes. You have to encode the expected behavior.
In other words, an eval is a formalized version of what the team already knows about the product.
That’s its strength. It’s also its limit.
Evals are excellent at protecting known knowledge. They are much weaker at discovering unknown problems.
The danger is developing inside the eval
This creates a subtle trap.
Writing another eval can be done alone in an afternoon.
Finding five great superusers is harder.
Recruiting them is harder. Getting them to use something unfinished is harder. Watching them struggle is harder. Interpreting ambiguous feedback is harder.
Accepting that the thing you spent two weeks building might be solving the wrong problem is much harder.
You can have an excellent eval score and a fundamentally broken product.
The division of labor
This is why I don’t think the future of AI product development is a choice between qualitative and quantitative methods.
They have different jobs.
Qualitative discovers. Quantitative protects.
Users expose needs, behaviors, workflows, expectations, and failure modes you didn’t anticipate. That expands the boundary of what the team knows.
Then you formalize the important discoveries. Turn them into test cases. Put them into the eval suite. Measure them continuously.
Now you can make the next 100 automated iterations without worrying that you’ve silently broken something you already solved.
The eval turns a discovery into a guarantee.
The loop becomes: build, observe, discover, formalize, evaluate, build again.
This is why I think evals are disproportionately powerful at two ends of the curve.
Early on, they take you from something that barely works to something that works reasonably well.
Later, they take you from something that works well to something reliable and protected against regression.
But in between, when you’re still discovering what actually matters, you need reality.
You need users. You need behavior. You need someone willing to say:
“Why would I use this?”
“Why does it work this way?”
“Why can’t it just do the thing?”
That’s the feedback an eval can’t invent for you.
The new competitive advantage is learning velocity
AI has made building dramatically faster. Everyone can produce another prototype.
The advantage now belongs to whoever learns from reality the fastest.
The best AI product teams will compress the distance between qualitative discovery and quantitative validation.
They’ll build something in the morning. Put it in front of users that afternoon. Watch what breaks. Fix it. Turn the important failure into an eval. Automate the regression.
Then get back in front of users and find out what they still don’t know.
That is the new product loop.
And in an era where anyone can build faster, the team that learns faster wins.
Evals protect yesterday’s knowledge. Users reveal tomorrow’s problems.
The best teams will need both.