Key takeaways
- AI removes the velocity bottleneck in weeks, not years, which means teams now drift from their original intent faster and with less warning.
- QA answers "does it work." A separate discipline has to answer "does it work for the reason it exists," and most teams have no system for the second question.
- Tying every evaluation run to a specific build turns product quality into a trendline instead of a one-off opinion about whether a release "feels better."
- Embedded and fractional product leaders create more durable value by building measurement systems than by making individual decisions while they are in the room.
- The next competitive edge in product work isn't shipping speed. It's the ability to prove, build by build, that speed hasn't detached the product from its intent.
I'm working with an AI-first company right now, and every week AI lets us ship something that used to take a month. That's not the scary part. The scary part is what all that speed exposed: velocity was never the bottleneck on this team, or on most teams I've worked with over 25 years. Direction was. When you can build in hours instead of weeks, it becomes incredibly easy to drift from what you originally set out to build, and you won't realize it until the product is already live. That gap has a name: drift. And AI just poured gasoline on it.
The bottleneck moved, it didn't disappear
For most of my career, the constraint on a product team was throughput. You had more ideas than engineering hours, so prioritization was the whole game. Say no to nine things, build the tenth well, ship it, learn, repeat. That constraint disciplined teams by default. You couldn't drift far because you couldn't move fast enough to drift far.
AI removed that constraint. Not partially, structurally. A team that used to ship one meaningful release a month can now ship one a week, sometimes one a day. The instinct is to treat this as pure upside. It isn't. When the cost of building drops to near zero, the cost of building the wrong thing also drops to near zero per instance, which means it happens far more often and compounds far faster. The teams that win the next cycle won't be the ones that ship fastest. They'll be the ones that can answer, build by build: is what we're building actually doing what it was designed to do?
QA checks the code. This checks the promise.
I want to be precise about the distinction here, because it's the whole argument. QA asks: does it work? Does the button fire the event, does the API return the right payload, does the page load without throwing an error. That's necessary and most teams do it reasonably well.
What almost no team does well is ask the second question: does it work for the reason it exists? A checkout flow can pass every QA test and still fail the customer if the redesign quietly optimized for conversion at the expense of trust. A dashboard can load fast and render correctly while no longer answering the question the user opened it to ask. QA verifies the mechanism. It says nothing about the intent. That's a product strategy problem wearing an engineering costume, and it's exactly the kind of gap the Product Assessment framework tries to surface when we evaluate how a PM reasons about outcomes versus outputs.
If AI speeds up building, it has to speed up measuring intent
Here's the bet I'm making, and it's the reason I spent a weekend building instead of resting: if AI lets us build products faster, it also has to let us measure product intent faster. Otherwise we're flying a faster plane with the same broken instruments. Speed without a corresponding increase in measurement discipline isn't progress, it's just a faster way to get lost.
So I built a system that evaluates the product from three angles teams usually keep separate. Experience: harsh 0 to 10 usability scoring, journey by journey, from the customer's perspective, not the team's. Performance: load times, reliability, build health on every run, no exceptions. Intent: does the live product actually accomplish the jobs users hired it to do, independent of whether it "feels" done.
Every run is tied to the build. That single design decision matters more than any of the three scores individually, because it turns the whole exercise into a trendline instead of a snapshot. Instead of debating whether a release "feels better" in a meeting, half of you nodding and half of you unconvinced, the team can see whether the product is moving closer to or further away from its intended experience, release over release. Synthetic personas run the journeys as brand-new customers on every build. The personas vary, their behavior varies, good days and bad days vary. The standard doesn't move. That consistency is the point.
The real leverage for an embedded product leader
I've spent enough of my career embedded, sometimes with weeks on an engagement rather than years, to know that my value can't just be the decisions I make while I'm in the room. Decisions age out the day I leave. What doesn't age out is a system that keeps making good product judgment repeatable after I'm gone. That's the shift AI is starting to enable for operators like me, and it changes what senior product leadership actually gets paid for.
It's not just a faster way to write code or generate features. It's becoming a force multiplier for product judgment, measurement, and alignment, three things that used to depend entirely on one experienced person's gut and were never written down anywhere the team could inspect. The next generation of product leaders won't win by using AI to build faster than the next team. They'll win by building the systems that keep everything they're building honest, and that's a distinct business acumen muscle, not just a technical one.
The question every team should be able to answer now
Speed without intent is drift. It's a plain statement but I've watched it play out at three different companies in three different ways: a loyalty product that optimized itself into irrelevance, a booking flow that got faster and less trustworthy in the same quarter, a sportsbook feature that shipped on schedule and solved a problem nobody had. None of those failed QA. All of them drifted.
If you're shipping faster than ever right now, and most teams are, ask yourself the only question that actually matters: can you still prove you're building what you set out to build? If you can't answer that with evidence, build by build, you don't have a speed advantage. You have a faster way to be wrong. This is also the kind of judgment gap that shows up clearly in a structured competency assessment: PMs who can articulate intent and measure against it consistently score differently than PMs who can only describe what they shipped.
Common questions
- What is product drift and why is AI making it worse?
- Product drift is the gap between what a product was designed to do and what it actually does after repeated iterations. It used to accumulate slowly because building was slow. AI collapses build cycles from weeks to hours, so the same drift that once took months to appear can now show up in days, often before anyone notices the product has moved away from its original intent.
- Is measuring product intent the same thing as QA or usability testing?
- No. QA verifies that the code functions as written. Usability testing checks whether users can complete a task. Measuring intent asks a different question: does the live product still accomplish the job it was originally designed to solve. A feature can pass QA and score well on usability while quietly serving a different purpose than the one it was built for.
- How should a fractional or embedded product leader use AI to create lasting value?
- Individual decisions made during an engagement disappear once the engagement ends. A measurement system that ties experience, performance, and intent scoring to every build keeps functioning after the leader leaves. That's the leverage: the judgment gets encoded into a repeatable process instead of staying locked in one person's head.
Next step
See which level your evidence actually supports
A free 10-minute adaptive interview. No account required. Evaluates your real decisions across 4 PM competency dimensions and maps them to your demonstrated level.