Earlier this year I wrote that AI was making many of the artefacts around Product Management cheaper to produce. PRDs, analysis, research synthesis, tickets, prototypes and presentations could all be created faster, so I expected Product value to move towards judgement, systems understanding and accountability.
At the time that was just a prediction, but over the last few months I have had a better chance to test it against how I actually use AI.
In my personal projects I went from effectively no portfolio of AI-built applications to around 15 in roughly eight weeks. At work I am building and deploying internal tools myself, while also using agents for code investigation, bulk analysis, persona testing and accessibility checks. I can investigate more, try more ideas and get much further into technical problems that I would previously have left alone because the effort was too high or it was outside of my existing skills
This extra capacity is useful but has also changed where I give my attention. What’s noticeable, is not that I can produce more, but how much more there is now to judge.
The first thing I had to relearn was when not to build
One evening I realised I had spent a significant amount of time planning 64 bespoke automated checkers before properly validating the principles behind them or looking closely enough at what already existed. I stopped and changed the order, before deciding what to build, I validated 189 principles against six established frameworks and used a deliberately sceptical agent to challenge the analysis.
I can’t say whether this removed most of the proposed checkers or materially reduced the eventual build, so that is not the lesson I take from this experience. The useful part was noticing how quickly cheap implementation had pulled me towards implementation before I had finished validating the need.
AI had made both things easier: building bespoke checks and doing much broader validation. I had simply reached for the first capability too quickly, since then I have become much more deliberate about validating the need and checking for reuse before bespoke automation starts gathering momentum.
That is a fairly ordinary Product lesson, but what changed here was the speed at which I managed to get the sequence wrong.
A decision record can look convincing even when no decision happened
A different problem appeared when Claude wrote plausible reasoning into a rules file as though a decision had already been made. The reasoning made sense to me, which made it easy to miss. I only noticed because I remembered that I had never actually made the decision and challenged why it had been documented without being raised with me.
I still want AI maintaining documentation, as keeping rules, rationale and project state current can be is slow and easy to neglect, particularly when several things are changing at once. The mistake was treating plausible reasoning as if it carried the authority of an actual decision.
So I changed the rule rather than removing the capability, AI could record a decision, but anything that required a named human decision-maker had to come back to that person first.
That has made me much more conscious of provenance. An AI-generated note does not need to contain an absurd hallucination to cause a problem. If it is plausible enough and appears in the right place, a future person or agent can reasonably treat it as settled.
I have written before about organisations losing the reasoning behind decisions over time in Decision Debt. Here the problem was almost the reverse, in that the record had gained reasoning for a decision that had never been made. Faster documentation is valuable, but only if it remains clear who actually decided what.
Testing became cheap
I have seen the same shift in evaluation. One of the larger evaluations of my Product Management skill involved 288 runs across two versions and two independent runner families, with the criteria hidden from the systems being tested and another model scoring the results. I would never have attempted this manually as it would have taken too long and been too complicated. AI made it practical to gather much broader evidence than a few hand-picked examples and to check whether a change actually improved behaviour rather than simply looking cleaner in the underlying files.
But the volume of evaluation did not remove the need to judge the evaluation itself. In a later gate, the system appeared to fail. When I looked at why this was, I realised the evaluator was over-specified. Changing the system to satisfy it would have meant optimising towards a bad measure, so I changed the evaluator instead. I kept the result because it was still informative, but did not treat it as a ‘release pass’.
Over time I cut the number of evaluation calls dramatically by carrying forward evidence that was still valid from previous runs and retesting only the things capable of changing the decision, this was the more useful lesson as it helped to optimise token usage and save time.
My overall aim with this, was not to maximise the amount of testing and make it comprehensive, it was to get enough trustworthy evidence to make the next decision.
I started this piece impressed by how much evaluation AI made possible, I have ended up more interested in whether a test is valid and whether its result should actually change anything.
AI can increase output without removing the wider work needed
There is evidence that the same pattern is appearing in software development more generally, AI can make producing a change faster, while leaving more work around deciding whether that change is valid, safe and ready to implement.
DORA’s March 2026 analysis describes AI speeding up initial code generation while some of this saved time is then spent on auditing and verification, it also reports higher AI adoption being associated with both greater delivery throughput and greater instability. By instability, DORA means deployments that cause problems in production or create unplanned rework, changes that need a hotfix, rollback or other remediation and deployments made specifically to fix user-facing bugs.
Black Duck’s 2026 survey of 831 software engineers and DevOps professionals found 92% reporting improved productivity and release velocity from AI coding assistants, while 90% also reported problems with AI-generated code. Manual review, security testing and code rework were the three largest bottlenecks they identified. Black Duck describes this as work being redistributed further down the software-development lifecycle rather than removed altogether.
There is counterevidence, METR’s controlled study of experienced open-source developers using early-2025 tools found that developers took 19% longer when AI was available, despite believing they had become faster. Its February 2026 follow-up suggests newer tools are probably providing more acceleration, but says selection effects and measurement problems make the size difficult to estimate reliably.
These findings do not cancel each other out, but they make me less interested in a productivity number. What matters for this article is that for my own use, AI has reduced the effort involved in attempting, investigating and producing enough to change my behaviour. Once that happens, some of the difficulty moves somewhere else.
Faster code creates more review, faster analysis makes provenance more important. Cheap tests make the quality of the test more important and better maintained documentation makes decision ownership easier to blur if it is not explicit.
The constraint is not always removed but instead it moves downstream.
Product has to know where its judgement ends
Another limit I have become more aware of is getting further into a discipline does not mean I suddenly own its judgement.
For example, if AI generates code for me, I do not become a security reviewer by trade. If it produces an accessibility assessment, I am not suddenly an accessibility specialist. The same applies to architecture, legal interpretation, research and data science. AI can give me more access to those activities and help me ask better questions, but that is different from having the professional judgement of somebody who specialises in them.
For Product, the decisions I still need to own are narrower, i.e. whether the problem is worth solving, what outcome matters, whether the evidence is strong enough to change direction, which trade-offs are acceptable and whether something should exist at all. Sometimes the right judgement is recognising that the next decision belongs to somebody else.
That boundary matters more to me now because AI makes it easy to produce outputs that look like specialist outputs and the appearance of competence has become cheaper and easier to mask as well.
Where I am putting friction back
I am not using AI less because of any of these experiences. If anything, I am using it more. I can build things that would previously have stayed as ideas, investigate far more of a problem before deciding what to do and test behaviour in ways I would not have had the time to attempt manually.
The change I am making is that I am being more deliberate about the points where speed should stop. Before building something bespoke, I want to know whether the need is validated and whether something established already solves enough of it. Before an AI tool records a decision, I want to know who actually made it and where it came from. Before accepting an evaluation result, I want to know whether the test itself is valid and before an irreversible action, I want the authority to be explicit.
These are questions I would normally have, the difference is I find myself having to be more explicit about what I would have previously taken for granted as good judgement. The key change is how quickly I reach these questions, how often and the frequency of applying good judgement.
Earlier this year I thought AI would make judgement a larger part of Product Management because it would automate more of the artefacts around the role and I still think that, but I had framed it mainly as substitution, e.g. AI produces more, so people spend more time on judgement.
My experience has been broader than this, the extra capacity is useful precisely because it lets me do more, but doing more also creates more decisions about what deserves to continue. It leads me to more questions on what evidence is good enough and what should be allowed to become real (or stay as a premise).
The practical change for me is not to slow down, it is to be much more deliberate about the points where speed should stop and a judgement needs to be made.


