Ways of working · essay

Spec-driven AI development doesn't have to reinforce the feature factory

Joni Lindgren Founder & Growth PM 22 min read

Since September 2025 I have built the things I need with spec-driven AI development. An agent turns a conversation into a spec, another builds it, and a third reviews it before it merges. I am not a developer, and for me this way of working has been revolutionary. It can also make a feature factory very efficient.

This week Matt Pocock’s skills arrived in a client’s repository ahead of an AI workshop. I had the same reaction to them that I had when I first installed superpowers, another set of skills for planning and building with Claude Code. Both sets are made for developers, and I have no comments on the development side. What I noticed is what they leave out: the outcome.

Pocock’s implement-spec skill states its goal as “the entire spec implemented on a single integration branch, with every ticket resolved the way the issue tracker closes work.” Superpowers checks its plans for “Spec coverage” and asks of every requirement: “Can you point to a task that implements it?” Both sets hold the code to the spec, and neither asks whether anyone behaves differently once the feature is live. That change in behaviour is what I mean by the outcome. Used that way, the agents count what ships, which is how a feature factory works.

Superpowers comes closest to the outcome: its brainstorming step asks for “the intended outcome, who it is for, and what success looks like.” But neither set has a place for the behaviour that should change, so nothing carries it from the spec into the build or the review.

Pocock already solves the same problem for architecture. His skills keep short decision records, and his format explains why: “These stop the next engineer from “fixing” something that was deliberate.” A decision record is short, versioned, and says why, so the next person, or the next agent, can read it before they change anything. Product decisions deserve the same treatment. What if a product decision record held the outcome definitions, the guardrails and the choices we have made for the product? Developer specs could reference it, and agents could look out for it while they plan, build and review. For scilla.studio’s benchmark tool, it could read like this:

  • User: product people and founders who want to compare their growth numbers.
  • Behaviour: they generate a diagnostic report of their numbers, and then do something with it: share it or download it.
  • Signal: how many of the people who generate a report go on to share or download it.
  • Guardrail: nobody needs an account to use it.

Generating the report is the aha moment. Sharing or downloading it shows that the report was worth something.

With the desired outcome on record (who the user is, which behaviour should change, how we will see it), AI can take it one step further:

  • Before the build, it can test every user story against the outcome and drop the ones that would change nothing.
  • During the build, it can add tracking for that behaviour in the same tickets, so the measurement ships with the feature.
  • After release, it can go back to the data on an agreed date and help validate whether the behaviour moved.

Without the record, the work ends when the code ships:

SpecCodeShip

With it, what we measure goes back into the decision:

Product decisionSpecCodeMeasureLearn

But isn’t the outcome the product manager’s job, settled before anyone writes a spec? It is, and these are engineering tools that make no claim to do product work. My hypothesis is that the outcome gets lost at that handover, because the agents build from the spec and the spec has no place for it. Recording the product decisions where the spec can reference them leaves the decisions with the product manager and puts them in front of the agents that build and measure.

There are limits, the same ones product development has without AI. Behaviour change can take weeks to show after a release. How do we document and version what we learn, and adapt our goals and hypotheses to it? And when the riskiest assumption is whether anyone wants the feature at all, a cheap test before the spec answers it better than a build followed by a measurement. Much of this is product operations: artefacts that need a home, and a process that is intertwined with the build process. Let’s not fumble the opportunity to build the system so that it helps us stay outcome-focused, instead of reinforcing the feature factory.

AI enablement for product teams

Keep up knowledge work with how fast development now moves. We help you build AI into how your product team works, from discovery to strategy, starting from your own context.

Talk about AI enablement →
Written by
Joni Lindgren
Founder & Growth PM · DM on LinkedIn
AI enablement for product teams Talk about AI enablement →