Pick one team, not the whole org
The selection criteria that actually predict whether a pilot survives contact, and why the most enthusiastic team is usually the wrong one to start with.
The pattern is so consistent it is almost funny. A platform group gets budget to fix the status-drift problem, and within a week the plan has become an org-wide rollout with a phased schedule, a communications plan, a champions network and a slide with eleven team names on it.
Three months later the slide has eleven teams on it and two of them have logged in.
The failure is not ambition. It is that an org-wide rollout has no mechanism for learning. When something goes wrong on team seven, you cannot tell whether it is a tooling problem, a convention problem or a team problem, because you changed all three at once across eleven groups. So the response is a rule, applied everywhere, to solve a problem you have not diagnosed. Do that four times and you have a configuration nobody understands and a reputation for breaking things.
One team, deliberately chosen, gives you a controlled experiment. That is worth more than three months of parallel progress, because parallel progress in the wrong direction is just distance from where you should be.
The criteria that actually matter
I have watched enough of these to have opinions about which team attributes predict a pilot that survives. Here are the ones that carry weight, roughly in order.
High agent usage. This is the strongest signal by some distance. The whole thesis rests on the gap between generation speed and record-keeping speed. On a team where agents are writing a substantial share of the changes, that gap is wide, painful and visible, and closing it produces an effect people can feel. On a team writing everything by hand, the gap is a nuisance, and your pilot will produce a small improvement that nobody notices. Pick the team where the pain is real, not the team that is easiest to schedule.
Work that lives in source control. Ground truth is read from branches, pull requests, reviews and merges. A team whose real output is configuration in a vendor console, dashboards in a BI tool, or infrastructure changed through a web UI has less signal to read, and the guarantees weaken proportionally. This is not a knock on those teams. It is a statement about where the evidence lives.
Reasonable volume. Enough events per fortnight that your shadow log means something. A team merging fewer than roughly twenty pull requests in two weeks will not generate enough data to calibrate against, and you will end up making decisions on a sample of eight.
A tech lead who will spend fifteen minutes a week on it. Not enthusiasm. Attention. The reviews during shadow mode are the entire human cost of this and they need someone who understands the branches. If the lead is drowning, the reviews will not happen and the pilot degrades into an unread log.
Stable enough to still exist in ninety days. If the team is mid-reorg, about to lose half its members, or in the middle of a platform migration, everything you learn will be confounded. Boring stability is an asset here.
A skeptic on the team. This is the counterintuitive one and I will defend it below.
Does not matter as much as people think
- Enthusiasm: absorbs rough edges silently, gives a false positive
- Codebase cleanliness: unrepresentative
- Seniority: compensates by hand, hides the signal
- Your preferred tracker: optimises for your convenience
Actually predicts survival
- High agent usage: the whole thesis rests on it
- Work that lands in source control
- Twenty merges a fortnight, or the log means nothing
- A tech lead who will spend fifteen minutes a week
- Still existing in ninety days
What does not matter as much as people think
Enthusiasm. The team that volunteers first is usually the team most tolerant of new tooling, which means they will absorb rough edges silently and give you a false positive. You will roll out to team two, who are normal, and discover everything the enthusiasts quietly worked around. A pilot’s job is to surface problems, and an audience that forgives problems is bad at that job.
Codebase cleanliness. Everyone wants to pilot on the tidy service. The tidy service is unrepresentative. More on this in the next article, which argues the case for brownfield directly.
Seniority of the team. Senior teams are often better at compensating for tooling gaps by hand, which again hides the thing you are trying to measure.
Whether they use your preferred tracker. The integrations exist across the common trackers. Choosing a team because they are on the tracker you know best optimises for your convenience rather than for the quality of the experiment.
The case for a skeptic
I want to make this argument properly because it sounds like contrarianism and it is not.
A pilot has two possible outputs: evidence that this works, and evidence that it does not. Both are valuable. Only one of them is easy to obtain from an enthusiastic team.
A skeptical senior engineer on the pilot team does three useful things. They find the failure modes fastest, because they are looking for them. They ask the questions your eventual rollout will face, early, when you can still change the answer. And if they are convinced at the end, that conversion is the single most persuasive artefact you will have when you go to team two, because everyone knows they were not predisposed to like it.
The failure mode to avoid is picking someone who is not skeptical but hostile: someone whose position is that the whole category is illegitimate. That is not a reviewer, that is an opponent, and you will spend the pilot litigating rather than learning. The distinction is whether they will change their mind on evidence. Ask them directly: what would you need to see to be convinced? If they can answer, they are a skeptic and you want them. If they cannot, they are not.
A pilot has two possible outputs: evidence that this works, and evidence that it does not. Both are valuable, and only one of them is easy to get from an enthusiastic team.
A skeptic finds the failure modes fastest because they are looking for them, and asks the questions your eventual rollout will face anyway, at a point where the answers are still cheap to change.
A concrete selection process
Here is how I would actually pick, in about two days.
Step one: list every team, and score each on five yes-or-no questions.
- Are agents writing a meaningful share of their changes?
- Does their work primarily land in source control you can read?
- Do they merge at least twenty pull requests a fortnight?
- Will their tech lead commit fifteen minutes a week for ten weeks?
- Will this team exist in its current form in three months?
Anything below four yeses is out. Do not negotiate with yourself about this.
Step two: for the remaining teams, look at the ticket lag and orphan rate. You measured these in the baseline step. Rank by ticket lag. The team with the biggest gap between merge and status transition is the team where a fix produces the most visible effect.
Step three: eliminate anyone under an immovable external deadline. A team shipping to a regulatory date in eight weeks will deprioritise the pilot the moment things get tight, and they should. That is correct behaviour and it makes them a bad pilot.
Step four: pick from what is left, preferring the team with a credible skeptic on it.
Step five: ask them. Not for permission, exactly, but you need one person on that team who actually wants to find out. A pilot imposed on a team with zero internal interest will produce compliance rather than data.
What to tell the team, and what not to
The framing of the ask matters more than most people expect. Two versions of the same sentence produce very different pilots.
The version that works: “We think your board lags reality by about four days and roughly a fifth of your merged work is not linked to anything. We want to run something for two weeks that writes nothing, just to see whether that is true and whether it is fixable. At the end you can tell us to stop.”
The version that does not: “You have been selected for the pilot of our new engineering efficiency programme.”
The first is an experiment with an exit. The second is a mandate with a survey at the end. Teams can tell the difference immediately and they respond accordingly.
Be specific about three commitments up front: the pilot writes nothing for the first cycle, every automated write once enabled captures the prior value and can be reverted, and the team can turn it off unilaterally without a meeting. All three of those are true, and saying them out loud removes most of the objection surface before it forms.
Resisting the pull to expand
Somewhere around week five, if things are going well, someone will suggest adding a second team since “it is basically working.” Do not, and here is the concrete reason rather than the principled one.
You are still changing variables on team one. You will lower a threshold, enable an action, adjust a convention. Every one of those changes needs a clean signal to evaluate against. Adding a team adds a second stream of noise and, worse, adds a second set of conventions that may conflict with the first, at exactly the moment when you are trying to work out whether a rule should be global or local.
Expand after the decision point, not before it. And when you do expand, expand to a team that is different from the pilot in one obvious way, not one that is similar. A second data point that resembles the first tells you almost nothing.
Where this breaks down
The one-team argument has real limits, and there are situations where it is the wrong call.
Some problems only appear at multi-team scale. Cross-team dependencies, portfolio rollups, epics that span groups, and shared repositories with several owning teams are all invisible in a single-team pilot. If the thing you are actually trying to fix is portfolio-level status drift across a programme, a one-team pilot will validate a mechanism that does not address your problem. In that case you need at least two interdependent teams in the pilot, and you accept the messier signal as the price of testing the real scenario.
A single team’s result may not generalise, and confident extrapolation from n equals one is a genuine risk. The pilot proves that this worked here. It does not prove it will work on the team with the different tracker, the legacy monolith and the manual QA gate. The honest position after a successful pilot is “we know one thing,” and the correct next move is a second pilot that is deliberately unlike the first, not a rollout plan with eleven names on it.
Sequential rollouts take longer, and sometimes that is decisive. If you have a genuine deadline, a compliance-driven date or a leadership window that closes, running teams one at a time may simply take more calendar than you have. Parallel is worse methodology and it is sometimes the only methodology available. If you must go parallel, at least stagger the starts by a fortnight so that team one’s shadow findings can inform team two’s setup.
And picking the team with the widest gap can select for dysfunction. The team with a four-day ticket lag and a thirty percent orphan rate may have those numbers because of process problems that no automation fixes. You will then be running a pilot inside a team that has an underlying management issue, and the results will be muddied by it. If the candidate team’s numbers look bad because the team is struggling rather than because it is fast, pick the second-worst instead.
The takeaway
One team, chosen on evidence rather than enthusiasm: high agent usage, work that lands in source control, enough volume to calibrate on, a lead with fifteen minutes a week, stability for a quarter, and preferably a skeptic in the room. Frame it as an experiment with an exit rather than a programme with a launch. Then hold the line against expanding until you have a decision.
The whole point of one team is that when something breaks you know exactly what broke it. That property is worth more than three months of parallel motion.
The next piece argues something that surprises people: once you have picked your team, you should point this at the messy legacy codebase rather than the clean new one, and most greenfield advice about agentic workflows actively misleads you.