How AI Is Changing Agency Reviews And Empanelment

Agency reviews exist to answer two questions: can this shop do the work, and can it do enough of it. AI has made both easier to answer and less useful to ask.
A review was always partly a capacity audit. Headcount, disciplines under one roof, offices in the right cities, a credentials deck showing the machine had run before. Procurement asked because the risk of hiring a shop that could not deliver at volume was real and expensive.
Why Agency Reviews Are Getting Harder To Score
Capacity now presents badly. A twelve-person shop can produce a response indistinguishable in polish from a hundred-person one, and the credentials deck no longer separates them either. The scoring rubric still has rows for scale, and those rows are increasingly measuring something that does not predict the outcome.
Meanwhile the volume of responses has risen, because responding got cheaper. More entries, less differentiation, same evaluation window. Evaluators who once read fifteen submissions carefully now skim forty, which quietly favours whoever writes to the rubric rather than whoever would do the work best.
Reviews are receiving more submissions that look more alike and mean less.
What Procurement Is Starting To Test Instead
Three shifts are visible. Live working sessions rather than presented decks, because a room reveals thinking that a document hides. Reference checks weighted more heavily than credentials. And questions about how a shop decides what not to do, which no tool answers on a shop's behalf.
What Empanelment Becomes
Rosters are likely to get shorter and more provisional. If a smaller shop can now be credible on a bigger brief, the argument for a long defensive panel weakens, and the argument for a tight one that gets revisited more often gets stronger.
That is uncomfortable for incumbents who were empanelled on scale, and it follows the same logic as will procurement judge agencies on speed not headcount.
