Not every company needs to lead on every dimension. A supplier that delivers more slowly but advises more accurately can allow that lag in delivery time as long as the customer weighs the advice more heavily than the wait. The question is not whether you win on all dimensions. The question is whether you know what weight each dimension actually has in the purchasing decision, and whether that weight holds up.
That last part is where it gets difficult. Weights are not fixed properties of a market. They shift as soon as the work behind a dimension changes in character.
A dimension weighs heavily when it is scarce: few suppliers can deliver it well, and the difference is noticeable to the customer. Delivery time long weighed heavily because planning was human work and mistakes were costly. As soon as a task is largely done by AI, with or without oversight, part of that scarcity disappears. What used to be a distinguishing capability becomes a baseline provision that most parties match within a foreseeable time.
That does not mean the dimension becomes unimportant. It means that the weight you assign to it today is based on a scarcity that may be disappearing. A benchmark that does not account for this weighs with outdated assumptions.
The work behind a dimension falls into parts: tasks AI can handle independently, tasks AI does with human oversight, and tasks that remain human work. This division determines whether a weight is stable or shifting.
Where a task is almost entirely taken over by AI, the weight of that dimension generally decreases, because most competitors can catch up within a comparable timeframe. Where oversight remains necessary, the difference between companies stays larger and more durable, because the quality of that oversight is itself a skill that not everyone builds up equally fast. Where the work remains human, little changes in the weight in the short term, although the underlying scarcity of skilled professionals itself may begin to shift.
This division differs per company and per sector, and also per company within the same sector. One company has already set up the takeover of a task with working controls; another is still starting with a pilot setup. That difference is measurable, not guesswork, and it is precisely what you find when you assess a competitor's AI maturity from the outside.
A weight in a benchmark is an estimate, not a measurement of the purchasing decision itself. The estimate rests on signals: what buyers in the market say they value, how deals are won and lost, how dimensions have shifted over the recent period. That is useful, but it is no guarantee that the same weighting still holds in a year's time, and it says nothing about an individual deal with specific conditions.
The outcome also says nothing when the evidence base is thin: if a score on a dimension rests on a single unsupported claim, or on outdated information about a competitor, then the weighting is only as good as the data it rests on. That is why every score comes with an evidence matrix: you can check what a score is based on and judge for yourself whether that evidence is sufficient for your situation. Where the evidence is weak, the score should be read as an indication, not as an established fact.
If a dimension weighs lightly and you score weakly on it, the question is not automatically how you fix that. Sometimes the right route is to acknowledge the weakness and strengthen the score on a more heavily weighted dimension; sometimes it is worth catching up anyway because the weight of that dimension is rising as more competitors hand the underlying task over to AI. What justifies addressing a lag and what you can leave alone depends on where the weight is moving, and that is a different question from what you do with a lag on a dimension.
The same applies in reverse, for a dimension where you currently stand strong: if the underlying work becomes largely takeoverable, the question is what you do when your strongest point becomes commonplace due to AI, and how long that lead still makes a difference is a matter to be answered separately via how long an AI lead holds up.
Which work in your own company can genuinely be taken over by AI, and which part will keep requiring oversight, is a question per task, not per role or department; the work scan from FTE TO AI answers that question at that level.
The weighting of a dimension is therefore not a fixed given, and a score without evidence is not a score to build on. What can be done: lay the claims on which you think you win side by side and check which ones can actually be defended with evidence. That is what the free dimension check does. You state where you think you win, and see which of those claims hold up, in the same way as with substantiating a claim about your own quality. The full benchmark, with the complete evidence matrix on your peer group, is under construction.