The common answer used to be: annually, or during a strategic reconsideration. That answer rested on an assumption that no longer holds everywhere: that the dimensions on which companies compete are themselves stable. Delivery time, response time, price consistency, quote turnaround time — these were outcomes of staffing and process, and those change slowly. An annual measurement therefore made sense.
When AI takes over the work behind a dimension, the dimension itself changes character. Delivery time that depended on how many planners were on staff becomes delivery time that depends on how well a planning system is set up. That can shift in months, not years. A comparison that keeps using the assumption of stability measures something that no longer exists.
Not every dimension shifts at the same speed, and that determines how often a comparison needs to be revisited.
For dimensions where AI can largely take over the task — planning, first-line response, quote calculation — the competitive position can flip in a short period. A competitor who organizes this work differently today can win within a few months on a dimension where it was still lagging last year. These dimensions call for a shorter cycle, not because the method dictates it, but because the underlying reality itself moves faster.
For dimensions where people work with AI support and provide oversight — advisory quality, complex case handling — the position shifts more gradually. The support changes the speed and consistency of the work, but the judgment remains with people, and that process doesn't change from month to month.
For dimensions that remain human work — relationship management, negotiation, building trust in a long sales cycle — the old rhythm is still largely intact. Here, an annual or biennial look is often sufficient, because the underlying labor doesn't fundamentally change shape.
The question "how often" therefore has no single answer. The answer is: per dimension, depending on which of the three categories that dimension currently falls into.
This is not a future scenario. In some companies the planning function is already largely automated and the delivery-time dimension shifts measurably; in other companies the same function still rests entirely with people and the dimension remains stable. The difference is not in the sector, but in the choices a company and its competitors have already made about which work they hand over to systems and which work they keep with people under oversight that approves or rejects. A benchmark that doesn't map that difference separately per dimension misses where the actual movement is.
That is also why why your customers see the difference before you do is relevant: customers often experience faster delivery or more consistent response months before an internal measurement picks it up. The comparison then lags behind the market, rather than leading it.
An evidence matrix makes scores traceable, but traceability is not the same as certainty. A number of limitations are inherent to the method, not to the execution.
Public evidence about a competitor is always a delayed reflection of the actual state. A claim on a website, a review, an annual report: these say something about a moment that has already passed. How recent and how direct the evidence is determines how much weight a score may be given; old evidence on a rapidly shifting dimension says little.
A score on a dimension on which you will never win also says nothing useful, unless that dimension carries weight in your customer's purchasing decision. How that weighting works, and why a low score there is not automatically a problem, is worked out in how do you weigh a dimension on which you will never win.
And a strong point today is no guarantee for tomorrow. If a claim rests on work that can now also be delivered by a system, the distinction disappears as soon as competitors introduce the same system. What that means for a strongest point that becomes commonplace through AI is on what do you do when your strongest point becomes commonplace through AI.
These limitations are not a reason not to build a benchmark. They are a reason to state, with every outcome, how old the evidence is and how sensitive the dimension is to shifting through automation. A score without those two pieces of information is difficult to interpret, even for a management team that wants to handle it sensibly.
How often a comparison needs to be revisited is therefore tied to a question that is not answered by the market but by the company itself: which work here can genuinely be taken over by AI, which work happens with oversight, and which work remains human work. That question is answered per task by the FTE TO AI work scan. Anyone who knows that also knows on which dimensions their own position can shift quickly and on which it cannot.
If any of these outcomes touches on decisions about personnel, separate legal requirements apply; this page and the underlying method offer no advice on that matter.
A peer group that is put together incorrectly renders every follow-up question about the repetition rhythm meaningless; see what is a peer group and how do you put one together for how that selection works. Anyone who wants to see where in the market there is still room that no one occupies will find the approach in what is a white space analysis.
The free dimension check is a first step: you name where you think you win, and see which of those claims can be defended with evidence. The full benchmark is under construction.