Put the conformance scores of 30 sites on one dashboard, and three things happen: leaders gloat, laggards rationalize, and operations teams start treating audit scheduling as a competitive game. Multi-site comparison is powerful and flammable in equal measure. The craft is in how the comparison is framed.
Fair baselines first
Comparable comparison demands comparable baselines. Sites in a comparison must run the same template family under the same scoring policy - mixing a 20-item retail checklist beside a 60-item industrial safety audit and ranking them together manufactures false heroes and false villains. Weight coverage by audit volume: a site audited four times a year should not be ranked as equal evidence to one audited weekly. When baselines are fair, the ranking that emerges is a fact about performance, not a product of measurement differences.
Stack the lenses, not just the totals
A single total-score column hides the diagnosis. The comparison view earns its keep through severity and trend lenses: critical count per site (the number management should actually be keyed to), finding aging by site (which site lives with its problems the longest?), section sub-scores alongside totals (two sites with equal totals often fail in different sections - one in safety, one in documentation). The lensed view converts "site A is worse" into "site A's kitchen documentation is the pattern."
Trend beats snapshot
The highest-integrity comparison emphasizes direction over position: same-site trend lines over the quarter, rather than a leaderboard of the current month. A site improving from 70 to 85 is a success story worth celebrating; a site holding a flat 92 while its aging climbs is the silent problem. Management attention should follow the trend divergence, not the raw rank.
The deployment discipline
Three practices keep comparison constructive: publish the comparison at the program level with the baseline rules visible; frame reviews as coaching conversations with the lagging site's own trend in context; and keep the corrective-action loop attached to the comparison - the lowest sites get the first support visits. When comparison feeds support rather than shame, the side-by-side view becomes a tool the field asks for, rather than a threat they dodge.
Run fair baselines, lens by severity and aging, emphasize trend over rank, and comparison becomes the mechanism that lifts the whole fleet together.