/aienm.

Quantifying Duplicate Work Eliminated by Shared AI Skills

Shared AI workflows cut duplicate rebuilds and recover millions in hidden productivity losses.

Staff Writer · · 11 min read
Cover illustration for “Quantifying Duplicate Work Eliminated by Shared AI Skills”
AI Productivity & ROI · July 28, 2026 · 11 min read · 2,421 words

There's a category of waste that never appears on any AI vendor invoice. Call it AI Execution Debt: prompts and workflows scattered across personal chat histories, Notion pages, and Slack threads instead of governed, discoverable assets. The next team that needs the same capability has no idea they exist, so they build from scratch and the clock resets.

Wipro's engineering team formalized this pattern in 2025 under the term "Prompt Technical Debt," describing the compounding maintenance cost and fragility that accumulates when speed takes priority over documentation and lifecycle ownership. Anyone who has watched a codebase degrade under deadline pressure recognizes the logic immediately: do it fast now, pay for it slowly, forever. The AI version is less visible but structurally identical.

What makes it worse is drift. Similar prompts built in isolation don't stay similar. They diverge as teams adjust and patch in response to local pressures, and what started as near-identical workflows eventually produce inconsistent outputs depending on where a user accesses them. Users notice. That loss of trust triggers manual compensation: copying context into prompts by hand, hardcoding summaries that should be dynamic, repeating setup steps across every new use case. Each of those behaviors is a micro-rebuild of something another team already solved, somewhere, and never told anyone about. I've sat in enough AI program reviews to recognize the pattern — usually when someone mentions their team "had to start over" on something another group built six months earlier.

Two distinct cost layers exist here, and conflating them is a common analytical mistake. One estimate puts the productivity tax from AI-generated rework at $186 per employee per month, reflecting downstream costs from low-quality AI outputs. That measures output quality. The upstream cost, the labor of building the same capability twice across two teams that never knew about each other, is separate, additive, and largely invisible to anyone not actively tracking it. These aren't the same number. They stack rather than cancel out.

Venn diagram: AI Execution Debt: Output Quality vs. Upstream Costs. Compares Output Quality Costs and Duplicate Build Costs; overlap: Compounding Debt.

The baseline productivity gains that duplicated effort is actively eroding

Before measuring what's lost to duplicate building, it helps to establish what's actually on the table. Federal Reserve Bank of St. Louis data from 2024 found that workers using generative AI saved 5.4% of their work hours, translating to a 1.1% workforce productivity increase. Real, if modest — and that's under reasonably functional conditions.

Microsoft and NBER research from 2025 found that Copilot users spent 3.6 fewer hours per week on email, a 31% reduction. But Microsoft's own research also shows employees need at least 11 weeks before realizing meaningful productivity improvement from AI tools. Organizations that cycle through redundant tool evaluations burn a significant fraction of that ramp window on setup rather than use. The productivity gain gets deferred, sometimes indefinitely, which is the kind of detail that never makes it into a vendor case study.

The developer data tells a more granular story. Sixty-six percent of developers cite "almost right, but not quite" AI solutions as their biggest daily time sink. That gap, between what a well-built shared skill could deliver and what a hastily assembled local prompt actually does, is the immediate practical value of reuse, expressed in daily frustration rather than annual savings projections. Meanwhile, only 19% of knowledge workers report clarity on what work should be done with AI. That's primarily a discoverability problem, not a training problem. When skills are ungoverned and invisible, employees can't find them, and the productivity gains never fully materialize. You can run all the AI literacy workshops you want; if the tools aren't findable, you're training people to navigate a library with no catalog.

These productivity numbers represent what's available. Duplicate rebuilding is what consumes them before they ever reach the income statement.

How organizations have put concrete numbers on eliminating duplicate AI work

The organizations that have most clearly captured reuse savings all did one thing first: they named and measured the redundant activity before designing any solution. That sequencing is what separates a documented case study from a press release.

QuantumBlack, McKinsey's AI unit, is the most thoroughly documented example at scale. Fifteen thousand-plus projects draw from a library of over 700 AI components built by more than 1,500 technology experts. Embedding the Brix MCP into GitHub Copilot reduced the time and steps required to discover and reuse assets by over 55%. That's not a license utilization metric. It's reuse expressed as time, a direct compression of the per-project setup cost that duplicate building would otherwise impose. For a global transportation client, QuantumBlack co-developed a centralized asset system that reduced time-to-value for new AI use cases from three months to one, a compression of more than 60%.

Freeport-McMoRan structured 60% of a reusable ML model for copper recovery prediction to transfer directly across sites, with only 40% customized per plant. Every additional deployment becomes structurally cheaper than the first, because the expensive part was already built.

Wyndham Hotels centralized its AI review workflows and cut brand-standards review time by 94%, saving between 40 and 80 hours per review cycle. Lumen Technologies quantified four hours of pre-call research per sales rep as a $50 million annual organizational drag, then used that figure to design Copilot integrations that compress the same task to 15 minutes. OneDigital found that $120,000 per year and 40% of its talent acquisition team's time were consumed by interview scheduling alone; after deploying conversational AI in 2024, scheduling costs dropped to near zero.

Each of these cases follows the same underlying logic: measure the redundant activity, compress it through shared infrastructure, verify the delta. The organizations that skipped the measurement step found it considerably harder to credibly claim the savings.

What AI Centers of Excellence reveal about reuse at the program level

A mature AI Center of Excellence is, structurally, an institutionalized answer to one question: how do we stop solving the same problems in parallel? The portfolio-level gains documented across CoE implementations include 30 to 40% reductions in AI project cycle times, 25 to 35% lower per-project costs, and three to five times faster skill development — gains from avoiding rebuilds, not from better models. That distinction matters more than it typically gets credit for.

The time-to-production math clarifies the stakes. Without a CoE, a typical project spends six to twelve weeks on tool evaluation and infrastructure setup before any real development begins. With a functioning shared registry and established architecture patterns, that drops to one to two weeks. Across 20 projects annually, that's 80 to 200 weeks recovered, the equivalent of one and a half to four person-years of capacity returned to productive work. For a company running $5 million in annual AI spending, the consolidation and efficiency gains available through this kind of centralization can reach $1.25 to $1.75 million.

PwC's research identifies centralized repositories of reusable AI assets as a consistent differentiator among top-performing companies, who are more than twice as likely to maintain them. PwC's Chief AI Officer frames the underlying principle as "solved once, for everyone": compliance controls, guardrails, and evaluation standards built once and applied across the portfolio, rather than reconstructed imperfectly by each team operating in isolation. Across studies, the average ROI for firms that successfully move AI from pilots to production-scale processes runs at roughly 1.7x, with cost savings of 26 to 31% across supply chain, finance, and people operations.

The CoE model doesn't produce these gains by being clever. It produces them by preventing the same work from happening twice. That's a less exciting story than a breakthrough model release, but it's the one that shows up in the financials.

The measurements that actually capture reuse value versus the ones that don't

The default AI scorecard — licenses purchased and seats activated — measures spending. It cannot detect whether five teams rebuilt the same skill or whether any published skill was ever reused. It's a purchasing metric dressed up as a value metric, and organizations that rely on it exclusively are flying blind on the question that actually matters.

The metrics that surface reuse value are different in kind. Prompt reuse rate tracks how often a published skill is called rather than rebuilt. Time saved, verified against a measured baseline rather than self-reported, provides an auditable delta. Active usage frequency reveals whether a shared skill is actually in circulation or quietly archived. Reduction in repetitive development work captures the upstream cost that output-quality metrics miss entirely.

Development velocity is the most direct proxy for duplicate elimination. Average time from concept to production, and reduction in setup time per new project, translate directly into labor hours recovered. The QuantumBlack 55% reduction in asset discovery and use steps is exactly this metric expressed in time.

Registry health requires its own set of indicators, and most organizations haven't defined them yet. How many published skills are actively called by more than one team? How often does a team that discovers an existing skill adopt it rather than build parallel? How quickly does a skill reach its second adopter after initial publication? That last number, the gap between publication and second adoption, is a leading indicator of both registry health and organizational trust in shared infrastructure. A skill that sits unused for six months after publication isn't a registry problem — it's a discoverability and culture problem, and measurement surfaces it.

The Lumen template is replicable: name the redundant activity, measure its current fully-loaded cost, design the shared solution, measure the delta. Before and after. That structure is what makes a claimed saving auditable rather than aspirational.

Why the savings are real but not automatic, and what undermines them

The savings are documented. They are also not self-executing, and pretending otherwise is how organizations end up with registries nobody uses and CoEs that issue quarterly reports nobody reads.

McKinsey found that 42% of companies abandoned most of their AI initiatives in 2025. Fragmented data and technology ecosystems, not model quality, were the primary cause of failure. An IBM global study of 2,000 CEOs in 2025 found that only 25% of AI initiatives delivered expected ROI, and only 16% scaled successfully. Those are the actual denominators against which projected savings must be held. Most AI programs run without rigorous reuse governance will fail to scale — that's the base rate, not pessimism.

The METR study from mid-2025 adds a counterintuitive data point worth sitting with: 16 experienced open-source developers took 19% longer on coding tasks when using AI tools, despite perceiving themselves to be 20% faster. Shared skills eliminate redundant building, but they cannot fix an adoption pattern in which expert users resist, misuse, or misjudge the tools. Those developers were moving at the speed of confidence rather than the speed of code. No registry fixes that, and it's worth being honest about the ceiling.

Accenture Federal Services put change management spend at nine times the technology cost for enterprise AI deployments as of May 2026. Nine times. A registry that nobody consults produces exactly zero reuse savings regardless of how well-architected it is. The human adoption cost must be modeled alongside the technical build cost, and it is structurally larger. Most organizations model the technology cost carefully and estimate the change management cost loosely, which is precisely backwards given those proportions. I've watched this mistake get made repeatedly, by smart people, on programs with real budget.

Brynjolfsson's study of 5,179 customer service agents provides further calibration: AI assistance improved weaker performers' output by up to 34%, while top performers saw minimal gains. A measurement framework that assumes uniform adoption will overestimate savings in departments where the highest performers dominate. Governance of the registry, active tracking of reuse, verification that claimed savings are realized rather than assumed: these must be somebody's explicit job responsibility, not a standing agenda item at a monthly steering committee.

How shared registries and MCP-based architectures make reuse the default path

A catalog is passive. Anyone who has watched a SharePoint site accumulate 4,000 documents that nobody searches understands the distinction viscerally. The structural goal of a shared registry isn't documentation; it's making reuse the path of least resistance, easier than rebuilding, so that the default behavior at the start of any new AI project is to check what exists rather than start from scratch.

An effective AI CoE functions simultaneously as hub and coach. The hub provides common building blocks: reference architectures, curated datasets, evaluation methods, governance frameworks. The coach helps individual teams identify where shared assets apply to their specific use case and supports the handoff from discovery to deployment. A hub without the coach produces underused assets, impressive in a catalog and invisible in practice. A coach without the hub produces duplicated advice, well-intentioned and structurally unable to scale.

The Model Context Protocol is enabling a new category of reuse that goes beyond static catalogs. When teams can publish, discover, and invoke best-in-class skills through a standardized interface without rebuilding the underlying infrastructure, the economics of reuse change fundamentally. QuantumBlack's Brix MCP is the live proof point: over 700 components, 10 million lines of codified best practices, surfaced through a discoverability layer that reduced asset discovery and use steps by 55%. That's an active reduction in per-project setup cost, measurable in time and labor at every deployment.

Agentic deployments at DXC Technology and Rimini Street reduced complex workflow cycle times by 30 to 50%, roughly three times the efficiency gains of rule-based automation applied to the same categories, per TechTarget reporting from April 2026. Composable, shared agent skills compound gains in ways that single-tool automation doesn't, because each skill reused eliminates not just the rebuild but also the integration work, the evaluation overhead, and the governance negotiation that would have accompanied it.

One enabling feature that rarely gets discussed: permissions inheritance. When a shared skill automatically inherits access controls from its connected data sources, any team that discovers and adopts it can use it immediately, without renegotiating data access from scratch. Across dozens of skills and dozens of teams, that latency reduction is itself a measurable gain. Not glamorous, but real and recurring.

The metric the whole architecture points toward is this: how many skills were built once and used by more than one team, and what was the fully-loaded cost of the skills that weren't. That number requires someone to own the registry, track reuse, and treat the answer as a performance metric rather than an administrative afterthought — no sophisticated measurement system required. Most organizations are not yet doing this. The ones that are building a meaningful and durable cost advantage over the ones that aren't.

Sources

  1. worklytics.co

More in AI Productivity & ROI