Most AI ROI questions get answered with the wrong numbers. Leaders look at how many people have a license and how many messages they send, see both rising, and conclude the investment is working. But seats and volume measure activity, not return. A company can grow usage every month and build no durable leverage if the best methods never become shared, approved assets.
The honest way to think about AI adoption ROI is to ask whether the company is getting structurally better at the work, not just busier with AI. That shifts the measurement from activity to reuse: which repeated tasks became approved, shipped company skills, and how much time each run saves the people who do the work.
Why seat counts are not ROI
Activity metrics are easy to collect and easy to misread. A high message count can hide ten people doing the same task ten different ways. A well liked method can produce inconsistent answers that quietly create rework downstream. And the task with the highest potential return may show little activity today simply because nobody has published it as a skill yet. None of that shows up in a seat count.
There is a second trap: self-reported time savings. Surveys record enthusiasm, not behavior. People overestimate what they use and underestimate what they abandon. The honest correction is not surveillance; it is estimating with the people who do the work. Sit with the team, time the task before and after the skill, and agree on how often it recurs. An estimate made with five practitioners beats a survey blasted at five hundred.
The inputs that actually drive ROI
Return on AI adoption is created by a small set of real inputs, and each one is countable inside your own teams. The unit that makes them countable is the company skill: a repeated task published by the person who does it best, approved by a team lead, and shipped to the whole team's Claude.
| Vanity metric | ROI signal it should be replaced with |
|---|---|
| Seats licensed | Repeated tasks that became approved company skills. |
| Messages sent | How often the repeated task recurs, agreed with the team. |
| Tools deployed | Methods that spread beyond their author into the whole team's Claude. |
| Demos given | Owner coverage and freshness: every skill owned and current. |
| Self-reported time saved | Time saved per run, timed with the people who do the work. |
A worked example: the weekly operating report
Take one task and follow the money. A RevOps manager spends 45 minutes every Monday assembling the weekly operating report: pipeline movement, renewals at risk, support load, and the three numbers the exec team always asks about. She has a Claude method that gets it done in 10 minutes, but it is hers alone; the four other team leads still build their versions by hand.
She publishes the method with knacks. It comes back drafted as a skill with examples and tests, her VP approves it, and it ships to the other leads' Claude with zero setup, with the company skill library as its home and her as the named owner.
Now the arithmetic is simple and honest about what it is. Five people, one report each per week, 35 minutes saved per run: roughly 12 hours a month. knacks does not track usage, so the volume side is not read off a dashboard; it is agreed with the five leads who build the report, each of whom can confirm the cadence in one sentence. The VP approved the skill on a Tuesday; by the following Monday every lead was building the report the same way. When the exec team changes what it asks for, she updates the skill once and everyone's report changes with it.
A simple way to frame the return
You do not need a complex model to start. For each skill: time saved per run, times how often the task recurs across the team, minus the modest cost of maintaining the skill. knacks does not track usage, so the volume in the equation is estimated with the people who do the work, not read off a surveillance dashboard. That is a feature: the team that does the task is the most reliable source for how often it happens. The point is not a precise figure on day one; it is estimating the same way every month, with the same people, so improvement is visible and dead weight is undeniable.
What to measure
Three checks per quarter do most of the work. First, count the methods that became approved, shipped skills: repeated tasks that used to live in one person's chat and now live in everyone's Claude. Second, check owner coverage: every skill should have a named owner who answers for it. Third, check freshness: when was each skill last updated, and does the changelog keep pace with pricing, product, and policy changes. All three come straight from the skill library, and none of them requires watching anyone work. No capture and no usage tracking, ever. knacks never sees chats, screens, or who runs what.
How knacks helps
knacks turns one person's AI method into an approved skill the whole team runs in Claude, and it runs the cycle as the knacks loop (Publish, Approve, Ship, Use, Improve). Anyone publishes repeated work in plain English; knacks drafts the skill with examples and tests; the team lead approves; the skill ships to the whole team's Claude, with the skill library as its home, stored as plain markdown in a GitHub repository your company owns; the team runs it in Claude on web, desktop, and Claude Code; and when the owner improves it, everyone is on the latest version immediately, with a one line changelog. Instead of reporting seats and messages, you report which repeated tasks became approved, shipped skills and the time saved estimates you made with the teams that run them. Next on the roadmap: tests that re-run on every new Claude model, so approved skills get checked when the model underneath them changes.
Measure AI by reuse, not seats.
Book a walkthrough and we will pick one repeated task where published reuse creates a return you can estimate and defend.
Book a walkthrough