Skip to main content

AI adoption you can measure

Most companies roll out AI assistants the same way: buy licenses, run a kickoff training, and hope. Three months later nobody can say who actually uses the tool, who is stuck, or whether it brings anything — the honest answer to "is it worth it?" is a feeling, not a number. The pain concentrates in three concrete problems. We show how to build a learning loop around AI adoption that attacks all three: it reads real usage telemetry, gives every developer a personal level and personalized advice, gives managers an aggregate view — and stays firmly on the right side of the line between measurement and surveillance.

Three real problems this solves — and who has them

1. You rolled out AI, but nobody can say what it brings

Who has it: an IT company with ~100 developers that rolled out an AI coding assistant for everyone. Licenses are paid, the kickoff training happened — and the only adoption signal management has is hallway anecdotes.

Adoption is nearly universal — 78% of organizations now use AI in at least one business function (McKinsey State of AI) — yet more than 80% report no material impact on enterprise earnings from generative AI, and only about 1% of leaders describe their rollout as mature (McKinsey). Tools in, impact invisible.

How we solve it: the assistant's real usage telemetry flows into one place, and a nightly batch turns it into per-developer statistics — prompts, tools and techniques used, activity over time — plus one aggregate picture for the whole company. "Is it worth it?" stops being a feeling: you see who works with it daily, who dropped off after week one, and how usage moves month over month.

2. One training, then a plateau

Who has it: the same company a quarter later. The enthusiasts sprinted ahead — and most of the team settled into using a powerful agent as a fancy autocomplete. Generic tips ("write better prompts!") help nobody, because every developer is stuck somewhere different.

Self-assessment makes it worse: in a randomized study, experienced developers using AI on familiar codebases estimated they were about 20% faster — while actually being 19% slower (METR). Without measurement, "we're doing great with AI" is exactly the kind of feeling that lies.

How we solve it: every night an AI coach reads each developer's actual usage and writes personal advice — what to try next, which habit to drop, which technique fits the work they actually do. The advice isn't fired blind: a second, independent pass evaluates every piece of advice and weak ones are rewritten before anyone sees them, and yesterday's advice is compared against what actually changed, so the coach keeps learning which recommendations work. Each developer also gets a level — apprentice, journeyman, master — computed from real signals, so progress has a name and a next step.

3. Measurement that doesn't become surveillance

Who has it: any company where "we'll measure how people use AI" raises an eyebrow — justifiably. Developer trust around AI tooling is already fragile — 46% of developers say they distrust the accuracy of AI output (Stack Overflow Developer Survey 2025) — and telemetry that reads like performance policing kills adoption faster than any missing feature.

How we solve it: privacy is designed in, not promised. The individual sees their own statistics, level and advice. The manager sees aggregate numbers and a coaching direction — never the content of prompts. Secrets (API keys, tokens) are redacted before any AI reads a prompt, and the methodology behind levels and scores is written down and approved by people before it counts. Measurement the team can read about openly is measurement the team will accept.

These are industry benchmark figures (McKinsey, METR, Stack Overflow). What your own adoption actually looks like shows up fast once telemetry is on the loop — that's what the free diagnostic maps.

The idea: a learning loop around AI adoption

All three problems share one cause: the rollout was an event, not a process. The fix is to wrap a controlled loop around how your team uses AI.

The goal is not to rank people or police prompts. The goal is that every developer gets a personal path to getting better, managers see whether the investment moves, and the system itself learns which advice actually helps.

Nothing heavy runs while people work. Everything is precomputed overnight; during the day developers and managers just open fresh, fast dashboards.

How it works

  • Signals — the AI assistant's telemetry (prompts, tools, techniques, sessions) lands in one store, tied to the user, with a 30-day window.
  • Nightly analysis — a batch computes per-developer statistics and a level from real signals, consistently for everyone at once.
  • Advice with quality control — AI writes the advice, an independent evaluation scores it, weak advice gets rewritten (generate → evaluate → improve). No raw first drafts reach people.
  • Learning from outcomes — yesterday's advice is checked against today's behavior; what worked shapes what's recommended next. The coach improves the same way the team does.
  • Human control — the levels methodology is approved by people before it counts, and managers get coaching directions, not transcripts.

Who sees what:

WhoWhat they see
Developerown statistics, level (apprentice → journeyman → master), personal advice
Managerteam levels and trends, "what to improve" per person — never prompt contents
Companyone adoption picture: active users, usage trends, where teams are stuck

Real numbers from the deployment described above (first weeks): 85 developers with history and 786 pieces of personal advice. Levels settled at 44% apprentices, 39% journeymen and 18% masters. And because the loop compares every piece of advice against the following week's behavior, it already knows which advice lands: the recommendation to try sub-agents was followed by real usage in 78% of measurable cases (21 of 27), broadening the palette of techniques in 30%, and planning mode in 23%. Early correlations, not proof of causation — the point is that the system knows these numbers at all, and tomorrow's advice already leans toward what works.

Who it makes sense for

This makes sense for companies where an AI tool is already rolled out — or about to be — to dozens of people, so the investment is real. Especially if:

  • licenses are bought but usage is a black box,
  • one training happened and nothing has changed since,
  • power users sprint ahead while the middle of the team stalls,
  • management asks "what does AI actually bring us?" and gets anecdotes,
  • measuring individuals raises justified privacy concerns.

The point isn't to score people. It's that getting better with AI stops being left to chance — each person knows their next step, and the company finally sees whether the investment moves.

The payback is concrete: the license spend you already committed starts compounding instead of stalling, the middle of the team moves instead of just the enthusiasts, and "is AI worth it?" gets a number instead of a shrug. One team with telemetry is enough to see it.

Measuring adoption goes hand in hand with the agents you actually deploy — see AI agents for business: what they actually do and where to start, and our AI solutions overview shows where both fit.

Want to see how your team actually uses AI? Get a free diagnostic — we map one team's real usage and show you exactly what we'd build, the impact and the cost. No obligation.