The question from management is not whether the team is happy with AI, but whether the money came back. Answering it requires a baseline taken before you start, adoption metrics kept separate from impact metrics, an honest formula, and the discipline to count anything you didn't measure as zero.

Short answer: Measuring the return on AI requires separating adoption (active users, frequency, depth) from impact (business metrics by department) and having a baseline from before the kickoff. Return is calculated from hours freed up times a reinvestment factor, plus incremental margin and avoided costs, minus total cost. Audit logs show events, not the content of chats.

In this guide
  1. What does measuring the return on AI training really mean?
  2. How do you build the baseline before you start?
  3. Which adoption metrics should you track?
  4. Which impact metrics apply to each department?
  5. What can you actually see in the admin tools, and what can't you?
  6. How is the return calculated? The formula, step by step
  7. A complete worked calculation (illustrative)
  8. What does management ask for, and in what format?
  9. How do you present the results on a single slide?
  10. Measurement traps that ruin a report
  11. How do you turn saved time into real value?
  12. An 8-question internal survey
  13. At what rhythm do you review all this?
  14. Tips that don't come in the manual
  15. What to do this week

What does measuring the return on AI training really mean?

It means two different things that almost everybody mixes up: adoption (how many people use it, how often and for what) and impact (what changed in the business because they used it). Adoption is measured in weeks and it's easy. Impact is measured in quarters and it's hard, because it requires having written down how things stood beforehand.

The confusion is costly in both directions. A report that only brings adoption —"80% of the team uses it every week"— doesn't answer management's only question, which is whether the money came back. And a report that only brings impact with no adoption —"we saved 900 hours"— doesn't survive the first uncomfortable question: who saved them, and how do you know?

There is a third category worth declaring from the outset: what you will not be able to measure. The quality of a proposal, the relief of someone who no longer does repetitive work, the judgment the team gained for catching an error. That gets collected through a survey and with examples, is presented as qualitative evidence, and is labeled as such. Mixing the qualitative with the quantitative in the same figure is the fastest way to lose credibility in a committee.

Anthropic's official documentation on evaluating results gives a criterion that works just as well inside a company: success criteria have to be specific, measurable, achievable and relevant, and they are best defined before you build anything (evaluations guide, accessed September 2026). Translated to your situation: if you can't write down today which number you're going to look at 90 days from now and where you're going to get it, you don't have a measurement plan yet.

How do you build the baseline before you start?

By picking few tasks, measuring them properly and writing them down before anyone touches the tool. It's the step people skip most often, and without it any later report ends up being a well-written anecdote.

A useful baseline has four data points per task: how long it takes (measured, not remembered), how many times a month it's done, who does it and what error or rework it generates. With that you can build any later calculation. Without it, you can't.

Baseline template (one row per task, three tasks per department)

Rule: if a data point has no source, it gets recorded as "not measured" and counts as zero in the final calculation. Never round upward.

Two field warnings. First: don't ask how long it takes, measure it. People's memory of their own timings is systematically optimistic on tasks they enjoy and pessimistic on the ones they don't. Timing five real runs takes two days and changes the entire report. Second: measure volume too, because almost all the value is there. Saving 20 minutes on something done twice a month isn't a business case; saving 2 minutes on something done 900 times a month is.

If you're about to launch a training program, this is week zero: the baseline is taken before the first session. The full calendar is in the post in this series on the eight-week training program, and the ten-question diagnostic it includes is the other half of the baseline.

Which adoption metrics should you track?

Five, and no more than that in the first quarter. More metrics don't give you more information: they give you more arguments about metrics.

MetricHow it's definedWhere it comes fromWarning sign
Weekly active usersPeople who used it at least one day in the last seven, over contracted seatsThe plan's admin dashboard, or a count reported by each departmentBelow 40% of seats after the first month
FrequencyActive days per person per monthSame source; supplement it with the internal surveyAn average below 4 days a month: that's anecdotal use
DepthNumber of distinct use cases per person and whether they work with their own documents or only with standalone chatSurvey and a review of the projects createdEveryone using it only to draft emails
BreadthDepartments with at least one case in production, not in testingCase inventory, reviewed monthlyOne department accounts for more than 70% of use
RetentionActive users in week 12 divided by active users in week 2Comparison of the two countsBelow 0.6: the novelty effect wore off and no habit was left

The most-used and least-useful metric is the number of messages or conversations. It rewards whoever types the most, not whoever gets things done; it rises on its own when someone is wrestling with a badly framed task and falls when the team learns to ask well on the first try. If management asks for it, present it as context and never as a success indicator.

Depth is the metric that best predicts real impact, and it's measured with two questions: does this person work with their own documents and projects, or only with a blank chat? How many distinct tasks did they handle this month? The jump in value happens when the answer to the first one is "yes". How to build that layer is covered in the post in this series on Projects and the knowledge base.

Want a hand applying this?

Free 30-minute working session. Uniamos is a remote AI automation agency (Austin and Mexico City) working over video calls and WhatsApp.

Book a free call →

Which impact metrics apply to each department?

The ones that already exist in your operation. The rule is not to invent new indicators for AI: you use the business's own, which already have history and are already discussed in committee.

DepartmentImpact metricData sourceHow often
SalesTime to prepare a proposal; proposals sent per rep; response time to an inbound requestCRM and send logMonthly
MarketingPieces published per month; cost per piece; hours of outside vendor workEditorial calendar and invoicesMonthly
Customer serviceFirst response time; queries resolved without escalation; reopened ticketsTicketing system or shared inboxWeekly
Administration and financeDays to monthly close; reworks due to errors; reconciliation hoursERP and close logMonthly
OperationsIncidents documented with an assigned cause; time to prepare reports and write-upsIncident systemMonthly
ProcurementHours spent comparing bids; lead time from request to purchase orderERPQuarterly
Human resourcesDays until a new hire's first contribution; time to fill a vacancyRecruiting logQuarterly

Three criteria for choosing yours. The data should already be recorded today with no extra work: if you have to set up a system in order to measure, the cost of measuring eats the savings. It should be debatable but not manipulable by whoever reports it. And it should have three months of prior history, so you can separate your program from seasonality; if there isn't any, measure a non-participating department in parallel.

A frequent situation at small and mid-sized companies in Mexico and Paraguay: the data exists, but it's scattered across spreadsheets, WhatsApp and the manager's head. There the first project isn't measuring AI, it's getting the dashboard in order; that groundwork is what we describe in dashboards and analytics. If the sales department has nowhere to pull volume from, a CRM with a free tier like HubSpot or a low-cost one like Zoho CRM solves the data-source problem before any AI tool does.

What can you actually see in the admin tools, and what can't you?

Less than management imagines and more than the team fears. It's worth knowing precisely, because this is the point where people over-promise most in a meeting.

ToolPlanWhat it showsWhat it does NOT show
Audit logsEnterprise onlyMore than 30 event types from the last 180 days: authentication, member additions and removals, project creation, deletion and visibility changes, file uploads, SSO and domain configuration changes. They are exported by the Owner or Primary Owner from the organization settings, and the download link expires in 24 hoursNeither the title nor the content of chats or projects: only their identifiers. It is no use for reading what someone wrote
Usage analyticsEnterpriseAggregate usage metrics for the organizationIt is not an individual productivity ranking or a surveillance tool
Compliance APIEnterprise and platform customersProgrammatic extraction of activity events, conversation data and file content, including Claude Code and Cowork. Only the Primary Owner can enable itIt is not designed for productivity metrics, but for compliance and retention. Its use must be stated in your internal policy
Spend controlsTeam and EnterpriseSpending limits per organization and per userIt attributes no value: it tells you how much is spent, not what was achieved
Account usage panelAll paid plansEach user's own consumption in their account settingsAnthropic does not publish absolute message counts per plan; limits are expressed as multiples and windows

Sources, all accessed in September 2026: audit logs, Compliance API, Enterprise plan, Team plan and roles and permissions. It's also worth reviewing custom retention —on Enterprise you can set your own retention periods for chats and projects with a minimum of 30 days, and by default data is retained indefinitely— and the policy on data use for training, which in commercial products is not to use your inputs or outputs by default.

The practical consequence: the admin dashboard gives you adoption, not impact. It tells you how many people log in and how often, not whether the proposal came out better or whether the monthly close went from eight days to five. That side you bring from your own systems.

And a rule for coexistence: announce in writing what is measured and what isn't before you measure anything. If the team suspects the logs are there to monitor conversations, usage drops within two weeks. How to write it up is in the post on use policy and governance. This is general guidance and not legal advice: consult your legal team.

How is the return calculated? The formula, step by step

With a simple formula and three built-in doses of honesty. The structure is that of any return on investment, but with one factor almost nobody includes and that is what separates a credible report from a brochure.

Return formula for an AI program

Return (%) = (Annualized benefit − Total annual cost) ÷ Total annual cost × 100

Annualized benefit = Hours freed up per year × Reinvestment factor × Fully loaded hourly cost + Margin on attributable incremental revenue + External costs avoided.

Hours freed up per year = (Time before − Time after) × Monthly volume × 12, summed across every measured task.

Reinvestment factor = the share of freed-up time that turns into valuable work. Between 0 and 1. If you don't know it, use 0.5 and declare it as an assumption.

Total annual cost = licenses + training hours valued + coordination hours + integrations or development + internal support.

Honesty number one: the reinvestment factor. Twenty minutes freed up are not automatically twenty minutes of value: they are if that person spends them on something that mattered and wasn't getting done. With a backlog of pending work the factor is high (0.7-0.9); when the freed-up time gets scattered, it's low (0.2-0.4). Ask about it in the survey and use the answer.

Honesty number two: what isn't measured counts as zero. If you didn't measure a department, don't estimate it. A report with three solid cases and eight blank departments is infinitely more defensible than one with eleven approximate figures.

Honesty number three: full costs. Licenses are the small part. With prices as of September 2026 —worth confirming on their site because they vary by region— the Team plan costs USD 25 per seat per month, 20 with annual billing (claude.com/pricing). Training hours for twenty people exceed that amount at almost any company. And if you also use the API to automate something, consumption is calculated separately at the per-million-token rates, and there prompt caching changes the bill a lot: a cache read costs a fraction of base input. If nobody on the team understands why one task costs more than another, the free tutorials on context and tokens at Claude Academy settle the question in half an hour.

A complete worked calculation (illustrative)

Everything that follows is an illustrative example with made-up figures to show the method. It is not a real case or a promise of results. Replace every number with your own.

Example company: a distributor with 40 people, 18 trained in the program, three departments with a measured baseline. Assumed average fully loaded hourly cost: USD 14.

Measured caseBeforeAfterMonthly volumeHours freed up per month
Prepare the standard proposal (sales)45 min18 min70 proposals31.5 h
Answer a frequent query (customer service)6 min4 min900 queries30 h
Variance report (administration)40 h/month26 h/month1 cycle14 h
Total75.5 h

Step 1. Hours freed up per year: 75.5 × 12 = 906 hours.

Step 2. Reinvestment factor. The internal survey says almost all of the sales time gets reinvested (there is a queue of customers who haven't been contacted) and the administration time gets scattered. A global 0.6 is applied and declared as an assumption: 906 × 0.6 = 543.6 effective hours.

Step 3. Annualized benefit: 543.6 × USD 14 = USD 7,610.

Step 4. Total first-year cost: licenses 18 × 25 × 12 = USD 5,400; training 18 people × 16 hours × USD 14 = USD 4,032; coordination 60 hours × USD 20 = USD 1,200. Total USD 10,632.

Step 5. First-year return: (7,610 − 10,632) ÷ 10,632 × 100 = −28%.

Step 6. Second year, with no training cost: (7,610 − 5,400) ÷ 5,400 × 100 = +41%.

This result is, deliberately, the one almost nobody shows you: a program that only shaves off minutes usually comes out negative in the first year and positive from the second year onward. Knowing that keeps you from promising what isn't going to happen.

Scenario B, the one that does change the calculation. Suppose the sales team, by responding faster, closes one additional deal a month worth USD 3,000 at a 25% margin: USD 750 a month, USD 9,000 a year. If you count that margin, you can't also count the 31.5 sales hours, because that would be the same benefit counted twice. The calculation becomes: 9,000 + (44 h/month × 12 × 0.6 × 14) = 9,000 + 4,435 = USD 13,435 in benefit against USD 10,632 in cost, that is +26% in the first year. The lesson is the usual one: the cases that pay for a program are not the ones that save minutes, they are the ones that change volume, lead time or external spend.

What does management ask for, and in what format?

Always the same six things, even if every committee phrases them differently. If you bring them ready, the meeting takes twenty minutes; if you don't, it takes an hour and ends with no decision.

  1. What it cost, all in. Licenses, hours and any development. One annualized figure with the breakdown underneath.
  2. How many people actually use it. Weekly actives over contracted seats, not people trained.
  3. What changed, with the before number next to it. Two or three cases, not twelve.
  4. What assumptions sit behind it. Reinvestment factor, hourly cost and what went unmeasured. Declaring it yourself before they ask changes the conversation.
  5. What risk was opened up and how it's covered. Data, vendor dependency, output quality.
  6. What decision you're asking for. Continue as is, expand to two more departments, stop. With an amount and a date.

On format: one slide, one appendix and no adjectives. Management doesn't need to understand how the tool works; it needs to decide whether to put in more money. And there's one question that always comes and is worth answering before it's asked: "does this let us cut headcount?". The honest answer at most small and mid-sized companies is no, not in the short term —what changes is the capacity to do more with the same team— and if the answer were different, that would have to be said too.

If the committee wants to see the detail of how the program under review was built, the natural appendix is the eight-week training plan; and if what's on the table is the next rollout phase, the phased framework is in the pillar post on rolling out Claude at your company.

How do you present the results on a single slide?

With six fixed blocks and no decorative charts. This structure works the same in a committee of five as in a board meeting.

Structure of the single results slide

Slide rules: no figure without a source; no percentage without the absolute number beside it; no other company's case; zero screenshots of conversations.

Two presentation details that change the outcome of the meeting. First: put the "what we didn't measure" block before the money block, not at the end. Whoever declares their limits first gets believed on everything else. Second: bring the appendix with the detail of each case and don't project it. If someone asks for depth, you have it; if not, you don't burn the committee's time.

If your company already runs on dashboards, the natural move is for these six blocks to stop being a slide and become a fixed view that updates itself. It saves rebuilding the same report every quarter and, along the way, forces you to freeze the definition of each metric.

Measurement traps that ruin a report

Seven of them, and all seven slip in on their own if nobody watches for them. Knowing them by name helps you catch them in the draft before management catches them in the meeting.

There is one additional, subtler trap that shows up in the second quarter: measuring what's easy to measure instead of what matters. The number of messages is easy; response time to a customer, not so much. If your report is full of activity metrics and empty of outcome metrics, the problem isn't the tool, it's the design of the measurement.

Want a hand applying this?

Free 30-minute working session. Uniamos is a remote AI automation agency (Austin and Mexico City) working over video calls and WhatsApp.

Book a free call →

How do you turn saved time into real value?

By deciding in advance what it's going to be used for. Freed-up time doesn't convert itself: if nobody assigns it, the day swallows it. This is the part of the job that belongs to the department heads, not to the trainer or the vendor.

Three approaches that work, in order of ease. The first is reassigning it to the backlog: if the sales team has forty accounts nobody has contacted, the freed-up hours have an obvious destination and the reinvestment factor approaches 1. The second is absorbing growth without hiring: same team, more volume; it's measured in orders, tickets or case files per person. The third is replacing outside spend: agency, design or translation hours that stop being invoiced, and it's the one that defends itself best in committee because it shows up as an invoice that no longer arrives.

What doesn't work is leaving it to individual discretion and waiting for a miracle. In practice, freed-up time with no destination gets divided between more meetings and more email, and the following quarter nobody can find the savings anywhere.

Reinvestment agreement by department (one line per case, signed in week zero)

When the destination for the time is "doing more of the same but better", it's worth being honest and booking it as a quality improvement, not as monetary savings. And when the destination is automating the rest of the process, you're no longer in training but in a process automation project, with its own budget and its own calculation.

An 8-question internal survey

It goes out in week 12 and is repeated every quarter. Eight questions, three minutes, responses not anonymous but used in aggregate and declared as such. It serves three purposes: estimating the reinvestment factor, surfacing cases you didn't know about, and anticipating attrition.

Quarterly adoption survey (copy and paste)

  1. In a normal month, how many days do you use it for work? (0 / 1-3 / 4-10 / 11-20 / more than 20)
  2. Which two specific tasks do you use it for most? (open)
  3. Do you work with your own documents or with department projects, or only with standalone conversations? (projects / individual documents / chat only)
  4. For those tasks, how long did they take you before and how long do they take you now? (two approximate figures)
  5. What have you used the time it frees up on? (backlog of pending work / more volume / new tasks / I don't notice it)
  6. How many times this month did you catch an error in an answer and correct it? (0 / 1-2 / 3-5 / more than 5)
  7. What stops you from using it more? (no time / I don't know how to apply it to my work / I don't trust it / access limitations / other)
  8. On a scale of 1 to 5, how much worse would your work get if it were taken away from you tomorrow?

How to use it. Question 5 gives you the reinvestment factor. Question 6 measures judgment, not failures: the higher it is, the better trained the team. Question 7 hands you next quarter's work plan. Question 8 is the best retention predictor there is and fits in a single row of the report.

Two rules to keep the survey honest. First: say explicitly that results are presented in aggregate and are not used to evaluate individuals; if that isn't believed, the answers turn diplomatic. Second: don't use question 4 as a direct input to the return calculation —it's self-reported— but as a clue for deciding which tasks deserve a properly timed measurement.

At what rhythm do you review all this?

At four different cadences with four different owners. The usual mistake is reviewing everything every month: it produces noise, premature decisions and fatigue.

CadenceWhoWhat gets reviewedWhat gets decided
Weekly (15 min)Coordinator and championsWho got left out, questions that are stuck, new casesOperational adjustments: access, one-off support
Monthly (45 min)Coordinator and department headsAdoption (actives, frequency), progress on the measured casesWhich case gets dropped and which gets expanded
Quarterly (90 min)ManagementThe single slide: adoption, impact, money, risksContinue, expand or stop, with an amount
AnnualManagement and ITContracted plan, underused licenses, data policy, retention and rolesRenewal, plan change, seat reallocation

The quarterly review is the only one that produces money decisions, which is why it's worth setting its dates on day one, in the calendar, with all four quarters loaded. A committee that gets convened when there's good news is not a tracking committee.

To prepare for the annual review it's worth first reading the Enterprise consumption guide and the article on usage credits on paid plans, which explain how spending behaves when the team goes past what the plan includes.

And there's one specific task that usually pays for the whole meeting: reviewing unused seats. People who changed roles, who left, or who never got started still occupy a license. Cross-check the seat list against the last quarter's actives and reassign. It's money you're already spending, and it can fund the next cohort of the program.

Tips that don't come in the manual

Ten things you learn from measuring several times that appear in no product guide.

  1. Measure three tasks, not thirty. A report with three timed cases always beats one with thirty estimates. Management doesn't buy coverage, it buys credibility.
  2. Time it in week zero and keep the sheet. With the date, the name of whoever measured and how. Six months from now, that sheet is the difference between a data point and an opinion.
  3. Declare the reinvestment factor before you calculate. If you pick it after seeing the result, it's no longer an assumption, it's an adjustment. Write it on the slide and let them argue about it: that argument earns you credibility.
  4. What isn't measured counts as zero. Turn this into a written rule of the report. It's the sentence that defuses "isn't this inflated?" before it's asked.
  5. Ask for volume before time. A 40-minute task done twice a month isn't a case; a 4-minute one done 900 times is. Volume decides where it's worth measuring.
  6. Use a control department. Leave a comparable team out of the program for the first quarter and compare. It costs nothing and solves the attribution problem in one stroke.
  7. Don't use the number of messages as a success indicator. It goes down when the team learns to ask well. If management asks for it, present it as context and explain why.
  8. Keep the adoption report separate from the money report. The first is monthly and operational; the second is quarterly and for the committee. Mixing them produces meetings where everything is discussed and nothing is decided.
  9. Review unused seats every quarter. It's the fastest saving of all and requires convincing nobody: if three licenses have been inactive for two months, reassign or cancel them.
  10. Keep two raw examples for every case. The before and after of a real, anonymized document. In a committee, two concrete examples do more than any percentage, and they carry the qualitative side that no table captures.
  11. Write down what went wrong too. A short section with errors caught and how they were fixed shows that the team has judgment. A report with no failures isn't believed.
  12. Freeze the definition of each metric in writing. "Active user" has to mean the same thing in January and in October. Half the apparent improvements between quarters are changes in definition.

What to do this week

Whether you're about to start or you've gone months without measuring, the kickoff is the same and it fits in one week.

  1. Monday. Pick three measurable tasks: high volume, data already recorded and a willing owner.
  2. Monday. Ask administration for the average fully loaded hourly cost of the roles involved.
  3. Tuesday. Time five real runs of each task. Not by asking: by measuring.
  4. Tuesday. Pull the monthly volume for the last three months from the CRM or the ERP and note the source.
  5. Wednesday. Define in writing the five adoption metrics and exactly what each one means.
  6. Wednesday. Check what you can actually see on your plan: if it isn't Enterprise, you won't have audit logs or organization analytics, and you'll be keeping the active-user count yourself.
  7. Thursday. Sign the reinvestment agreement with each department head: where the freed-up hours are going.
  8. Thursday. Schedule the year's four quarterly reviews with management, with a date and a time.
  9. Friday. Prepare the blank slide with the six blocks. Having it empty from day one keeps you from improvising in week 12.
  10. Friday. Write down and communicate what is going to be measured and what isn't. Before you measure anything.

If you want to stress-test the method with someone who has set it up before, at Uniamos we work on it alongside the training side in AI training with Claude, and you can write to us through contact. And if reviewing the cases shows you that what's missing isn't training but tidying up the process, the catalog by function in the post on use cases by department is a good place to look first.

Frequently asked questions

How long do you have to wait before measuring the return?

Ninety days at a minimum. Before week 12 everything you see is contaminated by the novelty effect. Adoption is measured from the first week, but the money calculation needs a closed quarter and a baseline from before the kickoff.

Do audit logs show what people write?

No. According to the official documentation, Enterprise audit logs cover more than 30 event types from the last 180 days, but they export neither the title nor the content of chats or projects, only their identifiers. For extracting content there is the Compliance API, which is a compliance tool and can only be enabled by the Primary Owner.

Which metrics can I see if I don't have an Enterprise plan?

Organization analytics and audit logs are Enterprise features. With Team you get spend controls and member administration, and each user sees their own consumption in their account settings. The active-user count is yours to keep, by department, complemented by the quarterly survey.

What reinvestment factor is reasonable?

If you haven't measured it, 0.5, declared as an assumption. Raise it to 0.7-0.9 when there is an obvious backlog of pending work, such as uncontacted accounts or a pile of tickets, and lower it to 0.2-0.4 when the freed-up time gets spread across small fragments throughout the day.

How do I avoid double counting?

By picking a single route for each benefit. If you count the margin from the additional sales a team produced, don't also count the hours that team saved: that's the same effect measured twice. Next to each case, write down which route you used and why.

Is it worth asking the team how much time they save?

It's worth it for direction, not for calculation. Self-reporting is optimistic in good faith. Use it to decide which tasks deserve a timed measurement and to estimate what the time gets reinvested in, which is something only the person living it knows.

What do I do if the return comes out negative in the first year?

Present it as it is and show the second year with no training cost. It's a normal result in programs that only shave off minutes. The useful conversation isn't how to dress it up, but which cases you would have to tackle to change volume, lead time or external spend, which is where the big return lives.

How do I measure it if the team was already using AI before the program?

By recording that prior use in the baseline, per person and per task. If you don't separate it out, you'll take credit for improvements that already existed. At companies where informal use is widespread, the program's first gain is usually control and consistency, not time, and it's worth saying so.

How often should the internal survey be repeated?

Every quarter, with the same eight questions and without changing their wording. The value lies in the comparison between quarters, and rewriting a single question is enough to lose it. If you need to add something, append it at the end and leave the originals untouched.

What about personal data when measuring adoption?

Always measure in aggregate and announce in writing what is being measured before you start. No adoption indicator needs to read conversations. This is general guidance and not legal advice: consult your legal team, and in Paraguay review the obligations of Law 7593 on personal data.

Related articles