Report 08 · Measurement series
The L&D Metrics That Matter
A working dictionary of the twelve L&D metrics worth computing, sorted into three tiers, with the formula, the honest reading, and the trap each one hides.
The short version
A headline metric measures a change in people or money, not a change in activity.Most L&D dashboards measure motion: enrollments, hours, completions, and smile-sheet scores. The field knows it. 72% of organizations track engagement as their learning-impact metric and 64% track retention, while few ever reach a business outcome, and completion-rate dashboards quietly reward the wrong design choices.
This guide sorts twelve metrics into three tiers (headline, diagnostic, and retired), each with its formula and its failure mode. The organizing rule is one line: a headline metric must measure a change in people or money, not a change in activity.
Evidence: 1 LinkedIn Learning (2025)
Organizations using engagement as their learning-impact measure.
1 LinkedIn Learning, 2025Companies using completion rate as their primary effectiveness metric.
3 Deloitte and Compliance Week, 2025Single-session content lost within a week without reinforcement, on the forgetting curve.
7 Ebbinghaus forgetting-curve and spacing-effect research, 2025Survey scores rose 12% and turnover rose 8% once a metric was tied to manager pay.
6 MIT Sloan Management Review, 2025What is inside
The question is no longer how much did our people train this year. It is whether the four honest numbers moved this month.
Motion is not outcome
Why the tiers exist, and the one rule that sorts them.Enrollments, hours, completions, and satisfaction scores all measure that something happened, not that anything changed. They are cheap to collect and comfortable to present, which is exactly why dashboards fill up with them. When 72% of organizations report engagement as their impact measure, the honest translation is that most L&D functions still cannot point to a change in capability or cost.
The fix is a filing system, not a new tool. Sort every metric you compute into three tiers by what it can honestly claim. Tier 1 is the headline four that belong on the exec slide. Tier 2 is the diagnostic six that explain why the headline four moved. Tier 3 is the pair you keep as telemetry and never present as achievement.
Evidence: 1 LinkedIn Learning (2025)
People or money, never activity
A headline metric must measure a change in people or a change in money. Active learner rate, verified skill delta, cost per active learner, and time to close a gap each pass that test. Hours and completions do not, because you can move them without anyone getting better or anything getting cheaper.
The three tiers at a glance
Freeze the definitions first. A trend is only trustworthy when the formula behind it has not changed.
Tier 1, the headline four
Four numbers, one slide, every month.These four survive contact with a CFO because each one is a ratio of something real. Compute them the same way every month and the trend becomes the product. Define active as completing at least one real unit of learning, never as a login: ghost logins flatter the number and poison everything computed from it.
Adoption, and the denominator under every other claim. Count someone active only when they completed at least one real unit. The trap is defining active as a login.
Capability change, the one number L&D alone can own. Report it per pivotal skill, averaged across the cohort. The trap is testing right after training, when recall is easy and predicts nothing. Check at 30 days or more.
Efficiency, honestly denominated. It falls as adoption rises, which makes it finance's favorite line. The trap is quietly switching to cost per seat when utilization is embarrassing.
Velocity, or how fast the organization converts a flagged gap into a verified capability. The trap is counting only the gaps you chose to close, which turns the metric into a highlight reel.
If you present only one thing to leadership, present verified skill delta. It is the only number that says people got better, not just busier.
Tier 2, the diagnostic six
The metrics that explain every move in the headline four.Diagnostics explain movement in the headline four. They belong in the L&D team's working view, not on the exec slide, because promoting them confuses activity with outcome. Completion rate is the clearest example: read per unit it is a useful content thermometer, but aggregated into a headline it just reports your format mix.
Evidence: 2 eLearning Industry (2025)
| Metric | Formula | What a bad trend tells you |
|---|---|---|
| Completion rate, per unit | completions ÷ starts | Content triage only. A unit under about 70% is too long, mistargeted, or badly timed. Traditional courses run 20–30% and micro formats about 80%, so the mix drives the number, not the quality. |
| Streak or frequency | median active days per week | Habit health. A fall here leads a fall in active learner rate by about a month, so it is your earliest warning light. |
| Retrieval accuracy | percent correct on spaced re-checks | Retention against the forgetting curve. Falling accuracy on 30-day re-checks means content is not resurfacing enough, a spacing bug rather than a people problem. |
| Manager ritual rate | teams running rituals ÷ all teams | The adoption engine. When this slips, learner activity follows within weeks, so coach managers rather than learners. |
| Coverage of pivotal skills | pivotal skills with 2 or more holders ÷ 10 | Risk posture. Every uncovered pivotal skill is a single-holder risk with a name attached, which feeds the exec conversation each quarter. |
| Seat utilization | active seats ÷ paid seats | Procurement hygiene. Below about 60%, renegotiate before renewing, and say so out loud, because volunteering the number buys credibility. |
Directional completion ranges are vendor-adjacent industry figures, not a single controlled study.
Evidence: 2 eLearning Industry (2025) / 4 Training Magazine and TalentLMS (2026)
A diagnostic visits, then goes home
A Tier 2 metric may appear on the exec slide only while it explains a Tier 1 move, for example when skill deltas dipped because ritual rate halved during a reorg. It presents the diagnosis and then goes back to the working view.
Tier 3, retire from headlines
Two numbers to keep as telemetry and never present as impact.Two of the most common headline metrics measure the wrong thing so reliably that they should come off the slide entirely. Keep both in the background for planning and content triage, but stop presenting either as a result.
Hours of learning logged
Hours measure sitting, and the sitting is compromised: 70% of employees multitask through training. Worse, hours reward the wrong design, because stretching content pads the metric. Keep hours for capacity planning and never present them as achievement.
Satisfaction averages, the smile sheets
Reaction scores correlate weakly with learning and behavior change, and people tend to rate fluent, easy sessions above the hard practice that actually works. Keep per-unit scores to catch broken content, because a sudden 2.1 means something, but drop the org-wide average, which means almost nothing.
Evidence: 4 Training Magazine and TalentLMS (2026)
Useful to run, wrong to headline
A metric earns headline space only when moving it requires someone to get better or something to get cheaper. Hours and smile sheets both move without either, which is exactly why they belong in telemetry and not on the executive slide.
A metric you can improve by making content longer is not a quality metric. It is a billing metric.
Name the denominator, compare cohorts
Two of the four rules for reading the numbers without fooling yourself.Name the denominator every time
Per-employee, per-learner, and per-active-learner can differ by more than 40% in a typical organization, so the same spend produces three different headline numbers. State the denominator in every axis label. That single discipline is what separates your deck from every vendor deck.
Compare cohorts, announced in advance
Pick trained against untrained on one operational metric, pre-register it, and hold the definition steady each quarter. Three quarters pointing the same way is a fundable pattern. A correlation you noticed afterwards is just a story.
Evidence: 5 ATD and Training Magazine (2026)
The denominator is not a footnote, it is the argument. A number that looks strong per active learner can look weak per employee, and a room that spots the switch stops trusting the rest of the slide. Choosing one denominator, labelling it, and keeping it is the cheapest credibility you can buy.
Evidence: 5 ATD and Training Magazine (2026)
Pick one denominator, label it on every chart, and keep it for a year. Consistency reads as honesty, and honesty is what gets the budget.
Guard against gaming and decay
The other two rules, and the reason a day-0 score proves nothing.Never pay anyone on a learning metric
When survey scores were tied to manager compensation, scores rose 12% while voluntary turnover rose 8%, so the metric stopped measuring reality the moment money was attached. Report team learning data to managers and bonus no one on it.
Evidence: 6 MIT Sloan Management Review (2025)
The second guard is against decay. A skill check on the last day of a module measures short-term memory, not capability, because without reinforcement up to 90% of a single session is gone within a week. The honest instrument is the 30-day spaced re-check, and conveniently it is the one a spaced platform produces for free.
Put the two guards together and the discipline is simple. Do not attach a learning number to anyone's pay, and do not claim a skill gain until it has survived a month. Both rules cost nothing and both remove the two easiest ways to fool yourself.
Evidence: 7 Ebbinghaus forgetting-curve and spacing-effect research (2025)
Claim only what survives 30 days, and compensate no one on a number you want to stay honest.
Is your dashboard honest?
Eight questions, a scoring key, and the one-page version.Check every statement that is true today.
| Score | Band | What it means |
|---|---|---|
| 0–2 | Activity theater | The dashboard flatters and informs no one. Start by renaming the denominators. |
| 3–4 | Half honest | Good instincts, soft definitions. Freeze the formulas and add the 30-day re-check. |
| 5–6 | Instrumented | The tiers exist. Add the pre-registered cohort pairing and hold the line on gaming. |
| 7–8 | Trusted source | Finance quotes your numbers unprompted, which is the endgame. |
Takeaways, the one-page version
Evidence: 5 ATD and Training Magazine (2026) / 6 MIT Sloan Management Review (2025)
This week, rebuild the exec slide around the headline four and add a denominator to every chart. This quarter, ship three identical monthly readouts and report your first time-to-close-gap median.
Sources and method
Every external numeric claim in this report points to one of these 2024 to 2026 sources. Forecasts and self-reported surveys are labelled so they are not mistaken for causal proof.
Workplace Learning Report 2025
Survey of 937 L&D and HR professionals and 679 learners. 72% track engagement and 64% track retention as their learning-impact measures, while few reach business-outcome measurement.
learning.linkedin.comPractitioner analyses of completion and impact
Practitioner reviews (McPheat 2025 and LMS metrics analyses) that completion does not equal impact, and that low per-unit completion signals content length or targeting problems.
elearningindustry.comCompliance training effectiveness survey
50% of companies use completion rate as their primary training-effectiveness measure.
complianceweek.com2025 Training Industry Report and TalentLMS 2026
70% of employees multitask during training. Traditional online course completion runs 20–30% versus about 80% for micro formats. Directional, vendor-adjacent industry figures.
trainingmag.comState of the Industry 2026 and 2025 methodology notes
Per-employee ($846) versus per-learner ($874) denominators, and per-active-learner differs again. Utilization gaps of more than 40% are typical.
td.org140-organization study on metrics and compensation
When survey scores were tied to manager compensation, scores rose 12% while voluntary turnover rose 8% (reported 2023).
sloanreview.mit.eduForgetting curve and spaced retrieval literature
Up to 90% of single-session content is lost within a week without reinforcement, and spaced retrieval is the honest measurement instrument.
en.wikipedia.org