What is the Kirkpatrick model?
The Kirkpatrick model is a four-level framework for evaluating training: Level 1 Reaction asks whether participants found the training engaging and relevant; Level 2 Learning asks whether knowledge or skill increased; Level 3 Behavior asks whether people apply it on the job; Level 4 Results asks whether an organizational outcome moved. Donald Kirkpatrick introduced the levels in 1959, and they remain the most widely used vocabulary in training evaluation.
The model is easy to recite and hard to run. The practitioner version of the problem: “We can quote our satisfaction scores from every session for three years, and I still can’t tell the board whether anyone works differently.” The levels are not the hard part — the record-keeping underneath them is, and that is what this page works through, level by level.
Key takeaways
- The four Kirkpatrick levels — Reaction, Learning, Behavior, Results — are a causal ladder: each level only means something read against the one below it.
- Most implementations stop at Levels 1–2, not from laziness but from infrastructure: Levels 3 and 4 require following the same person for months, and anonymous surveys cannot.
- Sopact calls the fix Continuous Kirkpatrick: the four levels run as one loop on one trainee record, under a persistent participant ID assigned at enrollment.
- The test that separates tools: ask to see one trainee’s enrollment baseline and their 90-day behavior evidence on the same screen, scored on the same rubric.
- New World Kirkpatrick, Phillips ROI, CIRO, and Brinkerhoff all refine the map — none of them removes the per-person join the upper levels demand.
What is Continuous Kirkpatrick?
Sopact calls the practice Continuous Kirkpatrick: the four Kirkpatrick levels run as one continuous loop on one trainee record — a persistent participant ID, assigned at enrollment, that carries the reaction pulse, the learning gain, the 90-day behavior evidence, and the tied result metric. Kirkpatrick supplies the level map; the Loop supplies the engine. The model stops being a year-end reporting framework and becomes the running structure of the program’s own data.
The differentiation is the data model, not the questionnaire. Form-centric survey tools treat every send as a fresh anonymous batch, so Level 1 and Level 3 can never land on the same person, and the ladder breaks between rungs however well the questions are written. A record-centric system never breaks the thread: the four levels become four views of one dataset. This page works the model itself; the full training lifecycle — needs assessment through delivery to ROI — is worked on training evaluation.
How the Kirkpatrick model evolved — and why practice never caught up
Donald Kirkpatrick published the four levels as journal articles in 1959 and consolidated them in Evaluating Training Programs (1994). In 2016, James and Wendy Kirkpatrick of Kirkpatrick Partners published the New World Kirkpatrick Model: plan backwards from Level 4, define leading indicators, and sustain Level 3 with Required Drivers — the reinforcement, monitoring, and accountability processes that keep new behavior alive after the course ends. The theory kept improving.
Tooling went the other way. The smile-sheet era (paper, then SurveyMonkey and Google Forms) made Level 1 free and made anonymous, disconnected responses the default. The LMS era (Cornerstone, Docebo, Moodle, TalentLMS) added completion rates and quiz scores — a Level 2 proxy that measures exposure and recall, not transfer. No mainstream tool owned Levels 3 and 4, because both require longitudinal, person-level records that neither a form tool nor an LMS keeps. The one test that separates the eras: ask to see one trainee’s enrollment baseline and their 90-day behavior evidence on the same screen, scored on the same rubric. A tool built on disconnected forms cannot show it.
The four levels of the Kirkpatrick model, worked one at a time
Each card below takes one Kirkpatrick level and shows three things: how it is commonly measured today, the exact point where that practice breaks, and the same level run on Sopact’s Loop — collected clean at the source on a persistent trainee ID, read on arrival, acted on in time. Question wording for every level lives in the banks on training evaluation survey questions; this page stays on what each level means and what it takes to measure it.
Level 1 is the easiest level to measure and the easiest to over-trust: decades of research find that satisfaction barely predicts learning or transfer.
Stage 1
Level 1 · Reaction
did it land?
TodayA smile sheet at the door · A 4.6/5 average reported as success · Comment boxes exported to a spreadsheet nobody reads⚠ The satisfaction average is the weakest signal in the model, yet for most programs it is the only level ever measured.
The Loop on this stage with Sopact
Collect — clean at the source
Two-minute pulse, same trainee IDRelevance and confidence itemsOne application-intent open-end
→ every source lands on one persistent ID
On arrival — read automatically
Intelligent Cell
Reads each comment on arrival and separates application intent (“I will use the framing script on my next escalation”) from politeness (“great session”).
Intelligent Row
Reaction lands on the same record that will hold learning, behavior, and results — Level 1 becomes the first point on a line, not the verdict.
Ask & act — the Assistant
“Which module produced politeness instead of application intent this week?”
→ Fix the weak module mid-course — the only moment Level 1 data is worth anything.
Level 2 is a subtraction, not a score — and a subtraction needs two measurements of the same person on the same instrument.
Stage 2
Level 2 · Learning
did they learn it?
TodayA post-quiz or self-rated confidence · Pre and post on different scales, if a pre exists at all · The cohort average reported as learning⚠ Without a baseline on the same instrument, a post-test average measures recall on one day — not gain.
The Loop on this stage with Sopact
Collect — clean at the source
Same scenario item as baselineSame rubric, pre and postConfidence re-check
→ every source lands on one persistent ID
On arrival — read automatically
Intelligent Cell
Scores each exit answer against the same rubric as the baseline and cites the exact phrase behind every score.
Intelligent Row
A learning gain per person, not per cohort: who moved, who did not, and what each no-gain trainee actually wrote.
Ask & act — the Assistant
“Show baseline and exit answers side by side for everyone below a one-point gain.”
→ Remediate named gaps before graduation, not in next year’s redesign.
Level 3 is the level the whole model exists for, and the one the tooling era quietly abandoned. The full treatment of the transfer problem is on behavior change after training.
Stage 3
Level 3 · Behavior
do they use it?
TodayA follow-up survey to a fresh anonymous link · A response rate under 20 percent, unmatchable to the cohort · Or nothing — the program ended, so the record did too⚠ Level 3 is where most Kirkpatrick implementations die: behavior lives 60–90 days after a record that no longer exists.
The Loop on this stage with Sopact
Collect — clean at the source
60/90-day follow-up, same IDManager or peer observationBarrier open-end
→ every source lands on one persistent ID
On arrival — read automatically
Intelligent Cell
Extracts on-the-job application evidence from each follow-up and quotes every transfer barrier in the graduate’s own words.
Intelligent Row
Baseline, learning gain, and 90-day application sit on one thread; non-appliers surface by name with their stated barrier.
Ask & act — the Assistant
“Which graduates show no application at 90 days, and which barrier repeats across the cohort?”
→ Run the refresher against the named barrier while the cohort is still reachable.
Level 4 closes the ladder. Which metrics belong at each level is covered in training metrics; the arithmetic that turns a Level 4 result into a percentage is training ROI.
Stage 4
Level 4 · Results
did it matter?
TodayAn ROI slide assembled at year-end · Staff asked to self-rate business impact · No join between the metric and who was trained⚠ Results claims fail on a missing join, not bad math: the metric was never connected to the individuals who were trained.
The Loop on this stage with Sopact
Collect — clean at the source
One tied operational metricTrained vs not-yet-trainedAttribution notes
→ every source lands on one persistent ID
On arrival — read automatically
Intelligent Cell
Joins the metric the organization already records — errors, retention, sales, ramp time — to each trained person’s thread.
Intelligent Row
Result, behavior, learning gain, and reaction on one record: a chain a board can interrogate link by link.
Ask & act — the Assistant
“Did the tied metric move for trained staff against baseline, and does the 90-day behavior evidence sit behind it?”
→ Report a Level 4 number you can trace, with the attribution limits stated plainly.
Kirkpatrick model examples
A worked Kirkpatrick model example, from a twelve-week communication-skills cohort run on Sopact Sense: at intake, participant P-1247’s baseline read “I freeze in meetings” (Level 1–2); his exit assessment showed a 34-point gain scored on the same rubric as his baseline (Level 2); a mid-cohort mentor interview and peer ratings supplied on-the-job evidence (Level 3); and by cohort end he had presented at an all-hands (Level 3 evidence a funder can quote). At the program level, the same records supported a Level 4 read: a regression across the cohort found mentor minutes predicted skill gains more strongly than LMS module completions — a finding that redirected budget, and one that is only computable when every level sits on one persistent ID.
The pattern compresses to any vertical. Customer-service onboarding: pulse per module (L1), scenario scored pre/post on one rubric (L2), QA-observed call behavior at 60 days (L3), escalation rate for the trained cohort against baseline (L4). Safety training: relevance pulse (L1), hazard-identification test (L2), supervisor spot-checks (L3), incident rate (L4). The instruments change; the ladder and the ID do not.
Why the Kirkpatrick model fails in practice
Most organizations that adopt the Kirkpatrick model measure Level 1, sample Level 2, and never reach Levels 3 and 4 — not because upper-level data matters less, but because it demands longitudinal infrastructure: the same person, followed for months, on one record. Level 1 can be measured in the room. Level 3 lives a quarter later, in the workplace, after the survey tool has forgotten everyone’s name. When the follow-up finally goes out on a fresh anonymous link, response rates collapse and nothing can be matched back — so the evaluation quietly retreats to the levels the tooling can hold.
The academic criticisms point at the same weakness from the other side. The standard criticisms of the Kirkpatrick model are that it implies a causal chain the levels do not guarantee — good reactions predict neither learning nor transfer; that it under-specifies how to measure behavior and results; that it evaluates after the fact instead of shaping design; and that it ignores context and inputs. All four criticisms are fair, and all four sharpen to the same operational point: the model names what to measure and stays silent on the record architecture that makes measuring it possible. Fix the architecture and the model works as intended; leave it broken and no refinement of the levels will save the evaluation.
How do you apply the Kirkpatrick model?
Apply the Kirkpatrick model in five steps: define the Level 4 result first and pick one operational metric for it; assign every trainee a persistent participant ID at enrollment; design one instrument per level, with the Level 2 baseline and exit scored on the same rubric; schedule the Level 3 follow-up at 60–90 days on the same ID; and read each level against the one below it, per person, not as cohort averages. Working backwards from Level 4 is the New World Kirkpatrick discipline; the persistent ID is what makes it executable by a program team instead of an evaluation department.
Step one usually exposes the real gap: if no operational metric can be named, the program has a design problem no survey will fix — start at training needs assessment instead. The per-level instruments themselves are a solved problem: the level-keyed banks on employee training survey questions supply the wording.
Kirkpatrick vs Phillips, CIRO, and Brinkerhoff
The Kirkpatrick model differs from its rivals by scope: Phillips ROI keeps Kirkpatrick’s four levels and adds a fifth that converts results to a return percentage; CIRO evaluates context and inputs before the training as well as outcomes after; and Brinkerhoff’s Success Case Method studies the most and least successful participants in depth instead of averaging everyone. The honest comparison is below — including where each one breaks.
Four evaluation models, honestly compared
| Model | What it adds | Where it breaks in practice |
|---|
| Kirkpatrick four levels (1959) | The shared vocabulary: Reaction, Learning, Behavior, Results — simple enough for every stakeholder | Assumes the upper levels get measured; most implementations stop at Levels 1–2 |
| New World Kirkpatrick (2016) | Plan backwards from Level 4; Required Drivers and leading indicators keep Level 3 alive | A planning discipline — still needs per-person data connected across months to execute |
| Phillips ROI | Level 5: net benefits over cost as a percentage, for the budget conversation | Inherits every join problem beneath it; isolating training’s effect is the hard part |
| CIRO | Evaluates context, inputs, and design — not just the aftermath | Thin on behavior; little guidance for measuring on-the-job transfer |
| Brinkerhoff Success Case Method | The why behind extreme outcomes, in participants’ own words | Success cases must be findable first — which requires per-person outcome data |
Model choice matters less than record architecture: every model’s upper levels depend on one person’s data connected across time. Which platforms can actually hold that thread is the job of training evaluation software.
A year-end evaluation tells you what happened. The Loop makes Levels 3 and 4 measurable.
Continuous Kirkpatrick is the Loop run on the level map: collect clean at the source so every wave lands on the same trainee record, analyze on arrival so reactions are coded and gains computed while the cohort still runs, improve in time so the weak module is fixed and the refresher aimed before the moment passes. Levels 3 and 4 stop being aspirations in a slide and become two more reads of a record that never died.
The same thread is what makes the claim defensible: a Level 4 number traces back through 90-day behavior, learning gain, and reaction to a trainee’s own words — the standard described in Loop traceability. When a funder asks “how do you know?”, the answer is a chain, not an assertion.
One method, three moves that never stop
1 · CollectClean at the source; all four levels land on one persistent trainee ID, from enrollment on.
2 · AnalyzeOn arrival; reactions coded, gains scored on one rubric, barriers quoted and cited.
3 · ImproveIn time to act; fix the module, remediate the gap, aim the refresher — this cohort.
Then the cycle runs again, a little sharper each cohort. Read the method: the Loop methodology →
Put the Kirkpatrick model to work this week
The fastest test of the model is one level of your real program run through it. Each prompt below pastes into Sopact Sense’s Assistant — one per Kirkpatrick level; the arrow above each links the Academy walkthrough with the expected output and tips.
Academy walkthrough → Apply the Kirkpatrick model to a survey
Map the four Kirkpatrick levels onto [PROGRAM]: for each level - Reaction, Learning, Behavior, Results - specify the instrument, the timing, the question set, and exactly what lands on the persistent participant ID. Then flag which level our current data cannot support, and why.
Academy walkthrough → Analyze pre, mid, and post survey data
Here are the pre and post responses for [COHORT]: [PASTE OR ATTACH]. Score both waves on the same rubric, compute each participant's Level 2 learning gain, list everyone below [THRESHOLD], and quote the baseline and exit answers behind each flagged gain.
Academy walkthrough → Measure behavior change after training
From our 90-day follow-ups for [COHORT]: [ATTACH]. Extract Level 3 evidence per graduate - is the skill applied on the job, how often, and with what result - then rank the transfer barriers by frequency and quote each barrier in the graduate's own words.
Academy walkthrough → Connect training to results
Join [OPERATIONAL METRIC] to the trained cohort's records: compare trained vs not-yet-trained staff against baseline, show whether the metric moved for the people with 90-day behavior evidence behind them, and state the attribution limits a board would probe.
Learn the how-to in the Academy
Each walkthrough is short and practical: what to do, the prompt to run, the output to expect, and the tips that keep it reliable.
Frequently asked questions
What are the four levels of the Kirkpatrick model?
Level 1 Reaction measures whether participants found the training engaging and relevant. Level 2 Learning measures whether knowledge or skill increased, pre to post. Level 3 Behavior measures whether people apply the skill on the job, typically read 60 to 90 days later. Level 4 Results measures whether an organizational metric moved. Sopact carries all four levels on one persistent participant ID, so each level reads against the last.
What is an example of the Kirkpatrick model in practice?
From a communication-skills cohort run on Sopact Sense: a participant’s intake baseline read “I freeze in meetings”; his exit assessment showed a 34-point gain on the same rubric (Level 2); mentor and peer evidence confirmed on-the-job change (Level 3), ending in an all-hands presentation; and across the cohort, mentor minutes predicted gains better than LMS completions — a Level 4 finding that redirected budget.
What are the criticisms of the Kirkpatrick model?
Four recur: the levels imply a causal chain they do not guarantee (good reactions predict neither learning nor transfer); the model under-specifies how to measure Levels 3 and 4; it evaluates after the fact rather than shaping design; and it ignores context and inputs. Sopact’s reading: every criticism sharpens to one operational gap — the upper levels demand longitudinal, person-level records — which is what Continuous Kirkpatrick fixes.
What is the New World Kirkpatrick Model?
The New World Kirkpatrick Model (James and Wendy Kirkpatrick, 2016) updates the original: plan backwards from Level 4 results, define leading indicators, and sustain Level 3 through Required Drivers — the reinforcement, monitoring, and accountability processes that keep new behavior alive. It is a planning discipline, and it still requires per-person data connected across months to execute; that record layer is the part Sopact Sense supplies.
What is the difference between the Kirkpatrick model and the Phillips ROI model?
Phillips keeps Kirkpatrick’s four levels and adds Level 5: net program benefits divided by program cost, expressed as a return percentage. The practical difference is effort — ROI requires isolating training’s effect and converting results to money. Choose Phillips when the budget conversation demands a percentage; either way, the Level 3–4 join problem underneath is identical, and Sopact’s persistent trainee ID is what makes both computable.
Why do most organizations stop at Level 1 and Level 2?
Because Levels 1 and 2 can be measured in the room, while Levels 3 and 4 require following the same person for months. Form-centric survey tools treat every send as a fresh anonymous batch, so the follow-up can never be matched back to the cohort. Sopact calls the structural fix Continuous Kirkpatrick: a persistent participant ID assigned at enrollment turns the four levels into four views of one dataset.
How do you measure Level 3 of the Kirkpatrick model?
Collect behavior evidence at 60 to 90 days on the same participant ID: a short follow-up asking for concrete application examples, a manager or peer observation, and one open-ended barrier question. Sopact Sense reads each follow-up on arrival, extracts the application evidence, and surfaces non-appliers by name with their stated barrier — so Level 3 triggers a refresher instead of a shrug.
How do you measure Level 4 of the Kirkpatrick model?
Pick one operational metric the organization already records — error rate, retention, sales, time to productivity — and join it to the trained group against baseline, ideally with a not-yet-trained comparison. A credible Level 4 claim keeps the Level 3 behavior evidence behind it and states its attribution limits plainly; on Sopact’s one-record thread, that chain is a query rather than a year-end reconstruction.
Is the Kirkpatrick model still relevant?
Yes — as the level map. Sixty-five years on, it remains training evaluation’s shared vocabulary, and no successor has replaced it. What changed is feasibility: persistent-ID records and AI that reads open-ended evidence on arrival make Levels 3 and 4 measurable by an ordinary program team. The model was never the barrier; the infrastructure was, and Continuous Kirkpatrick is Sopact’s name for closing that gap.
What is Continuous Kirkpatrick?
Continuous Kirkpatrick is Sopact’s name for running the four Kirkpatrick levels as one loop on one trainee record: a persistent participant ID carrying the reaction pulse, the learning gain scored on one rubric, the 90-day behavior evidence, and the tied result metric. Kirkpatrick supplies the level map; the Loop supplies the cadence, so evaluation happens while the program runs — not after it ends.
Next: run the model across the whole program on training evaluation, or pick up the per-level wording on training evaluation survey questions.
Four levels, one record
L1ReactionPulse coded on a persistent ID
L2LearningGain per person, same rubric
L3Behavior90-day evidence, no matching step
L4ResultsOne tied metric, a traceable chain
Continuous Kirkpatrick: all four levels on one trainee record.