Kirkpatrick Model: 4 Levels of Training Evaluation Explained
The Kirkpatrick four-level training evaluation model — reaction, learning, behavior, results. Definitions, examples, and sample questions per level for 2026.
The Kirkpatrick model is a four-level framework for evaluating training. Level 1 Reaction asks whether participants found the session relevant. Level 2 Learning asks whether knowledge or skill increased. Level 3 Behavior asks whether people apply it on the job. Level 4 Results asks whether an organizational outcome moved.
Donald Kirkpatrick introduced the levels in 1959 and they are still the shared vocabulary of training evaluation. If you run a program, you already know the map. What you probably do not have is a way to reach the top of it. The levels are not the difficult part. The record underneath them is, and that is what this page is about.
A seven-minute walkthrough of the four levels running on one connected learner record, from the enrollment baseline through the behavior evidence to the tied result metric.
Key takeaways
Levels 1 and 2 can be measured in the room. Levels 3 and 4 cannot, because they live weeks and months after the program, in the workplace.
Sopact calls the buying test the same-screen test: ask any platform to show one trainee's enrollment baseline and their later behavior evidence on one screen, scored on the same rubric.
Level 3 measurement should begin two to four weeks after training, not at ninety days. Waiting until the quarter is over means the finding arrives after the moment to act on it.
Completion is not change. A participant who finishes with no measured delta is a flag, not a success, and a cohort average will hide them.
Phillips ROI, CIRO, Brinkerhoff, LTEM, and Anderson each improve the map. None removes the per-person join the upper levels demand.
Why most Kirkpatrick evaluations never reach Levels 3 and 4
Your funders ask what changed, and your data ends at graduation. The reason most Kirkpatrick implementations stall is boring and structural. The baseline survey lives in one spreadsheet. The exit survey lives in another. The follow-up never quite happens. When the report is due, someone rebuilds the whole story by hand, and the participants who needed support mid-program were only visible in hindsight.
Run the timeline of a typical ten-week cohort and the failure is easy to see. Week one, the intake form captures a skills assessment and a confidence score, and files them in the intake sheet. Week five, a facilitator notices someone struggling and mentions it in a meeting, where it attaches to nothing. Week ten, the exit survey goes out as a new form in a new file, so matching it back to baseline means manual name-matching. Day ninety, the follow-up that would answer whether any of it lasted does not go out at all, because there is no longer a list of who to send it to. The journey was happening. The record was not.
This is why Level 3 is where most evaluations die. Behavior lives a quarter after a record that no longer exists, and when the follow-up finally goes out on a fresh anonymous link, response rates collapse and nothing can be matched back to the cohort. The evaluation quietly retreats to the levels the tooling can hold, and the program reports satisfaction for the third year running.
How do you choose a tool that can measure Kirkpatrick Levels 3 and 4?
Sopact calls it the same-screen test: ask any training platform to show one trainee's enrollment baseline and their behavior evidence from months later on the same screen, scored on the same rubric. A system that assigns a persistent participant ID at enrollment passes. A system built on disconnected form submissions cannot pass it at any price.
The distinction is the data model rather than the questionnaire. A form-centric tool treats every send as a fresh anonymous batch, so Level 1 and Level 3 never land on the same person and the ladder breaks between rungs however well the questions are written. A record-centric system never breaks the thread, so the four levels become four views of one dataset. It is the missing ID, not the missing field, that silently breaks the story.
The tooling history explains why so few products pass. The smile-sheet era, on paper and then on SurveyMonkey and Google Forms, made Level 1 nearly free and made anonymous responses the default. The LMS era, with Cornerstone, Docebo, Moodle, and TalentLMS, added completion rates and quiz scores, which measure exposure and recall rather than transfer. Neither era produced a mainstream tool that owned Levels 3 and 4, because both levels require longitudinal person-level records that a form tool discards and an LMS never kept. Run the same-screen test on your current stack before you design another instrument.
The four levels of the Kirkpatrick model: reaction, learning, behavior, results
Kirkpatrick Level 1: Reaction
Level 1 asks whether participants found the training relevant and worth their time, collected at session end as a short pulse. It is the easiest level to measure and the easiest to over-trust, because satisfaction barely predicts learning or transfer. A 4.6 out of 5 tells you the room was comfortable. The useful read is not the average at all, it is the difference between a comment expressing politeness and a comment expressing application intent.
Kirkpatrick Level 2: Learning
Level 2 asks whether knowledge, skill, or confidence increased, measured as a pre-and-post pair around the training. The per-person delta is the learning gain. Level 2 is a subtraction, and a subtraction needs the same person measured twice on the same instrument. A post-test average with no baseline measures recall on one day. A baseline on a 1-to-10 confidence scale and an exit on a 0-to-100 skills scale cannot be paired at all, and that drift is usually discovered a year later when the claim is due. Fix the scale at intake and treat it as a contract with every wave that follows.
Kirkpatrick Level 3: Behavior
Level 3 measurement should begin two to four weeks after training and continue over the following months, not start at ninety days. The guidance from Kirkpatrick Partners is explicit that leaving too long a gap between learning and measurement risks losing the opportunity to adjust the program while it can still be adjusted. In workforce training programs this is the difference between a Level 3 finding that redirects the next cohort and one that arrives as a post-mortem. A credible Level 3 read combines a self-rated application scale, a manager or peer verification of the same behavior, and one open question naming what is blocking application. Treat self-report on its own as amber rather than green: a graduate saying they use a skill is evidence, and a supervisor confirming it is proof. The transfer problem is worked in depth on behavior change after training.
Kirkpatrick Level 4: Results
To measure Level 4, pick one operational metric the organization already records, baseline it before the first session, and join it to the list of people who were actually trained, with a not-yet-trained comparison group where the program design allows one. Error rate, retention, promotion rate, sales, time to productivity, incident rate. The instrument at Level 4 is mostly not a survey question, it is that join. Level 4 claims usually fail on a missing join rather than bad arithmetic, because the metric was never connected to the individuals who went through the training. Which metrics belong at which level is covered on training metrics, and the arithmetic that converts a Level 4 movement into a return percentage is on training ROI.
Learner record · L-2024-118Illustrative · sample data
One persistent Learner ID. Four Kirkpatrick levels writing to the same record, each one readable against the one below it.
L1 · Session end
ReactionRelevance 4/5, confidence-to-apply 2 of 5. Comment coded as application intent, not politeness.
L2 · Pre / post
LearningSkills assessment 41% at intake, 78% at exit. Same rubric both waves, so the 37-point gain is computed rather than estimated.
L3 · Week 2–90
BehaviorSelf-reported weekly use, confirmed by supervisor. Amber becomes green only when a second source agrees.
L4 · Day 180
ResultsEmployment confirmed and wage recorded, pulled from the employer HRIS rather than a phone-call campaign.
Nothing here was re-matched or re-keyed. The ID assigned at enrollment is what makes it one record instead of four surveys.
How do you apply the Kirkpatrick model?
Apply the Kirkpatrick model in five steps: define the Level 4 result first and pick one operational metric for it; assign every trainee a persistent participant ID at enrollment; design one instrument per level, with the Level 2 baseline and exit scored on the same rubric; open the Level 3 follow-up two to four weeks after training and keep it running through ninety days on the same ID; and read each level against the one below it, per person, not as cohort averages.
Step one is the step that exposes the real gap. If no operational metric can be named, the program has a design problem no survey will fix, and the honest starting point is a training needs assessment instead. Step three is the step most teams skip, and skipping it is what produces mismatched scales and an unusable pre-post pair. Question wording is a solved problem, with level-keyed banks on employee training survey questions.
In practice the four levels map onto the four moments a program already has. Run your current cohort against this table and mark each row green, amber, or red before designing anything new.
Four program moments, four Kirkpatrick levels
Program moment
Level
What is collected
What it proves
Baseline week one
L2 baseline
Skills assessment, confidence rating, goals, on the scale every later wave reuses
A real starting point, so change is computed rather than reconstructed
Midpoint week five
L1 + early L2
Relevance pulse, the same confidence item repeated, one open blocker question
Who needs support while it still matters, flagged against their own baseline
Exit week ten
L2 gain
Every baseline measure repeated identically, plus one attribution question
A delta per learner, and the completers whose delta never moved
Follow-up week two to day 180
L3 + L4
Application evidence, manager verification, the tied operational metric
Whether it transferred, whether it lasted, and whether anything moved
Two rules keep the plan honest. Reuse the same IDs and the same codebook across every wave, so waves stack into a trend instead of piling up as separate studies. And state the evidence limit plainly: if the last measurement was at ninety days, the program has evidence to ninety days and no further, because sustained is a measurement rather than a mood.
Progress view · digital skills cohort 6Week 5 of 10
LearnerCohortProgram
Each learner is read against their own baseline, not the cohort average. Two are flagged because applied confidence is lagging knowledge growth.
L-118Skills 41 → 63
Support
L-121Skills 44 → 71
On track
L-127Skills 39 → 58
Support
Flag · week 5. L-118 and L-127 both cite the same module. Suggested: paired practice before week 7. The flag is a prompt for the team, not an automated action.
A worked Kirkpatrick example: one cohort, four levels
A worked Kirkpatrick model example from a ten-week digital-skills cohort of 23 participants: module-close pulses at Level 1, a skills assessment and confidence rating paired at intake and exit on identical scales at Level 2, application evidence collected from week two through day ninety at Level 3, and employment outcomes as the tied Level 4 result. Mean skill assessment moved from 41 percent at baseline to 78 percent at exit, an average delta of 34 points across the cohort. Confidence moved from 2.1 to 4.2 on a five-point scale.
The Level 3 story is the one that changed the program. At week five, one participant's applied confidence was lagging their knowledge growth, and the same pattern appeared in a second participant citing the same module. Paired practice sessions were added before week seven. At exit that first participant wrote: "I ran the whole client setup on my own this week. Two months ago I would not have tried." At day ninety, 18 of the 23 graduates reported weekly use of the trained skills, and the employment outcomes traced to individual follow-up responses rather than to a cohort estimate.
Two things about that report matter more than the numbers. It names its denominator, so 18 of 23 is stated as 18 of 23 rather than as 78 percent of an unspecified population. And every figure opens back to the participant records behind it, which is what lets a funder interrogate the claim link by link instead of taking it on trust. The pattern compresses to any vertical. Customer-service onboarding runs a pulse per module, a scenario scored pre and post on one rubric, QA-observed call behavior from week two, and escalation rate for the trained cohort against baseline. Safety training runs a relevance pulse, a hazard-identification test, supervisor spot-checks, and incident rate. Sales training, leadership development, workforce development, contractor and ethics training all follow the same shape. The instruments change. The ladder and the ID do not. Workforce training programs carry the heaviest Level 3 burden of any vertical, because the outcome that matters, whether the skill is still being used in a job months later, sits furthest from the classroom.
Day ninety. Levels 3 and 4 live here, and this is the evidence most programs never capture, because the record ended at graduation. Illustrative cohort, synthetic sample data.
How do you connect training to business results at Kirkpatrick Level 4?
Connect training to results by naming one metric the organization already tracks before the program starts, keeping the Level 3 behavior evidence attached to the people behind the movement, and stating the attribution limits plainly rather than claiming the metric moved because of the training. A Level 4 number with no Level 3 evidence underneath it is a coincidence with a chart. A Level 4 number that traces down through verified application, a learning gain scored on one rubric, and a trainee's own words is a chain a board can interrogate.
Source trail · Level 4 claim4 connected records
"18 of 23 graduates were using the trained skills at day ninety."
Every figure opens back to the people behind it. Here is what sits under that one sentence.
L4
Employment and wage recordPulled from the employer HRIS at day 180, joined on the Learner ID. No phone-call campaign.
L3
Supervisor verificationWeekly use of the trained skill confirmed by a second source, not self-report alone.
L2
Learning gain on one rubric41% to 78%, same instrument at intake and exit, so the delta is computed.
L1
Reaction pulse, codedApplication intent at session end, the first point on the line rather than the verdict.
"I ran the whole client setup on my own this week. Two months ago I would not have tried."Exit response · Learner L-118 · the bottom of the chain
The honest framing is contribution rather than causation. Observed change is what the data shows. Causal impact is a stronger claim that needs a comparison group or an explicit argument about what else could account for the movement. Programs lose credibility by overclaiming at Level 4 far more often than by underclaiming, and a report that grades its own findings, from strongly supported down to not yet measured, survives scrutiny that a confident summary does not.
Kirkpatrick vs Phillips ROI, CIRO, Brinkerhoff, LTEM, and Anderson
The Kirkpatrick model differs from its alternatives by scope: Phillips ROI keeps the four levels and adds a fifth that converts results to a return percentage; CIRO evaluates context and inputs before the training as well as outcomes after it; Brinkerhoff's Success Case Method studies the most and least successful participants in depth instead of averaging everyone; LTEM replaces the four levels with eight tiers that separate attendance and recall from actual task competence; and Anderson's Value of Learning Model starts from strategic alignment rather than from the course.
Six evaluation models, honestly compared
Model
What it adds
Where it breaks in practice
Kirkpatrick four levels 1959
The shared vocabulary of Reaction, Learning, Behavior, Results, simple enough for every stakeholder
Assumes the upper levels get measured; most implementations stop at Levels 1 and 2
New World Kirkpatrick 2016
Plan backwards from Level 4; Required Drivers and leading indicators keep Level 3 alive after the course
A planning discipline that still needs per-person data connected across months to execute
Phillips ROI adds Level 5
Net benefits over cost as a percentage, for the budget conversation
Inherits every join problem beneath it; isolating the training effect is the hard part
CIRO
Evaluates context, inputs, and design rather than only the aftermath
Thin on behavior, with little guidance for measuring on-the-job transfer
Brinkerhoff Success Case Method
The why behind extreme outcomes, in the participants own words
Success cases have to be findable first, which requires per-person outcome data
LTEM Thalheimer
Eight tiers that refuse to count attendance or recall as learning, forcing a task-competence standard
More demanding instruments at every tier, and no help with the longitudinal record they assume
Anderson Value of Learning
Starts from strategic alignment, asking whether the training should exist before asking whether it worked
Organization-level rather than learner-level, so it cannot answer who changed
The four recurring criticisms of the Kirkpatrick model are that the levels imply a causal chain the model does not guarantee, that it under-specifies how to measure Levels 3 and 4, that it evaluates after the fact instead of shaping design, and that it ignores context and inputs. All four are fair. All four sharpen to one operational point: the model names what to measure and stays silent on the record architecture that makes measuring it possible. That is why choosing a different model rarely fixes anything on its own, since every alternative's upper levels depend on the same per-person thread. Which platforms actually hold that thread is the subject of training evaluation software, and the full lifecycle from needs assessment through delivery to reporting is worked on training evaluation.
Making Kirkpatrick Levels 3 and 4 continuous with the Loop
A year-end evaluation tells you what happened. Running the level map on a live record instead is the Loop. Collect clean at the source, so every wave lands on the same trainee record instead of arriving as a new anonymous batch. Analyze on arrival, so reaction comments are themed and learning gains computed while the cohort is still running. Improve in time, so the weak module gets fixed and the refresher aimed at a named barrier before the moment passes. The funder report stops being a quarterly reconstruction and becomes the participant journeys, rolled up.
The division of labour matters as much as the method. Sopact keeps the record, connects the stages, computes the deltas, and flags the patterns. Deciding who gets support, what changes next cohort, and what goes to the funder stays with the people who run the program. And when the claim is challenged, a Level 4 number traces back through verified behavior evidence, a learning gain on one rubric, and a trainee's own words, which is the standard described in Loop traceability.
One method, three moves that never stop
1 · CollectClean at the source. All four levels land on one persistent trainee ID, from enrollment onward.
2 · AnalyzeOn arrival. Reactions themed, gains scored on one rubric, barriers quoted and cited.
3 · ImproveIn time to act. Fix the module, remediate the gap, aim the refresher, this cohort.
Then the cycle runs again, a little sharper each cohort. Read the method: the Loop methodology →
Kirkpatrick model walkthroughs, level by level
One walkthrough per level: what to collect, how to read it, and what to do when a level comes back weak.
What are the four levels of the Kirkpatrick model?
Level 1 Reaction measures whether participants found the training engaging and relevant. Level 2 Learning measures whether knowledge or skill increased from before to after. Level 3 Behavior measures whether people apply the skill on the job. Level 4 Results measures whether an organizational metric moved. Sopact carries all four levels on one persistent participant ID, so each level can be read against the one below it rather than as four separate studies.
When should you measure Level 3 of the Kirkpatrick model?
Begin two to four weeks after the training and continue over the following months, rather than waiting until ninety days. Kirkpatrick Partners' guidance is that too long a gap between learning and measurement risks losing the opportunity to adjust the program. Collect a self-rated application scale, a manager or peer verification of the same behavior, and one open question naming the blocker. Sopact treats self-report alone as amber rather than green, and surfaces non-appliers by name with their stated barrier.
How do you measure Level 4 of the Kirkpatrick model?
Pick one operational metric the organization already records, such as error rate, retention, promotion rate, sales, or time to productivity, baseline it before the first session, and join it to the trained group with a not-yet-trained comparison where the design allows. A credible Level 4 claim keeps the Level 3 behavior evidence behind it and names its denominator. On a record that passes Sopact's same-screen test, that chain is a query rather than a year-end reconstruction.
What is the same-screen test?
The same-screen test is Sopact's name for the one question that separates training evaluation tools: ask the platform to show one trainee's enrollment baseline and their behavior evidence from months later on the same screen, scored on the same rubric. A system that assigns a persistent participant ID at enrollment passes. A system built on disconnected form submissions cannot pass it, whatever its feature list claims.
What is the difference between the Kirkpatrick model and the Phillips ROI model?
Phillips keeps Kirkpatrick's four levels and adds Level 5, which divides net program benefits by program cost and expresses the answer as a return percentage. The practical difference is effort, because ROI requires isolating training's effect and converting results into money. Choose Phillips when the budget conversation demands a percentage. Either way the Level 3 and Level 4 join problem underneath is identical, and Sopact's persistent participant ID is what makes both computable.
What are the alternatives to the Kirkpatrick model?
The main alternatives are Phillips ROI, which adds a fifth monetized level; CIRO, which evaluates context and inputs before the program; Brinkerhoff's Success Case Method, which studies extreme outcomes rather than averages; LTEM, which replaces four levels with eight tiers that refuse to count attendance or recall as learning; and Anderson's Value of Learning Model, which starts from strategic alignment. Sopact's position is that every alternative's upper levels depend on the same per-person record, so the model choice matters less than whether the record exists.
What are the criticisms of the Kirkpatrick model?
Four criticisms recur: the levels imply a causal chain the model does not guarantee, since good reactions predict neither learning nor transfer; the model under-specifies how to measure Levels 3 and 4; it evaluates after the fact rather than shaping design; and it ignores context and inputs. All four are fair, and Sopact reads all four as pointing at one operational gap, which is that the upper levels demand longitudinal person-level records that most training tooling never kept.
What is the New World Kirkpatrick Model?
The New World Kirkpatrick Model, published in 2016 by James D. Kirkpatrick and Wendy Kayser Kirkpatrick in Kirkpatrick's Four Levels of Training Evaluation (ATD Press), updates the original in three ways: plan backwards from Level 4 results, define leading indicators that move before the final metric, and sustain Level 3 with Required Drivers, the reinforcement, monitoring, and accountability processes that keep new behavior alive. It is a planning discipline, and it still needs per-person data connected across months to execute, which is the layer Sopact supplies.
Why do most organizations stop at Level 1 and Level 2?
Because Levels 1 and 2 can be measured in the room, while Levels 3 and 4 live weeks and months later in the workplace. Form-centric survey tools treat every send as a fresh anonymous batch, so the follow-up cannot be matched back to the cohort and response rates collapse. Sopact's structural fix is a persistent participant ID assigned at enrollment, which turns the four levels into four views of one dataset and lets any platform be checked with the same-screen test.
What is a Kirkpatrick model example?
An illustrative ten-week digital-skills cohort of 23 participants shows the full ladder: module-close reaction pulses at Level 1; a skills assessment moving from 41 percent at baseline to 78 percent at exit on identical scales at Level 2, an average delta of 34 points; application evidence collected from week two through day ninety at Level 3, with 18 of the 23 graduates reporting weekly use; and employment outcomes as the tied Level 4 result. Sopact names the denominator on every figure and opens each one back to the participant records behind it.