The Kirkpatrick model evaluates learning through four connected levels: reaction, learning, behaviour and results. Its main contribution is to challenge the assumption that attendance, completion or satisfaction proves that a programme worked. Learners may enjoy a course without changing practice; they may gain knowledge without having the opportunity to apply it; and improved business results may be influenced by factors beyond learning.
The model is most useful when evaluation is designed before delivery. It helps people professionals identify what evidence is needed, what work conditions support transfer and how to make proportionate claims about impact. It should not be used as a mandatory checklist or a promise of perfect causation.
The four levels
| Level | Core question | Example evidence | Limitation if used alone |
| 1. Reaction | Did participants find the learning relevant, accessible and engaging? | Feedback, confidence, perceived relevance and access barriers | Positive reactions do not prove learning or behaviour change |
| 2. Learning | Did participants gain required knowledge, skill, confidence or judgement? | Knowledge checks, demonstrations, simulations and work samples | Learning in a controlled setting may not transfer to work |
| 3. Behaviour | Are participants applying capability in their work context? | Observation, manager feedback, work-quality evidence and self-reflection | Behaviour may depend on systems, workload and manager support |
| 4. Results | Did relevant organisational or stakeholder outcomes improve? | Quality, service, safety, productivity, retention or customer outcomes | Many variables may influence the result |
The levels are related, but they are not automatic steps. A strong reaction may support motivation, but it does not cause business impact by itself. Evaluation must be matched to the intervention’s scale, risk and intended outcome.
Begin with the result and work backwards
A practical way to design evaluation is to start with the performance or organisational outcome. What should be different if the learning succeeds? Then identify the behaviour required, the learning needed and the experience that supports participation.
| Backward-design question | Example: complaint-handling development |
| Result | Residents receive clearer, more timely responses and repeat complaints reduce |
| Behaviour | Advisers take ownership, use escalation routes and explain next steps accurately |
| Learning | Advisers practise case diagnosis, policy judgement and difficult conversations |
| Reaction | Learners consider scenarios realistic, can access the programme and know why it matters |
This approach connects evaluation to learning needs analysis. If the initial problem diagnosis is weak, later evaluation will also be weak because there is no valid outcome to assess.
Level 1: reaction with purpose
Reaction measures whether learning was relevant, accessible and usable. It should go beyond “Did you like the course?” Useful questions explore whether the content reflected real work, whether the pace and format were accessible, whether practice felt meaningful and what barriers participants expect when applying learning.
| Weak reaction question | Stronger question |
| “How satisfied were you?” | “How closely did the scenarios reflect decisions you face at work?” |
| “Was the facilitator good?” | “What will make it easier or harder to apply this learning in the next month?” |
| “Would you recommend this session?” | “Which part of the programme requires more practice or clarification?” |
Reaction evidence can identify design and access problems early. It should be reviewed promptly, particularly for new programmes or large-scale change.
Level 2: learning and demonstration
Learning evaluation should test the capability that matters. A knowledge quiz may be suitable for factual understanding, but judgement, communication or technical skill often requires a simulation, observed practice, case analysis or work sample. Pre- and post-assessment can show change, but should be used fairly and proportionately.
At fictional Glenmoor Utilities, supervisors attend a programme on incident escalation. Instead of relying only on a quiz, participants work through realistic scenarios, identify escalation thresholds and explain their rationale. Facilitators record whether the decision is accurate, timely and aligned with policy. This produces richer evidence than completion data alone.
Level 3: behaviour and transfer
Behaviour is often the most difficult level because work conditions influence transfer. People may know what to do but lack time, authority, manager support, system access or reinforcement. Evaluation should therefore consider both participant behaviour and the environment in which it occurs.
| Transfer condition | Evaluation question | Possible support |
| Manager reinforcement | Do managers observe, discuss and support new practice? | Manager briefing, observation prompts and coaching |
| Opportunity | Do participants encounter relevant tasks soon after learning? | Planned practice, stretch assignment or simulation refresh |
| Tools and process | Do systems enable the intended behaviour? | Job aids, system changes and accessible escalation routes |
| Social norms | Is the behaviour encouraged by peers and leaders? | Peer review, calibration and recognition |
| Workload | Do people have time to apply learning safely? | Capacity planning and prioritisation |
Glenmoor samples incident records and asks managers to observe escalation conversations over the following months. It also checks whether supervisors can access the escalation system during field work. If behaviour does not change, the answer may be a process improvement rather than more training.
✅ Get Your Assignment Done by Experts
Level 4: results and attribution
Results should be relevant to the original purpose. Glenmoor might track incident response time, quality of escalation records, repeat safety events and regulator feedback. But it should avoid claiming that learning alone caused a change, especially if staffing, systems or policy also changed.
| Result measure | Attribution question |
| Faster incident response | Did new digital tools or staffing changes also affect the result? |
| Fewer repeat errors | Was there a wider process redesign at the same time? |
| Improved customer outcomes | Did the affected employees have opportunity to apply the new behaviour? |
| Reduced risk cost | Are estimates based on credible assumptions and a relevant baseline? |
Use comparison groups, baseline trends, stakeholder insight and documented assumptions where feasible. This connects to return on investment in people interventions, which considers financial, operational, workforce and ethical outcomes together.
Strengths and limits of the Kirkpatrick model
| Strength | Limitation | Practical response |
| Moves attention beyond attendance and satisfaction | Can be treated as a rigid sequence | Select proportionate evidence for the decision |
| Highlights behaviour and results | Results are hard to attribute to learning alone | Combine data, qualitative insight and context |
| Encourages evaluation planning | May overlook structural barriers to performance | Assess transfer conditions and work design |
| Provides a shared language | Can lead to over-measurement | Focus on the few outcomes that matter |
The model is most valuable when it raises better questions rather than becomes a compliance form. Not every intervention requires a full four-level evaluation. A short information update may need only reaction and knowledge evidence; a high-cost leadership programme or safety-critical capability change may justify deeper behaviour and results analysis.
Workplace application: Glenmoor Utilities
Glenmoor’s programme achieves high reaction scores and improved scenario performance, but field observations show limited change in escalation behaviour. Investigation reveals that supervisors still need manager approval before using the new process and mobile signal is unreliable at several sites. The learning design is improved, but the more important action is to revise decision rights and provide offline system access. The evaluation prevents Glenmoor from blaming participants for a transfer failure caused by the work system.
Evaluation design, evidence governance and improvement
Evaluation evidence should be governed before launch. Glenmoor can identify the evaluation owner, the measures, data sources, baseline, review dates, confidentiality safeguards and the decisions that findings will inform. This prevents post-programme reporting from selecting only positive outcomes. It also clarifies where learning data will be combined with performance, customer or safety data and who may access that information.
Use evaluation as a learning cycle. If reaction is low because scenarios do not reflect work, revise the design. If learning is high but behaviour is low, inspect manager reinforcement, authority, tools and workload. If behaviour changes but results do not, test whether the selected result measure is valid or whether other system factors dominate. The purpose is not to prove that a programme succeeded; it is to improve capability and decide whether to continue, adapt, scale or stop the intervention.
Governance and inclusive evaluation
Evaluation should consider who can access learning and who benefits. Segment participation, reaction, learning and behaviour data where appropriate to identify whether shift workers, remote employees, part-time staff or other groups face barriers. Protect confidentiality, especially with small cohorts, and avoid using learning data punitively where it is intended for development.
Set evaluation ownership before launch. Programme leads may collect learning evidence; managers may support behaviour observation; operational leaders may own results measures; and people teams should ensure data quality, fairness and interpretation. Review findings with participants and stakeholders so that learning programmes are improved transparently.
Reporting evaluation evidence responsibly
Present results with their limitations. Glenmoor should distinguish participant feedback, demonstrated learning, observed behaviour and organisational trends rather than collapse them into a claim that the programme “worked”. This allows executives to see where the evidence is strongest, where transfer is blocked and what action is required next. It also protects participants from unfair conclusions when a wider system constraint prevents behaviour change.
Frequently asked questions
What are the four levels of the Kirkpatrick model?
Reaction, learning, behaviour and results. They provide a framework for considering whether learning was relevant, understood, applied and connected to meaningful outcomes.
Is participant satisfaction enough to show that training worked?
No. Satisfaction can be useful feedback, but it does not show whether learners gained capability, changed behaviour or improved results.
Does every programme need all four Kirkpatrick levels?
No. Use evaluation proportionately. The more costly, strategic or high-risk the intervention, the stronger the case for evaluating behaviour and results.
How does Kirkpatrick connect to ADDIE?
ADDIE provides a design process; Kirkpatrick provides an evaluation structure. Evaluation needs should be considered during ADDIE’s analysis and design stages, not added at the end.
References
Kirkpatrick, J.D. and Kirkpatrick, W.K. (2016) Kirkpatrick’s four levels of training evaluation. Alexandria, VA: ATD Press.
CIPD (2025) Learning needs analysis. Available at: https://www.cipd.org/en/knowledge/factsheets/learning-needs-factsheet/ (Accessed: 24 August 2026).
Salas, E., Tannenbaum, S.I., Kraiger, K. and Smith-Jentsch, K.A. (2012) ‘The science of training and development in organisations’, Psychological Science in the Public Interest, 13(2), pp. 74–101.