The Kirkpatrick model evaluates learning through four connected levels: reaction, learning, behaviour and results. Its main contribution is to challenge the assumption that attendance, completion or satisfaction proves that a programme worked. Learners may enjoy a course without changing practice; they may gain knowledge without having the opportunity to apply it; and improved business results may be influenced by factors beyond learning.

The model is most useful when evaluation is designed before delivery. It helps people professionals identify what evidence is needed, what work conditions support transfer and how to make proportionate claims about impact. It should not be used as a mandatory checklist or a promise of perfect causation.

The four levels

LevelCore questionExample evidenceLimitation if used alone
1. ReactionDid participants find the learning relevant, accessible and engaging?Feedback, confidence, perceived relevance and access barriersPositive reactions do not prove learning or behaviour change
2. LearningDid participants gain required knowledge, skill, confidence or judgement?Knowledge checks, demonstrations, simulations and work samplesLearning in a controlled setting may not transfer to work
3. BehaviourAre participants applying capability in their work context?Observation, manager feedback, work-quality evidence and self-reflectionBehaviour may depend on systems, workload and manager support
4. ResultsDid relevant organisational or stakeholder outcomes improve?Quality, service, safety, productivity, retention or customer outcomesMany variables may influence the result

The levels are related, but they are not automatic steps. A strong reaction may support motivation, but it does not cause business impact by itself. Evaluation must be matched to the intervention’s scale, risk and intended outcome.

Begin with the result and work backwards

A practical way to design evaluation is to start with the performance or organisational outcome. What should be different if the learning succeeds? Then identify the behaviour required, the learning needed and the experience that supports participation.

Backward-design questionExample: complaint-handling development
ResultResidents receive clearer, more timely responses and repeat complaints reduce
BehaviourAdvisers take ownership, use escalation routes and explain next steps accurately
LearningAdvisers practise case diagnosis, policy judgement and difficult conversations
ReactionLearners consider scenarios realistic, can access the programme and know why it matters

This approach connects evaluation to learning needs analysis. If the initial problem diagnosis is weak, later evaluation will also be weak because there is no valid outcome to assess.

Level 1: reaction with purpose

Reaction measures whether learning was relevant, accessible and usable. It should go beyond “Did you like the course?” Useful questions explore whether the content reflected real work, whether the pace and format were accessible, whether practice felt meaningful and what barriers participants expect when applying learning.

Weak reaction questionStronger question
“How satisfied were you?”“How closely did the scenarios reflect decisions you face at work?”
“Was the facilitator good?”“What will make it easier or harder to apply this learning in the next month?”
“Would you recommend this session?”“Which part of the programme requires more practice or clarification?”

Reaction evidence can identify design and access problems early. It should be reviewed promptly, particularly for new programmes or large-scale change.

Level 2: learning and demonstration

Learning evaluation should test the capability that matters. A knowledge quiz may be suitable for factual understanding, but judgement, communication or technical skill often requires a simulation, observed practice, case analysis or work sample. Pre- and post-assessment can show change, but should be used fairly and proportionately.

At fictional Glenmoor Utilities, supervisors attend a programme on incident escalation. Instead of relying only on a quiz, participants work through realistic scenarios, identify escalation thresholds and explain their rationale. Facilitators record whether the decision is accurate, timely and aligned with policy. This produces richer evidence than completion data alone.

Level 3: behaviour and transfer

Behaviour is often the most difficult level because work conditions influence transfer. People may know what to do but lack time, authority, manager support, system access or reinforcement. Evaluation should therefore consider both participant behaviour and the environment in which it occurs.

Transfer conditionEvaluation questionPossible support
Manager reinforcementDo managers observe, discuss and support new practice?Manager briefing, observation prompts and coaching
OpportunityDo participants encounter relevant tasks soon after learning?Planned practice, stretch assignment or simulation refresh
Tools and processDo systems enable the intended behaviour?Job aids, system changes and accessible escalation routes
Social normsIs the behaviour encouraged by peers and leaders?Peer review, calibration and recognition
WorkloadDo people have time to apply learning safely?Capacity planning and prioritisation

Glenmoor samples incident records and asks managers to observe escalation conversations over the following months. It also checks whether supervisors can access the escalation system during field work. If behaviour does not change, the answer may be a process improvement rather than more training.

Get Your Assignment Done by Experts

Free · 60 seconds

Reading this before you write? Make sure your draft hits the criteria.

Answer 3 questions and we’ll tell you exactly where you need support.

Find out

Level 4: results and attribution

Results should be relevant to the original purpose. Glenmoor might track incident response time, quality of escalation records, repeat safety events and regulator feedback. But it should avoid claiming that learning alone caused a change, especially if staffing, systems or policy also changed.

Result measureAttribution question
Faster incident responseDid new digital tools or staffing changes also affect the result?
Fewer repeat errorsWas there a wider process redesign at the same time?
Improved customer outcomesDid the affected employees have opportunity to apply the new behaviour?
Reduced risk costAre estimates based on credible assumptions and a relevant baseline?

Use comparison groups, baseline trends, stakeholder insight and documented assumptions where feasible. This connects to return on investment in people interventions, which considers financial, operational, workforce and ethical outcomes together.

Strengths and limits of the Kirkpatrick model

StrengthLimitationPractical response
Moves attention beyond attendance and satisfactionCan be treated as a rigid sequenceSelect proportionate evidence for the decision
Highlights behaviour and resultsResults are hard to attribute to learning aloneCombine data, qualitative insight and context
Encourages evaluation planningMay overlook structural barriers to performanceAssess transfer conditions and work design
Provides a shared languageCan lead to over-measurementFocus on the few outcomes that matter

The model is most valuable when it raises better questions rather than becomes a compliance form. Not every intervention requires a full four-level evaluation. A short information update may need only reaction and knowledge evidence; a high-cost leadership programme or safety-critical capability change may justify deeper behaviour and results analysis.

Workplace application: Glenmoor Utilities

Glenmoor’s programme achieves high reaction scores and improved scenario performance, but field observations show limited change in escalation behaviour. Investigation reveals that supervisors still need manager approval before using the new process and mobile signal is unreliable at several sites. The learning design is improved, but the more important action is to revise decision rights and provide offline system access. The evaluation prevents Glenmoor from blaming participants for a transfer failure caused by the work system.

Evaluation design, evidence governance and improvement

Evaluation evidence should be governed before launch. Glenmoor can identify the evaluation owner, the measures, data sources, baseline, review dates, confidentiality safeguards and the decisions that findings will inform. This prevents post-programme reporting from selecting only positive outcomes. It also clarifies where learning data will be combined with performance, customer or safety data and who may access that information.

Use evaluation as a learning cycle. If reaction is low because scenarios do not reflect work, revise the design. If learning is high but behaviour is low, inspect manager reinforcement, authority, tools and workload. If behaviour changes but results do not, test whether the selected result measure is valid or whether other system factors dominate. The purpose is not to prove that a programme succeeded; it is to improve capability and decide whether to continue, adapt, scale or stop the intervention.

Governance and inclusive evaluation

Evaluation should consider who can access learning and who benefits. Segment participation, reaction, learning and behaviour data where appropriate to identify whether shift workers, remote employees, part-time staff or other groups face barriers. Protect confidentiality, especially with small cohorts, and avoid using learning data punitively where it is intended for development.

Set evaluation ownership before launch. Programme leads may collect learning evidence; managers may support behaviour observation; operational leaders may own results measures; and people teams should ensure data quality, fairness and interpretation. Review findings with participants and stakeholders so that learning programmes are improved transparently.

Reporting evaluation evidence responsibly

Present results with their limitations. Glenmoor should distinguish participant feedback, demonstrated learning, observed behaviour and organisational trends rather than collapse them into a claim that the programme “worked”. This allows executives to see where the evidence is strongest, where transfer is blocked and what action is required next. It also protects participants from unfair conclusions when a wider system constraint prevents behaviour change.

Frequently asked questions

What are the four levels of the Kirkpatrick model?

Reaction, learning, behaviour and results. They provide a framework for considering whether learning was relevant, understood, applied and connected to meaningful outcomes.

Is participant satisfaction enough to show that training worked?

No. Satisfaction can be useful feedback, but it does not show whether learners gained capability, changed behaviour or improved results.

Does every programme need all four Kirkpatrick levels?

No. Use evaluation proportionately. The more costly, strategic or high-risk the intervention, the stronger the case for evaluating behaviour and results.

How does Kirkpatrick connect to ADDIE?

ADDIE provides a design process; Kirkpatrick provides an evaluation structure. Evaluation needs should be considered during ADDIE’s analysis and design stages, not added at the end.

References

Kirkpatrick, J.D. and Kirkpatrick, W.K. (2016) Kirkpatrick’s four levels of training evaluation. Alexandria, VA: ATD Press.

CIPD (2025) Learning needs analysis. Available at: https://www.cipd.org/en/knowledge/factsheets/learning-needs-factsheet/ (Accessed: 24 August 2026).

Salas, E., Tannenbaum, S.I., Kraiger, K. and Smith-Jentsch, K.A. (2012) ‘The science of training and development in organisations’, Psychological Science in the Public Interest, 13(2), pp. 74–101.