The skills day went well. Thirty nurses rotated through the stations, the evaluations came back glowing, and everyone signed the attendance sheet. Three months later, the director asks a simple question at the quality meeting: did it work? You realize you can show that people came and that they enjoyed it. You cannot show that anything on the units changed.
That gap is common in healthcare education, and it is not a sign of bad teaching. It is a sign that measurement was not planned before the training started. Evaluation added afterward almost always defaults to what is easiest to collect: attendance and satisfaction. Those numbers matter, but they answer a different question from the one leaders are asking.
This guide is for nurse educators, clinical managers and staff development leaders. It explains the four levels of Kirkpatrick's evaluation model, what data to collect at each level in a nursing setting, how competency validation fits in, and how to build a measurement plan before the first session. It complements our article on measuring the return on investment of clinical training, which deals with converting results into money; this one is about knowing whether the training worked at all.
Donald Kirkpatrick set out his approach in a series of articles titled "Techniques for Evaluating Training Programs," published in the Journal of the American Society of Training Directors in 1959 and 1960. It remains the most widely used framework for evaluating workplace training. The four levels are:
The levels form a chain. Good reactions make learning more likely, learning makes behavior change possible, and behavior change is what produces results. But each link can break. Nurses can enjoy a session and learn little; they can pass a skills check and never change practice, because the unit's equipment, workload or culture works against it.
The updated version of the model, which James and Wendy Kirkpatrick call the New World Kirkpatrick Model, makes two points that matter for nursing. First, plan backward: start with the Level 4 result you want, decide which Level 3 behaviors would produce it, and only then design the learning. Second, Level 3 rarely happens on its own. It needs what they call required drivers: reinforcement, monitoring, accountability and support, such as job aids, coaching and follow-up from charge nurses and managers.
Reaction data is easy to collect and easy to dismiss, but a well-designed Level 1 survey tells you things you need. The key is asking about relevance and intent, not only enjoyment. "Did you like the session?" tells you little. "How relevant was this to your current patient population?" and "How confident are you that you can use this on your next shift?" tell you much more.
Keep it short, collect it before people leave, and include one open question such as "What will get in the way of using this on your unit?" The answers to that question are often the most valuable data in the whole evaluation, because they predict where Level 3 will break.
Level 2 asks whether participants actually gained what the training intended. In nursing education this usually means a combination of knowledge tests, observed skills performance and confidence ratings.
A pre-test and post-test on the same objectives shows knowledge gain; without a pre-test, a high post-test score may only show that the group already knew the material. For skills, use a structured checklist with observable criteria, applied by a trained evaluator. Confidence ratings are worth collecting too, but treat them with care: confidence and competence are related, not identical.
If the training uses simulation, the INACSL Healthcare Simulation Standards of Best Practice include a standard on the Evaluation of Learning and Performance that distinguishes formative evaluation (feedback to help the learner improve), summative evaluation (a judgment at the end of a period of learning) and high-stakes evaluation (an assessment with major consequences, such as progression or employment decisions). Decide in advance which one you are doing, and tell participants. A debrief that suddenly becomes a pass or fail decision damages trust; our article on why simulation debriefing matters explains why.
Level 3 is where training either reaches patients or does not, and it is the level most often skipped. It requires going to the unit after the training, usually weeks to months later, and looking for the behavior you trained.
Good Level 3 data in nursing includes direct observation with the same checklist used in training, documentation audits, compliance audits the unit already runs, structured feedback from preceptors and charge nurses, and self-reports collected with specific questions ("In the last two weeks, how many times did you use the new handover format?"). Choose two or three observable critical behaviors, not twenty. If you trained a handover structure, the critical behavior is that nurses use it at shift change, and you can observe that directly.
When Level 3 results disappoint, look at the required drivers before blaming the training. Was there a job aid at the point of care? Did charge nurses reinforce the new behavior? Did the equipment or the electronic record support it? Often the training worked and the environment undid it.
Level 4 looks for the organizational outcomes the training was meant to influence: quality and safety indicators, incident reports, patient experience measures, time to independent practice for new hires, first-year retention, or audit findings. Many of these are already tracked by your quality or HR departments, so start by asking what exists rather than building new collection.
Be honest about attribution. A drop in an outcome measure after training may come from the training, from a staffing change, from a new device, or from chance. Compare against a baseline period, look for a comparison unit that was not trained yet if you can, and present Level 4 findings as evidence of contribution, not proof of cause. Leaders tend to trust a modest, well-argued claim far more than a dramatic one.
Because results take time, the New World model also recommends tracking leading indicators: short-term signs that critical behaviors are on track and results are likely to follow. A rise in audit compliance in the first month is a leading indicator for a quality outcome measured at six months.
Competency validation and training evaluation overlap but are not the same thing. Competency validation answers a question about an individual nurse: can this person perform this skill to the standard required? Training evaluation answers a question about a program: did this training change knowledge, practice and outcomes across the group?
A validated skills checklist is often your best Level 2 tool, and the same checklist used on the unit later becomes a Level 3 tool. That is a strong reason to design your validation methods and your evaluation plan together. Accrediting bodies such as The Joint Commission expect hospitals to assess and document staff competence, so this data is usually being collected anyway; the question is whether anyone aggregates it to judge the training. Our guides to building a competency validation program and planning a skills fair cover the validation side in detail.
A measurement plan fits on one page. Write it when you design the training, agree it with the manager who asked for the training, and set dates for each data collection point. The table below shows the structure, with examples for a training program on a structured handover format.
| Level | Question | Data to collect | Method and timing | Owner |
|---|---|---|---|---|
| 1 Reaction | Did nurses find it relevant and usable? | Relevance, intent to use, expected barriers | Short survey before leaving the session | Educator |
| 2 Learning | Can they perform the handover to standard? | Pre and post knowledge check; skills checklist score; confidence rating | Pre-test at start; observed role play at end | Educator, trained evaluators |
| 3 Behavior | Do they use it at shift change? | Observed handovers against the same checklist; preceptor feedback | Unit observation at about 30 and 90 days | Unit educator, charge nurses |
| Leading indicators | Is practice heading the right way? | Audit compliance trend; staff-reported use | Monthly, from existing audits | Manager |
| 4 Results | Did the outcome the training targeted change? | The quality or safety measure agreed with leadership, against baseline | Quarterly for two to four quarters | Quality department, manager |
| Required drivers | Is the environment supporting the behavior? | Job aid in place; reinforcement in huddles; barriers reported | Checked at each Level 3 visit | Manager, educator |
Three rules keep the plan realistic. Measure fewer things well rather than many things badly. Use data that already exists wherever you can. And collect a baseline before training starts, because without one, even a real improvement is hard to show.
What are the four levels of the Kirkpatrick model? Reaction (how participants responded to the training), Learning (whether they gained the intended knowledge, skills and confidence), Behavior (whether they apply it on the job) and Results (whether the organization sees the intended outcomes). Donald Kirkpatrick first described them in a series of articles published in 1959 and 1960.
How do you measure behavior change after nurse training? Choose two or three observable critical behaviors, then look for them on the unit weeks to months after training using direct observation with the same checklist used in the session, documentation or compliance audits, and structured feedback from preceptors and charge nurses. Check whether job aids and reinforcement were in place at the same time.
Is competency validation the same as training evaluation? No. Competency validation judges whether an individual nurse can perform a skill to standard. Training evaluation judges whether a program changed knowledge, practice and outcomes across a group. They work best designed together, because a validated skills checklist can serve as evidence at both Level 2 and Level 3.
How long after training should you measure results? Reaction and learning are measured during or at the end of the session. Behavior is usually checked weeks to months later, once nurses have had chances to use the skill. Results often take several months or quarters to show, which is why tracking leading indicators such as audit compliance in the meantime is useful.
Can you prove that training caused an improvement in outcomes? Rarely with certainty, because staffing, equipment and patient mix also change. A baseline, a comparison unit where possible, and evidence that the trained behavior actually changed on the unit let you make a credible case that training contributed. Present it that way.
Wahero designs facility training with evaluation built in: clear objectives, validated skills checklists, and follow-up that looks for change on the unit, not only in the classroom. Our services for healthcare facilities include on-site simulation and skills training brought to your facility, and our simulation lab page explains how scenario-based training is structured. If you are planning a program and want the measurement plan designed alongside it, talk to our team.
Educational content only
This material is published by Wahero Health Institute for professional education and is not individual medical advice, a care protocol, or a substitute for clinical judgment. Always follow your facility's policies, your state's nurse practice act, and your own scope of practice, and confirm medication doses against a current authoritative reference before administration. See our Terms of Use.