Before You Measure Outcomes, Build the Framework

I was recently asked to stand up a working group focused on institutional effectiveness and outcomes. The request was simple on the surface: help the organization get better at understanding whether its work was making a difference.
That kind of request can quickly turn into a hunt for metrics. How many courses did we deliver? How many people did we train? How many reports did we produce? How many systems did we deploy?
Those questions matter, but they are not enough.
A colleague had already done useful work through a thoughtful white paper. It raised the right issues and gave us a starting point. As I read it, though, I kept coming back to one concern: if we jumped straight into measures, every program would define success in its own way. We would collect a lot of numbers, but we might not learn much.
Before we measured outcomes, we needed to build the framework.

The temptation to count what is easiest
Most organizations are already measuring something. The problem is that they often measure what is easiest to count.
That usually means activity counts and output counts.
A training team can count courses delivered. A policy team can count reports produced. A technology team can count systems deployed. A support team can count the number of people served. These numbers are useful for management, staffing, budgeting, and workload analysis.
They do not, by themselves, prove that anything improved.
A course delivered does not mean learning happened. A person trained does not mean behavior changed. A report produced does not mean leaders made better decisions. A system deployed does not mean users adopted it or that the mission became more effective.
This is where institutional effectiveness gets difficult. The closer we get to meaningful change, the harder the measurement becomes. The data may be less clean. The timeline may be longer. Many factors may affect the result. The program may only control part of the chain.
That does not mean we should avoid outcomes. It means we should be clear about what we are measuring and what we can honestly claim.
Activities, outputs, outcomes, and impact are not the same thing
One of the first things I proposed was that the working group establish shared terms. Not because definitions are academic, but because people can sit in the same meeting using the same words to mean different things.
When that happens, the measurement conversation becomes muddy fast.
Here is how I think about the chain.
Term | Plain-language meaning | Example |
Activities | What we do | Deliver a course, conduct an assessment, hold a workshop, build a tool |
Outputs | What we produce | Number of graduates, reports, recommendations, products, deployed systems |
Outcomes | What changes because of the work | Improved capability, better decisions, changed behavior, stronger performance |
Impact | The broader long-term result | A sustained institutional change or mission-level improvement |
The distinction matters because each level answers a different question.
Activities answer, “What did we do?”
Outputs answer, “What did we produce?”
Outcomes answer, “What changed?”
Impact answers, “What larger difference did the change support over time?”
I have seen programs get described as successful because they were busy. The team delivered. The schedule was met. Products were completed. Stakeholders participated. All of that may be true and valuable.
But busyness is not the same as effectiveness.
If we want evidence-based decision making, we have to connect activities outputs outcomes in a way that shows a credible line of reasoning. We need to explain how the work is expected to create change, what evidence would show that change, and what data we trust enough to use.
A framework gives people a common language
Without a common framework, measurement becomes a collection of local practices. One office tracks volume. Another tracks customer feedback. Another tracks completion rates. Another builds a dashboard. Each may be doing reasonable work, but the organization cannot easily compare, learn, or make decisions across programs.
That is why I suggested starting with a small cross-functional working group.
Not a large committee. Not a standing bureaucracy. A small group with enough range to understand the work from different angles.
The purpose would be to build the basic architecture:
Common definitions
A shared understanding of activities, outputs, outcomes, impact, indicators, evidence, baselines, targets, and assumptions.
Standard templates
A simple way for programs to describe their logic, expected outcomes, indicators, data sources, and evidence gaps.
Evidence expectations
Guidance on what counts as credible evidence for different types of outcomes.
A maturity model
A way to understand where each program is starting and what kind of work comes next.
The goal was not to make every program look the same. The goal was to make the conversation consistent enough that leaders could understand what they were seeing.
A shared evaluation framework helps answer basic executive questions. Are we clear on the outcomes? Do we know how we would observe progress? Do we have data? Is the data reliable? Are we learning from it? Are we changing decisions because of it?

The maturity model matters because programs are not starting from the same place
One of the biggest mistakes in outcomes work is assuming every program is ready for the same assignment.
They are not.
Some programs may still be defining what their real outcomes are. They may have strong activities and committed teams, but no clear statement of what should change as a result of their work.
Others may know their outcomes but lack indicators. They can describe the intended change, but they have not yet decided what they would measure.
Some may have indicators but weak data. They know what they want to track, but the data is inconsistent, incomplete, or hard to access.
A few may already have outcome indicators, baseline data, regular reporting, and a habit of using evidence to adjust what they do.
Those programs should not all receive the same task.
If the framework is too advanced, early-stage programs will perform compliance. They will fill in boxes without real clarity. If the framework is too basic, mature programs will feel slowed down and may lose momentum.
A maturity model helps avoid both problems.
It allows the organization to say, “Here is the common path, but the next step depends on where you are.”
A simple maturity model might include stages like these:
Maturity stage | What it looks like | The right next task |
Emerging | Activities are clear, but outcomes are vague | Define intended outcomes and key assumptions |
Developing | Outcomes are stated, but measures are limited | Identify indicators and possible data sources |
Established | Indicators exist, but data quality varies | Improve data quality and evidence standards |
Advanced | Data is trusted and reviewed regularly | Use evidence to adjust decisions and test assumptions |
This is not about labeling programs as good or bad. It is about giving each program work that is useful.
A program that is still defining outcomes should not be pressured to produce a polished dashboard. A program with strong data should not be asked to sit through basic terminology for months. The framework should be consistent, but the work assigned should match each group’s maturity.
Trusted data comes before confident claims
At some point in any outcomes conversation, someone will ask, “What does the data show?”
That is the right question, but it has to be followed by another one: “Do we trust the data?”
Real evidence starts with data we trust.
That sentence became a central idea for me.
Trusted data does not mean perfect data. Perfect data rarely exists. Trusted data means we understand where it came from, how it was collected, what it includes, what it excludes, and what limits it has.
If a program counts participants, do we know who qualifies as a participant? If a system tracks completion, do we know whether completion reflects actual use or just access? If a survey reports satisfaction, do we know who responded and who did not? If a report claims improvement, do we know the baseline?
Data quality is not a technical side issue. It affects whether leaders can make sound judgments.
Poor data can create false confidence. It can make a program look more effective than it is. It can also hide real progress because the evidence is scattered or inconsistent.
That is why evidence expectations matter. Programs need to understand what kind of data fits the outcome they are claiming.
Some outcomes may need quantitative indicators. Others may require qualitative evidence, such as structured interviews, document review, expert assessment, or before-and-after comparisons. In many cases, the strongest picture comes from multiple sources.
The working group’s role would not be to demand one kind of evidence for every situation. It would help define what credible evidence looks like for different claims.

Contribution is different from attribution
Another issue we needed to address early was contribution versus attribution.
Organizations often want to know whether a program caused a result. Sometimes that is possible to show. Many times, especially in complex institutional work, it is not.
A program may contribute to a broader result without being able to claim it caused that result.
That distinction matters.
If several programs, leaders, partners, policies, and external conditions all shape an outcome, it would be misleading for one program to take full credit. At the same time, it would be equally misleading to say the program had no value just because it cannot prove sole causation.
Contribution asks a more practical question: did the program play a credible role in supporting the observed change?
To answer that, a program needs a clear theory of how its work connects to the outcome. It needs evidence that its activities happened, that outputs reached the intended audience, that intermediate changes occurred, and that those changes align with the broader result.
For example, a leadership development program may not be able to prove that it caused an organization-wide performance improvement. Many other factors may have influenced that result. But it may be able to show that participants applied specific practices, teams changed certain routines, decision cycles improved, or managers reported better use of shared processes.
That is a contribution story. It is more honest than an attribution claim the evidence cannot support.
Good institutional assessment should make room for both types of evidence. When attribution is possible, pursue it. When it is not, build a credible contribution case and be clear about the limits.
The working group had to be small and practical
I proposed a small cross-functional working group because this work needed range without becoming unwieldy.
Outcomes measurement touches strategy, operations, data, policy, programs, reporting, and leadership decisions. No single function owns the whole picture. If the group is too narrow, the framework will miss important realities. If it is too large, progress will slow and the work will become performative.
The group I had in mind would focus on practical products, not endless discussion.
The first products would likely include:
A shared glossary
A simple logic model or outcome mapping template
A standard way to describe indicators and data sources
Evidence criteria for different kinds of outcome claims
A maturity model with clear stages
Guidance on how programs should identify their next step
The point was to create enough structure for consistency, but not so much structure that programs spent more time feeding the framework than improving their work.
That balance is hard. Too little structure produces confusion. Too much structure produces compliance behavior. People learn how to satisfy the template, even if the evidence is weak.
The right framework should help responsible people think clearly.
It should make the work easier to explain, not harder. It should support better questions. It should help leaders see where outcomes are clear, where evidence is strong, where assumptions are untested, and where decisions need to change.

The real purpose is better decisions
Outcomes work can easily become a reporting exercise. Once that happens, the energy goes into producing the artifact. The report gets polished. The dashboard gets updated. The briefing gets delivered.
But the real question is whether anyone makes a better decision because of the evidence.
If the evidence shows that an activity is not producing the expected change, do we adjust it? If one program has stronger results than another, do we learn from it? If data quality is too weak to support a claim, do we fix the data or stop making the claim? If a program has matured, do we ask more advanced questions?
Institutional effectiveness should create a learning loop.
That loop starts with clarity. It moves through evidence. It ends with decisions.
For me, the working group was never about creating a measurement machine. It was about helping the organization become more disciplined in how it understands its own performance.
That requires humility. It means admitting that a full calendar is not proof of value. It means separating what we did from what changed. It means being careful about claims. It means giving programs room to mature without lowering the standard for evidence.
The framework comes first because it protects the integrity of everything that follows.
If definitions are unclear, measures will be inconsistent. If maturity is ignored, programs will be given the wrong tasks. If data is not trusted, evidence will not persuade. If contribution and attribution are confused, claims will either be inflated or unfairly dismissed.
The work begins with a few basic questions and the discipline to answer them honestly:
What are we doing?
What are we producing?
What is changing?
What evidence supports that conclusion?
What should we do differently?



Comments