Theory of Change as a Data Model: A Practitioner Note
Most nonprofit data systems track activities — people served, sessions delivered, dollars disbursed. Almost none track the logical chain connecting those activities to the outcome the funder actually cares about. Here's how to build a system that does.
Every nonprofit program has a theory of change somewhere — usually in a grant proposal, a strategy deck, or a program director's head — describing the logical chain from activity to outcome: this training leads to this skill, which leads to this behavior change, which leads to this measured impact. The problem is that this chain almost never exists as a data model. The organization's actual systems track activities (attendance, sessions delivered, dollars disbursed) and, separately, track outcomes (surveys, follow-up assessments), with no structural link between them — which means the theory of change lives in a document, while the data that could validate or challenge it lives in spreadsheets that were never designed to talk to that document at all.
This gap is why so much monitoring-and-evaluation reporting ends up being activity counts dressed up as impact evidence — "we served 4,000 families" is an activity count, not evidence that the theory of change is working, and funders increasingly know the difference, even when a report doesn't make it easy to see.
What it means to model the theory of change, not just the activities
The practical version of this is a small set of explicitly linked entities:
Activity — what was delivered, to whom, when. This is what most systems already capture reasonably well.
Intermediate outcome — the specific, measurable change the theory of change claims should follow from the activity, defined precisely enough to check, not described in the aspirational language a grant proposal uses. "Improved financial literacy" isn't checkable. "Score increase on a specific 12-item assessment, administered at a specific interval post-activity" is.
Causal link, stated as a claim, not assumed. The connection between activity and intermediate outcome is a hypothesis the organization is implicitly making — and modeling it explicitly, as a link the data can support or fail to support, is what turns "we believe this program works" into something checkable rather than something merely asserted.
Long-horizon outcome — the actual impact the funder cares about, usually measurable only much later than the program activity itself, and usually the hardest of the four to instrument, which is exactly why it gets skipped most often, and exactly why skipping it is the most costly gap.
Why this is worth the modeling effort for a resource-constrained organization
The honest objection here is real: this is more upfront data work than most nonprofit teams have capacity for, and the instinct to skip it in favor of just reporting activity counts is not laziness, it's realism about limited staff time. The case for doing it anyway is specific, not general: the organizations that build even a lightweight version of this model spend dramatically less time, at reporting season, reconstructing the story a funder wants from disconnected spreadsheets — because the connective structure was built once, at the data layer, instead of being manually reassembled by a program officer every time a report is due. The upfront cost is real; it is also a cost paid once, while the reporting-season reconstruction cost repeats every cycle for as long as the gap exists.
It also changes what the organization can learn about itself. A system that only tracks activities can tell you what you did. A system that tracks the causal chain, even imperfectly, can tell you where the chain is actually breaking — whether it's the activity-to-intermediate-outcome link that's weak, or the intermediate-to-long-horizon link — which is a fundamentally more useful diagnostic than "the long-horizon outcome wasn't as strong as hoped," because it tells the program team what to actually change, rather than just that something needs to.

