The New Jersey Coalition Against Sexual Assault (NJCASA) acknowledges that prevention programs need evaluation. Without it, it is difficult to know whether a program reached the intended participants, was delivered as planned, or contributed to meaningful change. The difficulty is that evaluation tools capture only the changes they were designed to detect. A program might measure or evaluate attendance, knowledge, attitudes, intentions, or self-reported behavior. It may not tell us whether participants trust the institution delivering the program, whether they feel responsible for one another, or whether they can challenge harmful behavior without losing their place in a peer group. These are the conditions that influence how people use what they have learned, yet they are often treated as background rather than part of the prevention effort itself. This ultimately creates a gap between what prevention practitioners understand about social change and what institutions can formally demonstrate.
Evidence-based approaches can create an additional tension. A program’s evidence of effectiveness is generally tied to a particular model, population, and set of implementation conditions. Organizations may therefore be expected to deliver it with fidelity, even when its assumptions (about relationships, identity, family, or community) do not reflect the people in the room. Adaptations that make a program more culturally or generationally relevant may be discouraged because they depart from the version that was evaluated. Fidelity can protect the elements that make an intervention effective, but when treated too rigidly, it can also limit responsiveness and innovation, particularly for communities that were not adequately represented in the original research.
This raises a larger question about which practices get recognized as evidence based in the first place. Every established intervention began as an untested idea, yet emerging community-led approaches often struggle to gain institutional support before they have accumulated the very evidence that funding would help them produce. Meanwhile, communities historically underserved by prevention systems may also be underrepresented in research and evaluation data. Treating only already-validated models as credible can therefore reproduce the limitations of the existing evidence base and narrow the space available to develop better approaches.
An important gap is seen in the evidence base. The CDC’s Sexual Violence Prevention Resource for Action includes strategies that address social norms, relationship skills, economic supports, workplace policies, school environments, and community-level risks. It also states that community and societal-level prevention strategies have received less evaluation than individual and relationship-level interventions. Relatively few evaluations have examined whether programs work similarly across racially, ethnically, and sexually diverse populations. Individual programs are usually easier to evaluate. They have identified participants, set activities, and defined start and end dates. Community change is harder to isolate. Institutions change slowly because they consist of individuals who move between schools, families, workplaces, congregations, neighborhoods, and online networks. A change in one setting may be contradicted in another. The effects may emerge after the funding period has ended. These problems do not make community change impossible to study, but they do limit what any single evaluation can claim.
While community-level change can be challenging to quantify, challenges exist with the metrics traditionally used to measure impact of prevention work at the individual level. Imagine that a school-based prevention program reports an increase in students’ knowledge of consent and greater confidence in bystander intervention. Those findings are useful, but they still leave important questions unanswered. Can students identify coercive behavior when the person responsible is well liked? Do they believe adults will respond fairly if they report it? Are Black students, disabled students, queer and transgender students, and students whose first language is not English equally likely to trust the reporting process? Can a student challenge a friend without being pushed out of the group? Does the school apply its policies consistently when the accused person is an athlete, a high-achieving student, or the child of an influential parent? A prevention program may change what students know without changing what they believe they can safely do. That distinction is easy to miss when knowledge and confidence are the primary outcomes.
Reported incidents present a similar problem. An increase in reports after a prevention initiative may indicate greater trust, improved awareness, easier access to reporting, an actual increase in incidents, or several of these at once. A decrease may reflect reduced harm, but it could also reflect fear, resignation, or loss of confidence in the institution. The number alone cannot tell us which explanation is correct. Interpretation requires knowledge of the setting, that is, how reports were handled, whether people experienced retaliation, whether the reporting process changed, and whether those most affected considered the institution credible.
Words such as “trust” and “belonging” are often used so broadly that they lose their meaning and intention. They become clearer when attached to observable situations. Trust may mean that a young person can ask a difficult question without expecting punishment or ridicule. Or it may mean that an organization admits when it mishandled a report and changes its procedures. Or it may mean that rules are enforced even when doing so is inconvenient for the institution. Belonging may mean that someone can disagree with a group without being expelled from it. It may mean that participants do not have to hide parts of their identities to be taken seriously. It may mean having some influence over the rules rather than simply being welcomed into rules made by others. These conditions can affect prevention. A person who trusts an institution may raise a concern earlier, someone who expects support from peers may be more willing to intervene, and participants who see their lives reflected accurately in a program may take its content more seriously.
None of these relationships is automatic. A close community can protect people and conceal abuse. A family may provide care while demanding silence. A faith community may meet material needs while protecting a respected leader. A peer group may challenge a stranger’s behavior but defend the same behavior from one of its own members. Community cohesion, therefore, is not necessarily evidence of safety. We need to know whom the group protects and what kind of behavior is excused. And what happens to a person who names harm? When it comes to evaluations, it is also important to remember that the evaluation process may be beholden to forces greater than itself. For example, founders establish reporting requirements, while program designers select outcomes and evaluators choose the instruments. Participants supply data, but they may have little influence over the questions or the interpretation.
Black feminist scholarship offers a useful challenge to this arrangement. Patricia Hill Collins (1989) argues that knowledge is shaped by lived experience, dialogue, care, and accountability. Applied to prevention evaluation, this raises practical questions: Who was involved in defining success? Whose interpretation is treated as credible? Who has the power to explain an unexpected finding? Who is accountable when an evaluation misrepresents the people it describes? Rather than rejecting established research methods, these questions probe us to examine how those methods are selected and used. The CDC’s 2024 Program Evaluation Framework now gives greater attention to these concerns. It recommends examining local history and power relationships, involving people affected by a program throughout the evaluation, using multiple sources of evidence, and interpreting findings collaboratively. It also distinguishes routine performance measurement from the broader work of evaluation. That distinction is important. Counting attendance tells us whether people came. It does not tell us why some people stopped coming. A survey can measure whether attitudes changed. It may not explain why the change appeared among one group but not another. And while an interview can provide that context, interviews also have their own limits.
Organizations sometimes treat qualitative methods as the answer to everything numbers leave out. They conduct interviews or ask participants to describe how a program affected them because stories can reveal experiences that a standardized survey misses. But they can also be extracted and simplified. Take, for instance, a situation where a participant gives complicated or critical feedback about a program. The organization has the power to select the portion that makes the program look effective so as to remove uncertainty and negate criticism, which now turns a critique into a testimonial. Inviting people to speak does not give them control over how their words are interpreted. To make community participation more substantive, the participants should have more influence over the questions and be able to review interpretations and challenge conclusions. Although meaningful community involvement, attention to cultural and political context, and recognition of power differences are central features of a robust evaluation approach, considerable inconsistencies in how these principles are practiced arise. As such, we should be cautious about claiming that an evaluation is participatory simply because it includes a focus group.
Funders have a legitimate interest in knowing how money was used and whether a program contributed to change. Accountability is necessary, especially in a field where the penalty for ineffective practices is continued human suffering. However, problems arise when funders limit the acceptable evidence, require the same indicators across different communities, or treat short-term results as a final judgment about a program’s value. Under those conditions, evaluation begins to influence the program it is supposed to study in various ways because programs learn which outcomes are likely to support renewal. If funders continue to administer evaluations the way they traditionally do using a narrow scope, then it may encourage programs to seek and replicate narrow outcomes. In other words, if evaluation processes favor activities that produce quick changes over sustained interventions, recruit participants who are easiest to reach and retain instead of the most vulnerable, or avoid questions that could produce difficult findings, then that narrow scope of work is what gets replicated. Work that requires time (e.g., earning trust, repairing a damaged institutional relationship, adapting an approach with community members, or responding to a problem that was not anticipated in the original proposal) can look inefficient within a short grant period.
This pressure has been documented. A 2023 study of funder-mandated performance metrics found that reporting requirements influenced how nonprofit staff defined client success, related to clients, understood client motivation, and delivered services. Research on charities dependent on grant funding has similarly found that funder-imposed measures can redirect attention from an organization’s mission toward the funder’s requirements, increasing the risk of mission drift (Henderson and Lambert, 2018). Therefore, the problem is larger than the reporting burden. Funding requirements can change what programs do and which forms of change they pursue.
The imbalance becomes especially clear when prevention programs are held responsible for conditions controlled by larger institutions. A school can fund a prevention curriculum while maintaining reporting procedures that students distrust. A workplace can require harassment training while protecting supervisors who produce strong financial results. A public agency can request evidence of community-level change through a short-term grant while treating housing instability, school exclusion, transportation barriers, or economic insecurity as outside the program’s scope. When the desired outcome does not appear, the community organization is often asked to explain what went wrong. The institutions surrounding the program are less often required to examine how their own policies affected the result. When local programs are expected to repair conditions they did not create and cannot control, this arrangement can produce a “convenient” division of responsibility. Funders define the period in which change should occur, the indicators through which change is measured, and the evidence required for continued support. If the results are inconclusive, the program may lose funding before it has had enough time to understand the finding and also reach the level of impact because the time frame of the grant is not long enough for the change it is supposed to help realize.
A null result may indicate that a program is ineffective. Or it may reflect a small sample, an unsuitable measure, weak implementation support, an unrealistic timeline, or institutional resistance outside the program’s authority. These explanations should not be treated as interchangeable. Funding decisions that make no distinction among them discourage honest reporting and weaken the field’s ability to learn. Practitioners then face an unavoidable choice. They can report uncertainty and risk being judged as unsuccessful, translate partial or unfinished change into confident claims, or redesign the program around what the funding system is prepared to recognize.
Prevention evaluation should help us decide whether a program deserves continued support, requires revision, or should end. It should also examine the conditions under which the program was expected to succeed. That requires attention to the funder’s role by asking a series of questions that make accountability more complete. For example, was the evaluation period appropriate for the kind of change being sought? Were the required indicators relevant to the community? Did the grant provide enough resources to collect credible evidence? Could the program report disappointing or ambiguous findings without placing its survival at immediate risk? Did the policies of the school, workplace, agency, or other partner support the outcomes the program was funded to produce? Current evaluation systems often place the greatest burden of proof on organizations with the least control over the conditions being measured, and often with the least number of resources to adequately measure those successes (e.g., database platforms, staffing shortages). Community programs must document change and explain uncertainty, all while justifying continued investment.
The institutions that set the terms of the work are usually treated as part of the background, when instead they should enter the evaluation. If a funder requires rapid evidence of a slow change, that affects the result. If an institution teaches intervention while punishing people who intervene, that affects the result. If a program is defunded whenever its findings complicate the expected hypothesis, the field loses both the program and the opportunity to learn from it. The final evaluation should therefore describe the program’s outcomes, the limits of the evidence, and the institutional conditions that shape both. Any process rigorous enough to judge the work of prevention practitioners should be rigorous enough to examine the decisions of those who fund, restrict, and evaluate that work.
Discussion Questions
Evaluation affects people differently depending on whether they design the measures, collect the data, participate in the program, or decide whether funding continues. A useful conversation must include those different positions and allow room to disagree with the argument made here. The following questions invite prevention workers, funders, and community members to examine their own role in deciding what counts as evidence and what follows from those decisions.
- Think of a change your prevention program considers important but struggles to document. What evidence might help you understand that change without reducing it to something easier (but less meaningful) to measure?
- If an evaluation shows no measurable change, what would you need to distinguish among an ineffective program, an unsuitable measure, an unrealistic timeline, inadequate support, and resistance from the surrounding institution? Who currently bears the consequences when that distinction is not made?
- What might your evaluation reveal if people who left the program, declined to participate, or criticized it had as much influence over its conclusions as those who completed it?
- When reporting requirements begin to influence which populations a program serves, which activities it chooses, or which findings it emphasizes, has accountability become a form of program control? How would practitioners and funders recognize that line?
- For funders: Can the organizations you support report uncertainty, disappointing findings, or unintended consequences without endangering future funding? How does your response to difficult findings inadvertently inform what grantees may disclose or omit the next time?
- Schools, workplaces, agencies, families, and community institutions often expect prevention programs to produce change. Which of their own policies or practices might undermine that change, and what could we learn if the institution itself were evaluated alongside the program?
Works Consulted
Basile, Kathleen C., et al. Sexual Violence Prevention Resource for Action: A Compilation of the Best Available Evidence. National Center for Injury Prevention and Control, Centers for Disease Control and Prevention, 2016. Revised Apr. 2025, https://www.cdc.gov/violenceprevention/pdf/SV-Prevention-Resource_508.pdf. Accessed 15 July 2026.
Collins, Patricia Hill. “The Social Construction of Black Feminist Thought.” Signs: Journal of Women in Culture and Society, vol. 14, no. 4, summer 1989, pp. 745–73. https://doi.org/10.1086/494543.
Henderson, Elisa, and Vicky Lambert. “Negotiating for Survival: Balancing Mission and Money.” The British Accounting Review, vol. 50, no. 2, Feb. 2018, pp. 185–98. https://doi.org/10.1016/j.bar.2017.12.001.
Iverson, Melissa, et al. “The Unintended Influence and Impact: Funder-Mandated Performance Metrics, Service Delivery, and Social Justice.” Human Service Organizations: Management, Leadership & Governance, vol. 47, no. 5, 2023, pp. 385–403. https://doi.org/10.1080/23303131.2023.2215270.
Kidder, Daniel P., et al. “CDC Program Evaluation Framework, 2024.” Morbidity and Mortality Weekly Report: Recommendations and Reports, vol. 73, no. 6, 26 Sept. 2024, pp. 1–37. https://doi.org/10.15585/mmwr.rr7306a1.
Kushnier, Lauren, et al. “Culturally Responsive Evaluation: A Scoping Review of the Evaluation Literature.” Evaluation and Program Planning, vol. 100, Oct. 2023, article 102322. https://doi.org/10.1016/j.evalprogplan.2023.102322.




