top of page

Welcome to the VBNN Digital Library

Unlock a Vast Knowledge Ecosystem

Featuring over 30,000 books, academic papers, illustrations, and expert insights—continuously updated to support your research and professional growth.

Welcome to our library!

Here, you will find an exclusive collection created 100% by our own faculty, meaning you will not find these resources anywhere else. Over the last 20 years, our team has written much more than what is currently online, and we are actively working to upload our complete back catalog. We update our platform regularly, so be sure to check back from time to time. If you ever need help finding a specific resource, you can always contact us!

Maximize Your Access

Log in to instantly view and download tailored resources directly aligned with your specific program and curriculum.

Ready to begin? Sign in above to explore your personalized dashboard.

Please note: Login is only possible using your institutional email address; otherwise, the system will not recognize your account.

VBNN Library AI

Introducing our fully integrated Library AI. Designed to support your research, you may submit inquiries in any language and receive precise, evidence-based responses drawn exclusively from our published scholarly articles and textbooks.

Search...

Latest Publications:

Search this site

Results found for empty search

  • Translational Science Fundamentals (Moving Discoveries from Bench to Bedside)

    Download the Book (PDF): Introduction In 1983, researchers publishing in the most prestigious basic-science journals in the world were confident about the clinical promise of their work. Two decades later, a team led by Despina Contopoulos-Ioannidis and John Ioannidis went back and checked. They identified 101 articles from leading journals, published between 1979 and 1983, whose authors had explicitly claimed that a discovery held real promise for prevention or treatment. By 2002, only 27 of those discoveries had been tested in a randomized trial. Five had led to a licensed product. One was in wide clinical use. The finding, published in the American Journal of Medicine in 2003, is not an indictment of basic science. Most of those papers were good science. It is a measurement of how much happens, or fails to happen, between a discovery and a patient. A second number circulates in almost every lecture on this subject. In 2000, Andrew Balas and Suzanne Boren estimated that it takes about seventeen years for research evidence to reach routine clinical practice, and that only about fourteen percent of original research ever does. The figure has been criticised, refined and occasionally mocked, most usefully by Zoë Morris, Steven Wooding and Jonathan Grant, whose 2011 review in the Journal of the Royal Society of Medicine was titled "The answer is 17 years, what is the question." Their point was that the time lag depends entirely on where one starts the clock and where one stops it, and that different studies measure different segments of a long and branching path. That point is the beginning of this book. What "translation" means Translational science is the study of how knowledge moves from one kind of work to another: from a laboratory observation to a first human dose, from a trial result to a clinical guideline, from a guideline to what actually happens in a clinic on a Tuesday afternoon, and from an effective intervention to a change in the health of a whole population. It is distinct from translational research, which is the work of moving any particular discovery along that path. The science asks why the path is so slow, so lossy and so uneven, and what can be done about it. Over the past two decades, the field has converged on a vocabulary of numbered stages. T0 is discovery: the identification of a mechanism, a target, a biomarker or a candidate intervention. T1 is the move into humans: first-in-human studies and the early trials that establish safety, dose and proof of mechanism. T2 establishes efficacy in patients and turns that evidence into guidance. T3 is the move from guidance into practice: implementation, dissemination and the study of why proven interventions are or are not used. T4 is the move from practice to population: the outcomes that follow at scale, and the policies that shape whether an intervention reaches everyone who could benefit. Different institutions draw these lines in slightly different places, and the history of how the lines came to be drawn is itself instructive. What matters more than the precise boundaries is the recognition that each stage asks a different question, demands a different kind of evidence, is carried out by different people in different institutions, and is funded and rewarded by different means. The argument of this book The central claim of what follows is simple to state and has consequences that take the rest of the book to work through. Discoveries do not stall in the pipeline so much as at its joints. The characteristic failure of translation is not that the science at any one stage is bad, though sometimes it is. It is that each stage is optimised for its own question and hands the next stage something it was never designed to use. A preclinical study optimised to publish a striking mechanism produces an effect size that no human trial will reproduce. A Phase I study optimised to find the highest tolerable dose hands forward a dose that may be wrong for the patients who will take it for years. An efficacy trial optimised for internal validity, with carefully selected participants and expert investigators, hands forward a result that community clinics cannot reproduce with their own patients and staff. An implementation programme optimised for uptake in one setting hands forward a model that no payer will fund and no legislature will mandate. At every joint, a question that the next stage needed answered was not asked, because asking it was nobody's job. The remedy, which the book develops stage by stage, is to design each stage with the next one's question in view. That means preclinical studies built to predict human outcomes rather than merely to demonstrate mechanism; first-in-human studies that collect the pharmacology later trials will need; efficacy trials that test interventions in forms that can actually be delivered; implementation studies that measure what policymakers will ask about; and policies designed so that their effects can be evaluated. It also means recognising that the pipeline is not a line but a loop. Observations at the bedside and in populations constantly send questions back towards the bench, and some of the most productive translation has run in that direction. How the book is organised Chapter 1 sets out the T0 to T4 framework, its origins and its limits, and argues for reading it as a map of handoffs rather than a sequence of stages. Chapter 2 examines the first and most notorious joint, between preclinical discovery and the decision to test in humans, where problems of reproducibility and model validity cause most candidate interventions to fail before or shortly after they reach a person. Chapter 3 is devoted to the first-in-human study itself: how starting doses are chosen, what went wrong in the most instructive Phase I disasters, and how dose-finding designs have evolved. Chapter 4 follows candidates through the long and costly work of establishing efficacy, where attrition is highest and where the choice of endpoints and trial designs determines whether a result will mean anything outside the trial. The second half of the book turns from products to people and systems. Chapter 5 addresses the gap between evidence and practice, and the implementation science that has grown up to study it. Chapter 6 follows interventions into communities, where the tension between fidelity and adaptation plays out and where sustainability is decided. Chapter 7 examines T4, the translation of evidence into public health policy and the population effects that follow. Chapter 8 considers the pipeline as a system: the flows that run backwards from practice to discovery, the de-implementation of practices that should never have been adopted, and the infrastructure that institutions have built to connect the stages. The conclusion draws out what the whole analysis implies for the people who fund, conduct and use translational research. A note on scope Translational science spans drugs, biologics, devices, diagnostics, behavioural interventions and policies. This book draws examples from all of them but leans on two families of cases: pharmaceutical development, where the early stages are best documented, and public health prevention, where the late stages are. It is written from a largely American and British vantage point, because those are the systems whose translational infrastructure is most thoroughly studied, but the problems it describes are general. Wherever a figure or a study is cited, it is a real one, named so that the reader can find it. Where the evidence is contested, the book says so. The reader need not be a scientist. Anyone who has wondered why a promising headline about a new treatment so rarely becomes a treatment, or why a proven intervention sits unused for a decade, is asking the questions this field exists to answer. The answers turn out to be less about any single failure of intelligence or effort than about the structure of the enterprise itself, and structures, unlike laws of nature, can be redesigned. Chapter 1. The Pipeline and Its Joints The word "translational" entered the vocabulary of biomedical policy in the early 2000s, and it arrived as a diagnosis. In 2003, Nancy Sung and colleagues, writing for the Clinical Research Roundtable convened by the US Institute of Medicine, published an analysis in JAMA of the central challenges facing the national clinical research enterprise. They identified two "translational blocks." The first lay between basic biomedical research and its application in human studies. The second lay between the results of clinical studies and their adoption in everyday practice and decision-making. The same year, Elias Zerhouni, then director of the US National Institutes of Health, launched the NIH Roadmap for Medical Research, which named "re-engineering the clinical research enterprise" as one of its three themes and set in motion the institutional changes that would later produce the Clinical and Translational Science Awards. The two-block model had the virtue of simplicity, and it captured something real. American biomedical research budgets had doubled between 1998 and 2003, and there was a growing unease that the flood of basic discovery was not producing a corresponding flow of new therapies or improvements in health. But the model also concealed a great deal. It grouped together, under "the second block," activities as different as running a randomised trial, writing a clinical guideline, persuading a clinician to follow it, and persuading a government to pay for it. As researchers from different disciplines began to look at the path in detail, the number of blocks began to multiply. From two blocks to five stages The first elaboration came from primary care. In 2007, John Westfall, James Mold and Lyle Fagnan published a short and influential piece in JAMA titled "Practice-based research—'Blue Highways' on the NIH Roadmap." Their argument was that the move from a successful trial to routine practice is not a single step. Trials are conducted largely in academic medical centres, with selected patients and expert staff. Most patients, however, are seen in community practices, and there is a distinct body of research, conducted in those practices, that asks whether and how trial findings work there. Westfall and colleagues labelled this T3 and called for practice-based research networks to carry it out. In 2008, Steven Woolf wrote in JAMA on "The meaning of translational research and why it matters." Woolf observed that the term was being used for two quite different enterprises. For laboratory scientists and the pharmaceutical industry, translation meant turning discoveries into drugs and devices, work that depended on molecular biology, animal models and early-phase trials. For health services researchers, translation meant ensuring that proven interventions actually reached patients, work that depended on epidemiology, behavioural science, organisational research and policy analysis. The two groups competed for the same money under the same label while needing entirely different expertise, infrastructure and timelines. Woolf's warning was that the first enterprise, being closer to the traditional interests of academic medicine, would tend to absorb the funds intended for both. The fullest version of the framework came from genomic medicine and epidemiology. In 2007, Muin Khoury and colleagues at the US Centers for Disease Control and Prevention, writing in Genetics in Medicine, described a continuum of translation research for genomic discoveries running from T1 through T4. In 2010, Khoury, Marta Gwinn and John Ioannidis extended it in the American Journal of Epidemiology in a paper titled "The emergence of translational epidemiology: from scientific discovery to population health impact." They added a T0 stage for the discovery research that precedes all application, and argued that epidemiology had roles to play at every stage, not just the last. The resulting five-stage scheme is now the most common reference point. As Table 1 sets out, each stage answers a distinct question, uses characteristic kinds of study, and ends with a handoff that the next stage depends on. Table 1. The five translational stages and their handoffs. Stage Central question Typical studies What it hands forward T0 Is there a mechanism or target worth pursuing? Laboratory, animal, genomic and observational discovery research A candidate intervention and a rationale T1 Is it safe and active in humans, and at what dose? First-in-human, Phase I, early Phase II, proof of mechanism A dose, a safety profile, early signals T2 Does it work in patients, and should it be recommended? Phase II and III trials, systematic reviews, guidelines Evidence of efficacy and a recommendation T3 Is it adopted and delivered well in real practice? Implementation, dissemination and effectiveness research A deliverable programme and its uptake T4 Does it improve population health, and for whom? Outcomes research, policy evaluation, surveillance Population impact and policy Other institutions draw the lines differently. The US National Center for Advancing Translational Sciences, established in December 2011, describes a spectrum from basic research through pre-clinical, clinical, clinical implementation and public health research, and deliberately depicts it without arrows so as to emphasise that each stage can feed any other. Some schemes fold Phase II and Phase III into T1; others place guideline development in T3. These are not trivial differences, because the labels determine which funders and which review panels consider an application. But they share the essential insight: the journey from discovery to population health has several distinct segments, and a problem that is solved in one segment can remain entirely unsolved in the next. Why the joints matter more than the stages It is tempting to read the T-stages as a production line, in which each station does its work and passes the product along. The metaphor misleads in two ways, and both matter for everything that follows. First, the stations are not run by the same organisation. On a factory line, a single management can see that the paint shop is producing parts that the assembly shop cannot use, and fix it. In translational science, T0 is conducted largely by academic laboratories funded by research councils and rewarded by publications in high-impact journals. T1 and T2 for drugs are conducted largely by companies, regulated by agencies such as the US Food and Drug Administration and the European Medicines Agency, and rewarded by market approval. T3 is conducted by health systems, professional societies and a relatively small community of implementation researchers, and is rewarded, if at all, by quality metrics and payment incentives. T4 is conducted by public health agencies and legislatures and is rewarded by political and electoral considerations. No single actor is responsible for the whole line, and none is rewarded for the quality of what it hands forward. Second, each stage defines success in its own terms, and those terms do not always serve the next stage. A preclinical paper succeeds if it is published and cited. A Phase I study succeeds if it identifies a dose that can be carried forward without unacceptable toxicity. An efficacy trial succeeds if it achieves statistical significance on its primary endpoint. An implementation programme succeeds if the intervention is adopted. A policy succeeds if it is enacted. Each of these definitions is reasonable in isolation. Each can be satisfied while failing the stage that follows. A published mechanism can be unreproducible. A tolerable dose can be far higher than necessary. A statistically significant result can depend on a population that looks nothing like the patients who will use the treatment. An adopted programme can be delivered so poorly that it has no effect. An enacted policy can go unevaluated. The consequence is that the most useful way to read the T-framework is as a map of handoffs, and to ask of each joint three questions. What does the receiving stage need to know? Who, at the sending stage, is responsible for finding it out? And what would it cost them to do so? When the answers are "a great deal," "nobody," and "more than they are rewarded for," the joint will leak. Much of the history of translational science since 2003 can be read as a series of attempts to change those answers, by building infrastructure that spans stages, by changing reporting and design standards, and by creating incentives for the sending stage to care about what happens downstream. The seventeen-year problem, measured properly The best-known summary statistic in the field, the seventeen-year lag estimated by Balas and Boren in 2000, was derived by adding together estimates of the time taken by successive steps: the delay from submission to publication, from publication to inclusion in reviews and textbooks, and from there to implementation. It is a composite of averages taken from different studies, which is why Morris, Wooding and Grant, in their 2011 review, found that estimates of time lags in the literature varied enormously and depended on how each study defined its start and end points. Their review is worth dwelling on because it demonstrates the value of the joint-by-joint view. Some lags are between stages, as when a positive trial result waits years for a guideline to incorporate it. Some are within stages, as when a trial takes years to recruit. Some are parallel, as when regulatory review and guideline development proceed simultaneously. Speeding up translation requires knowing which of these dominates for a given kind of intervention, and the answer is different for a cancer drug, a surgical technique, a vaccine and a behavioural programme. There is no single number because there is no single path. What the detailed studies do show, consistently, is that loss is at least as important as delay. The Contopoulos-Ioannidis study described in the introduction found that of 101 highly promising discoveries, only one reached wide clinical use in twenty years. That is not a story of slow progress along a path; it is a story of nearly everything falling off. At the other end of the pipeline, Elizabeth McGlynn and colleagues at RAND, in a 2003 study in the New England Journal of Medicine based on medical records and telephone interviews with adults in twelve US metropolitan areas, found that participants received about 55 percent of the care recommended for their conditions. The failure there is not that the evidence was slow to arrive. It had arrived, been summarised, and been turned into quality indicators. It was simply not being acted upon about half the time. One discovery, five joints The value of the handoff view is clearest when a single intervention is followed all the way along the path. The human papillomavirus vaccine is one of the few for which the whole journey, from laboratory discovery to measured population effect, is now documented, and each of its joints taught a different lesson. At T0, the decisive work was the identification of specific papillomavirus types in cervical cancer tissue. Harald zur Hausen's group in Germany reported HPV16 in 1983 and HPV18 in 1984, against a prevailing view that a herpesvirus was the more likely culprit; the work later earned him a share of the 2008 Nobel Prize in Physiology or Medicine. Establishing that a virus was a necessary cause of most cervical cancers turned a disease of uncertain origin into a potentially preventable infection. But a causal virus is not a vaccine. The critical technical step came in the early 1990s, when Ian Frazer and Jian Zhou at the University of Queensland, and independently groups at the US National Cancer Institute and elsewhere, showed that the viral L1 capsid protein could assemble itself into virus-like particles that carried no genetic material yet provoked a strong antibody response. At T1 and T2, the question became whether those particles were safe, immunogenic and protective in people. A problem arose that recurs in many prevention programmes: the outcome that mattered, invasive cancer, takes decades to develop. The trials therefore used high-grade precancerous lesions as their endpoint, which was defensible biologically but meant that the claim to prevent cancer rested, at licensure, on a surrogate. The quadrivalent vaccine was approved in the United States in June 2006 on the strength of trials such as FUTURE II, reported in the New England Journal of Medicine in 2007. At T3, the programme met a different kind of obstacle. The vaccine worked best if given before sexual debut, which meant vaccinating children and young adolescents against a sexually transmitted infection, an idea that provoked resistance in some countries. Delivery systems mattered enormously: Australia, which began a publicly funded school-based programme in 2007, achieved high coverage quickly, while countries relying on opportunistic vaccination in primary care lagged for years. The same vaccine, with the same efficacy, reached very different fractions of its intended population depending on how it was delivered. At T4, the long-awaited population evidence eventually arrived. A Swedish registry study by Jiayao Lei and colleagues in the New England Journal of Medicine in 2020, covering more than 1.6 million girls and women, found that those vaccinated before the age of 17 had an incidence of invasive cervical cancer close to 90 percent lower than those who were not vaccinated. Scottish data published in 2024 found no cases of invasive cervical cancer among women who had been fully immunised at 12 or 13 in the routine programme. These findings in turn fed policy: the World Health Organization launched a global strategy to eliminate cervical cancer as a public health problem in 2020, and in 2022 its advisory group concluded that a single dose could offer protection comparable to two or three, a change that dramatically lowers the cost of reaching low-income countries. Each joint in this story required a different kind of expertise and was owned by a different set of institutions. At each one, the intervention could have stalled, and in some countries it did. None of the difficulties at T3 or T4 could have been solved by better science at T0. That is the practical meaning of saying that translation fails at the joints. Two directions of travel One further limitation of the linear picture needs to be stated at the outset, because it will recur throughout the book. Knowledge does not flow only from bench to bedside. It also flows back. Some of the most consequential discoveries in medicine began with an observation in patients that sent researchers back to the laboratory. The recognition that a particular chromosomal abnormality was present in the white cells of patients with chronic myeloid leukaemia, first described by Peter Nowell and David Hungerford in Philadelphia in 1960, began at the microscope with patient samples. It took decades of laboratory work to identify the fusion gene responsible and the abnormal enzyme it produced, and then a targeted drug, before imatinib came back to patients at the end of the 1990s. Clinical trials that fail are another source of reverse flow: a well-designed negative trial can reveal that the mechanism studied in animals is not the one operating in humans, and send the question back to T0. Population surveillance at T4 can reveal an unexpected safety signal, a disparity in benefit, or an unanticipated effect that generates new hypotheses for the laboratory. This is why NCATS removed the arrows. The practical significance is that a good translational system does not merely push discoveries forward. It also has channels, and people, whose job is to carry questions backwards, and it treats the information in failures as seriously as the information in successes. Chapter 8 returns to this in detail. What the framework is for A framework of this kind earns its keep if it helps people make better decisions. For a laboratory scientist, the T-framework is a reminder that a discovery's value depends on questions that will be asked at later stages, and that a study designed with those questions in mind is worth more than one designed only to be published. For a clinical investigator, it is a reminder that the trial population, the endpoints and the form of the intervention all determine whether the result will be usable in practice. For a funder, it is a way of seeing where the pipeline is thinnest: historically, far more money has gone to T0 than to T3 and T4 combined, although the precise balance is hard to measure because funders classify their portfolios differently. For a policymaker, it is a way of recognising that evidence of efficacy is not evidence of population effect, and that the absence of population evidence may reflect the absence of anyone paid to generate it. The chapters that follow take the joints in order. The first, between discovery and the decision to test in humans, is where the greatest number of candidates are lost and where the scientific problems are, in some ways, the most fundamental. It is also the joint where the incentives of the sending stage, academic discovery science, are least aligned with the needs of the receiving one. Chapter 2. The First Joint: Why Promising Discoveries Fail to Travel In 2012, C. Glenn Begley, who had spent a decade as head of global cancer research at the biotechnology company Amgen, and Lee Ellis of the MD Anderson Cancer Center published a short comment in Nature that has been cited thousands of times since. Over the preceding decade, Begley's team had attempted to confirm the findings of 53 papers they considered landmark studies in preclinical cancer research, before committing resources to drug programmes built on them. They were able to confirm the scientific findings in only six, about eleven percent. A year earlier, Florian Prinz and colleagues at Bayer had reported in Nature Reviews Drug Discovery that in-house attempts to reproduce published data on potential drug targets had matched the published results in only about a fifth to a quarter of projects. These industry reports had limitations. The companies did not publish which papers they had tested or exactly how, so the claims themselves could not be checked, which is an irony the critics did not fail to note. But they were soon joined by more transparent evidence. The Reproducibility Project: Cancer Biology, a collaboration between the Center for Open Science and Science Exchange, set out to repeat key experiments from high-profile cancer biology papers published between 2010 and 2012. It reported its final results in eLife in 2021. Of the 193 experiments originally selected, the team was able to complete only 50, from 23 papers, in large part because the original papers did not report enough methodological detail and the original authors could not or would not supply it. Among the experiments that were completed, the replication effect sizes were, at the median, about 85 percent smaller than those originally reported. For translational science, these findings describe the condition of the first joint. When a company, a funder or an academic team decides to take a discovery towards human testing, it is relying on a body of preclinical evidence that is, on this evidence, often considerably weaker than it appears. The decision to move to T1 is expensive, and it is taken on the basis of information that the sending stage was not designed to make reliable. Where the weakness comes from It would be comforting to attribute the problem to misconduct, but fraud accounts for a small part of it. Most of the weakness has ordinary causes that operate on honest researchers. The first is design. Preclinical experiments, especially in animals, have historically been small, often unrandomised, and rarely blinded. A series of systematic reviews led by Malcolm Macleod, Emily Sena and colleagues in the CAMARADES collaboration, looking across animal studies of stroke and other conditions, found that studies which did not report randomisation or blinded outcome assessment tended to report larger treatment effects than those which did. That pattern is exactly what one would expect if unconscious bias in allocating animals or scoring outcomes inflates apparent effects. The second is publication bias. Sena and colleagues, in a 2010 paper in PLoS Biology, analysed data from animal studies of stroke and estimated that publication bias alone accounted for roughly a third of the apparent efficacy reported in that literature. Studies with null results are less likely to be written up, less likely to be accepted, and less likely to be cited, so the visible evidence is a selected sample of the more favourable results. The third is analytic flexibility. When an experiment can be analysed in several defensible ways, and the researcher chooses among them after seeing the data, the chance of finding a statistically significant result by chance rises well above the nominal five percent. John Ioannidis's much-discussed 2005 essay in PLoS Medicine, "Why most published research findings are false," made the general argument that small studies, small effects, flexible designs and fields with many competing teams all reduce the probability that a published positive finding is true. Preclinical biology has historically combined all of these. The fourth is the reward structure. A laboratory scientist is rewarded for novelty. A striking result in a high-impact journal can secure grants, promotions and a career. A careful confirmation of someone else's finding, or a well-conducted negative result, rarely does. The sending stage at the first joint is therefore systematically incentivised to produce exactly the kind of evidence the receiving stage finds least reliable: novel, dramatic, underpowered and unconfirmed. Leonard Freedman, Iain Cockburn and Timothy Simcoe attempted to put a price on the consequences. Their 2015 paper in PLoS Biology, drawing on published estimates of irreproducibility rates, suggested that about half of US preclinical research might be irreproducible and that the cost of that irreproducible research was on the order of 28 billion dollars a year. The authors were explicit that the estimate was rough, and the underlying irreproducibility rates are themselves contested. The order of magnitude, however, gives a sense of what is at stake at this joint. The validity problem Even perfectly reproducible preclinical findings can fail to translate, because a reproducible result in a model is not the same as a true result in humans. This is the problem of external or predictive validity, and it is at its most acute with animal models of complex human diseases. Stroke provides the most studied example. Victoria O'Collins and colleagues, in a 2006 systematic review in the Annals of Neurology, catalogued 1,026 experimental treatments for acute ischaemic stroke that had been tested in animals. Of these, 114 had been tested in patients. Only thrombolysis with tissue plasminogen activator had established itself as effective, and aspirin had a modest role. One compound, the free-radical trapping agent NXY-059, became a cautionary tale: it met the field's own recommended criteria for preclinical evidence, produced a marginally positive result in the SAINT I trial, and then showed no benefit in the larger SAINT II trial reported in the New England Journal of Medicine in 2007. Subsequent analyses of its preclinical record found that the animal studies had been small, had often used young, otherwise healthy animals rather than the elderly, hypertensive and diabetic population that suffers most strokes, and had often given treatment soon after the stroke rather than hours later as happens in clinical practice. Sepsis is another. In 2013, Junhee Seok and colleagues, in a large collaborative study in the Proceedings of the National Academy of Sciences, compared genomic responses to trauma, burns and endotoxin in humans and in the mouse models commonly used to study them, and reported that the mouse responses correlated poorly with the human ones. The paper was widely read as an indictment of mouse models of inflammation. A reanalysis of the same data by Keizo Takao and Tsuyoshi Miyakawa, published in the same journal in 2015, reached the opposite conclusion, finding substantial correspondence when the analysis focused on genes that changed significantly in both species. The dispute is instructive in itself. Whether a model is "valid" depends on which features of the disease one needs it to reproduce, and that can only be judged in relation to the specific therapeutic question being asked. Alzheimer's disease offers perhaps the starkest record. Jeffrey Cummings, Travis Morstorf and Kate Zhong, in a 2014 analysis in Alzheimer's Research & Therapy, examined drug-development programmes registered between 2002 and 2012 and calculated a failure rate of 99.6 percent. Many candidates had cleared amyloid or improved cognition in transgenic mice that carried human familial mutations, a form of the disease accounting for a small minority of patients. The eventual approval of antibodies such as lecanemab, which produced modest slowing of decline in large trials, came after two decades of failures, and the debate over how much benefit those drugs provide is still active. The general lesson is that a model's validity has several components that are easy to conflate. Face validity is whether the model looks like the disease. Construct validity is whether it is produced by the same causal mechanism. Predictive validity is whether interventions that work in the model work in patients. A model can have high face validity and poor predictive validity, as some stroke models appear to, or can be useful for one question and useless for another. The first joint leaks badly when models are chosen for convenience and familiarity rather than for their demonstrated ability to predict the human outcome that the next stage will measure. The valley of death To the scientific problems at the first joint must be added an economic one. A discovery made in an academic laboratory is typically far from being a product. Before it can be tested in humans, a candidate drug must be optimised for potency and selectivity, formulated, manufactured to a standard fit for human use, and put through a package of safety pharmacology and toxicology studies designed to satisfy regulators. These activities are expensive, unglamorous, and poorly suited to academic funding, which rewards hypotheses and publications rather than development. Yet they are too risky, at this early stage, to attract most commercial investment. The resulting gap has been called the valley of death since at least the early 2000s, and Declan Butler's 2008 feature in Nature, "Translational research: crossing the valley of death," helped to fix the phrase. Many public programmes since have been designed specifically to bridge it. In the United States, NCATS runs programmes that provide academic investigators with access to industrial-style drug-development expertise, and its Therapeutics for Rare and Neglected Diseases programme was created to advance candidates that no company was likely to pursue. Several universities have established their own drug-discovery units. Venture philanthropy, in which patient foundations fund early development directly, has also played an important role, most famously when the Cystic Fibrosis Foundation invested in the research that led to the drug ivacaftor. The valley of death is not only about money. It is also about knowledge. A laboratory scientist who has discovered a promising target may not know what regulators will require before a first-in-human study, what properties will make a molecule a viable drug, or what the clinical community needs to see before it will enrol patients. The bridging programmes that have worked best supply that knowledge as much as they supply money. Designing the first joint to hold Over the past fifteen years, the response to these problems has taken several forms, and they share a common logic: they make the sending stage accountable for what the receiving stage needs. The first is reporting standards. The ARRIVE guidelines (Animal Research: Reporting of In Vivo Experiments), first published by Carol Kilkenny and colleagues in PLoS Biology in 2010 and revised as ARRIVE 2.0 by Nathalie Percie du Sert and colleagues in 2020, set out what information an animal study needs to report for readers to judge its reliability, including sample-size calculation, randomisation, blinding and the handling of excluded animals. A 2012 paper in Nature by Story Landis and colleagues, arising from a US National Institute of Neurological Disorders and Stroke workshop, called for a core set of transparent-reporting standards along similar lines. Journals and funders have adopted these standards unevenly, and adherence has lagged endorsement, but the direction is clear. The second is funder requirements. From 2016, the NIH began requiring grant applicants to address the rigour of the prior research on which their proposals rested, the rigour of their own designs, the consideration of sex as a biological variable, and the authentication of key resources such as cell lines and antibodies. The sex-as-a-biological-variable policy, announced by Janine Clayton and Francis Collins in Nature in 2014, responded to evidence that preclinical research had relied disproportionately on male animals and cells, so that effects in females were often simply unknown when a product entered human testing. The third is preregistration and confirmatory design. Borrowing from clinical trials, some preclinical researchers now register their hypotheses and analysis plans before collecting data, which removes the analytic flexibility that inflates effect sizes. A more ambitious approach is the multicentre preclinical randomised trial, in which a promising intervention is tested in parallel in several independent laboratories using a common protocol. A 2015 study by Gemma Llovera and colleagues in Science Translational Medicine, which tested an anti-CD49d antibody in experimental stroke across six European centres, showed both that such trials are feasible and that they can overturn findings from single laboratories: the effect was seen in one stroke model but not in another. The fourth is a change in the models themselves. Human-derived systems, including organoids grown from patient stem cells, microphysiological "organ-on-a-chip" devices, and computational models, promise in some applications to predict human responses better than animals do. Regulators have begun to accommodate them. The FDA Modernization Act 2.0, signed into US law in December 2022, removed the statutory requirement that a drug be tested in animals before human trials, allowing sponsors to use other methods where they are adequate. In April 2025, the FDA announced a plan to phase out animal testing requirements for monoclonal antibodies and other drugs, in favour of what it called new approach methodologies. How quickly this changes practice remains to be seen, and for many questions, especially those involving the whole-body interaction of a drug with multiple organs over time, no current alternative fully substitutes for an animal study. One further development deserves separate mention because it reverses the usual direction of evidence at this joint. Instead of starting with a mechanism in a model and hoping it applies to people, researchers increasingly start with evidence from people. Large genetic studies can show that individuals who naturally carry variants disabling a particular gene have lower rates of a disease, which is a kind of natural experiment on what would happen if a drug inhibited that gene's product. The development of PCSK9 inhibitors for lowering cholesterol followed this path: families with gain-of-function mutations in PCSK9 had very high cholesterol, and people with loss-of-function variants had low cholesterol and fewer heart attacks, before any drug had been tested. Matthew Nelson and colleagues at GlaxoSmithKline, in a 2015 analysis in Nature Genetics, found that drug targets with this kind of human genetic support were roughly twice as likely to succeed in development as those without it, and a later analysis by Eric Minikel and colleagues in Nature in 2024 reached a similar conclusion with more data. Human genetic evidence does not replace laboratory work, but it answers, before the first joint is crossed, a question that animal models answer poorly: does modulating this target matter in humans at all? Finally, some of the most important changes are in how decisions are made at the joint. Several pharmaceutical companies now routinely attempt to reproduce key academic findings internally before investing in a programme. AstraZeneca, after a review of its own pipeline published by David Cook and colleagues in Nature Reviews Drug Discovery in 2014, adopted what it called the "5R" framework, asking whether a project had the right target, right tissue, right safety, right patient and right commercial potential, and reported that its later success rates improved. The right patient criterion is especially significant: it asks, at the earliest stage, which people the drug will eventually be tested in, and whether the preclinical evidence speaks to them. That is exactly the kind of downstream question that the first joint has historically neglected. What the first joint hands forward When the first joint works, what passes across it is not merely a molecule and a hypothesis. It is a candidate with reproducible evidence of activity in models chosen for their relevance, an understanding of how it is absorbed, distributed and eliminated, a toxicology package that identifies the organs most at risk, and a biomarker or other measure that can show, in the first human studies, whether the drug is engaging its target. It is also a clear statement of uncertainty: what the preclinical evidence does not show, and therefore what the first human studies will need to find out. Most candidates, even after all this, will fail in humans. But the goal at the first joint is not to eliminate failure; it is to make failure informative and to make it happen early, when it is cheap. A candidate that fails in Phase I because it does not engage its target in humans has taught the field something. A candidate that fails in Phase III because the animal effect was an artefact of unblinded scoring has wasted years and, more seriously, exposed patients to risk for no scientific return. The next chapter follows the candidate across the joint, into the first human beings who will receive it. Chapter 3. Crossing Into Humans: The Logic and Risk of Phase I On the morning of 13 March 2006, eight healthy young men were dosed in a private clinical research unit at Northwick Park Hospital in north-west London. They were the first human beings to receive TGN1412, a monoclonal antibody designed to activate a subset of immune cells by binding to a receptor called CD28. Six received the drug and two received a placebo. The doses were given at intervals of about ten minutes. Within ninety minutes, the six who had received the active drug began to experience severe headache, back pain and fever. Within hours, they had developed what Ganesh Suntharalingam and colleagues, reporting the case in the New England Journal of Medicine later that year, described as a cytokine storm: a massive release of inflammatory signalling molecules leading to multi-organ failure. All six were admitted to intensive care. All survived, but some suffered lasting harm. The dose they had received, 0.1 milligrams per kilogram of body weight, was about five hundred times lower than the highest dose that had caused no adverse effects in cynomolgus monkeys. By the conventional rules for selecting a first human dose, it was cautious. The problem was that the conventional rules assumed that toxicity in animals would predict toxicity in humans, and for this drug, acting on this target, it did not. Subsequent investigation suggested that the immune cells most responsible for the reaction in humans did not express CD28 in the same way in the monkeys used for testing. TGN1412 is the single most influential event in the modern history of first-in-human research, and it illustrates precisely what Phase I is for. It is the stage at which the entire preclinical case is put to its first real test, and where the gap between models and humans becomes, for the first time, directly visible in a living person. What Phase I is trying to learn The classical purpose of a Phase I study is to establish the safety and tolerability of a new intervention in humans, to characterise how the body handles it (its pharmacokinetics) and what it does to the body (its pharmacodynamics), and to identify a dose or range of doses for further study. For most drugs outside oncology, these studies are conducted in healthy volunteers, usually in specialised units, with single ascending doses given to successive small cohorts, followed by multiple ascending doses given over several days. In oncology, the logic is different. Cancer drugs have historically been cytotoxic, damaging healthy cells as well as tumours, and it would be unethical to expose healthy people to them. Oncology Phase I trials are therefore conducted in patients with advanced cancer who have usually exhausted standard options. That changes the ethical calculus: participants may hope for benefit, and the studies have a therapeutic as well as a scientific character. It also changes the design logic, because for cytotoxic drugs the working assumption was that more drug meant more effect, so the goal was to find the highest dose patients could tolerate, known as the maximum tolerated dose. The regulatory scaffolding for Phase I is broadly similar across major jurisdictions. In the United States, a sponsor must submit an Investigational New Drug application to the FDA, including preclinical pharmacology and toxicology data, manufacturing information and a clinical protocol, and may proceed after thirty days unless the agency places the study on hold. In the European Union, clinical trials are now authorised through a single application under the Clinical Trials Regulation, which became applicable in January 2022. In both systems, an independent ethics committee or institutional review board must also approve the protocol, and participants must give informed consent. Choosing the first dose The most consequential single decision in a first-in-human study is the starting dose. Too low, and many cohorts of volunteers will be exposed to a drug at doses that cannot possibly do anything, wasting time and money and raising its own ethical questions. Too high, and the first cohort may be harmed. The traditional approach, codified by the FDA in a 2005 guidance on estimating the maximum safe starting dose in healthy volunteers, begins with the no observed adverse effect level (NOAEL) in the most appropriate animal species, converts it into a human equivalent dose using scaling factors based on body surface area, and then divides by a safety factor, by default ten, to give a maximum recommended starting dose. This approach rests on toxicity: it asks how much drug can be given before something goes wrong in animals. After TGN1412, the UK government convened an Expert Scientific Group, chaired by Gordon Duff, whose report in December 2006 recommended that for high-risk agents, particularly those acting on the immune system with a novel mechanism, the starting dose should instead be based on the minimal anticipated biological effect level, or MABEL. This approach asks not how much drug causes toxicity but how much drug is needed to produce any pharmacological effect at all, using all available data on how the drug binds to its target and what fraction of receptors it occupies at a given concentration. For TGN1412, a MABEL-based starting dose would have been very much lower than the one used. The European Medicines Agency issued a guideline on first-in-human trials for potential high-risk medicinal products in 2007, incorporating MABEL, and substantially revised it in 2017. The other lessons of TGN1412 concerned the conduct of the study. All six active participants had been dosed before any had shown symptoms, because the dosing interval was shorter than the time the reaction took to develop. The revised guidance now emphasises sentinel dosing, in which one or two participants receive the drug first and are observed for an appropriate period before the rest of the cohort is dosed, and dosing intervals chosen on the basis of the drug's expected pharmacology. It also emphasises that studies of high-risk agents be conducted in units with immediate access to intensive care. The second warning: BIA 10-2474 In January 2016, a Phase I trial in Rennes, France, of BIA 10-2474, a drug that inhibits the enzyme fatty acid amide hydrolase, was stopped after one participant was declared brain dead and several others were hospitalised with neurological damage. The affected participants were in a multiple-ascending-dose cohort receiving 50 milligrams a day, the highest dose tested. Unlike TGN1412, the drug had not caused any comparable effect in animals at the doses tested, and other drugs in the same class had been given to humans without serious harm. Investigations by the French medicines agency and an expert committee identified several contributing concerns: the steep escalation between the previous cohort and the one in which harm occurred, doses far above those needed to fully inhibit the target enzyme, and the possibility that the drug had off-target effects on other enzymes in the brain. The episode reinforced a lesson that TGN1412 had already taught. The data that matter for dose escalation are not just whether any adverse effects have been seen, but whether each further increase in dose is actually needed. If the drug has already fully engaged its target at a lower dose, going higher adds risk without adding any therapeutic information. The 2017 revision of the EMA guideline, which was prepared with the Rennes events in mind, emphasises integrating pharmacokinetic, pharmacodynamic and target-engagement data into escalation decisions. Rules for climbing the dose ladder In oncology, where Phase I studies are conducted in patients and the aim has traditionally been to find the maximum tolerated dose, the design of dose escalation has been a subject of intense statistical debate. The approaches in widest use differ in how they decide whether to increase, hold or reduce the dose for the next cohort, and in how efficiently they identify the right dose, as Table 2 summarises. Table 2. Common dose-escalation designs in early oncology trials. Design How escalation is decided Main strength Main weakness 3+3 Fixed rules based on toxicities in cohorts of three Simple, familiar, needs no statistician at the bedside Often misidentifies the target dose; treats many patients at low doses Continual reassessment method (CRM) A statistical model of dose and toxicity, updated after each cohort Uses all accumulated data; identifies the target dose more accurately Requires modelling expertise; perceived as opaque Bayesian optimal interval (BOIN) Compares the observed toxicity rate at the current dose with pre-set boundaries Nearly as accurate as model-based designs; simple to run Still focused on toxicity rather than benefit The 3+3 design, which dates in essence from the 1970s and 1980s, treats three patients at a dose; if none experiences a dose-limiting toxicity, the next three receive a higher dose; if one does, three more are added at the same dose; if two or more do, escalation stops and the dose below is usually declared the maximum tolerated dose. It is transparent and easy to run. It is also, according to a large body of simulation studies, poor at identifying the dose it aims to find, and it treats a large fraction of patients at doses well below any likely to be effective. The continual reassessment method, introduced by John O'Quigley, Margaret Pepe and Lloyd Fisher in Biometrics in 1990, fits a statistical model relating dose to the probability of toxicity and updates it after each patient or cohort, recommending the dose whose estimated toxicity is closest to a target rate. It uses information more efficiently and more patients are treated near the target dose. Its adoption was slowed by concerns, some justified in its earliest versions, that the model could escalate too aggressively, and by the practical need for statistical support throughout the trial. The Bayesian optimal interval design, described by Suyu Liu and Ying Yuan in 2015, occupies a middle ground: it uses pre-calculated decision boundaries that can be written down in a table at the start of the trial, so that it is as easy to run as the 3+3 design while performing much closer to model-based methods. When the maximum is not the optimum All three designs share an assumption inherited from cytotoxic chemotherapy: that the right dose is the highest dose patients can tolerate. For many modern cancer drugs, that assumption is wrong. Targeted therapies that inhibit a specific enzyme may achieve full target inhibition at doses well below those that cause intolerable side effects. Immunotherapies may have flat dose-response relationships across a wide range. For such drugs, pushing to the maximum tolerated dose adds toxicity without adding benefit, and patients who take the drug for months or years may suffer side effects that lead them to reduce doses or stop treatment altogether. The FDA's Oncology Center of Excellence launched Project Optimus in 2021 to address this, and in 2024 published final guidance on optimising the dosage of oncology drugs. It asks sponsors to compare multiple doses, often in randomised fashion, before committing to a registrational trial, and to characterise the relationships between dose, exposure, activity and safety. A frequently cited illustration is sotorasib, a drug targeting a mutated form of the KRAS protein, which received accelerated approval in May 2021 at a dose of 960 milligrams daily. As a condition of approval, the FDA required a post-marketing trial comparing that dose with a quarter of it, reflecting uncertainty about whether the lower dose would be as effective with fewer side effects. This is a case where the handoff at the second joint, from Phase I to later development, had been designed around the wrong question. A Phase I design optimised to find the highest tolerable dose was handing forward a dose that later stages, and patients, did not need. The correction required changing what Phase I was asked to deliver. The ethics of being first The ethical issues in Phase I differ between healthy-volunteer studies and patient studies, but both concern the relationship between risk and benefit when the participant is the first human being to take a drug. For healthy volunteers, there is no prospect of direct benefit, and participation is typically paid. The concern is whether payment induces people to accept risks they would otherwise refuse, and whether the population of habitual paid volunteers, often economically marginal, bears a disproportionate share of the burden of drug development. The empirical record on harm is, on the whole, reassuring. Ezekiel Emanuel and colleagues, analysing data from healthy-volunteer Phase I studies conducted by a large pharmaceutical company in a 2015 paper in the BMJ, found that serious adverse events were rare and most adverse events were mild. The rarity of disasters such as TGN1412 and BIA 10-2474 is part of what makes them so instructive. For patients in oncology Phase I trials, the concern is different: whether patients understand that the primary purpose of the trial is to find a dose, not to treat them, and whether the prospect of benefit is realistic. Elizabeth Horstmann and colleagues, in a 2005 analysis in the New England Journal of Medicine of 460 Phase I oncology trials conducted between 1991 and 2002, found an overall response rate of 10.6 percent and a rate of death attributed to toxicity of 0.49 percent, with response rates higher in trials that included an established anticancer agent. More recent analyses have reported higher response rates in the era of targeted therapy. These data have been used to argue that Phase I participation offers more potential benefit than had often been assumed, and that describing these trials as purely non-therapeutic in consent discussions is inaccurate. They have also been used to argue that the true chance of benefit varies so widely by drug and setting that honest consent requires trial-specific information, not general reassurance. A different response to the ethics of first-in-human exposure is to reduce what needs to be learned at full dose. The FDA's 2006 guidance on exploratory Investigational New Drug studies allowed very limited early human studies, sometimes called Phase 0, in which subtherapeutic microdoses are given to a small number of people to learn about how a drug is distributed or whether it reaches its target. Such studies cannot establish safety or efficacy, but they can end a programme early, at low risk, if the drug behaves in humans very differently from how it behaved in animals. What Phase I should hand forward The standard output of Phase I, a recommended dose and a list of observed adverse events, is necessary but not sufficient. The next stage needs to know not just what dose can be tolerated but whether the drug reached its target, whether it engaged it, and whether that engagement produced the expected biological effect. These three questions, sometimes called the three pillars of survival in drug development, were articulated in a 2012 analysis by Paul Morgan and colleagues at Pfizer in Drug Discovery Today. Reviewing the company's Phase II programmes, they found that many failures were in programmes where it had never been established that the drug had reached and engaged its target at the doses tested. Without that information, a negative Phase II result cannot distinguish between a drug that does not work and a drug that was never given a fair chance. A Phase I study designed with the next stage in view therefore measures pharmacodynamic biomarkers alongside safety, explores more than one candidate dose, and characterises how exposure varies between people. It hands forward not just a number but an understanding of the dose-exposure-response relationship on which every later decision will depend. That understanding becomes indispensable in the stage the next chapter examines, where most candidates that survive Phase I will fail. Hashtags: #TranslationalScienceFundamentals #TranslationalScience #BenchToBedside #TranslationalResearch #T0ToT4 #PreclinicalResearch #FirstInHumanStudies #PhaseITrials #ClinicalTrials #ImplementationScience #DisseminationScience #PopulationHealth #Reproducibility #PredictiveValidity #ValleyOfDeath #DrugDevelopment #DoseOptimization #Pharmacokinetics #Pharmacodynamics #TargetEngagement #EvidenceToPractice #ImplementationResearch #TranslationalMedicine #ClinicalResearch #FutureOfTranslationalScience

  • The Peer Review Ecosystem (Ethics, Editorial Roles, and Constructive Critique)

    Download the Book (PDF): Introduction The first review most scientists write arrives without ceremony. An email from an editor they have never met, a manuscript title that sits somewhere near their own work, a deadline two or three weeks away, and a link to a submission system with a text box and a drop-down menu of recommendations. There is rarely any training attached. The graduate student or postdoctoral researcher who clicks "accept invitation" is expected to know, somehow, what the editor wants, what the authors deserve, what counts as a fatal flaw and what counts as a matter of taste, and how to say all of it in prose that will be read by strangers who have spent years on the work in question. Most early-career reviewers fill this gap by imitation. They recall the reviews they have received themselves, the sharp ones that stung and the lazy ones that missed the point, and they write something in between. Some become harsh, on the theory that rigour and severity are the same thing. Some become timid, afraid that a junior person has no standing to criticise a senior laboratory. Many produce long lists of small objections, because small objections are easy to find and feel like diligence. Very few are ever told whether their reviews were any good. This book is written for that reviewer. It treats peer review not as a single act of judgement but as an ecosystem: a set of roles, each with distinct obligations, that together produce the decisions which shape what enters the scientific record. Authors, reviewers, handling editors, editors-in-chief, publishers, research integrity officers, readers who comment after publication, and increasingly the platforms that publish reviews openly all occupy positions in this system. A reviewer who understands only their own position will misjudge what their report is for. A reviewer who understands the whole system can write a report that does its job. The argument The controlling claim of this book is simple to state and harder to practise. The ethics of peer review and the craft of peer review are the same subject. Fairness, confidentiality, disclosure of conflicts, and honesty about the limits of one's own expertise are usually taught, when they are taught at all, as a compliance layer: rules to observe before the real intellectual work begins. That framing is wrong. A review distorted by an undeclared rivalry is not an ethical failure sitting beside a technically excellent assessment; it is a bad assessment. A review that hides its uncertainty behind confident language misleads the editor about the evidence. A review that is contemptuous in tone gets its substantive points ignored, which means the flaws it found are more likely to survive into print. Constructive critique is not critique with the edges sanded off. It is critique that is accurate about the work, accurate about the reviewer's own position, and aimed at a decision someone else must make. Three commitments follow from that claim, and they run through every chapter. The first is that a review is written for two audiences with different needs. The editor needs help making a decision: is this work sound, is it important enough for this venue, and what would it take to fix? The authors need help improving the work, whatever the decision. A report that serves only one of these audiences is half a report. The second is that the reviewer is a witness, not a judge. Reviewers advise; editors decide. This matters practically, because it changes how a recommendation should be framed, and ethically, because it limits what a reviewer is entitled to do when something looks wrong. A reviewer who suspects manipulated data is not the investigator, the prosecutor, or the jury. They are the person who noticed, and their obligation is to report what they noticed clearly and to the right person. The third is that peer review is a human process with known and measurable failure modes. Reviewers disagree with each other far more than most scientists assume. They miss errors deliberately planted in manuscripts. They are swayed by the prestige of authors and institutions. They are, at times, cruel. None of this is a reason for cynicism, but all of it is a reason for method. A reviewer who knows where the process tends to fail can build habits that guard against those failures in their own work. What the chapters do The first two chapters establish the system. Chapter 1 asks what peer review is for, tracing how a practice that most scientists assume is ancient became standard only in the second half of the twentieth century, and summarising what the evidence says about its reliability. Chapter 2 maps the editorial system: who does what between submission and decision, what editors actually need from reviewers, and how a reviewer's report is used once it leaves their hands. Chapters 3 and 4 are about the core craft. Chapter 3 sets out a method for reading a manuscript critically, from the first pass that establishes what the paper claims to the detailed interrogation of design, analysis, and reporting. Chapter 4 turns to writing: how to structure a report, how to separate essential problems from preferences, how to phrase criticism so that it is heard, and how to make a recommendation that respects the editor's role. Chapters 5, 6, and 7 address the ethical terrain directly, though each treats ethics as part of accurate assessment. Chapter 5 covers conflicts of interest and confidentiality, including the newer question of whether a reviewer may feed a confidential manuscript to a generative AI tool. Chapter 6 examines implicit bias: what the experimental evidence shows about status, gender, and confirmation effects in review, and what an individual reviewer can do about biases they cannot see directly. Chapter 7 deals with disputes over data, from suspected image manipulation to authors who refuse to share the data behind their claims, and sets out what a reviewer should and should not do when something looks wrong. Chapter 8 looks at the changing shape of the system itself. Open identities, published reports, reviewed preprints, and portable reviews that travel between journals are no longer experiments at the margins. Several prominent journals now publish reviewer reports routinely, and at least one major life-sciences journal has abandoned accept-or-reject decisions after review altogether. These models change what a review is, who reads it, and what it is worth to the person who wrote it. The conclusion draws these threads into an argument about what an early-career reviewer should actually do differently, and about which problems in the system remain open. What this book leaves out The book concentrates on manuscript review for journals and preprint review platforms, because that is where most early-career scientists begin. Grant review shares many of the same ethical principles, and the discussion of conflicts, bias, and confidentiality applies there with little change, but the mechanics of study sections and funding panels differ enough that they are treated only in passing. The book also does not attempt to cover every discipline's conventions. Its examples come mostly from the life, physical, and social sciences, where the journal article remains the main unit of communication. Reviewers in fields where conference proceedings dominate, such as much of computer science, will find the principles transferable, and some of the most instructive evidence about bias in review comes from exactly those venues. Finally, the book does not offer templates to be filled in. Checklists have a place, and several appear in the chapters where they help, but a review assembled from a checklist without judgement reads like one. The aim is to give a new reviewer a clear enough picture of the system, and of their own role within it, that they can make good judgements in situations no checklist anticipated. A note on standing Early-career scientists often ask whether they have the standing to review at all, particularly when the authors are senior. The answer is that standing in peer review comes from expertise in the specific question at hand, not from seniority. An editor who invites a postdoctoral researcher usually does so because that person has recently worked with the precise technique, dataset, or model system the manuscript depends on, and often knows its pitfalls better than anyone else in the field. The appropriate response to that invitation is neither deference nor bravado but candour: say what you can assess with confidence, say what lies outside your competence, and do the work carefully. That candour is the thread that connects everything that follows. It is what makes a review useful to an editor, fair to authors, and worth the hours it takes to write. Chapter 1. What Peer Review Is For Ask a room of scientists when peer review began and many will point to the seventeenth century. In 1665 Henry Oldenburg, secretary of the Royal Society of London, launched the Philosophical Transactions, and the journal is often described as the origin of the practice. The description is misleading in an instructive way. Oldenburg edited the Transactions as a private venture, and he selected material largely on his own judgement and through his correspondence. The Royal Society did not take formal responsibility for the journal until 1752, and the system of sending papers to members for written reports developed gradually through the nineteenth century. Even then, it was one practice among several. Many journals were run by editors who decided alone or with a small circle of trusted advisers. The historian Melinda Baldwin has shown how recent the modern expectation really is. Nature, for most of its history after its founding in 1869, relied heavily on editorial judgement, and it did not require external refereeing for all research papers until the 1970s. Across much of science, universal external review became standard only in the decades after the Second World War, driven by the explosive growth of research funding, the resulting flood of submissions, and a growing need for scientific institutions to demonstrate accountability to the governments that paid for them. Peer review, in Baldwin's account, became a marker of scientific legitimacy in public life at roughly the same time it became routine inside journals. A well-known episode from 1936 captures the transition. Albert Einstein and Nathan Rosen submitted a paper on gravitational waves to Physical Review, arguing that such waves might not exist. The editor, John Tate, sent it to a referee, whose anonymous report identified a serious error. Einstein, accustomed to the German journals where editors published the work of eminent authors without external review, was indignant that the manuscript had been shown to a colleague before publication. He withdrew it and published elsewhere. The historian Daniel Kennefick later established that the referee was the cosmologist Howard Percy Robertson, and that Robertson subsequently helped Einstein's assistant understand the problem. The published version, when it appeared in the Journal of the Franklin Institute, reached a quite different conclusion from the original. The referee had been right. The story is often told as a joke at Einstein's expense. It is better read as a reminder that the practice now taken for granted was contested within living memory, and that its authority rests less on tradition than on whether it actually improves the work that passes through it. The functions a review serves Peer review is asked to do several jobs at once, and much confusion about how to review well comes from failing to separate them. At least four can be distinguished. The first is quality control: checking that the methods are sound, the analyses appropriate, and the conclusions supported by the evidence. This is the function most scientists have in mind when they speak of peer review, and it is the one a reviewer is best placed to perform, because it draws directly on technical expertise. The second is selection: judging whether the work is important, novel, or interesting enough for the particular venue. A methodologically flawless study may still be a poor fit for a journal that publishes only work of broad significance. Selection judgements are more subjective than quality judgements, more dependent on the journal's editorial policy, and more vulnerable to bias. Many journals now ask reviewers to separate the two explicitly, and some, notably PLOS ONE since its launch in 2006 and a number of other "soundness-only" journals since, have removed the importance criterion altogether, asking reviewers to judge whether the work is rigorous rather than whether it is exciting. The third is improvement. Reviews routinely make manuscripts better: they catch errors, suggest additional controls or analyses, point to missing literature, and force authors to state their claims more precisely. Surveys of authors consistently find that most believe their published papers were improved by review, even when they resented the process. The improvement function operates regardless of the decision. A rejected paper, revised in light of thoughtful reports, often goes on to a better version at another journal. The fourth is certification. Publication in a peer-reviewed venue signals to readers, hiring committees, funders, journalists, and courts that the work has passed some form of expert scrutiny. This is the function that gives peer review its public weight, and it is also the one most often overstated. Review certifies that a small number of experts, working with limited time and without access to the raw data, found no fatal problem with the manuscript as presented. It does not certify that the findings are true. A reviewer's report can serve all four functions, but not equally. It serves quality control and improvement best when it is specific and technical. It serves selection best when the reviewer states their view of the work's significance separately from their view of its soundness, so that the editor can weigh each against the journal's own priorities. And it serves certification best when it is honest about what the reviewer could and could not check. What the evidence says about reliability For a practice so central to science, peer review was subjected to systematic study remarkably late. The International Congress on Peer Review and Scientific Publication, first held in Chicago in 1989 under the sponsorship of JAMA, began to change that, and a substantial body of research now exists. Its findings are sobering, and an early-career reviewer should know them. Reviewers agree with each other far less than intuition suggests. A 2010 meta-analysis by Lutz Bornmann, Rüdiger Mutz, and Hans-Dieter Daniel, published in PLOS ONE, pooled studies of inter-reviewer agreement on journal manuscripts and found levels of agreement that were low by any conventional standard: reviewers assessing the same manuscript agreed only modestly better than would be expected by chance. Earlier, Peter Rothwell and Christopher Martyn had examined reviews of submissions to two clinical neuroscience journals and conference abstracts, reporting in Brain in 2000 that agreement between reviewers on whether to publish was little better than chance. Grant review shows similar patterns. Low agreement is not in itself proof that review is broken. Editors deliberately choose reviewers with different expertise, and two reviewers who assess different aspects of a paper will naturally reach different overall views. But low agreement does mean that any single review is a noisy signal, and that the recommendation at the bottom of a report is far less informative than the reasoning above it. That is one of the most practical lessons in this book: editors can combine and weigh reasoning, but they cannot do much with a bare verdict. Reviewers also miss errors. In a study published in JAMA in 1998, Fiona Godlee, Catharine Gale, and Christopher Martyn took a paper that had already been accepted by the BMJ, introduced eight deliberate weaknesses in design, analysis, and interpretation, and sent it to hundreds of reviewers. On average, reviewers identified only about two of the eight. A later study led by Sara Schroter, published in the Journal of the Royal Society of Medicine in 2008, inserted nine major errors into three papers and sent them to more than six hundred BMJ reviewers; the average reviewer detected fewer than three of the major errors, and training produced only modest improvement. In emergency medicine, William Baxt and colleagues reported in 1998 on a fictitious manuscript seeded with errors and sent to all reviewers of the Annals of Emergency Medicine; many recommended acceptance despite fundamental flaws. These studies share a design that flatters no one: a manuscript deliberately constructed to be flawed, sent to reviewers who had no reason to expect it. They probably overstate how often errors survive in the ordinary course of review, where several reviewers and an editor scrutinise the work, and where authors are usually not trying to hide anything. But they establish beyond reasonable doubt that individual reviewers, working under normal conditions, miss a large share of serious problems. Systematic reviews of whether editorial peer review improves the quality of published research have reached cautious conclusions. A Cochrane review led by Tom Jefferson in 2007 found little rigorous evidence either way, largely because few studies had been designed to answer the question well. Later work has found that reporting quality tends to improve between submission and publication, but the size of the effect varies and the mechanism is not always clear. What review cannot do The limitations of peer review follow from its structure, and they are worth stating plainly, because a reviewer who expects the process to do something it cannot will review badly. Review cannot detect fraud reliably. A reviewer sees a manuscript, not a laboratory. If an author fabricates data competently, presents internally consistent results, and describes plausible methods, there is usually no way for a reviewer to know. The major fraud cases of the past quarter-century, from the fabricated organic transistor results of Jan Hendrik Schön at Bell Labs, exposed in 2002, to the fabricated stem-cell lines of Hwang Woo-suk, whose papers in Science were retracted in 2006, passed through review at the most selective journals in the world. They were uncovered by readers, colleagues, and whistleblowers after publication, often because someone noticed duplicated figures or impossible consistency across supposedly independent experiments. Chapter 7 returns to what a reviewer can do when something looks wrong. The honest starting point is that review is not designed as a fraud detector and should not be judged as one. Review cannot replicate. The reviewer's evidence is the manuscript and whatever supplementary material the authors provide. Even where data and code are available, reviewers rarely have the time to rerun analyses, and almost never have the resources to repeat experiments. Review assesses whether a claim is plausibly supported by the evidence presented, not whether the claim will hold up. Review cannot compensate for a poor question. A reviewer can point out that a study is underpowered, that a comparison group is inappropriate, or that a conclusion overreaches. They cannot make an uninteresting study interesting, and a report that tries to redesign the whole project is usually of little use to anyone. Review cannot be neutral. Reviewers are human, drawn from the same communities as the authors, with the same rivalries, loyalties, theoretical commitments, and unconscious associations. Chapter 6 examines the evidence on bias in detail. The point here is that a system built on human judgement inherits human failings, and the question is not whether bias exists but how it can be reduced and made visible. Who relies on the verdict The weight carried by the phrase "peer-reviewed" extends far beyond the journal. Systematic reviewers and meta-analysts often restrict their searches to peer-reviewed literature, so a study that passes review becomes a data point in a pooled estimate that may inform a clinical guideline or a regulatory decision. Science journalists use publication in a reviewed journal as a threshold for coverage. Hiring, promotion, and funding committees count reviewed publications, and frequently weigh them by the selectivity of the venue. The law relies on it too. In Daubert v. Merrell Dow Pharmaceuticals, decided by the United States Supreme Court in 1993, the Court set out factors that federal judges may consider when deciding whether expert scientific testimony is admissible. Whether a theory or technique has been subjected to peer review and publication is one of them. The Court was careful to say that publication is not a requirement and that review is an imperfect filter, but the decision nonetheless made peer review a formal consideration in legal proceedings affecting product liability, criminal forensics, and environmental regulation. The COVID-19 pandemic showed both how much the certification function matters and how easily it can be bypassed or undermined. Preprints, which appear without review, became a main channel of scientific communication during 2020, and many were covered by the press before any expert had examined them. At the same time, peer review itself failed conspicuously in at least one high-profile case. In June 2020 The Lancet and the New England Journal of Medicine retracted papers based on a hospital database supplied by the company Surgisphere, after independent researchers raised questions that the company could not answer about where its data had come from. The papers had passed review at two of the most selective medical journals in the world. The data behind them could not be verified. For a reviewer, the lesson of this chain of reliance is not that every report carries the fate of public health. It is that the verdict a reviewer contributes to does not stay inside the journal. Downstream users rarely read the reviews; they see only the outcome. A reviewer who waves a paper through with a careless "minor revision" has, in effect, lent their expertise to every later use of that paper, without the users having any way to know how much scrutiny was really applied. That is one reason, discussed in Chapter 8, why the movement to publish reviews alongside papers has gathered strength: it lets downstream readers see what the certification consisted of. Why it is still worth doing well Given all this, it would be easy to conclude that peer review is theatre, and some critics have said as much. The conclusion does not follow. The evidence shows that review is noisy and fallible, not that it is useless. Manuscripts are routinely improved by it; errors are routinely caught; overstated claims are routinely moderated. Most working scientists can point to a review that saved them from publishing something wrong. And the alternatives that have been proposed, from publishing everything and letting readers sort it out to relying on metrics of attention, have failure modes of their own. The more realistic reform agenda, explored in Chapter 8, keeps expert review but changes when it happens, who can read it, and how reviewers are credited. What the evidence does suggest is that the quality of peer review depends heavily on the quality of individual reviews, and that individual reviews vary enormously. A careful, specific, well-reasoned report can steer a paper and an editor toward a sound outcome. A careless or hostile one can do real damage, delaying good work, letting weak work through, or discouraging a young author from submitting again. The difference between the two lies largely in the habits of the person writing, and those habits can be learned. That is where this book begins its practical work. Before turning to how to read and write a review, however, a reviewer needs to understand the system their report enters: who reads it, who decides, and what they need. That is the subject of the next chapter. A working definition It helps to end with a definition that reflects what the evidence and history suggest, rather than what the ceremony of "peer-reviewed" implies. Peer review is a structured request for expert advice, made by an editor to a small number of specialists, about whether a piece of work is sound, whether it suits a particular venue, and how it might be improved, delivered on the basis of the evidence the authors have chosen to present, under time pressure, and with the knowledge that the final decision belongs to someone else. Each clause of that definition carries an obligation. Structured means the reviewer should organise their advice so it can be used. Expert means the reviewer should confine confident judgements to what they actually know. Advice means the reviewer is not the decision-maker. The evidence the authors have chosen to present means the reviewer should notice what is missing as well as what is there. Under time pressure means the reviewer should prioritise, spending their limited attention where it matters most. The chapters that follow are, in a sense, an extended commentary on these obligations. Chapter 2. The Editorial System: Who Decides What A reviewer's report is one input into a decision made by someone else, and it travels through a system that most reviewers never see. Understanding that system is not administrative trivia. It determines what a report should contain, how its recommendation will be read, which remarks the authors will see, and what happens when a reviewer raises a concern about something other than the science. This chapter follows a manuscript from submission to decision, describing the people it passes through and what each needs. The journey of a manuscript A manuscript arrives through an online submission system, usually one of a small number of commercial platforms that most journals license. Before any scientist looks at it, editorial staff typically perform technical checks: that the required files are present, that the word count and formatting are within limits, that ethics approvals and data availability statements have been provided, that author details and conflict of interest declarations are complete. Many publishers now also run automated screening at this stage, including similarity checks for text overlap with published work and, increasingly, tools designed to detect manipulated images or patterns associated with commercially produced fraudulent papers. The manuscript then reaches an editor. At large journals this may be a professional editor, a full-time employee with a doctorate who no longer runs a laboratory; at most society and specialist journals it is an academic editor, a working scientist who handles manuscripts alongside research and teaching. The first decision this editor makes is whether to send the paper for review at all. Desk rejection, the return of a manuscript without external review, is common at selective journals, where a large majority of submissions may be declined at this stage. The reasons are usually scope, perceived significance, or obvious problems of quality or presentation. Desk rejection spares authors weeks of waiting and spares reviewers from assessing work the journal would never publish, which is why, uncomfortable as it is, it is generally considered good editorial practice when done promptly and with a brief explanation. If the paper goes forward, an editor identifies reviewers. At many journals an editor-in-chief assigns the manuscript to a handling editor, sometimes called an associate or academic editor, who has relevant expertise and manages the rest of the process. The handling editor searches for reviewers using their own networks, the manuscript's reference list, databases of past reviewers, and, at some journals, algorithmic reviewer-suggestion tools. Authors are often invited to suggest or exclude reviewers. Editors vary in how they treat these suggestions; some avoid author-suggested reviewers entirely, particularly since investigations in the mid-2010s revealed peer-review rings in which authors supplied contact details for fake reviewer accounts they controlled, leading to the retraction of scores of papers across several publishers. Finding reviewers is frequently the slowest step. Editors routinely send many invitations for each acceptance, and reviewer fatigue, the concentration of reviewing burden on a relatively small pool of willing and reliable people, is a widely recognised problem. This is one reason editors are often glad to invite early-career researchers, whose expertise in current techniques is frequently more precise than that of senior investigators and whose availability is sometimes greater. Once reviews are in, the handling editor reads them alongside the manuscript and reaches a decision, or makes a recommendation to the editor-in-chief who makes it. The decision letter goes to the authors with the reviewers' comments attached. If revisions are requested, the revised manuscript may return to the same reviewers, to a subset of them, or only to the editor, depending on the extent of the changes and the journal's policy. The roles and their obligations Each position in this system carries distinct responsibilities, and several of the most common problems in peer review arise when someone acts outside their role: a reviewer who treats their report as a final verdict, an editor who forwards reports without reading them, an author who treats review as an adversarial negotiation. The main roles and their core duties are set out in Table 1. Table 1. Principal roles in journal peer review and their core obligations. Role Primary responsibility Key ethical duties Typical failure Author Report the work fully and accurately Honest data, disclosure of conflicts, credit to contributors Overclaiming; withholding data Handling editor Manage review and reach or recommend a decision Choose qualified, unconflicted reviewers; weigh reports critically Treating reviews as votes Editor-in-chief Set policy and take final responsibility Consistency, appeals, handling misconduct Favouring prominent authors Reviewer Advise on soundness, significance, and improvement Confidentiality, disclosure, fairness, candour about expertise Harshness; scope creep Publisher and staff Operate systems and integrity checks Screening, record-keeping, corrections Opaque processes The table compresses a good deal, and two rows deserve fuller comment. The handling editor is the person a reviewer is really writing for. They are usually an expert in the broad field, but not necessarily in the specific technique or subfield of the manuscript; that is why they sought reviewers. They will read two or three reports, often of very different lengths, tones, and recommendations, and must reconcile them. What helps them most is a report that makes its reasoning visible, distinguishes the problems that would change the conclusions from those that would merely improve the presentation, and is candid about the reviewer's own confidence. What helps them least is a verdict without reasons, or a list of forty objections of undifferentiated weight. Good editors do not treat reviews as votes. If two reviewers recommend minor revision and a third identifies a fundamental flaw that the others missed, a careful editor will weigh the substance of the third report, not count the recommendations. Equally, if one reviewer recommends rejection on grounds that are clearly a matter of taste or theoretical allegiance, the editor may discount that recommendation. This is why the reasoning in a review matters more than the recommendation. Reviewers sometimes feel frustrated when an editor's decision does not follow their advice, but the system is designed that way: the reviewer advises on the evidence they have, and the editor decides with the benefit of all the evidence. The editor-in-chief holds ultimate responsibility for what the journal publishes and for its policies, and is usually the person who handles appeals and serious concerns about misconduct. Reviewers rarely deal with the editor-in-chief directly, but it is useful to know that the handling editor is not the last line of escalation. If a reviewer's serious concern is ignored, the editor-in-chief, and beyond them the publisher's research integrity team, are the appropriate next contacts. What the editor's invitation is asking A review invitation is a specific request, and reading it carefully is the first act of good reviewing. Most invitations include the title and abstract, a deadline, and sometimes a note from the editor about what they would like the reviewer to focus on. That note matters. An editor who writes "we would particularly value your assessment of the single-cell analysis" is telling the reviewer that someone else will cover the physiology. A reviewer who ignores the request and writes a general report may duplicate another reviewer's work while leaving the question the editor actually needed answered. Three questions should be answered honestly before accepting. First, is this within my competence? A reviewer does not need to be expert in every aspect of a manuscript, but they should be able to assess a substantial part of it with confidence, and they should be willing to say which parts they cannot assess. Accepting a review for which one is not qualified and then writing confidently about it is a quiet form of misrepresentation. Second, do I have a conflict of interest? Chapter 5 treats this in detail. At this stage, the question is whether any relationship with the authors, the work, or its subject would lead a reasonable observer to doubt the reviewer's impartiality. If so, the reviewer should either decline or disclose the relationship to the editor and let them decide. Third, can I meet the deadline? Late reviews are one of the most common complaints from both editors and authors, and they cost authors real time at stages of their careers when time matters greatly. It is better to decline promptly, perhaps suggesting a qualified colleague, than to accept and then delay. If circumstances change after accepting, a short message to the editor asking for an extension is courteous and usually granted. Early-career researchers sometimes receive review requests indirectly, when a senior colleague asks them to help with a review the senior colleague has been invited to write. This practice, sometimes called ghostwriting of reviews, is widespread. A survey led by Gary McDowell and colleagues, published in eLife in 2019, found that many early-career researchers had co-reviewed with a principal investigator, and that a large share of those had not been named to the editor. The authors argued that this deprives junior researchers of credit and misleads editors about who actually assessed the work. Most journals permit co-reviewing, but they expect the invited reviewer to obtain permission before sharing a confidential manuscript and to name the co-reviewer. An early-career researcher asked to help should ask whether the editor has been told; a principal investigator who asks for help should tell the editor and ensure the junior colleague is credited. What happens to a report Once submitted, a report is usually divided into two parts: comments for the authors and confidential comments for the editor. The authors will see the first part verbatim, sometimes lightly edited by the editor to remove anything inappropriate. They will not see the confidential comments. The division is useful but frequently misused. Confidential comments are appropriate for information the editor needs but the authors should not see: a note about the reviewer's own limits of expertise, a concern about possible duplicate publication or data manipulation that needs investigation before the authors are approached, a candid view of whether the work reaches the journal's bar for significance, or an explanation that the reviewer has a relationship with the authors that they judged not to be disqualifying. They are not appropriate for delivering criticisms that the reviewer is unwilling to make openly. A reviewer who writes a mild report to the authors and a damning note to the editor leaves the authors unable to respond to the real reasons for rejection. The general principle is that any substantive criticism of the science should appear in the comments to authors, where it can be answered. Most journals send each reviewer the other reviewers' reports and the decision letter after a decision is made. Reading these is one of the best ways for a new reviewer to calibrate. It shows what others noticed that one missed, how different reviewers weighted the same problems, and how the editor combined the advice. Where a reviewer's view differed sharply from the decision, it is worth considering whether the editor saw something in the other reports that justified the difference. If the manuscript is revised, the reviewer may be asked to assess the revision. The task then changes. The question is whether the authors have adequately responded to the points raised, not whether the reviewer can find new objections. Raising entirely new issues at the second round, when they could have been raised at the first, is a common source of author frustration and prolongs review without clear benefit. New problems introduced by the revision itself are fair game, and so are serious problems that were genuinely missed the first time, but the reviewer should acknowledge that they are new. Decisions and their meaning Journals use a small set of decision categories, though the labels vary. Accept is rare after first review. Minor revision means the paper is fundamentally sound and needs only limited changes, often without further external review. Major revision means substantial problems exist but appear fixable; the revised paper will usually be re-reviewed. Reject and resubmit, used by some journals, means the problems are serious enough that a substantially new manuscript is required, though the journal is willing to consider it. Reject means the journal will not consider the work further, whether because of fundamental flaws or insufficient fit or significance. A reviewer's recommendation should correspond to these meanings, and the report should make clear which category the reviewer has in mind and why. It is particularly helpful to state what would be required to change the recommendation. "I would support publication if the authors can show that the effect survives correction for multiple comparisons and holds in the second cohort" gives an editor far more to work with than "major revision". Reviewers should also remember that selection criteria vary between journals and that their recommendation is specific to the venue. A paper that is too incremental for a broad-scope journal may be an excellent contribution to a specialist one. A good report makes this explicit, separating the reviewer's view of the work's soundness, which should not change from journal to journal, from their view of its fit, which should. Variations on the standard workflow Not every manuscript follows the path just described, and a reviewer should recognise the variants, because each changes what the report is for. Registered Reports move the main round of review before the results exist. Introduced at the journal Cortex in 2013, under the editorial leadership of the psychologist Chris Chambers, and now offered by hundreds of journals, the format asks authors to submit their introduction, hypotheses, methods, and analysis plan before collecting data. Reviewers assess this Stage 1 protocol on the importance of the question and the rigour of the design. If the protocol is accepted in principle, the journal commits to publishing the final paper regardless of whether the results support the hypotheses, provided the authors follow the approved plan. At Stage 2, reviewers check that the plan was followed and that the conclusions follow from the results. For a reviewer, the shift is substantial. At Stage 1 there is no result to be impressed or disappointed by, so the assessment rests entirely on whether the design can answer the question. At Stage 2 the reviewer's freedom is deliberately restricted: they are not invited to reject the paper because the findings are unexciting or unexpected. The format was designed to counter publication bias, and it works only if reviewers respect its rules. Cascading or transfer review allows a manuscript rejected by one journal to be passed, with its reviews, to another journal from the same publisher, or in some arrangements to a different publisher altogether. Reviewers are often asked at the time of review whether they consent to their report being transferred. Consent costs little and can save authors months, but it means a report should be written so that it remains useful to an editor at a different venue; a report whose substance is simply "not important enough for this journal" transfers poorly. Special issues and conference proceedings often run on compressed timelines, with guest editors who may be less experienced than a journal's regular editorial board. In recent years several publishers have retracted large numbers of papers from special issues after discovering that guest editors had been impersonated or that review had been compromised. A reviewer invited to a special issue should apply exactly the same standards as for a regular submission and should be alert, as any reviewer should, to signs that a manuscript has been produced by a paper mill, a topic returned to in Chapter 7. Appeals and disputes Authors who believe a decision was wrong may appeal. The Committee on Publication Ethics, known as COPE, an organisation founded in 1997 whose guidance most major publishers follow, recommends that journals have a clear appeals process. Appeals are usually handled by the editor-in-chief or a senior editor, who may seek a further opinion. Reviewers are occasionally asked to respond to an author's rebuttal. When that happens, the right approach is to read the rebuttal on its merits, concede points where the authors are right, and explain clearly where the original concern still stands. A reviewer who treats an appeal as a personal challenge has forgotten that their role was always advisory. The editorial system, then, is designed to distribute responsibility. Authors are responsible for the honesty and completeness of what they submit. Reviewers are responsible for accurate, candid advice within their competence. Editors are responsible for choosing reviewers well, weighing their advice critically, and making decisions they can defend. When each role is performed well, the system can reach good decisions despite the noise in any single review. When a reviewer understands that their report is one part of this structure, they can write it to be used, which is the subject of the next two chapters. Chapter 3. Reading a Manuscript Critically Most weak reviews fail before a word is written. The reviewer reads the manuscript once, from beginning to end, marking objections as they arise, and then assembles those objections into a report. The result is predictable: a long list of comments of wildly varying importance, heavily weighted toward the introduction and early methods where the reviewer's attention was freshest, often missing the one problem in the analysis that actually determines whether the conclusions hold. Linear reading is how people read papers for pleasure or for background. It is not how to assess one. A more reliable approach treats reading as a sequence of distinct passes, each with its own purpose. The first establishes what the paper claims. The second tests whether the design and analysis can support those claims. The third examines the details: reporting, figures, statistics, data, and the fit between what was done and what is said. The passes need not be rigid, and experienced reviewers blend them, but separating them at first builds the habit of asking the important questions before the easy ones. The first pass: what is being claimed? Before assessing whether a paper is right, a reviewer must know precisely what it says. This sounds trivial and is not. Many manuscripts contain several claims of different strength, and the abstract, the discussion, and the title frequently state them differently. A title may announce that a gene "controls" a behaviour while the results show a correlation in one mouse strain; an abstract may report a "significant reduction" while the discussion concedes that the effect was small and confined to a subgroup. The first pass should therefore be quick and focused on the architecture of the argument. Read the title, abstract, and final paragraphs of the introduction, where the authors usually state their aims. Look at every figure and table, with their legends, without reading the results text. Then read the discussion's opening and closing paragraphs. At the end of this pass, a reviewer should be able to write down, in two or three sentences, the central claim of the paper, the key evidence offered for it, and the type of study that produced that evidence. It is worth actually writing this summary. It becomes the opening of the review, where it serves an important function discussed in the next chapter: it shows the authors and the editor that the reviewer has understood the work, and it exposes any misunderstanding before it contaminates the rest of the assessment. It also serves the reviewer. If the central claim cannot be stated clearly after a careful first pass, that is itself a finding: the manuscript is unclear about what it is arguing, and saying so is one of the most useful things a report can do. The first pass should also identify the claim's type. Is it causal, or descriptive, or predictive? Is it a claim about a mechanism, a population, a method's performance, or the existence of a phenomenon? The type determines what evidence is required. A causal claim from observational data needs a credible strategy for addressing confounding; a claim about a new method's superiority needs a fair comparison against current alternatives on relevant benchmarks; a claim about a population needs a sample that represents it. Much of the substance of a good review consists of matching the claim to the evidence its type requires. The second pass: can the design support the claim? With the claim in hand, the reviewer reads the methods and results closely, asking a single overarching question: if everything reported here was done exactly as described, would it justify the conclusion? This question sets aside, for now, whether things were done as described. It concerns the logic of the study. Several recurring problems are worth checking for explicitly, because they are common and because they are easy to miss when reading in the authors' frame. Alternative explanations. For every main result, ask what else could have produced it. In experimental work this usually means asking whether the controls rule out the obvious alternatives: whether a knockdown effect might be off-target, whether a drug effect might be due to the vehicle, whether a behavioural difference might reflect a motor deficit rather than a cognitive one. In observational work it means asking about confounding, selection, and reverse causation. A reviewer who can name a specific, plausible alternative explanation that the design does not exclude has found something important. A reviewer who says only that "other factors may be involved" has not. Comparison and baseline. Every claim of an effect is a claim about a difference, and the choice of comparison determines what the difference means. Is the control group appropriate? Is a new method compared with the current best alternative or with a straw man? Is a change measured against a baseline that could itself have shifted? Sample and power. Is the sample large enough to detect effects of the size claimed, or to make a null result informative? Small studies that report large effects deserve particular scrutiny, because they are disproportionately likely to be false positives or inflated estimates, a point made forcefully by John Ioannidis and by Katherine Button and colleagues, whose 2013 analysis in Nature Reviews Neuroscience found the median statistical power of neuroscience studies to be low. Where sample sizes are not justified by a power calculation or equivalent reasoning, a reviewer may reasonably ask for one. Independence of observations. A frequent and consequential error is to treat non-independent measurements as independent: multiple cells from the same animal, repeated measures from the same participant, several samples from the same culture. Doing so inflates apparent sample sizes and produces spuriously small p-values. It is worth checking what the unit of analysis is and whether it matches the unit of replication. Analytical flexibility. Were the analyses planned in advance, or chosen after seeing the data? Were outcomes, covariates, or exclusion criteria selected in ways that might have favoured the reported result? In clinical trials, comparison with the registered protocol is often revealing; in other fields, preregistration is increasingly common and should be checked when it exists. Where there is no preregistration, the reviewer can still ask whether the results are robust to reasonable alternative analytical choices. Generalisation. Do the conclusions extend beyond the conditions actually tested? Findings in one cell line, one species, one population, or one dataset are frequently described as if they held generally. A reviewer should ask the authors to match the scope of their claims to the scope of their evidence. By the end of the second pass, a reviewer should know whether there are any problems that, if not addressed, would mean the central conclusion is unsupported. These are the major concerns. They may be few, and there may be none. Identifying them correctly is the most valuable thing a reviewer does. The third pass: details, reporting, and data The third pass is where most reviewers spend most of their time, which is why it belongs last. Here the reviewer checks whether things were done and reported as they should be: whether statistical tests are appropriate and correctly described, whether figures display what the text says, whether the numbers in different places agree, whether the methods are described in enough detail to be reproduced, and whether data and code are available as the journal requires. Reporting guidelines are a practical aid here. Over the past three decades, groups of methodologists and editors have produced checklists specifying what should be reported for common study designs, and many journals require authors to complete the relevant checklist on submission. A reviewer does not need to audit every item, but knowing the relevant guideline helps identify omissions quickly. The most widely used are summarised in Table 2; the EQUATOR Network, an international initiative founded in 2008 to improve the reliability of health research reporting, maintains a searchable library of several hundred more. Table 2. Widely used reporting guidelines and the study types they cover. Guideline Study type Issued by or associated with Useful for checking CONSORT Randomised controlled trials CONSORT Group Randomisation, allocation, flow of participants PRISMA Systematic reviews and meta-analyses PRISMA Group Search strategy, selection, risk of bias STROBE Observational epidemiology STROBE Initiative Confounding, selection, missing data ARRIVE Animal research NC3Rs (UK) Randomisation, blinding, sample size STARD Diagnostic accuracy studies STARD Group Reference standard, spectrum of patients The guidelines matter to review for a reason beyond completeness. Their items encode the specific ways each type of study tends to go wrong. ARRIVE asks whether animals were randomly allocated and whether outcome assessors were blinded because unblinded, unrandomised animal studies have repeatedly been shown to report larger effects. CONSORT asks for a participant flow diagram because attrition that differs between arms can create an apparent treatment effect. Reading a manuscript against the relevant guideline is, in effect, reading it with the accumulated experience of the methodologists who studied that design's failures. Figures deserve particular attention in the third pass. Check that axes are labelled and scaled honestly, that error bars are defined, that the sample size for each panel is stated, and that representative images are accompanied by quantification. Look for panels that seem too clean or that resemble each other in ways they should not; Chapter 7 discusses what to do if something appears duplicated or altered. Check that individual data points are shown where sample sizes are small, since bar charts of means can conceal very different distributions. Numbers should be checked for internal consistency. Do the sample sizes in the methods match those in the figure legends and tables? Do percentages add up? Are reported test statistics, degrees of freedom, and p-values consistent with one another? Tools exist to help with the last question. The program statcheck, developed by Michèle Nuijten and colleagues, recomputes p-values from reported test statistics in APA-formatted text, and the GRIM test, proposed by Nick Brown and James Heathers in 2016, checks whether reported means are arithmetically possible given the sample size and the granularity of the underlying data. Inconsistencies found this way are usually typographical, but they can indicate deeper problems, and in either case they should be corrected. Finally, the reviewer should check what the data and code availability statement says, and whether it is true. A statement that data are "available on reasonable request" is weaker than a deposit in a public repository, and studies have found that such requests are frequently unanswered. Many funders and journals now require deposit where ethically possible. If the journal's policy requires data to be available and it is not, or if a link leads nowhere, the reviewer should say so. If the data are available and the reviewer has the time and competence, looking at them, even briefly, can reveal problems invisible in the manuscript. The three passes in practice A hypothetical example shows how the passes change what a reviewer notices. Suppose a manuscript reports that a dietary supplement improves memory in older adults. Its title says the supplement "enhances cognitive function in ageing". The abstract reports a randomised trial of 120 participants over twelve weeks and a significant improvement on a word-recall task. A linear reader might begin by objecting to the introduction's selective citations, move on to request more detail about the supplement's formulation, question the choice of font in a figure, and eventually, somewhere in the results, notice that the primary outcome listed in the trial registry was a composite cognitive score, not word recall. That last observation, buried as comment seventeen of twenty-two, is the one that matters. The first pass, done properly, would have surfaced it early. Writing down the central claim forces the question of which outcome supports it, and the type of claim, a causal effect from a randomised trial, directs the reviewer immediately to CONSORT's concerns: was the primary outcome prespecified, and was it the one reported? The second pass would then ask whether word recall was among several secondary outcomes, whether the composite score showed any effect, whether multiple comparisons were accounted for, and whether a twelve-week improvement on one task justifies a title claiming enhanced "cognitive function" in general. It would also ask about blinding, since a supplement with a distinctive taste might allow participants to guess their allocation, and about attrition, since older participants who experience side effects may drop out unevenly between arms. Only in the third pass would the reviewer check the formulation details, the consistency of the numbers across tables, and the figure design. Those checks are still worth doing. But the report built from these passes leads with the outcome-switching question and the overreach in the title, which are the problems that determine whether the conclusion stands, and relegates the rest to minor comments. The editor reading it knows at once what the decision turns on. The authors know what they must address, and they may well have a good answer: perhaps the registry was amended before unblinding, for documented reasons. A review that asks the right question clearly allows that answer to come out. Reviewing within one's competence Few reviewers are expert in every method a modern manuscript uses. A single paper in cell biology may combine imaging, genomics, animal behaviour, and sophisticated statistics; a paper in ecology may combine field sampling, remote sensing, and Bayesian modelling. The reviewer's obligation is not to be omniscient but to be clear about where their competence ends. The practical rule is to review confidently what one knows, to review cautiously what one partly knows, and to state plainly what one cannot assess. "I am not able to evaluate the phylogenetic methods and would recommend that the editor seek a specialist opinion on them" is a useful sentence. It tells the editor precisely where a gap exists in the assessment, and editors frequently act on it. By contrast, a reviewer who comments confidently on methods they do not understand may mislead the editor, burden the authors with inappropriate requests, or miss real problems while inventing false ones. Statistics deserve a special word. Many reviewers who are expert in their experimental domain are less secure in statistical analysis, and statistical problems are among the most common serious flaws in published research. A reviewer who is uncertain whether an analysis is appropriate should say so and suggest statistical review. Several journals employ statistical reviewers or editors for exactly this reason, and an editor who is told that a statistical question needs specialist attention can obtain it. Time, attention, and proportion A thorough review of a substantial manuscript takes time, typically several hours spread over more than one sitting. There is no merit in taking longer than necessary, but there is real cost in taking less than the work requires. The distribution of time matters as much as its total. A useful rough allocation is to spend a modest share on the first pass, the largest share on the second, and the remainder on the third, adjusting for the paper's complexity. Reviewers who spend most of their time correcting typographical errors and requesting additional citations have allocated their attention poorly, however long they spent. It is also worth leaving time between reading and writing. First reactions to a manuscript are often stronger than considered ones, in both directions. A day's distance frequently turns an objection that seemed decisive into one that is real but minor, or reveals that an apparently convincing result depends on an assumption not stated. The reviewer who writes immediately after a single read is likely to produce a report that reflects their mood as much as the manuscript. Reading with charity The final principle of critical reading is one that sounds soft but is methodological: read the manuscript as the strongest version of what the authors are trying to say. When a passage is ambiguous, consider the most reasonable interpretation before assuming the worst one. When a control seems missing, check whether it appears in the supplementary material or whether a different control addresses the same concern. When a result seems implausible, ask what would have to be true for it to be correct. Charitable reading is not credulity. Its purpose is to ensure that criticisms, when they come, are aimed at what the authors actually did rather than at a misreading. A report that attacks a straw man wastes the authors' time, undermines the reviewer's credibility with the editor, and often leaves the real weaknesses untouched. A report that engages with the strongest version of the argument, and still finds problems, is one that the authors must take seriously. With the reading done, the reviewer knows what the paper claims, whether the design can support it, and where the details fall short. The next task is to turn that knowledge into a report that an editor can use and authors can act on. Hashtags: #ThePeerReviewEcosystem #PeerReview #ScholarlyPeerReview #EditorialEthics #ResearchIntegrity #ConstructiveCritique #EditorialRoles #ReviewerResponsibilities #HandlingEditors #EditorsInChief #ReviewerEthics #ConflictOfInterest #Confidentiality #ReviewerBias #ImplicitBias #PublicationEthics #COPE #CriticalManuscriptReview #ResearchQuality #ScientificPublishing #OpenPeerReview #RegisteredReports #EditorialDecisionMaking #ResearchTransparency #FutureOfPeerReview

  • Scientific Integrity and Whistleblowing (Investigating Image Manipulation and Fraud)

    Download the Book (PDF): Introduction In the spring of 2022 a neurologist at Vanderbilt University named Matthew Schrag, working on a separate question entirely, began looking closely at the Western blots in a group of papers on amyloid beta, the protein at the centre of the dominant theory of Alzheimer's disease. Among them was a paper published in Nature in 2006 that described a specific assembly of the protein, called Aβ56, and reported that it impaired memory in rats. The paper had been cited thousands of times. Schrag's concerns, set out in a dossier and later investigated by the journalist Charles Piller for Science, centred on bands in the blots that appeared to have been duplicated, spliced and erased. In June 2024 Nature* retracted the paper. The retraction notice referred to signs of excessive manipulation, including splicing, duplication and the use of an eraser tool, and recorded that the first author, Sylvain Lesné, did not agree with the retraction while his coauthors did. Three features of that episode are worth holding onto, because they recur throughout this book. The first is that the evidence was on the page for sixteen years. Nothing new had to be discovered in a laboratory; someone simply had to look at the published figures with trained eyes and the right software. The second is that the person who looked was not a colleague of the authors, not a reviewer, not an editor, and not an institutional official. He was an outsider who took on the work in his own time. The third is that the distance between the first credible alarm and the formal correction of the record was measured in years, and it was closed only after sustained attention from a major news outlet. That pattern, visible evidence, outsider detection, slow correction, is the subject of this book. Over the past fifteen years the detection of fabricated and falsified research has been transformed. What was once almost entirely a matter of chance, a junior researcher noticing something odd at the bench and deciding to risk a career by reporting it, has become in large part a forensic discipline practised on the published record itself. Image analysts can find a duplicated panel in a figure from a paper published two decades ago. Statisticians can show that a set of reported means could not have arisen from integer data, or that baseline variables in a randomised trial are too similar, or too different, to be the product of chance. Linguists and bibliometricians can flag the fingerprints of a paper mill in the phrasing of an abstract. None of this requires access to anyone's laboratory notebooks. It requires only the paper, the method, and the will to apply it. The controlling argument This booklet argues one thing. Detection of research fraud has outrun correction. The forensic methods for finding fabricated images, impossible numbers and manufactured papers are now good enough, cheap enough and scalable enough that the bottleneck in scientific integrity is no longer finding problems but acting on them. Journals, universities and funders were built for a world in which misconduct was rare, detected from the inside, and handled confidentially one case at a time. They are now facing a world in which problems are detected from the outside, in public, at industrial scale, and in which some of the misconduct is itself industrial. The institutions have not caught up, and the consequence is that the scientific record is being corrected largely by volunteers, at personal cost, on a timescale that bears no relation to the speed at which bad work enters it. It follows that the most important reforms are not better detection tools, though those help. They are changes to what happens after a problem is found: who is obliged to respond, how quickly, on what standard of evidence, with what protection for those who raise concerns, and with what separation between correcting a paper and judging a person. The final chapters make that case in detail. What this book covers The first chapter sets out what research misconduct is, in the legal and regulatory sense, and what the best evidence says about how common it is. It separates fabrication and falsification from the much larger territory of questionable research practice and honest error, because the forensic methods discussed later behave very differently across those categories. The next two chapters deal with images. Chapter 2 describes the forms that image manipulation takes in published figures, from the simple reuse of a panel to the splicing, cloning and erasing of bands in a blot, and explains why a visible duplication is evidence of a problem but not, on its own, evidence of intent. Chapter 3 turns to the tools: the software that now screens submissions at several major publishers, the older techniques of contrast adjustment and error-level analysis, and the looming difficulty posed by images generated or altered with machine learning, which leave few of the traces that current methods look for. Chapter 4 is about numbers. It explains in plain terms how a set of forensic statistical tests work, from the arithmetic consistency checks known as GRIM and SPRITE, through the analysis of baseline balance in randomised trials pioneered by the anaesthetist John Carlisle, to the examination of digit patterns and spreadsheet metadata. Each test is described with what it needs, what it can show and where it breaks down. Chapter 5 moves from individuals to organisations. Paper mills, businesses that manufacture and sell authorship of fraudulent manuscripts, have become the largest single source of fabricated research by volume. The chapter describes how they operate, how they were discovered, how they are detected, and why the scale of the problem, documented in 2025 in a large study in the Proceedings of the National Academy of Sciences, has changed what integrity work has to mean. Chapters 6 and 7 look at people and institutions. Chapter 6 is about whistleblowers, both the insiders who report colleagues and the outsiders, the so-called sleuths, who scrutinise the literature in public. It examines what they have achieved and what it has cost them. Chapter 7 follows an allegation through the formal machinery of institutional inquiry and investigation and journal correction, and identifies where that machinery stalls. Chapter 8 sets out a programme of reform grounded in the evidence of the previous chapters. A note on names and cases This is a book about real cases, and some of them involve named people. The standard applied throughout is simple. Where a finding of misconduct has been made by a competent body, whether a government office, a university investigation or a court, it is reported as that body's finding and attributed to it. Where a paper has been retracted, the retraction is reported along with, where relevant, the reasons the journal gave. Where allegations have been made but not adjudicated, or where the person concerned disputes them, that is said plainly. Several of the people discussed have denied wrongdoing, and some have litigation under way. A book about the discipline of evidence has no business running ahead of the evidence about individuals, and this one does not try. The same discipline applies in the other direction. A duplicated image or an impossible mean is a fact about a document. It is not, by itself, a fact about the character of any person who signed it. Much of what follows is an argument that institutions should learn to act firmly on the first kind of fact without waiting for the second, which is slower, harder and properly surrounded by due process. Keeping those two things apart is not a courtesy to the accused. It is the only way to correct the record at the speed the record now requires. Chapter 1. What Counts as Misconduct, and How Much of It There Is Every investigation of a published paper begins with a question that sounds simpler than it is: what, exactly, is wrong? A paper can be wrong because its authors made a mistake in a calculation. It can be wrong because they analysed their data in a dozen ways and reported the one that worked. It can be wrong because an image was assembled carelessly from the wrong folder. And it can be wrong because someone invented the numbers. Only the last of these, and a few close relatives, is research misconduct in the formal sense. The distinctions matter enormously, both for what the forensic methods described later in this book can show and for what institutions are entitled to do once they have shown it. The narrow definition and why it is narrow In the United States the working definition comes from a federal policy issued by the Office of Science and Technology Policy in 2000 and adopted across research agencies. Research misconduct is fabrication, falsification or plagiarism in proposing, performing or reviewing research, or in reporting research results. Fabrication is making up data or results and recording or reporting them. Falsification is manipulating research materials, equipment or processes, or changing or omitting data or results, such that the research is not accurately represented in the record. Plagiarism is the appropriation of another person's ideas, processes, results or words without giving appropriate credit. The policy states explicitly that research misconduct does not include honest error or differences of opinion. A finding of misconduct under this framework requires three things: a significant departure from the accepted practices of the relevant research community; that the misconduct was committed intentionally, knowingly or recklessly; and that the allegation is proven by a preponderance of the evidence. For research funded by the US Public Health Service, which includes the National Institutes of Health, these rules are codified in the federal regulations at 42 CFR Part 93 and overseen by the Office of Research Integrity, usually called ORI. Those regulations were substantially revised in a final rule published in September 2024, which took effect at the start of 2025 and gave institutions until the start of 2026 to bring their own policies into compliance. The revision clarified definitions, permitted institutions to handle multiple respondents and allegations more efficiently, and adjusted procedures around confidentiality and the scope of investigations, but it left the core triad of fabrication, falsification and plagiarism untouched. The narrowness is deliberate. A finding of misconduct can end a career, bar a scientist from federal funding and in rare cases lead to criminal prosecution. Regulators have therefore confined it to conduct that is plainly dishonest rather than merely poor. Other jurisdictions draw the line in slightly different places. Several European codes, including the European Code of Conduct for Research Integrity published by ALLEA, the federation of European academies, use the same fabrication, falsification and plagiarism core but surround it with a larger list of unacceptable practices, such as manipulating authorship, withholding data without justification, or misrepresenting the achievements of others. Some national systems, including those in Denmark and Sweden, have statutory definitions with their own nuances. But the idea of a hard core surrounded by a softer, larger periphery is almost universal. For the forensic investigator this has a practical consequence. The methods described in this book detect anomalies in documents. A duplicated band in a blot, a mean that cannot be produced by the stated sample, a set of baseline characteristics that are implausibly balanced: each is an observation about a paper. None of them, taken alone, establishes intent, which is a fact about a person's state of mind. The step from anomaly to misconduct is taken, when it is taken at all, by an institution that can interview people, collect raw data and notebooks, and weigh explanations. The investigator's job is to establish that the published record cannot be relied on. That is a different and much lower bar, and one of the central claims of this book is that institutions should be willing to act on it without waiting for the higher one. The larger territory: questionable practice and honest error Around the narrow core lies a large territory of conduct that distorts the record without falling within the legal definition. It is usually called questionable research practice. The canonical list includes failing to report all of a study's outcome measures, deciding whether to collect more data after looking at the results, selectively reporting studies or conditions that worked, excluding data points after seeing their effect on the result, rounding p-values down across a significance threshold, and presenting an unexpected finding as though it had been predicted all along. The distinction between these practices and falsification is sometimes a matter of degree rather than kind. Dropping an outlier with a documented rationale is standard practice. Dropping every observation that weakens an effect until the effect crosses a threshold, and then failing to mention that anything was dropped, begins to look like changing or omitting data such that the research is not accurately represented, which is the regulatory definition of falsification. Where the line falls in a given case depends on facts that the published paper usually does not reveal. Honest error is different again. Laboratories are busy places, figure assembly is tedious, and large projects generate thousands of image files with similar names. Many duplicated images in the literature are the result of someone dragging the wrong file into a figure panel. Many statistical inconsistencies are typographical. A mature integrity system needs to be able to correct these errors quickly and without stigma, precisely so that it can reserve its heavier machinery for the cases that warrant it. As later chapters show, the absence of a fast, stigma-free route to correction is one of the reasons errors and fraud are so often tangled together in practice. Authors resist correcting an honest mistake because a correction looks like an admission, and institutions hesitate to demand one because it might. The grey zones around the core Between the hard core of fabrication and falsification and the soft periphery of questionable practice lie several categories that matter to forensic investigators because they generate many of the anomalies found in the literature. The first is duplicate publication, in which substantially the same data are published more than once without disclosure, sometimes with different framing or in a different language. It inflates the apparent weight of evidence, which matters when trials are pooled in meta-analyses, since the same patients are counted twice. Closely related is the reuse of figures from one's own earlier papers to represent new experiments. When an image published in 2015 as showing one cell line reappears in 2019 as showing another, the second use is not self-plagiarism in any harmless sense. It is a misrepresentation of what the 2019 experiment produced, and so falls under falsification if done knowingly. The second is authorship manipulation: adding people who contributed nothing, omitting people who did, or, as Chapter 5 describes, selling places on the author list. Authorship matters for integrity not only because it misallocates credit but because it determines who is responsible when something goes wrong. A paper whose listed authors cannot explain how a figure was made has a responsibility gap that makes investigation harder. The third is manipulation of the review and citation system itself: fabricated reviewer identities, reviews written by the authors under false names, and coordinated citation of one another's work to inflate metrics. These practices do not falsify data directly, but they defeat the mechanisms meant to catch falsified data, and several publishers have retracted large batches of papers on the ground that peer review was compromised. What unites these grey zones is that each leaves a documentary trace, in duplicated text, repeated images, reviewer email addresses or citation patterns, that can be found without access to anyone's laboratory. That is why forensic reading of the literature has become so productive: much of what corrupts the record is visible in the record. How common is it? Estimating the prevalence of something people have strong reasons to conceal is inherently difficult. Three kinds of evidence exist: surveys in which researchers report their own behaviour and what they have witnessed in others, audits of samples of published papers, and the record of retractions. Each has predictable biases, and each tells part of the story. The most cited survey evidence is a meta-analysis by Daniele Fanelli, published in PLoS ONE in 2009, which pooled the results of eighteen surveys. On average about 2 percent of scientists admitted to having fabricated, falsified or modified data or results at least once, and up to about a third admitted to other questionable practices. When asked about colleagues rather than themselves, respondents reported much higher figures: around 14 percent said they had observed colleagues falsifying, and up to 72 percent reported questionable practices by others. The gap between self-report and report about others is itself informative. It may reflect underreporting of one's own behaviour, the fact that one bad actor can be observed by many colleagues, or both. A more recent and larger estimate came from the Dutch National Survey on Research Integrity, whose results were published in PLOS ONE in 2022 by Gowri Gopalakrishna and colleagues. The survey used a technique designed to encourage honest answers about sensitive behaviour and drew on nearly seven thousand respondents across all disciplines in the Netherlands. It found that 4.3 percent of respondents reported having fabricated data and 4.2 percent reported falsification in the previous three years, with 8.3 percent reporting at least one of the two. Just over half reported engaging frequently in at least one questionable research practice. Rates of fabrication and falsification were highest among PhD candidates and junior researchers and in the life and medical sciences. Earlier work in psychology by Leslie John, George Loewenstein and Drazen Prelec, published in Psychological Science in 2012, surveyed around two thousand academic psychologists in the United States and found that questionable practices such as failing to report all dependent measures and deciding to collect more data after checking significance were admitted by large proportions of respondents. That study helped trigger what became known as the replication crisis in psychology, and it is important to the argument here because it showed that the boundary between ordinary bad habit and falsification is heavily populated. The main surveys are summarised in Table 1, which is intended to convey the order of magnitude rather than a precise rate, since the samples, time windows and question wording differ. Table 1. Selected survey estimates of fabrication, falsification and questionable practice. Study Population Fabrication or falsification Questionable practices Fanelli (2009), PLoS ONE Meta-analysis of 18 surveys About 2% admitted at least once; about 14% observed in colleagues Up to about 34% admitted; up to 72% observed in colleagues John, Loewenstein and Prelec (2012), Psychological Science About 2,000 US psychologists Low admission rates High admission rates for several practices Gopalakrishna et al. (2022), PLOS ONE About 6,800 researchers in the Netherlands 8.3% reported at least one in previous three years 51.3% reported at least one frequently Audits of published papers give a different view. The most important for this book is the study by Elisabeth Bik, Arturo Casadevall and Ferric Fang, published in mBio in 2016, which examined images in more than twenty thousand papers from forty biomedical journals and found that about 3.8 percent contained a problematic duplication. Chapter 2 examines that study in detail. In clinical research, the anaesthetist John Carlisle's analysis of trials submitted to the journal Anaesthesia, discussed in Chapter 4, found false data in a striking proportion of trials for which individual patient data were available. Audits of this kind do not measure intent, but they measure something arguably more important: the fraction of the record that cannot be trusted as it stands. What retractions tell us, and what they do not The retraction record is the most visible measure and the most misleading. A retraction happens only when a problem has been detected, reported, investigated and acted on, and each of those steps filters out cases. Retraction rates therefore measure the efficiency of the correction system at least as much as the prevalence of the underlying problem. The best-known analysis of why papers are retracted is by Ferric Fang, R. Grant Steen and Arturo Casadevall, published in the Proceedings of the National Academy of Sciences in 2012. They examined more than two thousand retracted biomedical and life-science articles indexed in PubMed and, by consulting secondary sources such as ORI findings and news reports as well as retraction notices, found that about two-thirds were attributable to misconduct, including fraud or suspected fraud, duplicate publication and plagiarism. Only about a fifth were attributable to error. The study's broader lesson was that retraction notices themselves were often uninformative or misleading about the reason, so that reading notices alone understated the role of misconduct. The absolute numbers have grown sharply. The Retraction Watch database, founded by the journalists Ivan Oransky and Adam Marcus and acquired in 2023 by Crossref, the nonprofit that registers digital object identifiers for scholarly works, which then made it openly available, now records tens of thousands of retractions. In 2023 more than ten thousand papers were retracted in a single year, a record driven largely by mass retractions at the publisher Hindawi, then owned by Wiley, of papers linked to paper mills and compromised special issues. Chapter 5 examines that episode. Even so, retractions remain a very small fraction of the literature, on the order of a few papers in every ten thousand, far below any plausible estimate of the fraction that contains fabricated or falsified material. The gap between those two numbers is the space in which this book operates. Why it matters It is tempting to treat fraud as a problem for the careers of the people involved, and to assume that the scientific process will absorb the damage because false results fail to replicate. The second assumption is weaker than it sounds. Replication is expensive and rarely attempted directly. A fabricated result in a crowded field can shape grant priorities, animal experiments and drug development for years before anyone establishes that it cannot be reproduced, and the failure to reproduce it is often attributed to technical differences rather than to the original being false. The costs are concrete. An analysis by Andrew Stern, Arturo Casadevall, R. Grant Steen and Ferric Fang, published in eLife in 2014, examined papers retracted as a result of research misconduct identified by ORI between 1992 and 2012 and estimated that they accounted for roughly 58 million dollars in direct NIH funding, a small fraction of the agency's budget but a substantial sum in absolute terms. That figure does not include the downstream work built on retracted findings, the patients enrolled in trials justified by them, or the careers of junior colleagues whose names appeared on compromised papers. In clinical medicine the stakes are higher still. When fabricated trials enter systematic reviews, they shift pooled estimates of treatment effects and therefore clinical guidelines. The case of the Japanese anaesthetist Yoshitaka Fujii, whose trials were examined statistically by John Carlisle in 2012 and subsequently investigated by the Japanese Society of Anesthesiologists, led to the retraction of more than a hundred and seventy papers, many of them trials of drugs for postoperative nausea. Those trials had been included in meta-analyses. A fabricated trial is not an isolated lie; it is a contaminant that spreads through every synthesis that includes it. That, in the end, is why the rest of this book is concerned less with punishing individuals than with the integrity of the record. The forensic methods that follow are ways of reading a document and asking whether it could have been produced by the process it describes. When the answer is no, the most urgent question is not who is to blame but who is relying on it, and how quickly they can be told. Chapter 2. The Anatomy of a Manipulated Image Most of the evidence in experimental biology is ultimately visual. A Western blot shows whether a protein is present and in roughly what quantity. A micrograph shows the shape of cells, the location of a fluorescent marker, the architecture of a tissue. A gel shows fragments of DNA sorted by size. A flow cytometry plot shows how a population of cells divides across two measured properties. In each case the published figure is offered as a faithful record of what the instrument recorded, and readers are invited to see the result for themselves. That invitation is what makes images such powerful evidence, and it is also what makes their manipulation such a direct form of deception. A falsified number in a table can be hidden among other numbers. A falsified image is placed at the centre of the argument. What an honest image looks like To understand manipulation it helps first to understand what an unmanipulated scientific image contains. Every digital image produced by a camera or scanner carries, in addition to the signal of interest, a great deal of incidental information: electronic noise from the sensor, variations in background intensity, dust on the lens or the scanner glass, scratches on a membrane, uneven staining, the particular granularity of film if the image was scanned from an X-ray film exposure. In a Western blot, the background around each band has a texture that is effectively unique. In a micrograph of cells, the exact arrangement of every cell in the field is unique, as is the pattern of debris between them. This incidental information is the forensic analyst's friend. Two independent experiments will never produce identical noise, identical dust and identical arrangements of cells. If two panels that are supposed to show different samples share the same background blemishes in the same relative positions, they came from the same original image, whatever their labels say. The signal itself can be similar between experiments; the noise cannot. Honest image processing is permitted and often necessary. The guidelines that shaped modern practice were published in the Journal of Cell Biology in 2004 by Mike Rossner, then the journal's managing editor, and Kenneth Yamada, in an article titled "What's in a picture? The temptation of image manipulation." Their core rules have since been adopted, with local variations, by most life-science journals. No specific feature within an image may be enhanced, obscured, moved, removed or introduced. Adjustments of brightness, contrast or colour balance are acceptable if they are applied to the whole image and do not obscure, eliminate or misrepresent any information present in the original, including the background. Grouping images from different parts of the same gel, or from different gels, fields or exposures, must be made explicit by the arrangement of the figure, for example with dividing lines, and in the legend. Nonlinear adjustments, such as changes to gamma settings, must be disclosed. These rules are simple, and they draw a clear line. Cropping a blot to show the relevant bands is acceptable. Cutting out a lane and closing the gap so that the result looks like a continuous gel, without marking the join, is not, because it conceals that the lanes were not run side by side. Raising the contrast of a whole image so a faint band becomes visible is acceptable if the background is raised too. Painting over a band that should not be there is not. The Journal of Cell Biology did more than publish guidelines; it began screening every accepted manuscript's figures before publication. Rossner later reported that over several years roughly a quarter of accepted manuscripts contained at least one figure that had to be remade because it had been manipulated in a way that violated the guidelines, most of them for reasons that did not affect interpretation, and that about one percent were found to contain manipulations that did affect the conclusions, and so had their acceptance revoked. Those figures established two things early: that guideline violations were common, and that most of them were not fraud. The taxonomy of problems The most influential attempt to classify image problems in the published literature, rather than in submissions, is the study by Elisabeth Bik, Arturo Casadevall and Ferric Fang in mBio in 2016. Bik, then a microbiologist working in industry, examined by eye the images in 20,621 papers published between 1995 and 2014 in forty biomedical journals. She flagged papers that appeared to contain inappropriately duplicated images, and each flagged paper was reviewed by her two coauthors, with a paper counted only if all three agreed. The final tally was 782 papers, about 3.8 percent. The rate varied greatly between journals, from well under one percent at the Journal of Cell Biology, which had screened images for years, to more than twelve percent at the International Journal of Oncology. The proportion of problematic papers rose markedly after about 2003, around the time digital image handling became routine. The study sorted problems into three categories, which have since become the working vocabulary of the field. As Table 2 sets out, the categories differ in what is duplicated, how, and how readily an innocent explanation is available. Table 2. Categories of inappropriate image duplication, after Bik, Casadevall and Fang (2016). Category What is seen Typical innocent explanation Weight as evidence of intent Simple duplication The same image used twice to represent different conditions Wrong file selected during figure assembly Low on its own Duplication with repositioning Overlapping or identical images shifted, rotated, flipped or rescaled Rarely plausible; repositioning requires deliberate action Moderate to high Duplication with alteration Regions cloned, stamped, spliced or erased within or between images Very rarely plausible High The first category is the most common and the most forgiving. A paper that shows the same loading-control blot under two different experiments, or the same micrograph labelled as two different treatments, may well be the product of a mistake. Laboratories store thousands of images, often with sequential or nearly identical file names, and a tired postdoctoral researcher assembling a figure late at night can easily pick the wrong one. Such errors still invalidate the figure, because one of the two panels does not show what it claims, but they do not in themselves suggest dishonesty. They call for a correction backed by the original data. The second category involves an extra step. If a micrograph appears in two figures but one version has been rotated by ninety degrees, or mirrored, or cropped to a different region that overlaps the first, someone performed an action beyond selecting a file. A rotated duplicate is harder to explain as a slip of the mouse, because it makes the duplicate harder to recognise, which is the effect one would expect if concealment were the aim. That is not proof: people rotate images for layout reasons, and overlapping fields of view can arise legitimately if the same slide was imaged twice and both images were then mislabelled. But the innocent explanations are narrower. The third category is the most serious. Here the image has been altered internally: a region has been copied from one part of an image to another, a band has been pasted into a lane, a patch of background has been cloned over something that was there, or pieces from different images have been combined without disclosure. Bik and her coauthors judged that at least half of the problematic papers in their sample showed features in the second or third categories, suggestive of deliberate manipulation. Alteration of this kind is very difficult to produce by accident. A cloned region in a histology image, where the same cluster of cells appears twice in one field, or a Western blot in which the background around one band is identical to the background around another, requires a person to select a region and paste it. The characteristic forms of alteration Within the third category, certain patterns recur so often that experienced analysts recognise them almost at a glance. Band splicing without disclosure. A Western blot is assembled from lanes that were not run on the same gel, or not adjacent on the same gel, and the joins are hidden. Tell-tale signs include a sharp vertical discontinuity in the background, a change in background texture or intensity from one lane to the next, or bands that sit at slightly different heights from their neighbours in a way inconsistent with a single run. Splicing is sometimes legitimate if disclosed, which is why the rules require dividing lines; undisclosed, it conceals the conditions under which the comparison was made. Band duplication. A single band, or a set of bands, appears more than once, sometimes in the same blot to represent different samples, sometimes in different blots in the same or different papers. The shapes of bands vary from run to run in small, characteristic ways, so identical shapes with identical surrounding noise indicate a common source. Erasure and background cloning. Something is removed from an image and the gap is filled with a patch of background copied from elsewhere, or with a uniform tone. Erased areas can be detected when the patch has a noise pattern identical to another region, or when it has no noise at all, which real images never lack. Adjusting the contrast of a suspect image to an extreme setting often reveals rectangular or irregular areas whose texture does not match their surroundings. The retraction notice for the 2006 Nature paper on Aβ*56 referred specifically to the use of an eraser tool, alongside splicing and duplication. Cloned cells and tissue. In micrographs, clusters of cells or features of tissue are duplicated within a single field, sometimes to make a population look denser or a treatment effect look stronger. Because the arrangement of cells in any real field is unique, a repeated constellation is a strong signal. Recycled images across papers. The same image appears in papers by the same group published years apart, representing different cell lines, different treatments or even different organisms. This pattern is especially important because it can only be detected by looking across papers, which neither reviewers nor editors usually do. Duplicated plots and spectra. The same logic applies beyond photographs. Flow cytometry dot plots, in which each dot represents a single cell, contain thousands of randomly positioned points; two plots with identical scatter are the same data. Spectra, chromatograms and other traces carry noise that functions like the background in a blot. A worked examination It helps to see how an experienced analyst approaches a single suspect figure. Consider a composite of the kind common in cell biology papers: a Western blot showing a protein of interest across six lanes representing a control and five treatments, above a second blot showing a loading control, such as actin, for the same six lanes. A reader has noticed that the bands in lanes two and five look alike. The analyst begins by obtaining the highest-resolution version of the figure available, usually from the publisher's website or the supplementary files, since images embedded in a downloaded article file are often compressed. She crops the two lanes and places them side by side at high magnification. The bands are similar in shape, but that alone means little; bands of the same protein often look alike. She therefore looks away from the bands, at the background. In the lanes in question there is a small dark speck just above and to the right of each band, at the same distance, and a faint diagonal streak running through both. Those features are incidental, and their repetition is the first real evidence of common origin. She then adjusts contrast and gamma across the whole figure. At extreme settings the background of the blot, which looked uniform, resolves into a mottled texture. Between lanes four and five a sharp vertical line appears, with slightly different background intensity on either side, suggesting that lane five was placed there from elsewhere. She overlays lane two on lane five with partial transparency and finds that the speck, the streak and the mottled texture align exactly, with no rotation. She checks the loading-control blot beneath and finds that it shows no such join, which means that the protein-of-interest blot and its loading control do not share a common history, even though the figure presents them as a matched pair. Finally, she searches for the same image elsewhere: in the other figures of the same paper, and, using image-comparison software, in earlier papers by the same group. She finds nothing further. Her report records what she found and nothing more: that lanes two and five of the upper blot appear to derive from the same original image, that there is a discontinuity consistent with undisclosed splicing between lanes four and five, and that the upper blot and loading control appear not to derive from a single membrane as the layout implies. She does not say that anyone committed fraud. She asks the authors and journal for the original, uncropped images. What happens next depends on whether those images exist and what they show. From image to inference The fact that an image has been duplicated or altered tells the reader that the figure does not show what it claims. That is a sufficient reason to correct or retract the paper if the original data cannot be produced. It is not, on its own, a finding about who did it or why, and the investigators who find these problems are generally careful to say so. Several considerations shape the inference. The first is whether the problem affects the conclusions. A duplicated loading control in a supplementary figure is less consequential than a duplicated result panel that constitutes the main evidence for the paper's claim. The second is pattern. A single duplication in a large paper is more consistent with error than a series of duplications across several papers from the same group, particularly if they span different journals, years and first authors. The third is the response. Authors who can promptly produce the original, unprocessed images showing what the figure should have contained are demonstrating that the underlying experiment was performed. Authors who cannot, or who produce new images that were evidently generated after the question was raised, are demonstrating something else. Some of the most consequential episodes in the history of the field turned on images. In 2005 the South Korean researcher Hwang Woo-suk and colleagues published in Science a paper claiming to have derived patient-specific embryonic stem cell lines by cloning. Within months, anonymous users of a Korean online forum for biologists had pointed out duplicated photographs of cell colonies in the paper's supplementary material. An investigation by Seoul National University concluded that the data in the 2005 paper, and in an earlier 2004 paper in Science, had been fabricated, and both papers were retracted. The image duplications did not by themselves prove fabrication, but they were the thread that, once pulled, unravelled the rest. The Stanford case of 2022 and 2023 followed a similar logic. Concerns about images in papers on which Marc Tessier-Lavigne, then Stanford's president, was an author had been posted on PubPeer, a website on which readers can comment on published papers, some years earlier, and were brought to wide attention by reporting by Theo Baker in the student newspaper The Stanford Daily. A scientific review panel commissioned by the university's board of trustees reported in July 2023. According to the report as published, the panel found no evidence that Tessier-Lavigne had himself manipulated data or had known of the manipulation when it occurred, but it identified repeated instances of manipulation of research data or subpar scientific practices in papers from laboratories he led, and found that he had not taken adequate steps to correct problems when they were raised. He announced his resignation as president and said he would retract or correct several papers; retractions of papers in Cell and Science followed. Both cases show the same structure. The image problem is the entry point. It establishes that something in the record is wrong. What it cannot establish, and what only an investigation with access to people and original data can, is how the wrong thing came to be there. The next chapter turns to the tools that now allow image problems to be found at a scale no individual could manage by eye, and to the new difficulty of images that were never photographed at all. Chapter 3. Screening at Scale: Software, Source Data and the Synthetic Image For most of the history of image forensics the principal instrument was the human eye. Elisabeth Bik's survey of more than twenty thousand papers, described in the previous chapter, was done largely by looking. So was much of the work of the volunteers who scrutinise the literature on PubPeer. The eye is remarkably good at this task once trained: it notices a repeated cluster of cells, a familiar blemish, a band whose shape it has seen before. But it does not scale. A single analyst can examine perhaps a few dozen papers carefully in a day. The world's journals publish several million papers a year. Between those numbers lies the case for automation, and over the past decade software has begun to close part of the gap. This chapter describes what the tools do, how they work in principle, what they miss, and the arrival of a problem they were not built to solve. The analyst's manual toolkit Before automated screening became widespread, image forensics relied on a small set of techniques that remain useful and that explain much of what the software now does. The simplest is aggressive adjustment of brightness, contrast and gamma. An image that looks clean at normal settings can reveal its history when pushed to extremes. Pasted regions often carry a slightly different background level or noise texture from their surroundings; at high contrast those differences become visible as rectangles, sharp edges or areas of unnatural smoothness. Applying a false-colour lookup table, which maps small differences in grey level to strongly contrasting colours, achieves a similar effect. The US Office of Research Integrity has for many years distributed a set of simple scripted actions, known as forensic droplets, for Adobe Photoshop that apply these kinds of transformations to an image so that an investigator can inspect it quickly. The open-source scientific image program ImageJ, and its widely used distribution Fiji, provide equivalent capabilities. A second technique is direct comparison. Two suspect panels are overlaid, one made semi-transparent or subtracted from the other. If they share an origin, the subtraction leaves almost nothing, or the overlay reveals perfect alignment of every speck of noise. If one has been rotated, flipped or rescaled, the analyst applies the inverse transformation before comparing. A third is error level analysis, which exploits the way JPEG compression works. When a JPEG image is saved, it is divided into blocks and compressed with some loss of information. Resaving an image at a known quality level and measuring how much each region changes can reveal areas that were compressed a different number of times from the rest, as can happen when content from another image is pasted in. Error level analysis is available in free online tools such as FotoForensics and Forensically, the latter of which also offers clone detection and noise analysis in a browser. It must be used with care. Many innocent operations, including resizing, format conversion and the publisher's own processing of figures, produce patterns that look suspicious to an untrained eye, and the technique has a history of being misread in public disputes. It is a pointer for closer inspection, not evidence in itself. The common thread in all of these methods is that they look for inconsistency: between two images that should differ but do not, or within one image whose parts should share a history but do not. That principle carries over directly into automated tools. How automated screening works Commercial image-integrity software emerged in the late 2010s and has spread quickly. The two products most widely used by publishers are ImageTwin, developed in Austria, and Proofig, developed in Israel. Several large publishers have also built tools of their own. The details of the algorithms are proprietary, but the general approach is well understood. The software first extracts every image from a manuscript, separating composite figures into their component panels. It then looks for matching regions, both within a single panel and across all panels in the paper. The matching relies on techniques from computer vision that identify distinctive local features in an image, such as corners, edges and textured patches, describe them in a way that is robust to rotation, scaling and changes in brightness, and then search for the same features elsewhere. Where a large number of features match in a consistent geometric arrangement, the software reports a probable duplication and shows the analyst the two regions side by side, often with the transformation that links them. This is essentially copy-move detection, and it is well suited to finding the second and third categories of the Bik taxonomy, repositioned and altered duplicates, which are exactly the ones a tired human is most likely to miss. The more powerful capability, and the one that has changed the economics of detection, is comparison across papers. A duplication within one manuscript can be found by comparing the manuscript to itself. The reuse of an image from a paper published five years earlier, perhaps in another journal, can only be found by comparing the manuscript against a large database of previously published figures. Some vendors have assembled such databases from openly available literature, which allows a submission to be checked against millions of earlier images. The coverage of these databases is limited by what the vendor can legally obtain and index, and a figure published behind a paywall, or in a journal that has not shared its content, may be invisible to them. This is one of several places where the integrity of the record depends on cooperation between publishers that commercial incentives do not naturally produce. Adoption has moved from pilots to policy. In January 2024 the editor-in-chief of the Science family of journals, Holden Thorp, announced in an editorial titled "Genuine images in 2024" that the journals would use Proofig to screen images in papers under consideration, having piloted it for several months with, in his words, clear evidence that problematic figures could be detected before publication. Other publishers have announced screening of all or most submissions, and the STM association's Integrity Hub, a shared infrastructure launched by publishers to detect problems across submissions, has worked on duplicate-image and duplicate-submission detection across participating publishers. What software misses and what it overcalls Automated screening is valuable, but its limits are important, because they shape what the rest of the integrity system must do. It produces false positives. Many legitimate scientific images contain repeating structures: the regular pattern of wells in a plate, the repeated motif of a crystal lattice, the similar appearance of many cells of the same type, scale bars and labels that recur across panels. A loading control may legitimately be shown more than once if the same blot was probed for several proteins and the paper says so. The software flags matches; a trained human must decide whether they matter. Publishers that have adopted these tools generally describe them as triage, with every flag reviewed by an editor or integrity specialist before any action is taken. It produces false negatives. Small duplicated regions, heavily processed images, low-resolution figures and duplications involving images outside the reference database can all escape detection. So can manipulation that does not involve duplication at all: a band that has been darkened, a background that has been uniformly smoothed, a quantification that does not match the image it accompanies. Most fundamentally, duplication-based software can only find a manipulation that leaves a copy behind. It detects reuse. It does not detect invention. An image that was fabricated from nothing, or taken from an experiment different from the one described but never published anywhere, contains no duplicate to find. That limitation was always there, but until recently it mattered less, because inventing a convincing scientific image from nothing was hard. It is no longer hard. The arithmetic of screening A further limit is statistical rather than technical, and it is often overlooked by those who expect software to settle cases. It concerns base rates. Suppose, purely for illustration, that a journal receives a thousand manuscripts containing images, and that 4 percent of them, forty papers, contain a genuine problem, a figure close to the rate Bik and her colleagues found in the published literature. Suppose also that the screening software is good: it detects 90 percent of genuine problems and wrongly flags only 5 percent of clean papers. The software will correctly flag 36 of the 40 problematic papers. It will also flag 5 percent of the 960 clean ones, which is 48 papers. Of the 84 papers flagged in total, then, more than half are clean. And four genuinely problematic papers pass unflagged. The numbers are hypothetical, but the structure is general. Whenever the thing being screened for is rare, even an accurate test produces many false positives relative to true ones. This is why a flag from screening software cannot by itself justify a rejection, still less an accusation, and why every serious user of these tools treats a flag as the start of a human review. It is also why the most useful output of a screening tool is not a verdict but a precise pointer to the regions of an image that match, which a human can then examine with the methods described earlier in this chapter. The same arithmetic runs in the other direction for the published literature, where screening is being applied retrospectively to millions of papers. A tool with a small false-positive rate, applied at that scale, will flag very large numbers of innocent papers in absolute terms. If those flags are posted publicly without review, they can cause real harm to authors who did nothing wrong. The responsible practice among analysts who use these tools on published work is to verify every flag by eye before raising it, which returns part of the bottleneck to human attention. The synthetic image problem Generative machine-learning models, first generative adversarial networks and more recently diffusion models, can produce photographs of faces, landscapes and objects that are difficult to distinguish from real ones. The same techniques can be applied to scientific images. A model trained on a large collection of Western blots can generate new blots, with plausible bands, plausible backgrounds and entirely novel noise. Researchers demonstrated several years ago that such synthetic blots could be produced and that experienced scientists had difficulty distinguishing them from real ones. Similar results have been shown for histology and microscopy images. A synthetic image defeats duplication detection by construction. Its noise is unique, because it was generated rather than copied. It will not match any image in any database. Techniques for detecting generated images exist, based on statistical regularities that current generators leave in the frequency content of an image or on inconsistencies in fine structure, but they are in an arms race with generators that improve continually, and a detector trained on one generation of models may fail on the next. There is already evidence that fabricated images of a more primitive kind were being produced at scale before generative models became widely available. In 2020 Bik and several other analysts who publish under pseudonyms, including those known as Smut Clyde, Morty and Tiger BB8, described a group of more than four hundred papers that shared a distinctive style of Western blot, with bands of a characteristic tadpole-like shape on similar backgrounds. The papers came from different hospitals and research groups, mostly in China, and the blots were not exact duplicates of one another. What linked them was a common appearance, suggesting that they had been produced by the same source rather than by the independent experiments the papers described. The group became known as the Tadpole paper mill. The episode is instructive because detection depended not on finding identical pixels but on recognising a shared manufacturing style, which is a far harder task for software and one that synthetic images will make harder still. Chapter 5 returns to paper mills in detail. From detection to provenance If the forensic analysis of published images is becoming less reliable as generation improves, the natural response is to stop relying on the published image alone and to demand the evidence behind it. This is a shift from detection to provenance: rather than asking whether an image looks manipulated, asking whether it can be traced back to an original acquisition. Several journals and publishers already require or request source data. EMBO Press was an early proponent of publishing the source data behind figures alongside the paper. Many journals now require uncropped, unprocessed versions of all blots and gels as supplementary files, so that readers can see what surrounded the displayed bands. Some funders and institutions require raw data to be retained for a set period and produced on request. These requirements raise the cost of fabrication because they require the fabricator to produce not one convincing image but a consistent set of originals, with plausible file formats, metadata and surrounding context. Stronger forms of provenance are technically feasible. Microscopes, gel imagers and plate readers could sign their output files cryptographically at the moment of acquisition, embedding a verifiable record of which instrument produced the image and when. Industry standards for content provenance, developed largely with news photography in mind, provide a model. Laboratory data management systems could record a hash of every raw file as it is created, making later substitution detectable. None of this is yet widespread in academic research, and all of it faces practical obstacles: older instruments, proprietary file formats, the cost of storage, and the reluctance of researchers to accept systems that feel like surveillance. But the direction is clear. As synthetic images make the published figure less trustworthy as evidence in itself, the evidential weight has to move to the chain of custody behind it. What the tools have changed The practical effect of automated screening has been to move a substantial share of image detection from after publication to before it. A duplication caught at submission costs the journal an email and costs the authors an embarrassing request for original data. The same duplication caught ten years after publication may cost a university a year-long investigation, a journal a contested retraction and a field years of work built on a false premise. Screening at submission is therefore enormously cost-effective, and the most important argument for it is not that it catches fraudsters but that it catches errors while they are still cheap to fix. Screening has also changed the incentives, though more slowly than one might hope. A researcher who knows that every image will be checked against millions of others has less reason to reuse one. A paper mill that knows its standard blots will be recognised has an incentive to switch to generated ones. That second effect is a reminder that detection alone cannot end the problem. Each improvement in screening moves the most determined fraud to the next method that screening cannot see. The published literature, meanwhile, still contains everything that was published before screening became routine, and the tools are now good enough to find much of it. The result is a large and growing backlog of flagged papers awaiting action from journals and institutions. The existence of that backlog, and the reasons it is not being cleared, are among the central concerns of the second half of this book. Before turning to them, the next chapter examines the other great body of forensic method: the analysis of numbers. Hashtags: #ScientificIntegrityAndWhistleblowing #ScientificIntegrity #ResearchMisconduct #ResearchFraud #Whistleblowing #ResearchIntegrity #ImageManipulation #ImageForensics #DataFalsification #DataFabrication #WesternBlotManipulation #ImageDuplication #ResearchWhistleblowers #ScientificSleuthing #PaperMills #PublicationEthics #ResearchTransparency #ForensicStatistics #GRIMTest #DataProvenance #SourceData #Retractions #ResearchAccountability #FraudDetection #FutureOfResearchIntegrity

  • Geospatial Analysis in R and Python (Remote Sensing, GIS, and Ecological Niche Modeling)

    Download the Book (PDF): Introduction In the late summer of 1854, John Snow walked the streets of Soho with a notebook, recording where people had died of cholera. The map he later published, with its stacked bars of deaths clustered around the Broad Street pump, has become the founding image of spatial epidemiology. It is usually told as a story about a map. It is better told as a story about decisions. Snow chose which deaths to count, how to assign each one to a household address, which water sources to mark, and how to measure proximity. He knew that straight-line distance was the wrong measure, so he also traced walking routes through the street network. Every one of those choices shaped what the map could show. The pump handle came off because the choices were good ones. A modern analyst facing the same kind of question has tools Snow could not have imagined. A disease surveillance team can pull geocoded case records from a health information system, overlay them on a population grid at 100-metre resolution, attach land surface temperature from a satellite that passes overhead every day, compute a local cluster statistic in a few seconds, and fit a Bayesian spatial model that borrows strength from neighbouring districts. An ecologist can download a few hundred thousand occurrence records for a mosquito species, stack them against nineteen bioclimatic variables, and project the species' suitable range onto climate scenarios for the end of the century before lunch. What has not changed is that every result still rests on choices about representation, scale and dependence. The difference is that software now makes most of those choices silently. A spatial join assigns each case to a polygon using a rule the analyst may never have read. A raster reprojection resamples values with an interpolation method set by default. A regression treats observations as independent unless told otherwise. A species distribution model draws ten thousand background points from wherever the extent happens to reach. In each case the code runs, the output looks plausible, and the error, if there is one, is invisible in the result. This booklet is built on a single claim: the craft of geospatial analysis consists in making explicit the choices the software would otherwise make for you. The tools in R and Python are now mature, fast, and largely interchangeable. Both ecosystems sit on the same small set of open-source libraries that read file formats, transform coordinates and compute geometric relationships. What separates a defensible analysis from a misleading one is rarely the language or the package. It is whether the analyst knew which coordinate reference system was in play, what spatial support each variable had, how neighbours were defined, how the observations were sampled, and how the model was validated against data that were not simply its own spatial neighbours. Who this is for The intended reader already does some quantitative work in R or Python and has an environmental or health question with a place in it. That reader might be an epidemiologist who has mapped incidence rates by district and wants to know whether the clustering is real, an ecologist who has been handed occurrence records and asked for a distribution map, a public health analyst estimating how many people live within a given distance of a contaminated site, or a graduate student who has inherited a pipeline built on packages that no longer install. The booklet assumes familiarity with ordinary regression and with writing a script. It does not assume training in geography, cartography or remote sensing. Because the reader works in one language and often collaborates with people who work in the other, the booklet treats R and Python side by side. The aim is not to teach either syntax in full. Package documentation does that better, and it changes every year. The aim is to explain what the operations do, where they can go wrong, and what the equivalent tool is called in each ecosystem, so that a reader can move between them without mistaking a difference in naming for a difference in method. What the booklet covers, and what it leaves out The chapters move from the representation of spatial data to its analysis. The first chapter describes the shared software engine beneath both ecosystems and the two data models, vector and raster, that everything else builds on. The second deals with coordinate reference systems, the most common source of quiet error in the field. The third and fourth cover the operations that construct analytical datasets: vector overlays, spatial joins and aggregation for exposure assessment, and raster algebra, resampling and zonal statistics. The fifth turns to satellite imagery, from the physics of reflectance through cloud masking, spectral indices and compositing to the accuracy assessment of classified maps. The last three chapters deal with inference. The sixth covers spatial autocorrelation: how neighbours are defined, what Moran's I and its local variants measure, and how to test for clusters of disease without being fooled by small populations. The seventh covers spatial regression, from the classical lag and error models of spatial econometrics to geographically weighted regression and the hierarchical Bayesian models now standard in small-area disease mapping. The eighth covers ecological niche and species distribution modelling, which has become central to the epidemiology of vector-borne and zoonotic disease, and which has its own well-documented ways of producing confident maps from biased data. Several substantial topics are left out, deliberately. Cartographic design is not covered, because the booklet is concerned with analysis rather than presentation. Point process models receive only brief mention. Network analysis, trajectory analysis of movement data, and deep learning for image segmentation are each large enough to deserve their own treatment and are named only where they touch the main argument. Commercial desktop GIS is not discussed at all. Where a topic is omitted, the reason is that treating it at booklet length would mean treating everything else more thinly. A note on software The R spatial ecosystem went through a significant transition between 2020 and 2023. The packages that had carried it for fifteen years, rgdal, rgeos and maptools, were retired from CRAN in October 2023, and the older raster package has been superseded by terra. Current practice rests on sf for vector data, terra and stars for rasters and data cubes, and spdep, spatialreg and related packages for spatial statistics. In Python, geopandas reached version 1.0 in 2024 and is built on shapely 2.0, whose vectorised geometry operations made large vector workflows practical. Raster work runs through rasterio, xarray and rioxarray, and spatial statistics through the PySAL family of packages. Any booklet that named function signatures in detail would be out of date before it was printed. Package and function names are given where they help a reader find the right tool, and the reader should expect small details to drift. The concepts beneath them, which are the subject of this booklet, do not. How to read it The chapters are ordered so that each builds on the last, and a reader new to the field should take them in sequence. A reader with a specific problem can go directly to the relevant chapter, though the second chapter on coordinate systems is worth reading by everyone, because its errors propagate into every later step. Throughout, the emphasis falls on reasoning rather than recipes. The recurring question is not "which function do I call?" but "what has this operation assumed, and is that assumption true of my data?" An analyst who asks that question at every step will produce work that holds up. One who does not will, sooner or later, produce a beautiful map of an artefact. Chapter 1. One Engine, Two Languages A newcomer to geospatial analysis often spends the first weeks deciding whether to learn it in R or in Python, as though the choice would determine what analyses are possible. It mostly does not. Beneath the surface syntax, both ecosystems call the same compiled libraries to read files, transform coordinates and compute geometric relationships. When an R user reads a shapefile with sf and a Python user reads the same file with geopandas, the bytes pass through the same code. When each reprojects the data, the same transformation pipeline runs. When each asks whether two polygons intersect, the same algorithm answers. Understanding that shared foundation is the first step toward understanding why results agree across languages when they should, and why they occasionally do not. The three libraries underneath Three open-source C and C++ libraries do most of the work in both ecosystems. GDAL, the Geospatial Data Abstraction Library, reads and writes spatial file formats. Its raster side handles GeoTIFF, NetCDF, HDF, JPEG2000 and more than a hundred other formats; its vector side, historically called OGR, handles shapefiles, GeoPackage, GeoJSON, PostGIS connections, KML and many others. GDAL also provides virtual file systems that let software read a file directly from a web server or cloud bucket without downloading it first, a capability that has reshaped remote sensing workflows and is taken up again in Chapter 5. PROJ transforms coordinates between reference systems. Since its sixth major version, released in 2019, PROJ has used a formal description of coordinate reference systems known as WKT2 and a database of transformations maintained from the EPSG registry. The change mattered because the older approach, which described systems with short text strings, silently dropped information about datums. Transformations that had been accurate to metres could be off by a hundred metres or more without any warning. Chapter 2 deals with this at length. GEOS, the Geometry Engine Open Source, is a C++ port of the Java Topology Suite. It computes predicates such as intersects, contains and touches, and operations such as buffers, unions, intersections and convex hulls, on planar geometries. It treats coordinates as if they lay on a flat plane. That assumption is harmless for projected data in metres over a modest area and wrong for longitude and latitude over a large one. For that reason sf in R, from version 1.0 onward, sends geometric operations on geographic coordinates to a different library, s2, which Google developed to compute on the sphere. An intersection or buffer on longitude and latitude in sf is therefore computed on a spherical model by default, while the same operation in geopandas is computed on the plane by GEOS unless the analyst projects the data first. This is one of the few places where the two ecosystems give different answers to what looks like the same call, and it is a direct consequence of which engine each hands the work to. The vector data model Vector data represent the world as discrete objects with explicit boundaries: points, lines and polygons, each carrying attributes. The standard that governs them is the Simple Features specification of the Open Geospatial Consortium, also published as ISO 19125. It defines seven core geometry types: point, linestring, polygon, and their multi-part variants multipoint, multilinestring and multipolygon, plus a geometry collection that can hold a mixture. Both sf and geopandas implement it directly, which is why a GeoPackage written in one reads cleanly in the other. In both ecosystems a vector dataset is a data frame with one special column holding geometries. In sf it is an sf object whose geometry column is a list of geometries with an attached coordinate reference system; in geopandas it is a GeoDataFrame with an active GeoSeries. Everything an analyst knows about data frames carries over: filtering rows, joining tables, grouping and summarising. The spatial operations are additional verbs that act on the geometry column. Simple Features has rules about validity that matter more than they appear to. A polygon's exterior ring must not cross itself; interior rings, which represent holes, must lie inside the exterior ring and must not overlap each other. Administrative boundary files, especially those digitised by hand or simplified for web display, frequently break these rules. An invalid polygon can make an intersection fail outright, or worse, return a result that is wrong without raising an error. Both ecosystems provide validity checks and repair functions, st_is_valid and st_make_valid in sf, is_valid and make_valid in geopandas. Running the check on every boundary file before analysis costs seconds and prevents a class of error that is otherwise very hard to trace. A second subtlety concerns topology. Simple Features stores each polygon independently. Two neighbouring districts share a boundary, but each stores its own copy of the vertices along it. If the file has been simplified, the copies may no longer coincide, leaving slivers of overlap or gaps between districts. This becomes important in Chapter 6, where the definition of which areas are neighbours depends on whether their boundaries touch. A gap of a few metres between two districts that obviously share a border will make them non-neighbours under a strict contiguity rule, and the resulting spatial weights will be wrong. The raster data model Raster data represent the world as a regular grid of cells, each holding a value. A satellite image is a raster; so is a digital elevation model, a gridded climate surface, or a population density layer. The grid is defined by its extent, its cell size, its number of rows and columns, and its coordinate reference system. Values are stored in bands, so a single raster file might hold one band of elevation or thirteen bands of multispectral reflectance. The raster model has two conceptual traps. The first concerns what a cell value means. In a digital elevation model, the value may represent the elevation at the cell centre. In a satellite image, it represents a weighted average of reflected light over roughly the cell's footprint, blurred by the sensor's point spread function. In a population grid, it represents a count of people within the cell. In a climate surface, it is an interpolated estimate from station data, with uncertainty that the file does not record. These are different kinds of quantities, and they behave differently when the grid is resampled or aggregated. Averaging population counts to a coarser grid is wrong; they must be summed. Summing temperatures is meaningless; they must be averaged. The software does not know which kind of quantity a layer holds, so the analyst must. The second trap concerns alignment. Two rasters can have the same cell size and coordinate reference system and still be offset from each other by half a cell, because their grid origins differ. Map algebra between them, such as dividing one by the other, requires that cells coincide. If they do not, one layer must be resampled onto the other's grid, and resampling changes values. Chapter 4 treats this in detail. In R, terra is the main package for raster data. It was written by Robert Hijmans as the successor to his earlier raster package, with a C++ core that makes it much faster and able to process files larger than memory by working through them in blocks. The stars package, developed by Edzer Pebesma, takes a different approach: it represents spatiotemporal arrays with any number of dimensions, which suits data cubes of many images over time. In Python, rasterio provides a thin, explicit interface to GDAL for reading and writing raster files, while xarray provides labelled multidimensional arrays, and rioxarray connects the two so that an xarray object knows its coordinate reference system and geotransform. Data cubes and the blurring of the models The distinction between vector and raster is less absolute than it once was. Satellite archives now deliver thousands of images of the same place over time, and analysts increasingly treat them as a single four-dimensional array indexed by band, time, and two spatial coordinates. This is the data cube model. In Python, xarray has become the standard container for it, often backed by dask so that computations are split into chunks and run lazily across many processors. In R, stars fills the same role, and the gdalcubes package builds regular cubes from irregular collections of images. Vector data can also be cubed. A set of district polygons with weekly case counts over five years is a vector data cube: one spatial dimension indexed by district and one temporal dimension indexed by week. stars supports this directly, and the representation is natural for spatiotemporal disease surveillance. The practical consequence is that the analyst should think less about which data model a file uses and more about what the dimensions of the problem are. A question about how vegetation greenness in the three months before an outbreak relates to case counts per district involves a raster cube, a vector cube, and an aggregation from one to the other. The software can perform each step, but only the analyst can say which aggregation respects the meaning of the variables. Reading the same data in both languages The ecosystems have converged enough that most operations have a direct counterpart in the other language, as Table 1 sets out. The table lists the main package for each task. Names change over time, but the pairing has been stable for several years. Table 1. Principal packages for common geospatial tasks in R and Python. Task R Python Vector data frames sf geopandas, shapely Raster files terra rasterio, rioxarray Data cubes stars, gdalcubes xarray, dask Coordinate transformation sf, terra (via PROJ) pyproj Spatial weights and autocorrelation spdep libpysal, esda Spatial regression spatialreg, GWmodel, R-INLA spreg, mgwr Zonal statistics exactextractr rasterstats, exactextract Species distribution models dismo, predicts, biomod2, ENMeval elapid, scikit-learn Several points about the pairing are worth stating. First, the R side of spatial regression and Bayesian disease mapping is considerably deeper than the Python side. The R-INLA package, which fits latent Gaussian models by integrated nested Laplace approximation, has no mature Python equivalent, and much of the small-area estimation literature publishes code in R. Second, the Python side is stronger in large-scale remote sensing, machine learning and cloud-native processing, because xarray, dask and the scientific Python machine learning stack were built with large arrays in mind. Third, the SDM tooling in R is more mature and more thoroughly tested, though the elapid package, described in the Journal of Open Source Software in 2023, provides a Python implementation of Maxent and related tools. When the two disagree Because both ecosystems use GDAL, PROJ and GEOS, they usually agree to many decimal places. When they do not, the cause is almost always one of four things. The first is the spherical versus planar distinction already described. A buffer of one degree around a point on geographic coordinates means different things on a sphere and on a plane, and the two ecosystems default to different ones. The second is library versions. sf and geopandas may be linked against different versions of PROJ, and a newer PROJ may use a more accurate transformation grid that an older one lacked. Both ecosystems report the versions of their linked libraries, through sf_extSoftVersion in R and geopandas.show_versions in Python, and those versions belong in any methods section that reports transformed coordinates. The third is default parameters. Resampling methods, the treatment of cells that are partly inside a polygon during zonal statistics, and the handling of missing values all have defaults, and the defaults differ between packages. A mean temperature by district computed in one package may differ from another because one counted only cells whose centres fell inside the district while the other weighted every cell by the fraction of its area inside. The fourth is the definition of neighbours in spatial statistics. Queen and rook contiguity, distance thresholds, the number of nearest neighbours and the row standardisation of weights are all choices, and the defaults in spdep and libpysal are not always identical. None of these differences is a bug. Each is a choice that one package made on the analyst's behalf. When results from the two languages disagree, the disagreement is useful: it reveals a choice that had not been made explicit. The remainder of this booklet is, in a sense, a catalogue of such choices and of how to make each one deliberately. File formats and why they matter The choice of file format is usually treated as a matter of convenience, but several formats in common use carry limitations that shape analysis. The shapefile, introduced by Esri in the early 1990s, remains the most widely distributed vector format, and it is a poor one. A single dataset is spread across at least three files, and often six or seven, which travel separately and are easily separated. Attribute field names are limited to ten characters, so a column named population_2020 becomes truncated on writing, sometimes colliding with another truncated name. Text encoding is ambiguous, which corrupts place names with accented characters. Individual component files cannot exceed two gigabytes. There is no way to store a proper null value in numeric fields, so missing values are often written as zero, which is catastrophic for a column of case counts. GeoPackage, an OGC standard built on SQLite, avoids all of these problems in a single file, and it should be the default for exchanging vector data. For large analytical datasets, GeoParquet, which stores geometries in the columnar Apache Parquet format, has become the fastest option in both ecosystems; geopandas reads and writes it natively, and sf can do so through GDAL or the sfarrow and geoarrow packages. FlatGeobuf fills a similar niche for streaming vector data over the web. For rasters, the GeoTIFF remains standard, and its Cloud Optimized variant, the COG, organises the file internally in tiles with overviews so that software can request only the portion it needs over a network. NetCDF and HDF5 are common for climate and atmospheric data with many time steps. Zarr, a chunked array format designed for cloud storage, is increasingly used for large data cubes and is read natively by xarray. The format affects not only speed but correctness: a GeoTIFF records a nodata value in its header, and software that ignores it will treat the fill value, often minus 9999 or zero, as a real measurement. Reproducibility as a design constraint Geospatial workflows are harder to reproduce than most statistical work. They depend on compiled libraries whose versions matter, on large input files that may be revised by their publishers, and on web services that change. A climate surface downloaded in one year may be silently updated the next; a satellite archive may be reprocessed to a new collection, changing every reflectance value slightly. The practical defences are straightforward. Record the versions of GDAL, PROJ and GEOS alongside package versions. Use environment managers, renv in R and conda or pixi in Python, which can pin compiled libraries as well as packages. Record the exact version or collection identifier of every remote dataset, and where possible the date it was retrieved. Store intermediate products in open formats such as GeoPackage for vectors and Cloud Optimized GeoTIFF or Zarr for rasters, so that a reviewer can inspect them without the original software. Container images, built with Docker or Apptainer, provide the strongest guarantee when a workflow must be rerun years later. These habits matter more in spatial work than elsewhere because spatial errors are hard to see. A misaligned raster or a mistransformed coordinate produces numbers that look entirely reasonable. The only protection is a workflow whose every step can be inspected and rerun, and whose choices are recorded rather than inherited from defaults. Chapter 2. Coordinates, Projections and the Geometry of Error Every spatial dataset makes a claim about where things are, and that claim is only meaningful relative to a coordinate reference system. Two numbers such as 51.5 and minus 0.12 identify a point in central London only if the reader knows that they are latitude and longitude, in that order, in degrees, on a particular model of the Earth. Change any of those assumptions and the same two numbers identify somewhere else, or nowhere. Coordinate reference systems are the least glamorous part of geospatial analysis and the most frequent source of error that nobody notices, because a mistransformed dataset still draws on a map and still produces statistics. It just produces the wrong ones. Datums, ellipsoids and the moving Earth A geographic coordinate reference system has three components: an ellipsoid, a datum, and a prime meridian. The ellipsoid is a mathematical surface that approximates the shape of the Earth, slightly flattened at the poles. The datum ties that ellipsoid to the physical Earth by fixing its position and orientation. The prime meridian, almost always Greenwich, sets where longitude zero lies. Different datums place the same ellipsoid in slightly different positions, so the same point has different coordinates under each. The differences are not trivial. In Great Britain, the historical Ordnance Survey datum OSGB36 differs from the global WGS84 by up to about 120 metres. In North America, coordinates under NAD27 can differ from modern systems by tens of metres or more. Even between modern datums the differences are measurable: NAD83 and WGS84 were nearly identical when defined in the 1980s but now differ by one to two metres across much of the continent, because the North American plate has moved and the two systems treat that movement differently. Plate motion is not a pedantic concern. The Australian plate moves north-east at about seven centimetres a year, and by the 2010s the country's official datum, GDA94, was out of step with satellite positioning by around 1.8 metres. Australia adopted GDA2020 to correct this. For most environmental and epidemiological work, errors of a metre or two are well within the uncertainty of the data. For high-resolution imagery with ten-metre pixels, or for linking survey plots to individual trees, they are not. WGS84, identified by the code EPSG 4326, is the datum used by GPS and by most global datasets. It is not a single fixed system but a series of realisations that have been updated several times to track the International Terrestrial Reference Frame. When a dataset says it is in WGS84, the realisation is usually unstated, and for most purposes that ambiguity is harmless. PROJ, in recent versions, will warn when a transformation involves a datum ensemble whose members differ by more than a threshold, and the warning should be read rather than suppressed. Projections and what they preserve A projected coordinate reference system flattens the curved surface onto a plane so that positions can be expressed in linear units, usually metres. No flat map can preserve all geometric properties of a sphere. Every projection trades off among area, shape, distance and direction. Equal-area projections preserve area: a square kilometre anywhere on the map covers a square kilometre on the ground. Albers equal-area conic and Lambert azimuthal equal-area are common choices for continental analysis; the Mollweide and equal Earth projections serve for global work. Conformal projections preserve local angles and therefore local shape, at the cost of distorting area. Mercator is the famous example, as is Transverse Mercator, on which the Universal Transverse Mercator system is based. Equidistant projections preserve distances from one or two points, or along certain lines. The practical rule is simple: choose the projection according to the calculation. If the analysis computes areas, densities or rates per square kilometre, it needs an equal-area projection. If it computes distances over a small area, a local conformal projection such as the appropriate UTM zone keeps distance errors below about one part in a thousand within the zone. If it computes distances over a large area, it should compute them geodesically on the ellipsoid rather than on any projection. Web Mercator, EPSG 3857, deserves particular warning. It is the projection of most online map tiles, and data exported from web tools often arrive in it. It is conformal and wildly non-equal-area: at 60 degrees latitude, areas are inflated by a factor of four relative to the equator, and the inflation grows without bound towards the poles. On a Web Mercator map Greenland looks roughly the size of Africa, though Africa is about fourteen times larger. Any area, density or buffer computed in Web Mercator coordinates is distorted in proportion to latitude. It is a display projection and should never be used for measurement. Geographic coordinates as a trap Much environmental data arrives in longitude and latitude, and a great deal of analysis is carried out on it without projection. That is sometimes fine and sometimes badly wrong. A degree of latitude is about 111 kilometres everywhere. A degree of longitude is about 111 kilometres at the equator and shrinks with the cosine of latitude: about 78 kilometres at 45 degrees and 56 kilometres at 60 degrees. A buffer of 0.1 degrees around a point is therefore a circle only at the equator; everywhere else it is an ellipse stretched north and south. A nearest-neighbour search on raw degrees in Scandinavia will systematically prefer neighbours to the east and west over those to the north and south. The same geometry affects rasters. A global climate grid with 30 arc-second cells, the resolution of WorldClim, has cells of about 0.86 square kilometres at the equator but only about 0.43 square kilometres at 60 degrees latitude. Any procedure that treats cells as equal, such as sampling background points uniformly by cell for a species distribution model or counting cells to estimate the area of suitable habitat, over-represents high latitudes. The correction is to weight by true cell area, which terra computes with cellSize, or to work in an equal-area grid. As noted in Chapter 1, sf now avoids many of these errors for vector operations by computing on the sphere through the s2 library whenever data are in geographic coordinates. Distances, areas and buffers are then computed correctly in metres. geopandas does not do this: its geometric operations are planar. Python users must either project the data before computing areas and distances or use pyproj's Geod class, which computes geodesic distances and areas on the ellipsoid using the algorithms Charles Karney published in 2013, accurate to within a few nanometres. The axis-order problem A particularly irritating source of error concerns the order of coordinates. Mathematicians and most GIS software write coordinates as x then y, which for geographic data means longitude then latitude. The EPSG registry, however, defines EPSG 4326 with latitude first, following geodetic convention. For years most software ignored the official order. Since GDAL 3 and PROJ 6, the libraries respect the authority's axis order by default, while offering an option to override it. The consequence is that the same file can be read with swapped coordinates depending on software and settings. sf uses traditional GIS order, longitude first. pyproj requires an always_xy argument set to true to guarantee longitude-first order when creating a transformer; without it, a transformation from EPSG 4326 expects latitude first. Web services following the OGC's WMS 1.3 and WFS 2.0 standards also honour authority order. A dataset with swapped axes places points at the mirror position across the line where latitude equals longitude, which for data in Africa or Europe often lands in the ocean and is easy to spot, but for data near that line can produce plausible-looking errors. The defence is a set of simple checks run on every point dataset as it enters an analysis. Plot the points over a base map. Confirm that longitudes fall between minus 180 and 180 and latitudes between minus 90 and 90. Check that points fall on land if they should, within the expected country if they should, and not at exactly zero, zero, a location in the Gulf of Guinea that is the default value in many databases and has become known as Null Island. Occurrence data for species distribution models, discussed in Chapter 8, suffer from all of these problems, and the CoordinateCleaner package in R was written specifically to detect them. Reprojecting vectors and rasters Reprojecting vector data is conceptually clean. Each vertex is transformed independently, and the geometry is reconstructed from the transformed vertices. The only subtleties are that straight lines between vertices in one projection are not straight in another, so long edges may need densifying before transformation, and that geometries crossing the antimeridian at 180 degrees longitude can wrap incorrectly. Reprojecting rasters is not clean. The grid of the output raster does not align with the grid of the input, so every output cell must be assigned a value computed from input cells that only partly overlap it. This is resampling, and the choice of method changes the data. Nearest-neighbour resampling assigns each output cell the value of the input cell closest to its centre. It preserves original values exactly and is the only correct choice for categorical data such as land cover classes, where averaging class codes would produce nonsense. It produces blocky artefacts and can duplicate or drop cells. Bilinear interpolation takes a distance-weighted average of the four nearest input cells and cubic convolution uses sixteen; both produce smoother surfaces suited to continuous variables such as temperature or elevation, but both invent values that were never measured and smooth away local extremes. Average, sum, mode, minimum and maximum methods aggregate all input cells that fall within each output cell, and are appropriate when moving from fine to coarse resolution. Population counts must be resampled with a sum, or with an area-weighted redistribution, never with bilinear interpolation, which does not preserve totals. Each reprojection degrades a raster slightly. Two reprojections degrade it more. The rule that follows is to reproject rasters as few times as possible, ideally once, and to prefer transforming vector data into the raster's coordinate system rather than the reverse. When a study requires many raster layers on a common grid, the analyst should define that grid explicitly, with its projection, origin, cell size and extent, and resample every layer onto it once, choosing the method layer by layer according to what the values represent. In terra this is done with project using a template raster; in Python, rioxarray's reproject_match performs the same operation. Heights and discrete global grids Two further reference questions arise often enough in environmental work to need a word. The first concerns height. Elevations are measured relative to a vertical datum, and there are two broad kinds. Ellipsoidal heights, which GPS receivers report directly, are measured from the mathematical ellipsoid. Orthometric heights, which is what most people mean by elevation above sea level, are measured from the geoid, an irregular surface that approximates mean sea level and departs from the ellipsoid by as much as about a hundred metres in either direction. The Shuttle Radar Topography Mission elevation model, for instance, reports heights relative to the EGM96 geoid, while the Copernicus global digital elevation model uses the later EGM2008. Combining GPS heights from field plots with an elevation model without converting between the two can produce elevation errors of tens of metres. For studies where altitude is itself an exposure, as in the epidemiology of malaria transmission in East African highlands, where small differences in elevation shift temperature enough to change mosquito development rates, such errors matter. The second concerns alternatives to conventional projections for global or multi-scale aggregation. Discrete global grid systems partition the Earth's surface into cells of roughly equal area at several nested resolutions. The most widely used is H3, a hexagonal hierarchical index released as open source by Uber in 2018, with bindings in both R and Python. Hexagons have the useful property that every neighbour is at the same distance from the cell centre, unlike squares, whose diagonal neighbours are farther than edge neighbours. That makes hexagonal grids attractive for aggregating point data such as case locations or species occurrences to a uniform lattice before computing autocorrelation or density. H3 cells are not exactly equal in area, and their boundaries do not nest perfectly across resolutions, but for many purposes they avoid both the latitude distortion of geographic grids and the need to choose a projection. The price is that most raster data do not come in hexagons, so the aggregation step must be done carefully and with the same attention to what each variable represents that applies to any resampling. Choosing a working system A defensible choice of coordinate reference system for a study usually follows from three questions. What calculations will be performed? Areas and densities need equal area. Local distances need a conformal projection or geodesic computation. Neighbour relations for autocorrelation need a system in which distance means the same thing in every direction. How large is the study area? A single city or district fits comfortably in a UTM zone or a national grid. A country spanning several UTM zones, or a continent, needs a conic or azimuthal projection centred on the region. A global analysis should either use an equal-area global projection or work on the ellipsoid directly. What are the native systems of the key rasters? If the most important raster, such as a satellite scene, is in a particular UTM zone, it is often better to bring everything else into that zone than to reproject the scene. For health data, a fourth consideration applies. Many national statistical agencies publish boundaries in a national system, and many health information systems store facility locations in WGS84. Merging them without explicit transformation is one of the most common errors in applied work. The fix is to declare the coordinate reference system of every dataset as it is read, transform to the working system immediately, and check the result visually. Positional accuracy of health data Epidemiological data carry positional error beyond anything introduced by projection. Case addresses are converted to coordinates by geocoding, which matches street addresses against a reference database. Match rates vary widely, and addresses that fail to match are not missing at random: rural addresses, informal settlements, new developments and addresses written in non-standard formats fail more often, so geocoded datasets systematically under-represent certain populations. Studies in the geocoding literature have repeatedly found that positional errors for matched addresses are larger in rural areas, where an address may be interpolated along a long road segment, than in dense urban grids. Privacy adds a further, deliberate layer of error. The Demographic and Health Surveys programme, which runs household surveys in many low- and middle-income countries, publishes cluster locations with random displacement: up to two kilometres for urban clusters and up to five kilometres for rural clusters, with one percent of rural clusters displaced by up to ten kilometres. The displacement is constrained to stay within the correct administrative area. Analysts who link DHS clusters to environmental rasters, such as distance to a health facility or mean vegetation index, must account for it, typically by extracting values over a buffer that matches the displacement radius rather than at the published point. Geomasking of this kind protects confidentiality, but it also biases any analysis that relates outcomes to fine-scale exposures, generally toward the null, because the exposure assigned to each cluster is a noisy version of the true one. The lesson generalises. The coordinates in a health dataset are an estimate, with an error structure that depends on how they were produced. That error structure belongs in the analysis plan, not in a limitations paragraph added afterwards. Chapter 3. Vector Operations and the Construction of Exposure Most spatial analyses in environmental health begin with a question of linkage. Which people live near the smelter? Which villages lie within the flood extent? What proportion of each district's population is within an hour's travel of a clinic? Which survey households drink from surface water in an area where the river is contaminated? None of these questions is answered by a statistical model. Each is answered by a sequence of geometric operations that construct a dataset, and the statistical model then takes that dataset as given. If the construction is wrong, the model faithfully analyses the wrong thing. This chapter concerns those constructions: spatial joins, buffers, overlays, distance calculations and aggregations. They are simple to perform and deceptively hard to perform well. The same few pitfalls recur across studies, and each one is a choice the software makes unless the analyst makes it first. Spatial joins and their predicates A spatial join attaches attributes from one layer to another based on a geometric relationship rather than a shared key. Joining case locations to district polygons, to find which district each case falls in, is the canonical example. In sf it is st_join; in geopandas it is sjoin. Both default to the intersects predicate: a case is joined to every district its geometry intersects. For points strictly inside polygons, intersects behaves as expected. Points exactly on a shared boundary intersect both polygons and are joined twice, silently duplicating the case. With coordinates rounded to four or five decimal places, as health records often are, points on boundaries are more common than one might expect. The fix depends on intent. Using the within predicate excludes boundary points entirely; assigning each point to a single polygon, by keeping the first match or the largest overlap, keeps them once. In sf, st_join has a largest argument that assigns each feature to the polygon it overlaps most. Whichever rule is chosen, the analyst should count rows before and after every join, because a join that changes the number of rows has almost always done something unintended. Joins between polygon layers are subtler still. Joining health facility catchments to administrative districts will match every catchment to every district it touches, including districts it overlaps by a few square metres because of boundary imprecision. The largest-overlap rule, or an explicit intersection followed by area weighting, is almost always what is wanted. A further hazard is that spatial joins are left joins by default in sf and inner joins by default in geopandas. A case that falls outside every district, perhaps because it was geocoded to the wrong side of a coastline, is retained with missing district attributes in one and dropped silently in the other. The unmatched records are exactly the ones worth inspecting. Buffers and proximity as exposure Proximity is the most common proxy for environmental exposure. Residents within 500 metres of a major road are classed as exposed to traffic pollution; households within five kilometres of a mine are classed as exposed to its emissions; children within a certain distance of a malaria vector breeding site are classed as at elevated risk. Buffers are easy to compute and hard to justify. The first question is the metric. A buffer drawn in geographic coordinates without spherical computation is not circular, as Chapter 2 explained. A buffer drawn in a projected system is circular in that projection, which is close enough over short distances in an appropriate projection and wrong in Web Mercator. The second question is the cut-off. A binary buffer asserts that exposure is uniform inside and absent outside, which is rarely true. Traffic-related air pollutants such as ultrafine particles and nitrogen dioxide decline steeply over the first few hundred metres from a busy road, then more slowly. A study that compares residents within and beyond a fixed buffer will find different effects depending on where the cut-off falls, and the temptation to try several cut-offs and report the one that gives the clearest result is a well-known route to false findings. Where the decay of exposure with distance is known from measurement, it is better to model it directly, with a continuous distance term or a dispersion model. Where it is not known, the cut-off should be fixed in advance on physical grounds and sensitivity to it reported. The third question is whether straight-line distance is the right measure at all. Snow understood in 1854 that walking distance through streets, not straight-line distance, determined which pump a household used. Modern equivalents are everywhere. Access to health care depends on travel time along roads and paths, not on distance as the crow flies. A village separated from a clinic by a river without a bridge is far from it regardless of the straight-line distance. Network distance along a road graph can be computed with sfnetworks or dodgr in R and with osmnx and networkx in Python, using road data from OpenStreetMap. Where roads are sparse, as across much of rural Africa and Asia, a friction surface approach is more realistic. Each raster cell is assigned a travel speed according to its land cover, road class and terrain, and a least-cost path algorithm computes the minimum travel time from every cell to the nearest destination. The Malaria Atlas Project published a global friction surface of this kind, with a map of travel time to cities for 2015 by Daniel Weiss and colleagues in Nature in 2018 and a follow-up mapping travel time to health care facilities in Nature Medicine in 2020. Both surfaces are freely available and have become standard covariates for access-related outcomes. The gdistance package in R and the scikit-image graph routines in Python provide the cost-distance computations needed to build such surfaces locally. Overlay and the transfer of attributes between units Environmental and health data rarely share geographic units. Cases are reported by health district, population by census tract, pollution by monitoring station or model grid, land cover by pixel. Bringing them together requires moving attributes from one set of units to another, and the method of transfer matters. The simplest approach is areal weighting. The source polygons are intersected with the target polygons, and each piece receives a share of the source attribute in proportion to its area. If a census tract with 4,000 residents is split evenly between two health districts, each district receives 2,000. This assumes that the attribute is spread uniformly over the source polygon, which for population is almost never true. A rural tract whose residents are concentrated in one village at its northern edge will have its population misallocated in proportion to area. Dasymetric mapping improves on areal weighting by redistributing counts according to ancillary information about where they are likely to be. For population, the ancillary data are usually land cover, building footprints or nighttime lights: residents are allocated only to cells classed as built-up, or in proportion to building volume. Gridded population products such as WorldPop, developed at the University of Southampton, and the Global Human Settlement Layer produced by the European Commission's Joint Research Centre, are themselves the output of dasymetric models, redistributing census counts to cells of about 100 metres using satellite-derived covariates. For most health applications in low- and middle-income countries, these products provide better denominators than areal weighting of census polygons, with the caveat that their accuracy is lowest precisely where census data are oldest and settlement patterns are changing fastest. Intensive variables, those expressed per unit such as rates, densities and concentrations, must be transferred differently from extensive variables such as counts. A rate cannot be split by area; it must be averaged, and the average should be weighted by population rather than by area if the rate refers to people. Many analyses go wrong by applying areal weighting to a rate or area-weighted averaging to a count. The modifiable areal unit problem Once data are aggregated to areas, results depend on the areas chosen. This is the modifiable areal unit problem, a name given to it by the geographer Stan Openshaw in a 1984 monograph, though the phenomenon had been observed decades earlier. It has two parts. The scale effect is that statistics change as units are aggregated into larger units. Correlations between variables tend to strengthen with aggregation, because aggregation averages out individual-level noise. A modest association between deprivation and hospital admission at the level of small census areas can become a strong one at the level of large districts. The zoning effect is that statistics change when the same area is divided into the same number of units along different boundaries. Openshaw demonstrated that by redrawing the boundaries of a fixed number of zones, correlations between two census variables could be driven across a very wide range. Electoral gerrymandering exploits the same property. The modifiable areal unit problem is closely related to the ecological fallacy, which is the error of inferring individual-level relationships from group-level data. The classic demonstration is W. S. Robinson's 1950 paper in the American Sociological Review, which showed that the correlation between being foreign-born and being literate in the United States had opposite signs at the level of states and at the level of individuals: states with more immigrants had higher literacy, though immigrants themselves were less often literate, because immigrants had settled in states with high literacy. There is no general solution. The honest responses are to analyse data at the finest resolution available, to state the unit of analysis as part of the finding rather than as an incidental detail, to repeat key analyses at more than one scale, and to avoid individual-level interpretation of area-level associations. In disease mapping, the hierarchical models of Chapter 7 reduce the instability that small units produce without abandoning them, which is one reason they have displaced crude rate maps. From points to surfaces Some analyses keep cases as points rather than aggregating them. Kernel density estimation converts a set of points into a smooth surface by placing a small bump, the kernel, over each point and summing. The result depends much more on the bandwidth, the width of the bump, than on the shape of the kernel. Too narrow a bandwidth yields a surface with a spike at every case; too wide a bandwidth smears everything into a single hill. Automatic bandwidth selectors exist, and they disagree with one another. A kernel density surface of cases is almost never what an epidemiologist wants on its own, because it mostly reflects where people live. The informative quantity is relative risk: the density of cases divided by the density of the population at risk. When individual controls are available, as in a case-control study, the ratio of a case density to a control density gives a spatially varying relative risk surface. Julia Kelsall and Peter Diggle set out this approach in 1995, and the sparr package in R implements it with adaptive bandwidths and tolerance contours that mark regions where the risk is significantly elevated. The point-pattern package spatstat, developed over many years by Adrian Baddeley, Rolf Turner and Ege Rubak, provides the broader toolkit for point processes, including Ripley's K function for testing whether points cluster more than expected under a given background intensity. Aggregation with denominators and uncertainty Almost every health rate has a numerator and a denominator from different sources, often at different spatial resolutions and different dates. Cases come from surveillance for the current year; the population denominator comes from a census five or ten years old, projected forward. If the projection is wrong for a particular district, perhaps because of migration after a conflict or displacement after a flood, that district's rate is wrong, and in a small district the error can be larger than any true variation in risk. Denominator uncertainty is rarely propagated into analysis, and it should be at least acknowledged. WorldPop and some other gridded population products now publish uncertainty surfaces alongside their estimates. Where denominators are known to be poor, a sensitivity analysis using alternative population sources, such as comparing census projections against a gridded product, reveals which findings depend on the choice. A worked construction: population within reach of care A common planning question shows how these operations fit together and where the choices lie. Suppose a ministry of health wants to know, for each district, what share of the population lives within one hour's travel of a facility offering emergency obstetric care. The inputs are a table of facilities with coordinates, a district boundary file, a gridded population surface, and a friction surface giving travel time per metre across each cell. The first step is cleaning the facility list: checking that coordinates fall inside the country, that none sit at zero latitude and zero longitude, and that facilities flagged as offering the service actually do. Facility registries are notoriously out of date, and in practice this step often takes longer than all the others combined. The second step is computing travel time. The friction surface is converted to a graph in which each cell connects to its eight neighbours, with edge costs equal to the time to cross between cell centres. A least-cost algorithm propagates outward from every facility at once and records, for each cell, the minimum time to reach any facility. In R, gdistance builds the transition matrix and computes accumulated cost; in Python, the minimum-cost-path routines in scikit-image do the same on a NumPy array. The result is a raster of travel time on the friction surface's grid. The third step is aligning the population raster to that grid. If the population grid is finer, it must be aggregated with a sum. If it is coarser, it must be disaggregated with care, since dividing a cell's population evenly among its sub-cells assumes uniformity. If the grids are offset, one must be resampled, again with a method that preserves totals. The fourth step is a zonal computation: for each district, sum the population in cells with travel time under sixty minutes and divide by total population in the district. The next chapter discusses how cells that straddle district boundaries should be treated, which for small districts can change the answer by several percentage points. Each step embeds a choice that could be defended differently. Walking speeds in the friction surface assume a mode of travel; a woman in labour may be travelling by motorcycle, by donkey cart or on foot. The facility list assumes that listed facilities are functioning. The population grid assumes that its dasymetric model placed people correctly. The one-hour threshold is conventional, not physiological. A well-reported analysis states each assumption, and a careful one shows how the district rankings change when the most uncertain of them are varied. In published work of this kind, the ranking of districts is usually more robust than the absolute percentages, and that is worth saying explicitly to the people who will use it. Privacy, confidentiality and the construction of data Linking health records to places raises obvious concerns about confidentiality. A map of individual case locations, even without names, can identify people in sparsely populated areas. Standard responses include aggregation to areas large enough that each contains a minimum number of people, random displacement as in the Demographic and Health Surveys, and more elaborate geomasking methods such as donut masking, which displaces each point by at least a minimum and at most a maximum distance so that the original location cannot be recovered by averaging. Every such method trades confidentiality for analytical precision. From the perspective of this booklet's argument, the point is that masked coordinates are a data construction with known properties. An analyst who knows the masking rule can model it; one who treats masked coordinates as exact will mis-estimate exposure. The same holds for every other construction in this chapter. Joins, buffers, overlays and aggregations each embed a rule, and the rule should be chosen, recorded and, where it matters, varied to show that the conclusion survives. Hashtags: #GeospatialAnalysisInRAndPython #GeospatialAnalysis #SpatialDataScience #RemoteSensing #GeographicInformationSystems #GIS #SpatialEpidemiology #SpatialStatistics #CoordinateReferenceSystems #MapProjections #VectorData #RasterData #SpatialJoins #RasterAnalysis #SpatialAutocorrelation #MoransI #SpatialRegression #BayesianSpatialModels #SatelliteImagery #LandSurfaceAnalysis #EcologicalNicheModeling #SpeciesDistributionModels #Maxent #GeospatialReproducibility #FutureOfGeospatialScience

  • Delphi Methodologies (Consensus Building Among Expert Panels in Policy and Health)

    Download the Book (PDF): Introduction Every year, thousands of published documents announce that a group of experts has agreed. A panel of clinicians agrees on which outcomes every trial in a disease should measure. A group of public health specialists agrees on what governments should do next in a pandemic. A committee of speech therapists, psychologists and paediatricians agrees on what to call children whose language does not develop as expected. Very often, the phrase that carries the weight is short: "consensus was reached using a modified Delphi process." Readers are invited to treat that phrase as a guarantee. They usually should not, at least not until they know what lies behind it. The Delphi method is a structured way of asking a panel of knowledgeable people the same questions several times, in writing, without letting them see who said what, and showing them between rounds how the group as a whole answered. It was devised at the RAND Corporation in the early 1950s to extract usable judgements from military and technical specialists about questions that no data could settle, and it has since spread into medicine, nursing, public health, education, environmental management, technology foresight and public policy. It is cheap, it can be run by email or on a web platform, it can bring together people on different continents who would never sit in one room, and it produces an output that looks like evidence: percentages, medians, lists of statements that crossed a threshold. That last property is both the method's great strength and its standing temptation. A Delphi study converts opinion into numbers. Those numbers can be honest summaries of a considered, independent judgement by the right people, or they can be the predictable result of a narrow panel, leading questions, feedback that nudged the waverers, a threshold chosen after the data came in, and the quiet departure of the people who disagreed. From the outside, the two can look identical. The published paper reports that 87 per cent of respondents agreed, and nothing on the page tells the reader which kind of 87 per cent it was. The argument of this book This book makes one argument and follows it through the life of a Delphi study. The consensus a Delphi study reports is not discovered; it is produced, by a chain of design decisions each of which could have been made differently. Who counts as an expert, how the questions are worded, which scale is used, what feedback panellists see, how many rounds are run, where the threshold sits, and how dropouts are handled all shape the final agreement at least as much as the underlying state of knowledge does. A Delphi result is therefore credible not because agreement was reached, but because the choices that produced the agreement were made in advance, for stated reasons, and disclosed in enough detail that a reader can tell agreement apart from conformity, attrition or design. This is not a counsel of despair. It is the reason the method can be done well. If consensus were simply out there waiting to be found, there would be little a researcher could do to improve a Delphi study beyond finding more experts. Because it is produced, every stage offers a lever, and the difference between a careful study and a careless one lies in how those levers are pulled. The chapters that follow are organised around those stages, and each asks the same two questions: what choices exist here, and how does each choice change the agreement that will eventually be reported? Why this matters now Delphi studies have never been more common or more consequential. Clinical guideline groups use them to fill the gaps where trials are absent. Core outcome sets, which determine what hundreds of future trials will measure, are routinely built on them. Reporting guidelines for research are developed through them, which gives the method a curious recursive authority: the rules that govern how evidence is reported are often themselves the product of expert consensus. During the COVID-19 pandemic, a Delphi study involving 386 experts from 112 countries and territories, published in Nature in November 2022, set out 41 consensus statements and 57 recommendations for ending the pandemic as a public health threat. Governments, professional bodies and funders now read Delphi outputs as a legitimate basis for action. At the same time, methodologists have spent two decades documenting how inconsistently the method is used. A systematic review by Ian Diamond and colleagues, published in the Journal of Clinical Epidemiology in 2014, examined 100 Delphi studies and found that only 72 of the 98 aiming at consensus defined what consensus meant; that the median percentage threshold, where one was given, was 75 per cent but ranged from 50 to 97; and that in 70 studies the process simply stopped after a fixed number of rounds rather than when consensus was reached. Reviews in health care quality indicators and in core outcome set development found similar variability. The response has been a wave of reporting guidance, including CREDES for palliative care in 2017, and, in 2024, both the ACCORD guideline for consensus methods in biomedicine and DELPHISTAR for Delphi studies across the social and health sciences. Reporting guidance helps readers see what was done. It does not tell researchers what to do, and it cannot by itself stop a well-reported study from being badly designed. That gap is where this book sits. What the book covers and what it leaves aside The first chapter sets out what the Delphi method is for, where it came from, and the family of variants that now shelter under the name, from the classical forecasting Delphi through the RAND/UCLA Appropriateness Method to the policy Delphi and the real-time online versions. It also argues, against a common habit, that Delphi is a method of last resort for questions that evidence cannot yet answer, not a shortcut around evidence that exists. The second chapter treats the protocol: the steering group, the evidence review that should precede the first questionnaire, and the decisions that must be written down before any panellist is contacted. The third addresses the panel, which is the single largest determinant of what a Delphi study will conclude: who qualifies as an expert, how different stakeholder groups should be represented, how large the panel needs to be, and how to recruit people who will stay. Chapters four and five follow the rounds. The fourth covers the first questionnaire, whether open or seeded, and the craft of writing items and choosing rating scales. The fifth covers what happens between rounds, where the method's distinctive mechanism, controlled feedback, does its work, and where it can do damage. Chapter six turns to the statistics of consensus: percentage agreement, medians and interquartile ranges, the RAND/UCLA definitions of disagreement, coefficients of variation, Kendall's W, and tests of stability between rounds. The chapter argues that no measure is correct in the abstract, and that the defensible choice is the one that matches the question and is fixed in advance. Chapter seven confronts the method's oldest criticism: that it manufactures agreement through social pressure, even without face-to-face contact. It examines groupthink, conformity, anchoring on feedback and selective attrition, and sets out the design choices that reduce each. Chapter eight takes the method into policy, where the goal is often not agreement at all but a clear map of disagreement, and where the legitimacy of the output depends on who was asked as much as on what they said. It ends with the reporting standards that allow outsiders to judge a finished study. The conclusion draws the argument together into a set of commitments for anyone who designs, commissions, or relies on a Delphi study. Some things are deliberately left out. The book does not survey every application domain, and it does not teach the statistics of forecasting accuracy in depth. It treats the nominal group technique and consensus conferences only where they illuminate a Delphi choice. It gives no software instructions, because platforms change faster than principles. Its examples are drawn mainly from health and public policy, because that is where the method is now most heavily used and most consequential, but the reasoning applies wherever a group of people with relevant knowledge must be asked to judge something the evidence cannot yet settle. A note on the reader The book is written for three kinds of reader. The first is the researcher about to run a Delphi study, often for the first time, often as part of a doctorate or a guideline project, and often with a steering group that has more enthusiasm than experience. The second is the commissioner: the guideline body, funder, professional society or ministry that will pay for a study and then act on its output. The third is the reader of Delphi results, including journal editors, reviewers, clinicians and officials, who needs to know which questions to ask of a paper before taking its conclusions seriously. All three share a need that the phrase "consensus was reached" does not meet. They need to know what kind of agreement they are looking at, how it was made, and how much weight it can bear. That is what the following chapters try to supply. Chapter 1. What Delphi Is For The Delphi method was born of a practical embarrassment. In the early years of the Cold War, the United States Air Force needed estimates about matters for which no data existed and none could be gathered: how an adversary would target American industry, how many weapons would be needed to achieve a given effect, when particular technologies would become feasible. The obvious source of such estimates was expert opinion. The obvious way to gather it was to put the experts in a room. And the obvious problem, familiar to anyone who has sat on a committee, was that the room distorted what the experts knew. The most senior or most voluble person set the terms of discussion. People defended positions they had stated publicly rather than revise them. Dissent was expensive. Agreement, when it came, often reflected the social dynamics of the meeting more than the balance of knowledge within it. Researchers at the RAND Corporation, most prominently Norman Dalkey and Olaf Helmer, set out to keep what was valuable about pooled judgement while removing what was corrosive about face-to-face deliberation. Their experiment, conducted for the Air Force in the early 1950s and published in the journal Management Science in 1963 under the title "An Experimental Application of the Delphi Method to the Use of Experts," asked a small panel to estimate, from the viewpoint of a Soviet strategic planner, the number of atomic bombs needed to reduce American munitions output by a specified amount. The experts never met. They answered questionnaires, received controlled information about the group's answers and about relevant factors raised by others, and answered again. Their estimates, initially spread across a very wide range, converged over the rounds. The name, a nod to the oracle at Delphi, was meant with some irony; the method's authors knew that nobody was consulting a god. Within a decade the technique had been turned to long-range forecasting of science and technology, most famously in the study reported by Theodore Gordon and Olaf Helmer in 1964, and had begun to spread into corporate planning, education and health. By 1975, when Harold Linstone and Murray Turoff edited the collection that became the field's standard reference, The Delphi Method: Techniques and Applications, the method had already diversified into forms its originators had not anticipated. Linstone and Turoff offered a definition broad enough to cover them all: Delphi is "a method for structuring a group communication process so that the process is effective in allowing a group of individuals, as a whole, to deal with a complex problem." The four defining features Beneath the variety, four features define the method, and each exists to solve a specific problem of face-to-face groups. The first is anonymity. Panellists do not know, at least during the rating process, who gave which answer. The point is not secrecy for its own sake but the removal of status cues. A junior clinician can disagree with a professor without the disagreement being noticed as such; a professor can change her mind without losing face. Anonymity in this sense is often called quasi-anonymity, because panellists may know who else is on the panel, and the organisers always know who said what. What matters is that responses are not attributed to individuals when fed back to the group. The second is iteration. The same questions, or refined versions of them, are put to the panel more than once. Iteration gives people the chance to reconsider in light of what others think and why, which is the mechanism by which Delphi hopes to improve on a single survey. A one-round expert survey is not a Delphi study, however carefully designed, because nobody has had a chance to respond to anyone else. The third is controlled feedback. Between rounds, the organisers tell panellists something about how the group answered: typically a measure of central tendency such as the median, some measure of spread, often the panellist's own previous answer for comparison, and sometimes the reasons others gave. Feedback is "controlled" because the organisers choose what to show, and those choices, as later chapters will argue, are among the most consequential in the entire design. The fourth is statistical aggregation of the group response. The output is not a negotiated form of words agreed by everyone in a room but a summary of individual ratings. Every panellist's judgement enters the final result, and the result can be expressed with a measure of how much the panel agreed. This is what allows a Delphi study to report that 82 per cent agreed with one statement and 54 per cent with another, and it is what gives Delphi outputs their appearance of quantitative rigour. Each of these features is a partial remedy, and each introduces its own difficulties. Anonymity removes status effects but also removes accountability; a panellist who need never defend a rating may give it less thought. Iteration allows reconsideration but also creates opportunities for fatigue and dropout. Controlled feedback provides information but also creates a reference point that can pull ratings towards it regardless of their merits. Statistical aggregation includes everyone but can conceal the difference between a panel that genuinely agrees and one whose members have simply stopped resisting. The rest of this book is, in effect, an extended account of how to use these features while guarding against their side effects. The family of variants Few Delphi studies today follow the procedure Dalkey and Helmer described. The name now covers a family of related designs that differ in their purpose, their first round, their use of meetings, and the kind of output they seek. The main members of the family differ on these points in ways worth setting side by side, as Table 1 does. Table 1. Main variants of the Delphi method and how they differ. Variant Primary purpose First round Face-to-face element Typical output Classical Delphi Forecasting or estimation Open questions None Converged estimates with spread Modified Delphi Agreement on statements or items Items drawn from evidence review Sometimes a final meeting List of items meeting a threshold RAND/UCLA Appropriateness Method Rating appropriateness of clinical indications Evidence summary plus rating Second-round panel meeting Indications classed as appropriate, uncertain or inappropriate Policy Delphi Exposing options and disagreement Open or seeded None or optional Map of positions and arguments Real-time or online Delphi Faster iteration at scale Seeded None Continuously updated group ratings The classical Delphi begins with open questions, uses the answers to build a structured questionnaire, and iterates until estimates stabilise. It remains common in technology foresight. The modified Delphi, now the dominant form in health research, replaces the open first round with a list of items derived from a systematic review, a qualitative study, or existing guidelines. The word "modified" is used so loosely that it tells a reader very little; it can mean a seeded first round, a final consensus meeting, a change of scale, or all three. A study that describes itself as a modified Delphi should always say what was modified and why. The RAND/UCLA Appropriateness Method, developed in the 1980s to study the overuse and underuse of medical procedures, is a particular and well-specified hybrid. Panels of seven to fifteen clinicians, with nine the traditional number, receive a review of the literature and rate hundreds of clinical scenarios on a nine-point scale of appropriateness. They then meet, discuss the areas of disagreement under a trained moderator, and rate again privately. Each scenario is classed as appropriate if the panel median falls between 7 and 9, uncertain between 4 and 6, and inappropriate between 1 and 3, with any scenario rated "with disagreement" treated as uncertain regardless of its median. The method's user's manual, published by RAND in 2001 under the lead authorship of Kathryn Fitch, remains one of the most precise descriptions of any consensus procedure. The policy Delphi, proposed by Murray Turoff in 1970, inverts the usual goal. Its purpose is not to reach agreement but to generate the strongest possible range of options and arguments on a contested policy question, and to identify where and why informed people disagree. Chapter eight returns to it in detail. Real-time Delphi, described by Theodore Gordon and Adam Pease in 2006, abandons discrete rounds. Panellists log in to a platform, see the current distribution of responses and the reasons given, enter or revise their own ratings, and can return as often as they like. The design trades the clean structure of rounds for speed and for a continuous exchange of reasons. The e-Delphi, now nearly universal, is simply a Delphi conducted through email or web survey software. It changes the logistics rather than the logic, although online administration makes very large panels feasible, and very large panels behave differently from small ones. When Delphi is the right tool The method's popularity has outrun its justification. Delphi exists to extract considered judgement where better sources of knowledge are unavailable. It is appropriate when three conditions hold together. The first condition is that the question cannot be answered adequately by existing evidence. If randomised trials have established that a treatment works, a panel of experts agreeing that it works adds nothing, and a panel disagreeing that it works should not override the trials. Delphi belongs in the gaps: questions of definition and terminology, choices about what to measure, judgements about appropriateness in populations the trials excluded, forecasts of future states, priorities among many plausible actions. The CATALISE study, which asked 59 experts from ten disciplines across six English-speaking countries to agree on how to identify children with language impairments, is a good example. The disagreement it addressed was not about data that could be collected; it was about categories, thresholds and words, which only the professional community could settle. The second condition is that relevant knowledge is distributed among people who cannot easily be brought together, or whose interaction would be distorted if they were. International panels, panels that mix patients with professionals, and panels on politically sensitive topics all benefit from written, anonymous exchange. The third condition is that the answer will be used for something that needs a defensible, transparent basis. A Delphi study is laborious. If a steering group could simply decide, and nobody would reasonably question its authority, a Delphi study may be ceremony rather than method. Delphi is the wrong tool when the question is empirical and answerable, when the aim is to legitimise a conclusion already reached, or when the relevant expertise is so narrow that the panel would consist of the steering group's collaborators. It is also the wrong tool when what is needed is deliberation in the full sense, a process in which people change their understanding through sustained argument. Written rounds allow the exchange of brief reasons; they are poor at the kind of back-and-forth that reveals a hidden assumption or produces a new synthesis. For some problems, a well-run meeting, or a Delphi combined with one, does better. Neighbouring methods Delphi is one of several formal consensus methods, and knowing its neighbours clarifies what it does and does not offer. The nominal group technique, described by André Delbecq and Andrew Van de Ven in the early 1970s and set out at length in their 1975 book with David Gustafson, Group Techniques for Program Planning, brings a small group of perhaps eight to twelve people into a room but constrains how they interact. Participants first write down ideas silently, then share them one at a time in round-robin fashion without debate, then discuss each idea for clarification, and finally vote or rank privately. The technique keeps face-to-face contact while blunting its worst effects, and it is fast: a single session can generate and prioritise a list. What it cannot do is include large or dispersed panels, and it cannot fully remove the influence of a forceful participant during the discussion phase. The consensus development conference, used for many years by the United States National Institutes of Health and adopted in various forms elsewhere, is a public event at which experts present evidence to a lay or mixed panel, which then deliberates in private and issues a statement. It resembles a jury more than a survey. Its strength is the depth of the evidence presentation and the visible, accountable nature of the panel's deliberation; its weakness is that it relies on a small number of panellists and on the dynamics of one closed meeting. Delphi differs from both in giving up deliberation almost entirely in exchange for scale, anonymity and a quantified output. That trade is often worth making, but it is a trade. A Delphi study cannot discover that its panel was using a key term in two different senses unless the free-text comments happen to reveal it. It cannot follow a promising line of argument in real time. Several modern designs therefore combine methods: a Delphi to establish where agreement and disagreement lie across a large panel, followed by a smaller meeting, sometimes run with nominal group rules, to resolve the items left in between. The RAND/UCLA method builds that combination into its core. When a study reports such a hybrid, the reader should ask which stage decided what, because a meeting at the end can overturn or quietly reframe what the Delphi rounds found. The evidence on whether it works It is reasonable to ask whether Delphi actually improves judgement. The best evidence comes from forecasting and estimation, where answers can later be checked. Gene Rowe and George Wright reviewed the experimental literature in the International Journal of Forecasting in 1999 and again in a 2001 chapter in Principles of Forecasting. Comparing Delphi with "staticised groups," in which individual judgements are simply averaged without feedback, they counted twelve studies favouring Delphi, two favouring staticised groups, and two ties. Comparing Delphi with unstructured interacting groups, they counted five studies favouring Delphi, one favouring the interacting group, and two ties. Their conclusion was cautiously positive: iteration with feedback tends to improve accuracy over a single round, and the structured process tends to beat free discussion. Two caveats matter. Many of those experiments used students answering almanac-style questions, which is some distance from senior clinicians judging the appropriateness of surgery. And accuracy is a meaningful criterion only when there is a true answer to be accurate about. Most health and policy Delphi studies ask questions about value, priority or definition, for which there is no later moment of truth. What the studies can be judged on is whether the process was fair, whether the panel was appropriate, and whether the reported consensus reflects considered judgement rather than artefact. These are the standards the rest of this book applies. Consensus as a product A final observation frames what follows. In the original RAND experiments, convergence of estimates was treated as a sign that the panel was approaching a better answer. That interpretation made some sense for estimation questions, where independent errors can cancel out and exchange of information can correct mistakes. It makes less sense for the kinds of questions modern Delphi studies ask. When a panel converges on the view that a particular outcome should be measured in every trial, the convergence may reflect shared recognition of the outcome's importance, or it may reflect the fact that the item was well worded, appeared early in the questionnaire, was supported by a high group median in the feedback, and was opposed mainly by panellists who dropped out after round one. The method does not itself distinguish these possibilities. Its designers must. That means treating every design decision as a decision about what kind of consensus the study is capable of producing, and making those decisions before the data can influence them. The next chapter begins at the point where that discipline starts: the protocol. Chapter 2. Deciding Before Asking: The Protocol A Delphi study is decided long before the first questionnaire goes out. By the time panellists are rating items, most of the determinants of the final result have already been fixed: the question being asked, the items on offer, the people doing the rating, the scale they will use, the feedback they will receive and the rule that will separate consensus from its absence. The protocol is where those determinants are chosen. It is also where they can be protected from the most insidious threat to a Delphi study's credibility, which is not bad faith but hindsight: the natural tendency, once results start to arrive, to adjust the rules until the results look sensible. This chapter treats the protocol as the study's first and most important act. It covers the steering group that designs the study, the definition of the question, the evidence review that should precede any rating, and the specific decisions that belong in writing before anyone is invited. The steering group and its power Every Delphi study has an organising group, whether it is called a steering committee, a management group, a working group or simply the research team. The label matters less than the recognition that this group holds more power over the outcome than any panellist. It decides what is asked and how. It interprets free-text comments and decides whether they justify new items. It chooses what feedback to show. It often has the final word in a closing meeting where items near the threshold are resolved. In many published studies, the steering group also writes the final statements, which means its members translate the panel's ratings into the form of words that will be cited. The composition of the steering group should therefore be treated with the same seriousness as the composition of the panel. A group drawn entirely from one institution, one discipline or one school of thought will tend, without any intention to mislead, to frame items in ways that reflect its own assumptions. A steering group for a core outcome set in a surgical condition that includes only surgeons is likely to produce an initial item list heavy with technical and complication-related outcomes and light on the outcomes patients care about, such as return to work or bodily appearance. Including a methodologist, at least one person from each major stakeholder group that will sit on the panel, and, where the topic concerns patients or the public, people with lived experience of the condition in the steering group itself, is the most effective single protection against framing bias. The steering group's conflicts of interest also matter. A Delphi study on the appropriate use of a device, organised by people who hold patents on it or receive consultancy payments from its manufacturer, may be methodologically impeccable and still not be trusted. The ACCORD reporting guideline for consensus methods, published in PLOS Medicine in January 2024, asks studies to report the roles of those who organised the process and their funding and conflicts of interest, precisely because the organisers' influence is so large. The best protection is to declare interests in the protocol and to design the process so that no single interested party controls a decision that matters. Defining the question A surprising number of Delphi studies are vague about what they are trying to decide. "To reach consensus on best practice in the management of X" is an aspiration, not a question. Before designing anything, the steering group needs to be able to state precisely what the panel is being asked to judge, for what purpose, and for which population or setting. A well-defined Delphi question has four elements. It specifies the object of judgement: outcomes to be measured, recommendations for practice, definitions of a condition, indicators of quality, priorities for research, forecasts of a future state. It specifies the criterion of judgement: importance, appropriateness, feasibility, desirability, likelihood, clarity. It specifies the scope: which population, which setting, which country or health system, which time horizon. And it specifies the use: what will happen to the output, and who will act on it. The criterion is where many studies go astray. Asking panellists whether a quality indicator is "important" conflates at least three distinct judgements: whether the thing it measures matters, whether the indicator validly measures it, and whether it is feasible to collect. A panellist might believe that timely pain relief in the emergency department matters enormously while doubting that "time from triage to first analgesic dose" measures it well and knowing that most departments cannot extract the data. A single rating of importance forces that panellist to average three different judgements in an unknown way. The RAND/UCLA tradition, which asks separately about validity and feasibility in some applications, is instructive here. Separate criteria mean separate ratings, which lengthens the questionnaire, but they produce results that can be interpreted. The scope needs equal care. An international panel asked about the appropriateness of an intervention without being told which health system to imagine will answer from its members' own contexts, and the resulting disagreement may reflect differences in resources rather than differences in judgement. If the question is meant to apply to high-income settings, say so. If it is meant to apply everywhere, consider asking panellists to rate for their own setting and analysing by region. The evidence comes first A Delphi study that asks experts to rate items without first showing them what is known is asking them to rely on memory, and memory is selective. Panellists will recall the studies they read, the cases they saw, and the conference presentations that impressed them. Different panellists will recall different things, and some of the resulting disagreement will be disagreement about facts that a literature review could have settled. For this reason most rigorous modified Delphi designs begin with a structured evidence review, and many distribute a summary of it to panellists. The RAND/UCLA method makes this explicit: panellists receive a literature review, often a substantial one, before they rate a single indication. In core outcome set development, the COMET Initiative's handbook, published by Paula Williamson and colleagues in Trials in 2017, recommends identifying candidate outcomes from a systematic review of the outcomes measured in existing trials, supplemented where possible by qualitative research with patients to capture outcomes that trials have ignored. The evidence review does two jobs. It generates candidate items, so that the first questionnaire reflects the full range of what has been proposed rather than what the steering group happened to think of. And it gives panellists a shared factual base, so that their ratings reflect judgement about what the evidence means rather than disagreement about what it says. The distinction matters because a Delphi study cannot resolve factual disputes; it can only record them. If one panellist believes a treatment reduces mortality and another believes it does not, iteration will at best produce a compromise between two readings of the literature, which is worse than either reading checked against the literature itself. The evidence summary is itself a source of influence. How it is written, what it emphasises and what it omits all shape the ratings that follow. A summary that describes one intervention as "well established" and another as "promising" has done some of the panel's work for it. The discipline is to present evidence neutrally, in a consistent format for each item, with the strength of evidence stated using an explicit system where one is available, and to have the summary checked by someone outside the steering group. What must be decided in advance The heart of the protocol is a list of decisions that must be made and written down before the data can influence them. Several reporting guidelines, including CREDES, published by Saskia Jünger and colleagues in Palliative Medicine in 2017, ACCORD, and DELPHISTAR, published by Marlen Niederberger and colleagues in PLOS ONE in August 2024, ask studies to report these decisions. The protocol is where they should first appear. Panel eligibility and structure. Who qualifies as an expert, by what criteria, and in which stakeholder groups. Whether groups will be analysed separately. The target size of each group and the minimum number of respondents required for a round to count. Recruitment. How potential panellists will be identified and approached, and what they will be told about the time commitment. The first round. Whether it is open, seeded, or both; where seeded items came from; whether panellists can propose new items; and the rule for adding them. Items and scale. The exact wording of the instructions and the rating scale, including the labels on each point and whether there is an "unable to judge" option. Feedback. What panellists will see between rounds: which statistics, whether their own previous rating, whether comments from others, and whether results will be split by stakeholder group. The consensus definition. The statistical rule for declaring consensus that an item should be included, consensus that it should be excluded, and the absence of consensus. The minimum number of respondents to whom the rule applies. Stopping. The maximum number of rounds, and whether the study stops when consensus is reached, when responses stabilise, or after a fixed number of rounds regardless. Item carry-forward. Which items return in later rounds. Some designs re-present every item; others drop items that have already reached consensus. The choice changes the dynamics of later rounds. Attrition. Whether panellists who miss a round are invited to the next; how non-response will be analysed; what response rate will be considered adequate. Any final meeting. Who attends, what it may decide, and how its decisions will be recorded and reported alongside the Delphi results. Analysis and reporting. How free-text comments will be analysed, and how items that did not reach consensus will be reported. Diamond and colleagues' review found that consensus was specified in advance, with a threshold, in only 42 of the studies that reported achieving it. Where the rule is chosen afterwards, the reader cannot tell whether a 70 per cent threshold was selected because it was principled or because it allowed a particular item through. The protocol removes that doubt, and it is only as good as the record of it. Publishing the protocol in a journal, registering it on a public platform such as the Open Science Framework, or, for core outcome sets, registering the study in the COMET database, creates a timestamped record that the study can later be checked against. A protocol in miniature It helps to see how these decisions fit together in a single, deliberately simple case. The following is a hypothetical design, constructed for illustration rather than drawn from any published study. Suppose a national nursing body wants to agree on a set of indicators for the quality of pressure-injury prevention in residential care homes. The steering group comprises two tissue-viability nurses, a care-home manager, a geriatrician, a health services researcher experienced in consensus methods, and two relatives of care-home residents. Its first act is to narrow the question. It is not asking which prevention practices work; that belongs to trials and systematic reviews. It is asking which measurable aspects of practice a regulator should monitor, judged on two separate criteria: whether the indicator reflects something that matters for residents, and whether a typical home could collect it without new systems. A scoping review of existing indicator sets and inspection frameworks produces an initial list of thirty-four candidate indicators, each with a one-paragraph evidence note written to a common template. The panel will comprise three groups: nurses working in care homes, specialists in tissue viability and wound care, and residents' relatives together with advocacy representatives. The protocol sets a target of at least twenty respondents per group and states that any group falling below twelve in a round will be reported descriptively but excluded from the consensus rule. The protocol fixes a nine-point scale for each criterion, labelled at the ends and middle, with an "unable to rate" option. It defines consensus for inclusion as at least 70 per cent of respondents in every group rating an indicator 7 to 9 on both criteria, and no more than 15 per cent rating it 1 to 3 on either. It defines consensus for exclusion symmetrically. It sets a maximum of three rounds, with every item not yet at consensus carried forward, and states that the process stops early if no item changes category between two consecutive rounds. Feedback will show the distribution of ratings for each group separately, the panellist's own previous rating, and a summary of reasons given, grouped by theme and stripped of identifying detail. Panellists who miss round two will still be invited to round three. A final online meeting, open to a sample of panellists from each group, may discuss items left without consensus but may not add or remove items that met either rule. None of these choices is uniquely correct. A different steering group might use a five-point scale, a higher threshold, or a single combined criterion. The point is that each choice has been made for a reason that can be stated, before any rating has been seen. When the results arrive, nobody will be tempted to wonder whether the threshold was chosen to include a favoured indicator, because the threshold was public before the indicator was rated. Allowing for the unforeseen Pre-specification is not rigidity. Delphi studies encounter surprises: an item that turns out to be ambiguous, a stakeholder group that responds far less than expected, a round-two comment that reveals a crucial missing option. A protocol can anticipate some of these by stating in advance how they will be handled. It might say, for example, that items flagged as unclear by more than a stated share of respondents will be reworded and re-presented, with the rewording reported; or that if a stakeholder group falls below a minimum number, its results will be reported descriptively but not used to determine consensus. When something happens that the protocol did not foresee, the honest course is to make the change, record why, and report it as a deviation. Deviations are normal. What damages credibility is the undisclosed deviation, the threshold quietly relaxed from 80 to 70 per cent in the final round, or the stakeholder group whose disagreement was reported only in a supplementary file. A reader who can see the deviation can judge it. A reader who cannot see it has been misled, even if the change was defensible. Ethics and consent Delphi panellists are research participants, and a protocol must address their interests. They need to know what they are agreeing to: how many rounds, how long each will take, how their responses will be used, whether they will be named in the publication, and whether they can withdraw. Many studies acknowledge panellists by name in the final paper, which rewards participation but has consequences. If panellists know they will be listed as contributors to a consensus statement, some may feel bound to its conclusions and reluctant to express dissent that might later be attributed to them, even though individual ratings remain anonymous. Offering panellists the choice to be acknowledged or not, and being clear that acknowledgement does not imply endorsement of every statement, helps. Patient and public panellists raise particular considerations. They may need plain-language versions of items, support with technical terms, and payment for their time on terms comparable to those offered to professionals. They may be rating items that touch on their own experience of illness. The protocol should state how these needs will be met. Failing to meet them is not just an ethical lapse; it is a methodological one, because a patient group that cannot understand the items or loses heart after round one will be underrepresented in the final result, and the consensus will tilt towards the professionals by default. The protocol as a public commitment It is tempting to think of the protocol as paperwork required by ethics committees and journals. That view misses its function. A Delphi study produces a consensus whose authority rests on a claim that the process was fair. The protocol is the evidence for that claim. It shows that the rules were set before anyone knew which items would pass them, that the panel was chosen by criteria rather than by acquaintance, and that the feedback and stopping rule were not adjusted to steer the outcome. A study that can point to a protocol written, registered and followed, with deviations disclosed, can say something that a study without one cannot: that its consensus was not engineered after the fact. Every other chapter in this book depends on that foundation. The most consequential decision the protocol records is who will be asked, and that is where the next chapter turns. Chapter 3. Who Counts as an Expert If a Delphi study's result could be predicted from a single piece of information, that information would be the composition of the panel. Items, scales and thresholds shape the edges of the outcome; the panel shapes its centre. A panel of surgeons and a panel of patients, asked the same questions about the same operation, will reach different conclusions, and both will be reported as consensus. A panel recruited through one professional society will tend to reproduce that society's orthodoxy. A panel that is ninety per cent from high-income countries will produce recommendations that assume high-income resources, however "international" the title claims it to be. This is not a flaw that better statistics can remove. It follows from what Delphi is: a structured way of aggregating the judgements of particular people. The judgements belong to those people, and the consensus is theirs. The question for the designer is therefore not how to find the objectively correct panel, since there is none, but how to choose a panel whose judgement the intended users of the result will have good reason to trust, and how to describe it so that readers can see whose consensus they are looking at. What expertise means in a Delphi study The word "expert" suggests a person with specialist knowledge, usually certified by qualifications, experience, and publication. That meaning is appropriate for some Delphi questions and misleading for others. For estimation and forecasting, the relevant expertise is knowledge that bears on the answer. A panel forecasting the adoption of a medical technology needs people who understand the technology, the regulatory pathway, health system purchasing and clinical practice. Formal credentials are a reasonable proxy, although the research of Philip Tetlock, summarised in his 2005 book Expert Political Judgment, is a standing reminder that credentialed experts are often poorly calibrated forecasters, and that breadth of perspective can matter more than depth of specialism for predicting complex events. For questions of value and priority, the relevant expertise is different. Deciding which outcomes matter most in a chronic disease is not a matter on which professors know more than patients. Patients have knowledge of what living with the condition is like, what they would trade for what, and which symptoms dominate daily life, that no amount of clinical training supplies. The core outcome set movement has taken this seriously for more than a decade, and the inclusion of patients as a distinct panel group is now standard in well-conducted core outcome set studies. The same logic applies to carers, to service users in social care, to affected communities in environmental policy, and to front-line staff whose knowledge of what is feasible is practical rather than academic. For questions of definition and terminology, expertise includes knowledge of how terms will be used. The CATALISE panel on children's language impairments deliberately spanned education, psychology, speech and language therapy, paediatrics, child psychiatry and other disciplines, because the terminology it proposed would have to work across all of those professions and in the services where they meet. A Delphi protocol should therefore define expertise in relation to the question, not in the abstract. For each stakeholder group, it should state what knowledge or experience qualifies a person for inclusion and how that will be checked. Typical criteria for professionals include a minimum number of years in relevant practice, current involvement in the relevant field, publications or leadership roles, or nomination by a professional body. Typical criteria for patients and carers include direct experience of the condition or service, sometimes within a stated time frame. The criteria should be specific enough that a reader could, in principle, apply them. Heterogeneity and its consequences Designers face a basic choice between homogeneous and heterogeneous panels. A homogeneous panel, drawn from one profession or perspective, will usually reach consensus faster and on more items, because its members share assumptions. A heterogeneous panel, spanning professions, sectors and lived experience, will agree on less, take longer, and produce a result that is more robust precisely because it has survived challenge from different angles. Rowe and Wright's review of forecasting evidence recommended heterogeneous panels, and the logic extends beyond forecasting. The value of pooling judgements comes from the partial independence of the errors in them. A panel of people who all trained in the same place, read the same journals and attend the same meetings shares its blind spots, and pooling their views multiplies confidence without adding information. A panel of people with different vantage points brings errors that partly cancel and knowledge that partly complements. Heterogeneity has costs. Items must be written so that all groups can understand them. Feedback must be designed so that the majority group does not swamp the minority. And the analysis must decide what consensus means when groups disagree. If 85 per cent of clinicians rate an outcome as critical and 40 per cent of patients do, has the panel reached consensus? Pooling the groups into a single percentage would hide the disagreement and let the result depend on how many of each were recruited. The better practice, now common in core outcome set work, is to analyse each group separately and to require consensus within every group, or at least to report group-level results alongside any pooled figure. Chapter six returns to the statistics. How big should a panel be? There is no statistical formula for Delphi panel size, because a Delphi panel is not a sample from which inferences are drawn about a population. It is the population whose judgement is being reported. The question of size is therefore a practical and substantive one, not a matter of power calculations. Published guidance and practice span a wide range. The RAND/UCLA method uses panels of seven to fifteen, with nine as its traditional number, reflecting its reliance on a face-to-face meeting in which everyone must be able to participate. Rowe and Wright suggested between five and twenty experts for forecasting panels. Core outcome set studies commonly recruit hundreds of panellists across several groups, and the COVID-19 study published in Nature in 2022 involved 386. The ACCORD guideline's own Delphi recruited 72. Several considerations bear on the choice. The first is the number of distinct perspectives that need representation. A panel with four stakeholder groups needs enough members in each to produce a stable percentage; with only six patients, one person changing their rating moves the patient agreement figure by more than sixteen percentage points, and a threshold of 70 per cent becomes a question of whether four or five people agree. The second is expected attrition, discussed below; a panel that begins at the minimum size will end below it. The third is the burden of the questionnaire; a very long item list with a very large panel produces an enormous analytic task and may reduce the quality of each rating. The fourth is credibility: a consensus statement meant to guide an international field may need to show broad participation to be taken seriously, even if a smaller panel would produce similar results. A useful rule is to decide, for each stakeholder group, the smallest number of respondents at which a percentage agreement would be meaningful, commonly around fifteen to twenty, then inflate the recruitment target to allow for the attrition the study expects. ACCORD's own design required agreement from at least 80 per cent of a minimum of twenty respondents, a helpful example of making the floor explicit. Finding and recruiting panellists How panellists are found determines who is found. Invitations sent through a single professional society reach its members, who are not all practitioners in the field and who may share its official positions. Invitations to authors of relevant papers reach researchers rather than practitioners. Snowball recruitment, in which initial panellists nominate others, reaches their networks, which may be homogeneous. Open calls through social media reach people who follow the organisers and are motivated enough to respond. A structured approach helps. Chitu Okoli and Suzanne Pawlowski, writing in Information & Management in 2004, described a "knowledge resource nomination worksheet" for Delphi recruitment. The designers first list the disciplines, organisations and types of knowledge the panel needs, then identify individuals in each category through literature, organisations and nominations, then rank and invite them category by category until each is filled. The value of such a procedure lies less in its details than in its effect: it forces the designer to decide what the panel should look like before seeing who agrees to join, and it makes gaps visible. Geographic and demographic balance deserve explicit attention. International panels in health research are routinely dominated by North America, Western Europe and Australasia, partly because recruitment runs through English-language journals and networks. If the result is meant to apply globally, recruitment should target low- and middle-income settings deliberately, and the final panel composition should be reported by region so that readers can judge how global the consensus is. The same applies to gender, career stage, practice setting and, for patient groups, to age, ethnicity and severity of condition. Conflicts of interest among panellists should be collected and reported. The aim is not usually to exclude people with interests, since in many fields those with the deepest knowledge also have commercial or professional stakes, but to allow readers to judge whether the panel as a whole was balanced. A panel on the appropriate use of a drug in which a majority receive funding from its manufacturer is not neutral, however anonymous its ratings. Attrition and why it matters so much Every Delphi study loses panellists between rounds. People are busy, questionnaires are long, and the second and third rounds offer less novelty than the first. Response rates of 70 to 80 per cent per round are often considered acceptable, and studies commonly report cumulative retention well below that. The ACCORD Delphi retained 58 of 72 panellists in its first round, 54 in its second and 51 in its third. DELPHISTAR's own Delphi study had 91 completers in its first round, 69 in its second and 56 in its third. Attrition would be merely an inconvenience if it were random. It rarely is. Those who drop out may differ systematically from those who stay: they may be less engaged with the topic, more pressed for time, less convinced of the study's value, or, most importantly for a consensus method, in disagreement with the emerging view. A panellist who sees in round two that her ratings are far from the group median has several options. She can hold her position, which requires explaining it; she can move toward the group; or she can stop responding. The third option is the least effortful. If dissenters leave at a higher rate than others, the remaining panel will show more agreement than the original one did, and the study will report rising consensus that is in fact rising homogeneity. The effect is hard to detect after the fact unless the study was designed to look for it. The protocol can require that the round-one ratings of those who later dropped out be compared with those of completers; if dropouts were systematically more likely to have rated against the eventual consensus, the reader should know. Studies can also report consensus figures using the original panel as denominator as well as the responding panel, which shows how much of the apparent agreement depends on who stayed. Retention can be improved. Clear expectations at recruitment, short questionnaires, prompt rounds with reasonable deadlines, personalised reminders, and feedback that shows panellists their contribution was read all help. So does inviting non-responders to later rounds rather than excluding them, which keeps the door open for those who missed a deadline without losing interest. Some studies offer modest payments or certificates of participation; for patient panellists, payment for time is increasingly regarded as a matter of fairness rather than incentive. Weighting, self-rated expertise and the organisers' own votes Once a heterogeneous panel is assembled, a question follows that many studies never ask aloud: should every panellist's rating count equally? Pooling ratings into a single percentage answers it implicitly, and the answer depends on recruitment rather than principle. If a study recruits 150 clinicians and 30 patients and pools them, clinicians carry five times the weight of patients, not because anyone decided they should but because they were easier to find. There are three defensible responses. The first is to analyse groups separately and require consensus in each, which gives each group a veto without weighting individuals. The second is to weight groups equally in a pooled figure, so that the patient percentage and the clinician percentage each count for half regardless of numbers. The third is to fix group sizes in advance and recruit to them, accepting pooled figures because the balance was designed. Each is reasonable. What is not reasonable is to let the ratio emerge from recruitment and then report a pooled percentage as if it were neutral. A related idea, explored since the early RAND work, is to weight panellists by their own assessment of their expertise on each item, or to exclude ratings from those who declare themselves unfamiliar with an item. The attraction is obvious: a cardiologist asked about a rehabilitation outcome may know less than a physiotherapist, and a self-rating allows the method to recognise that. The difficulty is that self-assessed expertise is only loosely related to actual knowledge, and confidence varies by personality, gender and professional culture as well as by competence. A more modest and more common practice is to offer an "unable to rate" or "outside my expertise" option and to exclude those responses from the denominator for that item, reporting how many were excluded. That lets panellists abstain honestly without inviting them to rank themselves. The steering group's own participation is a final issue. In many studies, members of the organising group also sit on the panel. There are arguments for this: they are often among the most knowledgeable people available, and excluding them may weaken the panel. But they have designed the items, they know the purpose of the study, and they may have views about what the result should be. At a minimum, a report should state whether organisers rated items and how many of them did. Where the panel is small, a steering group that makes up a fifth of it can move an item across a threshold on its own. Some studies exclude organisers from rating altogether, or analyse the results with and without them, which is the clearest way to show that the consensus did not depend on them. Reporting the panel Because the panel so largely determines the result, the description of the panel is among the most important parts of any Delphi report. Readers should be able to see how many were invited and how many responded in each round, broken down by stakeholder group; the qualifying criteria and how they were checked; the geographic distribution; the relevant demographic and professional characteristics; conflicts of interest; and the characteristics of those who dropped out compared with those who stayed. A reader armed with that information can make the judgement that the phrase "an international panel of experts" does not permit. She can ask whether these are the people whose agreement should guide her decision, and whether what they agreed was likely to have been shaped by who was in the room and who left it. That judgement cannot be delegated to the method. The next chapter turns to the other side of the encounter: what the panel is actually asked. Hashtags: #DelphiMethodologies #DelphiMethod #ExpertConsensus #ConsensusBuilding #ExpertPanels #PolicyDelphi #HealthDelphi #ModifiedDelphi #ClassicalDelphi #RANDUCLAAppropriatenessMethod #ExpertJudgement #ConsensusMethods #ControlledFeedback #IterativeRounds #PanelAnonymity #ConsensusThresholds #StakeholderRepresentation #ExpertSelection #PanelAttrition #ConsensusMeasurement #KendallsW #CoreOutcomeSets #ConsensusGuidelines #PolicyConsensus #FutureOfDelphiResearch

  • Structural Equation Modeling (Path Analysis, Latent Constructs, and Fit Indices)

    Download the Book (PDF): Introduction Every structural equation model is an argument about where covariance comes from. That sentence is the whole of this book compressed, and most of the trouble researchers get into with the method comes from forgetting it. Consider what actually happens when a researcher fits a model in AMOS or lavaan. She has a data set: perhaps four hundred teachers who answered twenty questionnaire items about workload, emotional exhaustion, collegial support, and intention to leave the profession. From those responses the software computes a covariance matrix, a square grid of numbers describing how each item moves with every other item. Twenty items produce 210 unique variances and covariances. That matrix is the evidence. Everything else, the latent constructs, the arrows, the standardized coefficients printed to three decimal places, the indirect effect with its bootstrapped confidence interval, is a hypothesis about how that matrix came to be. The researcher's diagram says, in effect: these four items share variance because they all reflect a common cause called exhaustion; these five share variance because they reflect support; exhaustion depends on workload and is buffered by support; intention to leave depends on exhaustion. Once the diagram is drawn and the parameters estimated, the model implies a covariance matrix of its own, the pattern of associations that would appear if the story were exactly true. Estimation chooses parameter values that make the implied matrix as close as possible to the observed one. Fit indices summarize how close it came. Modification indices point to where it fell short. Seen this way, several things that confuse newcomers become clearer. A model can fit well and still be wrong, because many different stories can generate the same covariance matrix. A model can fit badly for reasons that have nothing to do with the structural hypotheses the researcher cares about, because most of the matrix's cells concern relationships among indicators of the same construct. A significant indirect effect does not show that mediation occurred, because covariances carry no information about temporal order. And a fit index crossing a published threshold does not certify anything, because the threshold was derived under conditions that may not resemble the researcher's model at all. Why the method is worth the effort None of this is an argument against structural equation modeling. It is an argument for using it with a clear head, because what the method offers is real and not easily obtained any other way. Its first gift is the separation of measurement from theory. Psychological, educational, and organizational research deals almost entirely in constructs that cannot be observed directly: self-efficacy, reading comprehension, transformational leadership, organizational commitment, anxiety. We observe responses to items, scores on tests, ratings by supervisors, and we treat them as imperfect windows onto the constructs. Ordinary regression on summed scale scores pretends the windows are clear. Measurement error in predictors then biases coefficients, usually toward zero for a single predictor and unpredictably when several correlated predictors are involved. A latent variable model estimates the error explicitly and relates the constructs to one another as if they had been measured without it. The gain is not cosmetic. Two constructs that correlate at .45 as observed scale scores can correlate at .60 or more once unreliability is removed, and a regression coefficient that looked trivial can look substantial. Its second gift is the ability to test a whole theory at once. A regression answers one question about one outcome. A structural model lets a researcher write down a system: workload affects exhaustion, exhaustion affects commitment and turnover intention, support moderates the first path. The system has testable consequences beyond its individual coefficients, because the paths it omits are claims too. A model that says workload affects turnover intention only through exhaustion predicts a particular pattern among the three constructs, and the data can contradict it. Its third gift is flexibility. The same framework handles confirmatory factor analysis, path analysis, mediation, multiple-group comparison, measurement invariance, latent growth curves, and latent interactions. Researchers who learn the logic once can apply it to a remarkable range of designs. Why the method goes wrong The same flexibility makes structural equation modeling easy to misuse. It is possible, with a few clicks in AMOS or a few lines of lavaan syntax, to produce output that looks authoritative and means very little. The recurring failures are well documented in the methodological literature and in any reviewer's memory. Researchers treat fit indices as pass marks, reporting that a model "fit well" because its CFI exceeded .95 and its RMSEA fell below .06, without asking whether those cutoffs were designed for models like theirs. They add correlated errors suggested by modification indices until the numbers improve, then report the final model as if it had been hypothesized in advance. They estimate mediation models on cross-sectional survey data and describe the resulting indirect effects in causal language. They ignore the fact that alternative models, with arrows pointing the other way, would fit exactly as well. They bundle measurement problems and structural problems into a single fit statistic and cannot tell which kind of misfit they have. These are not failures of software. AMOS and lavaan both compute what they are asked to compute, correctly. They are failures of understanding what the numbers are answers to. What this book does This book is built around a single claim: that a structural equation model is a set of explicit hypotheses about covariance, and that every practical decision in the method, from how constructs are specified to how fit is judged to how misfit is repaired, becomes clearer once the researcher keeps that claim in view. The chapters follow the order in which a careful analysis actually proceeds. The first chapter develops the core logic: how a diagram implies a covariance matrix, what identification means, and why degrees of freedom are the currency of testability. The second treats path analysis with observed variables, the historical root of the method, and shows both what it can do and why measurement error limits it. The third and fourth chapters turn to latent constructs: first the conceptual work of specifying a measurement model and judging its reliability and validity, then the practical work of confirmatory factor analysis, including estimation choices for non-normal and ordinal data, missing data, and measurement invariance. The fifth and sixth chapters address the structural questions that draw most researchers to the method in the first place: mediation, moderation, and their combination. They take seriously both the statistical machinery and the causal assumptions the machinery cannot supply. The seventh chapter explains what each of the common fit indices measures, where the famous cutoffs came from, and why methodologists have spent a quarter century warning against using them mechanically. The eighth chapter is about diagnosis: how to locate the source of misfit using residuals and modification indices, how to respecify without deceiving oneself, and how to report what was done so that readers can judge it. Throughout, the book shows how the ideas map onto the two programs most widely used in the target fields. AMOS, distributed by IBM alongside SPSS, is built around drawing path diagrams and remains common in management, marketing, and education departments. lavaan, an open-source package for R created by Yves Rosseel and described in the Journal of Statistical Software in 2012, is built around a compact model syntax and has become the default in much of psychology. Neither is presented as superior. They estimate the same models by the same methods, and a researcher who understands the model can move between them. Where they differ in defaults or capabilities, the difference is noted, because defaults quietly shape results. The reader assumed here knows ordinary regression and the idea of a correlation, has perhaps run a factor analysis, and wants to use structural equation modeling responsibly in a thesis, a grant, or a journal article. No matrix algebra is required, though a few simple formulas appear where they illuminate what a statistic is doing. The goal is not to replace the standard textbooks, several of which are listed at the end, but to give a reader the judgment that makes those textbooks usable: a sense of what questions the method can answer, what it cannot, and how to tell the difference when the output is on the screen. Chapter 1. Covariance as Evidence The single most useful habit a structural equation modeler can develop is to look at a model diagram and ask: what pattern of associations does this picture predict? Everything in the method follows from the answer. Estimation is the search for parameter values that make the prediction match the data. Testing is the comparison of the prediction with the data. Identification is the question of whether the data contain enough information to pin the parameters down. This chapter builds that habit from the ground up. From arrows to associations Begin with the smallest interesting case. Suppose a researcher believes that parental involvement (X) raises a student's academic self-concept (M), which in turn raises end-of-year achievement (Y). She measures all three directly, say as standardized scale scores, and draws a chain: X points to M, M points to Y. There is no arrow from X to Y. In standardized form, the model has two path coefficients, call them a for X to M and b for M to Y. The rules for reading associations off such a diagram were worked out by the geneticist Sewall Wright in a series of papers beginning around 1918 and set out in his 1921 paper "Correlation and Causation." Wright's tracing rules say that the correlation between any two variables equals the sum, over every permissible route connecting them, of the products of the coefficients along each route. A permissible route may go backward along arrows and then forward, but not forward and then backward, and it may pass through each variable only once. Apply the rules. The correlation between X and M is a, since there is one route. The correlation between M and Y is b. The correlation between X and Y is the product a × b, since the only route runs from X through M to Y. So the model predicts three correlations from two parameters, and the third is not free: it must equal the product of the other two. That is a testable claim. If the observed correlation of X with M is .40 and of M with Y is .50, the model insists that X and Y correlate at .20. If they actually correlate at .35, something is wrong with the story. Perhaps parental involvement also affects achievement directly, through homework help or school choice, in addition to its effect through self-concept. The missing arrow was a hypothesis, and the data can reject it. Now add the direct path c′ from X to Y. The model has three parameters and three correlations to explain. It will reproduce them exactly, whatever the data are. The model is no longer wrong in any detectable way, which also means it is no longer testable. It has become a reparameterization of the correlation matrix, a way of redescribing the data rather than making a claim about them. This small example contains the three ideas that organize the rest of the method. A model implies a pattern of associations. The pattern is testable only to the extent that the model has fewer free parameters than there are associations to explain. And the omitted arrows, not the drawn ones, are where the testable content lives. The model-implied covariance matrix Real models are larger, and they work with covariances rather than correlations, so that variables retain their original scales. But the logic is identical. Given a model and a set of values for its parameters, one can compute the covariance matrix the model implies. Methodologists write this as Σ(θ), the covariance matrix Σ expressed as a function of the parameter vector θ. The observed sample covariance matrix is S. Estimation finds the values θ̂ that make Σ(θ̂) as close as possible to S, according to some definition of "close." The most common definition is maximum likelihood. Under the assumption that the observed variables follow a multivariate normal distribution, maximum likelihood chooses the parameter values under which the observed data would have been most probable. In practice, the software minimizes a discrepancy function, usually written F, which is zero when the implied and observed matrices are identical and grows as they diverge. Both AMOS and lavaan use maximum likelihood by default for continuous data. Other estimators, discussed in Chapter 4, use different definitions of closeness that are better suited to non-normal or ordinal data. What matters conceptually is that the observed covariance matrix, along with the means if the model includes them, is the complete evidence base. With complete data and a normal-theory estimator, two data sets with identical covariance matrices and means will produce identical parameter estimates, identical standard errors, and identical fit statistics, regardless of whether they came from a randomized experiment, a longitudinal panel, or a single survey administered on one afternoon. The method does not know how the data were collected. Any causal interpretation must come from the design and from the researcher's substantive knowledge, not from the covariances themselves. This point is easy to agree with and hard to remember. It will return in every later chapter. Counting information: degrees of freedom With p observed variables, the covariance matrix contains p(p + 1)/2 unique elements: p variances on the diagonal and p(p − 1)/2 covariances off it. Five variables give 15 elements; ten give 55; twenty give 210. If the model also estimates means, add p more. The model's degrees of freedom are the number of unique elements minus the number of freely estimated parameters. When degrees of freedom are positive, the model is overidentified: it makes more predictions than it has parameters, so it can fail. When they are zero, the model is just-identified: it reproduces the data perfectly and cannot be tested as a whole, though its individual parameters can still be estimated and tested. When they are negative, the model is underidentified: there are more unknowns than pieces of information, and no unique solution exists. Degrees of freedom are the currency of testability. Each one represents a constraint the model imposes on the data, a place where it says "this covariance must equal that combination of parameters," and each constraint is an opportunity for the data to disagree. A researcher who adds a correlated error term to improve fit spends a degree of freedom. The fit improves, necessarily, but the model now says less. A model with many degrees of freedom that still fits well has survived many tests. A model with one or two has survived very few. This is why the phrase "the model fit the data" must always be read alongside the degrees of freedom. A just-identified path model always fits perfectly and proves nothing by doing so. A confirmatory factor model with 200 degrees of freedom and acceptable fit has passed a demanding examination, though, as later chapters will show, not necessarily the examination the researcher cared about. Identification Counting is necessary but not sufficient. A model can have positive degrees of freedom overall and still contain parameters that the data cannot determine. Identification asks whether each free parameter has a unique value implied by the population covariance matrix. If two different parameter values would produce the same implied matrix, no amount of data could tell them apart, and the parameter is not identified. The most common identification problem in practice concerns the scale of latent variables, which is taken up in Chapter 3. A latent variable has no natural units. Is exhaustion measured on a scale from zero to one, or zero to a hundred? The data cannot say. So the modeler must fix the scale, either by setting one factor loading to 1.0, which gives the latent variable the units of that indicator, or by setting the latent variance to 1.0, which standardizes it. AMOS does the first automatically when a latent variable is drawn with its indicators; lavaan's cfa() and sem() functions also fix the first loading of each factor to one by default, and the argument std.lv = TRUE switches to fixing the variances instead. The two choices produce identical fit and equivalent solutions; they differ only in the units in which the estimates are expressed. Other identification problems are subtler. A factor with only two indicators is not identified on its own; it can be estimated only if it correlates with other factors in the model, and even then the solution may be unstable. A single-indicator factor requires the researcher to fix the indicator's error variance to some value, usually derived from a known reliability. Nonrecursive models, which contain feedback loops or reciprocal paths, require instrumental variables that affect one variable in the loop but not the other, a condition that must be justified substantively. Models with correlated errors between indicators that load on different factors can become unidentified in ways that no simple rule detects. Two practical warnings follow. First, software will sometimes produce estimates for a model that is empirically underidentified: formally identified, but with data that leave some parameter almost undetermined. Symptoms include enormous standard errors, warnings about a non-positive-definite information matrix, and estimates that swing wildly when the model is changed slightly. lavaan issues explicit warnings in such cases; AMOS reports that the model is probably unidentified and may suggest constraints. These warnings deserve to be read rather than suppressed. Second, identification is a property of the model and the population, not of any particular estimation run. A researcher who is unsure whether a novel model is identified can fit it to a covariance matrix generated from known parameter values and check whether the software recovers them. Writing the model down A structural equation model can be expressed as a diagram, as a set of equations, or as program syntax. The three are equivalent, and fluency in moving among them protects against errors. The diagram conventions are nearly universal. Observed variables appear as rectangles, latent variables as ovals or circles. A single-headed arrow represents a directional effect, a regression coefficient or factor loading. A double-headed curved arrow represents a covariance between exogenous variables or between error terms. Each endogenous variable, meaning any variable with an arrow pointing into it, has an error or disturbance term, usually drawn as a small circle with an arrow into the variable. In AMOS the diagram is the model: the researcher draws it on a canvas with the program's tools, and AMOS translates the drawing into equations. This makes the method visually accessible, but it also means that omissions are easy to miss. An error term left undrawn is an error variance fixed to zero, a strong and usually unintended claim. lavaan takes the opposite route. The model is typed as a short block of text, and the diagram, if wanted, is drawn afterward by a separate package such as semPlot. The syntax uses a handful of operators that map directly onto diagram elements, and Table 1 sets out the correspondence along with how the same element is created in AMOS. Table 1. Core model elements in diagrams, lavaan syntax, and AMOS. Element Diagram lavaan operator Example AMOS action Factor loading Arrow from oval to rectangle =~ exhaust =~ e1 + e2 + e3 Draw latent variable with indicators Regression path Arrow between variables ~ quit ~ exhaust + support Draw single-headed arrow Covariance or variance Curved two-headed arrow ~~ e1 ~~ e2 Draw double-headed arrow Intercept or mean Not usually drawn ~ 1 quit ~ 1 Enable means and intercepts Defined parameter Not drawn := ind := a*b User-defined estimand plugin Parameter label Label on arrow name* exhaust ~ a*workload Name the parameter in its properties One practical consequence of the difference in interfaces is that lavaan makes certain defaults invisible in a different way than AMOS does. When the sem() function is used, lavaan automatically estimates residual variances for all observed and latent endogenous variables and, by default, allows exogenous latent variables to covary. An AMOS user must draw the curved arrows between exogenous latent variables explicitly; an omitted one is a covariance fixed at zero. Researchers moving between programs should always check the full parameter list in the output, which both programs provide, rather than trusting that the model estimated is the model intended. What the estimates mean Once estimated, the model yields several kinds of numbers, and it pays to be precise about what each represents. Unstandardized coefficients are in the original units. A path of 0.35 from workload to exhaustion, with both measured on five-point scales, means that a one-point increase in workload is associated with a 0.35-point increase in expected exhaustion, holding the other predictors of exhaustion constant. Standard errors and significance tests are computed for the unstandardized estimates. Standardized coefficients rescale everything to unit variance. They make paths comparable within a model, and for factor loadings they indicate how strongly each indicator reflects its factor. But standardized coefficients depend on the variances in the particular sample, so they are poor guides for comparing groups or studies whose variances differ. When a researcher wants to say that an effect is stronger among novice teachers than experienced ones, the comparison should be made on unstandardized coefficients with the latent scales anchored equivalently in both groups, a requirement discussed under measurement invariance in Chapter 4. The squared multiple correlation, R², for each endogenous variable gives the proportion of its variance accounted for by its predictors in the model. For an indicator of a latent factor, R² is the proportion of that item's variance attributable to the factor, a quantity sometimes called item reliability. Error variances are estimated parameters, and they should be inspected. A negative error variance, known as a Heywood case, is impossible in a population and signals a problem: a misspecified model, an indicator that is nearly redundant with another, a factor with too few indicators, or a small sample. Heywood cases are not to be patched silently by fixing the variance to a small positive number; they are symptoms, and the underlying cause needs to be found. A model is an argument It may seem that this chapter has dwelt on elementary mechanics. But the mechanics carry a conceptual point that shapes everything later. A structural equation model makes claims of two kinds. It claims that certain parameters are nonzero, and those claims are tested by the significance of individual estimates. It also claims that certain parameters are zero, the omitted arrows, and those claims are tested by overall fit. Researchers tend to focus on the first kind, because the paths they drew represent their hypotheses. But the second kind is where the model earns its right to be called confirmatory. A model in which every variable is connected to every other says nothing that the correlation matrix did not already say. The model is also an argument about direction, and here the covariance matrix is silent. The chain from parental involvement to self-concept to achievement implies exactly the same correlations as the chain running in reverse, from achievement to self-concept to involvement, and as a model in which self-concept is a common cause of the other two. All three models have one degree of freedom and all three predict that the correlation between the endpoints equals the product of the other two. They are equivalent models, observationally indistinguishable, and no fit statistic can choose among them. The choice must rest on theory, on the design of the study, and on knowledge that lies outside the data. Chapter 8 returns to equivalent models, because they are the most frequently ignored threat to interpretation in published structural equation modeling. The remaining chapters take these ideas into increasingly realistic territory: first path models with observed variables, then latent constructs, then the structural relationships among them, then the judgment of fit and the diagnosis of failure. In each, the question to hold in mind is the one this chapter began with. What pattern of associations does this model predict, and what, exactly, would count against it? Chapter 2. Path Analysis with Observed Variables Structural equation modeling began without latent variables. Sewall Wright developed path analysis to study inheritance in guinea pigs and the determinants of birth weight and bone size in livestock, using measured variables and diagrams of hypothesized influences. The method passed into economics as simultaneous equation modeling and into sociology in the 1960s, where Otis Dudley Duncan's work on status attainment made path diagrams a standard way of representing how family background, education, and occupation link across a life. Only in the early 1970s, with Karl Jöreskog's development of the LISREL model and program, were path analysis and factor analysis joined into the general framework now called structural equation modeling. Path analysis with observed variables remains useful in its own right. Many research questions in education and management involve variables measured well enough, or by single indicators unavoidably, that a latent variable model would add complexity without benefit: grade point average, years of experience, firm size, number of absences, test scores from standardized instruments with published reliabilities. More important for this book, path analysis is where the logic of direct, indirect, and total effects is easiest to see, and where the limits of observed-variable models become clear enough to motivate everything that follows. Systems of regressions A path model is a set of regression equations estimated simultaneously. Each endogenous variable is regressed on the variables that point into it. Consider a model from organizational research: transformational leadership (L) is hypothesized to raise employees' psychological empowerment (E), which in turn raises both job satisfaction (S) and discretionary effort, often labelled organizational citizenship behavior (C). Satisfaction also affects citizenship behavior directly. Leadership has no direct path to satisfaction or citizenship behavior. Written as equations, the model says E depends on L; S depends on E; and C depends on E and S. Each equation has its own disturbance term, representing everything that affects the outcome but is not in the model. If the disturbances are assumed to be uncorrelated, and the model is recursive, meaning there are no feedback loops and all causal flow runs one way, the path coefficients could be estimated by running three separate ordinary regressions. The estimates would be the same as those from maximum likelihood estimation of the system. What the simultaneous approach adds is the overall test: the model omits two paths, L to S and L to C, and it assumes that the disturbances are uncorrelated. Four observed variables provide ten variances and covariances; the model estimates one exogenous variance, four path coefficients, and three disturbance variances, eight parameters in all, leaving two degrees of freedom, one for each omitted path. The chi-square test with two degrees of freedom asks whether those omissions are consistent with the data. This is worth dwelling on because it illustrates what "confirmatory" means. The researcher's theory is that leadership works entirely through empowerment. The two missing arrows are the theory. If the data show that leadership predicts satisfaction even after empowerment is accounted for, the chi-square will register it, and the modification index for the L to S path will be large. The model has made a falsifiable claim. Direct, indirect, and total effects Path models distinguish among three kinds of effect, and the distinction is the foundation of mediation analysis. The direct effect of one variable on another is the coefficient on the arrow connecting them, if there is one. In the leadership model, the direct effect of empowerment on citizenship behavior is the coefficient on the E to C path. An indirect effect runs through one or more intervening variables and equals the product of the coefficients along the route. Empowerment affects citizenship behavior indirectly through satisfaction: the E to S coefficient multiplied by the S to C coefficient. Leadership affects citizenship behavior through two indirect routes: L to E to C, and L to E to S to C. Each has its own product. The total effect is the sum of the direct effect and all indirect effects. The total effect of leadership on citizenship behavior, in this model, is the sum of the two indirect products, since there is no direct path. Both AMOS and lavaan compute these quantities. In AMOS, the "Indirect, direct and total effects" option in the Analysis Properties dialog produces tables of all three; to test a specific indirect effect when a variable has several, AMOS requires a user-defined estimand, since its default indirect effect sums all routes. In lavaan, the researcher labels the relevant paths and defines each indirect effect explicitly with the := operator, for example writing ind1 := a*b after labelling the paths a and b. The explicit approach has a virtue: it forces the researcher to state which indirect effect is of interest, rather than accepting an aggregate that may combine theoretically distinct mechanisms. The decomposition of effects is algebraically straightforward. Its interpretation is not, and Chapter 5 is devoted to the difference. For now, note that the words "direct" and "indirect" are relative to the model. A direct effect is whatever part of the association is not carried by the mediators included in the model. Add another mediator, and part of the former direct effect becomes indirect. The direct effect is a residual, not a mechanism. Assumptions that the diagram hides A path diagram looks like a description of how the world works. What it actually encodes is a set of statistical assumptions, some of them strong, that the researcher must be prepared to defend. The first is that the disturbances of different endogenous variables are uncorrelated unless the model says otherwise. This assumption means that no omitted variable affects two endogenous variables at once. In the leadership model, if an unmeasured factor such as an employee's general positive affect raises both empowerment and satisfaction, then the disturbances of E and S are correlated, and the estimated E to S path absorbs a spurious component. The coefficient will be too large, and the indirect effect of leadership through satisfaction will be overstated. This is the omitted-confounder problem that afflicts all observational research, but path diagrams make it easy to forget. The diagram shows arrows the researcher believes in; it does not show the unmeasured common causes the researcher has assumed away. A useful discipline, when drawing any path model, is to ask of every pair of endogenous variables: is there anything that could cause both of these that is not in the model? If the answer is yes, and it usually is, the estimates between them should be regarded as upper or lower bounds, not as effects. The second assumption is correct functional form: the relationships are linear and additive unless the model specifies otherwise. A path coefficient summarizes the average linear association. If empowerment raises satisfaction strongly at low levels and hardly at all at high levels, the coefficient will describe neither regime well. The third assumption is correct direction. The arrow from empowerment to satisfaction could as easily run the other way: satisfied employees may feel more empowered. With cross-sectional data, as Chapter 1 showed, models differing only in the direction of such arrows can fit identically. The fourth assumption, and the one that motivates latent variable modeling, is that the variables are measured without error. The problem of measurement error In ordinary regression, error in the outcome variable is harmless to the coefficients; it inflates standard errors but does not bias estimates. Error in a predictor is another matter. Classical measurement error in a single predictor attenuates its coefficient toward zero, by a factor equal to the predictor's reliability. If empowerment is measured with reliability .70, the estimated effect of empowerment on satisfaction will be, on average, about 70 percent of its true size. With several correlated predictors, the damage is less predictable. Suppose satisfaction and empowerment both predict citizenship behavior, and empowerment is measured with more error than satisfaction. The regression will understate the role of empowerment, and because the two predictors are correlated, some of the variance that truly belongs to empowerment will be credited to satisfaction. The coefficient for the better-measured predictor can be biased upward. In mediation models this has a specific consequence that has been demonstrated repeatedly: measurement error in the mediator attenuates the estimated indirect effect and inflates the estimated direct effect. A researcher who finds "partial mediation" may be looking at complete mediation observed through an unreliable measure. The standard reliability coefficient for scale scores, Cronbach's alpha, does not help as much as its ubiquity suggests. Alpha estimates reliability correctly only when all items measure the construct equally well, an assumption called tau-equivalence that rarely holds; otherwise it underestimates reliability of the sum score, and it can be distorted by correlated errors. More fundamentally, knowing the reliability does not remove the error from the analysis. It only quantifies it. A small calculation shows how large the distortion can be. Suppose that, in the population, empowerment correlates .50 with satisfaction and .30 with citizenship behavior, and satisfaction correlates .40 with citizenship behavior, with all three measured perfectly. Regressing citizenship behavior on empowerment and satisfaction gives standardized coefficients of about 0.13 for empowerment and 0.33 for satisfaction. Now suppose empowerment is measured with reliability .60, a figure not unusual for short scales, while satisfaction is measured with reliability .90. The observed correlations shrink: empowerment with satisfaction becomes about .37, and empowerment with citizenship behavior about .23. Rerunning the regression on these attenuated correlations gives roughly 0.11 for empowerment and 0.34 for satisfaction. The poorly measured predictor has lost about a fifth of its coefficient, and the well-measured one has gained slightly despite its own small measurement error, even though nothing about the true relationships has changed. In a larger model with more correlated predictors, the redistribution can be larger and can even change signs. There are two remedies within the observed-variable framework. The first is to correct the path model for known unreliability by specifying each scale score as the single indicator of a latent variable, fixing the indicator's loading to one and its error variance to (1 − reliability) × observed variance. This is a legitimate technique, particularly when a construct is measured by a well-established instrument whose reliability is known from large samples, and it is easily specified in either program. In lavaan, for a scale score with observed variance 0.80 and reliability .85, the error variance would be fixed at 0.12 by writing a line such as emp_score ~~ 0.12*emp_score alongside a single-indicator factor definition. In AMOS, the error variance is fixed by entering the value in the parameter's properties. The second remedy is to model the items directly, allowing the software to estimate each item's relationship to the construct and each item's error. That is the subject of Chapters 3 and 4, and it is the step that makes structural equation modeling distinctive. Nonrecursive models and reciprocal effects Most published path models are recursive. But some theories genuinely involve reciprocal influence: job satisfaction and job performance may each affect the other; students' motivation and achievement may reinforce one another. A model with a two-way arrow structure, satisfaction pointing to performance and performance pointing to satisfaction, is called nonrecursive. Such models are not identified without additional information. The standard solution uses instrumental variables: for each variable in the loop, at least one predictor that affects it directly but has no direct effect on the other variable in the loop. For the satisfaction and performance loop, one might use pay fairness as an instrument that affects satisfaction but not performance directly, and cognitive ability as an instrument that affects performance but not satisfaction directly. The quality of the estimates depends entirely on the validity of those exclusion claims, and they are often hard to defend. Weak instruments, those only modestly correlated with the variable they are meant to predict, produce unstable estimates with very large standard errors. In practice, reciprocal effects are better studied with longitudinal data, in which the influence of satisfaction at time one on performance at time two, and of performance at time one on satisfaction at time two, can be estimated in a cross-lagged panel model. Even these models have been criticized in recent years because the traditional cross-lagged panel model confounds stable between-person differences with within-person change. Hamaker, Kuiper, and Grasman's 2015 paper in Psychological Methods proposed the random-intercept cross-lagged panel model, which separates the two by giving each person a stable trait component, and it is now widely used for questions about reciprocal processes within individuals. Its lesson generalizes: a path coefficient in a panel model answers a question about some level of analysis, and the researcher must know which. Sample size and estimation in path models Path models with observed variables are relatively undemanding of sample size compared with latent variable models, since they estimate fewer parameters. A common heuristic calls for at least ten to twenty cases per estimated parameter, but heuristics of this kind have weak foundations. Sample size requirements depend on the size of the effects, the reliability of the measures, the number of parameters, the distribution of the data, and the purpose of the analysis. A simulation study by Wolf, Harrington, Clark, and Miller, published in Educational and Psychological Measurement in 2013, found that required sample sizes for structural equation models ranged from as few as 30 to more than 400 depending on such features, and that no single rule of thumb held across conditions. The more defensible approach is a power analysis for the specific parameters or fit test of interest, which can be carried out by simulation in lavaan or with dedicated tools, and which forces the researcher to state expected effect sizes in advance. Maximum likelihood assumes multivariate normality. With observed variables that are skewed, such as counts of absences or measures of rare behaviors, the parameter estimates remain consistent but the standard errors and chi-square statistic can be badly wrong. Robust corrections, discussed in Chapter 4, address this and are available in both programs, though they are easier to request in lavaan. AMOS's primary alternative for non-normal data is bootstrapping, which it implements well and which is valuable particularly for indirect effects, whose sampling distributions are not normal even when the variables are. What path analysis teaches Path analysis with observed variables is, in one sense, a set of regressions with a diagram attached. Its enduring value is conceptual. It teaches that a model is defined as much by its omitted paths as by its included ones; that effects decompose into direct and indirect components that depend on what else is in the model; that every arrow embeds assumptions about confounding and direction that the data cannot check; and that measurement error, in predictors and mediators especially, distorts exactly the coefficients researchers most want to interpret. That last lesson leads directly to latent variables. If the quantities of interest are constructs rather than scale scores, and the scale scores contain error, then the model should represent the constructs and their measurement explicitly. Doing so requires a measurement model, and specifying a measurement model well is harder, and more consequential, than most researchers expect. Chapter 3. Latent Constructs and the Measurement Model A latent variable is a claim that something unobserved explains why certain observed things go together. When a researcher says that eight questionnaire items measure teacher self-efficacy, the statistical content of that statement is precise: the eight items covary because each responds to the same underlying quantity, and once that quantity is held constant, the items should no longer covary at all. The latent variable is the common cause that accounts for the shared variance. What remains in each item after the common cause is removed is unique variance, a mixture of random error and item-specific content that has nothing to do with self-efficacy. That definition is worth stating plainly because the language of latent variables tempts researchers into a looser view, in which a construct is whatever a set of items happens to be labelled. The looser view produces measurement models that fit poorly for reasons the researcher cannot diagnose, and structural estimates that mean something different from what the construct's name suggests. The discipline of the measurement model begins with taking the common-cause claim seriously. The reflective model and its implications The standard latent variable model in structural equation modeling is the reflective, or common factor, model. Each indicator is written as a linear function of the factor plus an error: the item score equals an intercept, plus a loading times the factor, plus a unique term. The arrows run from the construct to the items. Variation in self-efficacy produces variation in the item responses, not the other way around. The model makes a strong and testable prediction, known as local independence: indicators of the same factor are uncorrelated once the factor is controlled. For a single factor with four indicators, this implies that the six covariances among the items are fully explained by four loadings and the factor variance. With the scale fixed, that is four free parameters to account for six covariances, and the model has two degrees of freedom. With three indicators the model is just-identified and cannot be tested in isolation. With two it is underidentified on its own. This is the source of the widely repeated advice that each factor should have at least three indicators, and preferably four or more: not a matter of reliability alone, but of whether the model's defining claim can be checked. Several consequences follow from the reflective model, and each is a useful test of whether it suits a given construct. Indicators should be interchangeable in principle. Dropping one indicator of a well-defined reflective construct should not change what the construct means, only how precisely it is measured. If removing an item would change the construct's content, the reflective model is probably wrong for it. Indicators should correlate positively with one another, after any reverse-coded items are recoded. If two items supposedly reflecting the same factor are uncorrelated, a single common cause cannot explain them. Indicators should relate to other variables in the same way, apart from differences in strength. If one self-efficacy item correlates strongly with years of experience and another not at all, the items are responding to different things. Formative constructs and composites Not every construct fits the reflective model. Socioeconomic status is often measured by income, education, and occupational prestige. It makes little sense to say that a person's latent status causes their income and their education; rather, income and education are components that together define status. A rise in income raises status without any expectation that education will follow. Such a construct is called formative, or, when the weights are fixed rather than estimated, a composite. Formative indicators need not correlate. Removing one changes the construct's meaning. Local independence does not apply. Forcing a formative construct into a reflective measurement model typically produces poor fit, low and uneven loadings, and a factor whose meaning drifts toward whichever indicator dominates the covariance structure. Formative constructs raise identification problems of their own in covariance-based structural equation modeling: a formatively measured latent variable is identified only if it emits paths to at least two other variables, and its meaning then depends partly on what those outcome variables are. A lively methodological debate over the past two decades has questioned whether formatively measured latent variables are coherent at all, with some methodologists arguing that such constructs are better treated simply as weighted composites of observed variables. The practical advice is conservative. If the construct is a sum or index of its components by definition, as with a checklist of stressful life events or a firm's portfolio of human resource practices, treat it as an observed composite and model it as such. Reserve latent variables for constructs whose indicators are plausibly effects of a common cause. A related tradition, partial least squares structural equation modeling, is widely used in management and information systems research and treats all constructs as weighted composites. It optimizes explained variance in outcome constructs rather than fit to the covariance matrix, and it has distinct software such as SmartPLS. Its relationship to covariance-based modeling has been a source of heated exchanges in the methods literature. This book concerns covariance-based modeling as implemented in AMOS and lavaan; researchers choosing between the approaches should recognize that they answer different questions and that the fit indices of later chapters do not transfer to partial least squares. Scaling the latent variable A latent variable has no inherent units, and the model cannot be estimated until some are assigned. Chapter 1 noted the two common choices. Fixing the first loading to one, the default in both AMOS and lavaan, gives the factor the metric of that marker indicator: a one-unit difference in the factor corresponds to a one-unit expected difference on the marker item. Fixing the factor variance to one, requested in lavaan with std.lv = TRUE, standardizes the factor and frees all loadings. The choice does not affect fit, but it affects interpretation and occasionally estimation. If the marker indicator is weakly related to the factor, the marker-variable approach gives a factor with a poorly defined scale and can produce estimation difficulties. In multiple-group models, fixing the factor variance to one in every group imposes equal variances, which may be untrue; the marker-variable approach, or fixing the variance only in a reference group, is then appropriate. A third option, effects coding, constrains the loadings of each factor to average one, so the factor takes on the average metric of its indicators; it is useful when a researcher wants latent means and variances expressed on the original response scale. Reliability, properly estimated The measurement model yields direct estimates of how well each indicator reflects its factor, and from these one can compute reliability for the construct as a whole. The standardized loading of an item, squared, is the proportion of its variance explained by the factor. An item with a standardized loading of .80 has 64 percent common variance; one with .50 has only 25 percent. Items with standardized loadings below about .40 contribute little and may be measuring something else, though whether to retain them depends on the construct's content and not only on the numbers. For the reliability of a composite, the model-based coefficient most often recommended is omega, associated with the psychometrician Roderick McDonald. For a unidimensional factor with uncorrelated errors, omega is the squared sum of the unstandardized loadings divided by that quantity plus the sum of the error variances. Unlike Cronbach's alpha, omega does not assume equal loadings. When loadings are unequal, as they almost always are, alpha underestimates the reliability of the sum score, though usually not by much; when errors are correlated, alpha can be misleading in either direction. Dunn, Baguley, and Brunsden made the case for moving from alpha to omega in the British Journal of Psychology in 2014, and the argument is now widely accepted. In lavaan, the companion package semTools provides reliability coefficients including omega from a fitted model through its compRelSEM() function. AMOS does not report omega directly; users compute it from the estimated loadings and error variances, which is simple arithmetic. A short example makes the difference concrete. Suppose four items measuring emotional exhaustion have standardized loadings of .85, .80, .60, and .45. Their error variances in standardized form are one minus the squared loadings: .28, .36, .64, and .80. The sum of loadings is 2.70, its square 7.29, and the sum of error variances 2.08, so omega is 7.29 divided by 9.37, about .78. Cronbach's alpha computed on the same items, assuming the model holds exactly, comes out close to .76. The difference is small here, as it usually is, but the two coefficients diverge more as loadings become more unequal. More instructive is what the loadings reveal that neither coefficient does. The fourth item shares only about 20 percent of its variance with the factor. It may be poorly worded, it may tap a different facet of exhaustion, or it may be responding to something else entirely. Omega and alpha both summarize the scale; the loadings show which item is carrying less of the construct and invite a look at its content. A related quantity, composite reliability, is computed from the standardized solution in the same way and appears frequently in management research. Along with it, researchers often report average variance extracted, the mean of the squared standardized loadings for a factor. Fornell and Larcker, in a 1981 paper in the Journal of Marketing Research, proposed that average variance extracted should exceed .50, meaning that the factor accounts for more of its indicators' variance than error does. That standard is widely cited and should be understood as a guideline, not a law; a construct measured by many moderately loading items can be reliable as a composite while falling short of it. Discriminant validity A measurement model with several factors makes a second kind of claim: that the constructs are distinct. If the estimated correlation between two latent factors approaches one, the data do not support treating them as different things, regardless of what they are called. Three approaches are common. The first compares a model in which the two factors are distinct with a model in which their correlation is fixed at one, or in which their indicators load on a single factor; a significant loss of fit when the constraint is imposed indicates that the constructs can be distinguished. With large samples, this test detects trivial differences, so it establishes that the factors are not identical but not that they are meaningfully distinct. The second, proposed by Fornell and Larcker, requires that each factor's average variance extracted exceed its squared correlation with any other factor. The logic is that a construct should share more variance with its own indicators than with another construct. The third, the heterotrait-monotrait ratio proposed by Henseler, Ringle, and Sarstedt in the Journal of the Academy of Marketing Science in 2015, compares the average correlation between items of different constructs with the average correlation among items of the same construct. Their simulations suggested that the Fornell-Larcker criterion often fails to detect a lack of discriminant validity, and that ratios above about .85 or .90 signal a problem. The heterotrait-monotrait ratio is available in semTools through the htmt() function. Whatever the criterion, the substantive question is whether two constructs that the theory treats as different are empirically separable in the measures used. In organizational research, constructs such as job satisfaction, affective commitment, and engagement are often correlated above .70 as latent variables, and a structural model that treats one as a cause of another may be relating a construct to a near-duplicate of itself. The measurement model is the place to discover this, before the structural paths are interpreted. Correlated errors and method variance The reflective model assumes that indicators' unique parts are uncorrelated. Real questionnaires violate this in predictable ways. Items with similar wording, such as two items both beginning with "I feel confident that," share something beyond the construct. Negatively worded items often share variance with each other that has more to do with response style than content. Items measured by the same method, such as self-report on the same occasion, share method variance that inflates their correlations with one another, including correlations across constructs. Podsakoff, MacKenzie, Lee, and Podsakoff reviewed these sources of common method bias in the Journal of Applied Psychology in 2003, and their paper remains the standard reference. Its central message is that procedural remedies, such as measuring predictor and outcome from different sources or at different times, are more effective than statistical ones. Among statistical remedies, adding a method factor on which all items load, alongside their substantive factors, is common but can create identification and interpretation problems. The popular Harman single-factor test, which checks whether one factor accounts for most of the variance, is widely regarded as an inadequate diagnostic. Correlated errors between specific items should be specified when there is an a priori reason: items that share a stem, repeated measurement of the same item across occasions, or items from the same subscale nested within a broader construct. In longitudinal models, the error of each item at time one should typically be allowed to correlate with the same item's error at later occasions; failing to do so misattributes item-specific stability to the construct. What should not be done is to add correlated errors because a modification index suggests them, with no substantive rationale. Chapter 8 returns to why. Parcels, higher-order factors, and bifactor models Three extensions of the basic measurement model come up often enough in applied work to need brief treatment. Item parceling combines items into small composites, such as averaging three items at a time, and uses the parcels as indicators. It reduces the number of parameters and often improves fit. Little and colleagues have defended parceling when the goal is to estimate structural relations among constructs whose unidimensionality is well established. Critics point out that parceling can hide misspecification in the item-level model, so that good fit with parcels is weak evidence that the items measure what they are said to. A reasonable position is that parceling is defensible for a well-validated unidimensional scale and indefensible when the measurement model is itself in question. A higher-order factor model treats first-order factors as indicators of a broader factor: for example, emotional exhaustion, depersonalization, and reduced accomplishment as facets of burnout. With three first-order factors, the second-order structure is just-identified and fits exactly as well as a model with three correlated first-order factors, so it cannot be tested against that model; with four or more it imposes testable constraints. A bifactor model lets every item load on a general factor and on one specific factor, with general and specific factors uncorrelated. It has become popular for questions about whether a multidimensional scale can be scored as a whole. Bifactor models tend to fit well, sometimes because their flexibility lets them absorb noise, and their specific factors are often weakly defined once the general factor is extracted. Indices such as omega hierarchical, which estimate how much of the total score's variance reflects the general factor, are more informative than fit comparisons for deciding whether a total score is meaningful. The measurement model as foundation Everything in the structural part of a model depends on the measurement part. The meaning of a latent variable is defined by its indicators and the pattern of their loadings. Structural paths are estimated between these latent variables, and their magnitudes depend on how much error the measurement model has removed. Discriminant validity problems in the measurement model become multicollinearity problems in the structural model. Misspecification anywhere tends to spread. This is why careful practice treats the measurement model as a hypothesis to be tested on its own before any structural paths are added, and why the next chapter concerns the confirmatory factor analysis that does the testing. Hashtags: #StructuralEquationModeling #SEM #PathAnalysis #LatentConstructs #LatentVariableModeling #CovarianceStructure #MeasurementModel #ConfirmatoryFactorAnalysis #CFA #ModelIdentification #DegreesOfFreedom #DirectEffects #IndirectEffects #MediationAnalysis #MeasurementError #ReflectiveMeasurement #FormativeConstructs #FactorLoadings #CompositeReliability #DiscriminantValidity #ModelFit #FitIndices #ModificationIndices #AMOS #FutureOfStructuralEquationModeling

  • Scientific Grant Budgeting and Financial Compliance (A PI's Field Guide)

    Download the Book (PDF): Introduction Most scientists meet grant accounting the way travellers meet customs officers: briefly, reluctantly, and with a vague sense that something could go wrong without their quite knowing what. The proposal is written in a rush of scientific ambition, the budget is assembled in the last seventy-two hours before the deadline by a department administrator working from a spreadsheet template, and the award notice arrives months later carrying a stack of terms and conditions that almost nobody reads end to end. Then the money starts to move, and for three or five years the principal investigator spends it on people, reagents, instruments and travel while an entirely separate machine inside the university records, allocates, reconciles and reports every dollar. For most PIs, most of the time, the two worlds never collide. The experiments run, the reports go in, the award closes. But when they do collide the consequences are out of all proportion to the sums involved. A postdoc's salary charged to the wrong grant for eight months becomes a cost transfer that an auditor flags as late and unexplained. A promise in the proposal that the department would "provide a technician at no cost to the project" becomes a cost-sharing commitment the university must document to the penny. A freezer bought in the last month of an award, for no reason anyone can articulate beyond "the money was there", becomes an unallowable cost the institution must refund. And at the far end of the spectrum, patterns of misstatement across many awards become False Claims Act settlements measured in millions of dollars, announced by the Department of Justice and reported in Science and Nature. This booklet is written for the scientist in the middle of that picture: a principal investigator, or someone about to become one, who wants to understand the financial architecture of sponsored research well enough to make good decisions and avoid preventable trouble, without becoming an accountant. It concentrates on federally funded research at American universities, because that is where the rules are most elaborate, most codified and most consequential, and because the principles carry over well to foundations, industry sponsors and research systems elsewhere. It uses the vocabulary your sponsored programs office uses, and explains it, so that the next conversation with that office goes faster and better. The argument of this book The central claim is simple, and everything else in the book is an elaboration of it: a grant budget is not an estimate; it is a set of representations the institution makes to the government, and nearly every compliance failure is a gap between what was represented and what the records later show. That framing changes how a PI should think about almost every financial decision. When you list a percentage of your effort in a proposal, you are representing that you will devote that effort. When you describe a cost as direct, you are representing that it can be specifically and consistently identified with this project. When you name a collaborator's institution as a subrecipient, you are representing that it will carry out part of the programmatic work under your institution's oversight. When you offer the department's contribution of a microscope or a technician, you are representing a commitment the university must later prove it kept. Auditors, inspectors general and, in the worst cases, federal prosecutors do not ask whether the science was good. They ask whether the representations were true. Seen this way, compliance stops being a thicket of arbitrary rules and becomes something more tractable: the discipline of making only promises you can keep, and keeping records that show you kept them. The rules themselves, which run to hundreds of pages across the federal Uniform Guidance, agency policy statements and institutional procedures, turn out to be mostly consistent applications of a few principles. Costs must be necessary for the project, reasonable in amount, allocable to the project in proportion to the benefit it receives, treated consistently with how the institution treats similar costs elsewhere, and documented. Most of what follows is a working out of those principles in the places where PIs actually make decisions. What the book covers The chapters follow the path of money through an award, from the rules that govern it to the audit that tests it. Chapter 1 sets out the regulatory framework: the federal Uniform Guidance at 2 CFR Part 200, the agency-specific overlays from NIH and NSF, the institution's own policies, and the four or five tests every cost must pass. It also explains why the 2024 revision of the Uniform Guidance, effective for awards made from October 2024, changed several thresholds that matter to PIs. Chapter 2 turns to direct costs: salaries and effort, fringe benefits, the NIH salary cap, supplies, travel, participant support, and the categories of expense that are ordinarily indirect but can sometimes be charged directly. The emphasis throughout is on building a budget you can defend line by line. Chapter 3 explains indirect costs, formally called facilities and administrative costs or F&A. It shows what the rate actually pays for, how it is negotiated, why it is applied to a modified base rather than to the whole budget, and why the attempt in 2025 to replace negotiated rates with a flat cap ended up in the federal courts. Chapter 4 deals with equipment: where the line between supplies and equipment falls, what depreciation means in a university context, who owns an instrument bought with grant funds, and how to avoid the classic end-of-award purchase that auditors love to question. Chapter 5 addresses cost sharing, the most misunderstood commitment in research finance, and explains why the most expensive words in many proposals are "at no cost to the sponsor". Chapter 6 covers subawards: how to tell a subrecipient from a vendor, what your institution must do to monitor the partner it pays, and how the rules for international collaboration changed at NIH in 2025. Chapter 7 follows the award through its active life: effort reporting, cost transfers, rebudgeting, no-cost extensions and closeout, which is where most small problems either get fixed or harden into findings. Chapter 8 examines audits and penalties: how the single audit works, what the Office of Inspector General looks for, how False Claims Act cases arise, and what the largest university settlements of the past two decades actually involved. The conclusion draws these threads into a practical stance for a working PI, and argues that the institutions and investigators who suffer least from audits are not those who know the most rules, but those whose budgets and records were designed from the outset to tell a single true story. A note on scope and currency Research finance is not static. The thresholds in the Uniform Guidance were revised in 2024. NIH's salary cap rises most years. The rules for foreign subawards at NIH were rebuilt in 2025. The fight over indirect cost rates that began in February 2025 produced appellate decisions in 2026 and a policy debate that is not over. This booklet describes the rules as they stand in 2026, flags the places where they have recently moved, and tries to explain the reasoning behind them so that a reader can interpret the next change rather than merely memorise the current number. Two cautions apply throughout. First, your institution's policies may be stricter than federal rules, and when they are, they govern. A university that sets its equipment capitalisation threshold at $5,000 applies that figure, not the federal ceiling of $10,000. Second, nothing here substitutes for the specific terms of your award. The notice of award, the program solicitation and the agency's policy statement override general principles where they conflict. The value of understanding the principles is that you will know which questions to ask, and of whom, before rather than after the money is spent. Chapter 1. The Rules Behind the Money A research grant looks, from the investigator's side, like a gift with strings attached: the agency gives money, the scientist does research, the strings are paperwork. Legally it is something else. A federal grant is a form of financial assistance governed by statute, by government-wide regulation, by agency policy and by the terms of the individual award, and the institution that accepts it takes on binding obligations about how every dollar will be spent and recorded. The principal investigator is not usually the legal recipient. The university is. But the PI is the person whose decisions generate almost every cost, and so the PI is, in practice, the first line of compliance. Understanding the structure of the rules is worth an hour of any investigator's time, because it answers a question that otherwise produces endless frustration: why does the sponsored programs office say no to things that seem obviously reasonable? Usually the answer is that the request collides with one of a small number of principles that sit at the top of the hierarchy, and that no amount of scientific justification can move. The layers of authority At the top sits federal law. Statutes create the agencies, appropriate their money and occasionally impose specific restrictions, such as the annual appropriations provision that limits how much of an individual's salary NIH grants may pay, or the provisions that have, since fiscal year 2018, barred the Department of Health and Human Services from unilaterally changing the way indirect cost rates are applied to NIH awards. Statutes are rarely read by PIs, but they explain why some rules are immovable: an agency cannot waive what Congress has required. Beneath statute sits the government-wide regulation that governs almost all federal grants to universities, nonprofits, states and local governments: Uniform Administrative Requirements, Cost Principles, and Audit Requirements for Federal Awards, codified at Title 2 of the Code of Federal Regulations, Part 200. Everyone in research administration calls it the Uniform Guidance. The Office of Management and Budget issued it in December 2013 to consolidate eight older circulars into a single framework, including Circular A-21, which had governed university cost principles since the 1950s, and Circular A-110, which set administrative requirements. It became effective for new awards in December 2014 and has been revised several times since, most substantially in a revision published in April 2024 and effective for awards and amendments made on or after 1 October 2024. The Uniform Guidance is divided into subparts. Subpart D covers post-award requirements such as financial management, property, procurement, subrecipient monitoring and closeout. Subpart E contains the cost principles, the rules about what can be charged and how. Subpart F contains the audit requirements. Several appendices matter to universities, most importantly Appendix III, which sets out how institutions of higher education identify and allocate indirect costs. Each federal agency then adopts the Uniform Guidance in its own part of the regulations and publishes policy documents that explain how it applies it. For biomedical researchers the central document is the NIH Grants Policy Statement, which is incorporated by reference into every NIH award. For NSF-funded researchers it is the Proposal and Award Policies and Procedures Guide, usually called the PAPPG, which governs both what goes into a proposal and how the award is managed. The Departments of Energy and Defense, NASA, the Department of Agriculture and others have their own equivalents. These documents do not override the Uniform Guidance, but they fill in details and exercise options the Uniform Guidance leaves open. NSF's prohibition on voluntary committed cost sharing, for example, and NIH's specific rules about the salary cap and cost transfers, live at this level. Next comes the individual award. The notice of award or grant agreement specifies the budget, the period of performance, the reporting schedule and any special conditions. A program solicitation may impose rules that apply only to that competition, such as a mandatory cost-share percentage or a cap on the number of months of senior personnel salary. The award terms are the most specific layer, and where they conflict with more general guidance, they generally control. Finally, the institution's own policies apply to every award it accepts. Universities are required to have written policies and procedures for many things the Uniform Guidance governs, and they are free to be stricter than federal rules. Many are. A university may require prior approval for purchases that federal rules would allow without it, or set a lower capitalisation threshold for equipment, or require monthly rather than annual reconciliation of grant ledgers. When a PI hears "the federal rules allow it but our policy doesn't", this is the layer speaking, and it binds just as firmly as the others. For contracts rather than grants, a different body of rules applies. Research contracts with federal agencies are governed by the Federal Acquisition Regulation, which has its own cost principles in Part 31 and its own compliance culture. Contracts are less common than grants for most academic scientists, but they are frequent in defence and energy research, and they bring stricter requirements around deliverables, invoicing and cost accounting. Where this booklet speaks of awards, it means grants and cooperative agreements unless it says otherwise. The tests every cost must pass The heart of the cost principles is a short section, 2 CFR 200.403, that sets out the factors affecting the allowability of costs. Read it once and most of the rest of Subpart E falls into place. To be allowable, a cost must be necessary and reasonable for the performance of the award and allocable to it. It must conform to any limitations or exclusions in the cost principles or the award. It must be consistent with policies that apply uniformly to both federally financed and other activities of the institution. It must be accorded consistent treatment, meaning that a cost cannot be charged directly to an award if a cost of the same kind in similar circumstances has been charged as indirect. It must be determined in accordance with generally accepted accounting principles. It must not be used to meet cost-sharing requirements on another federal award. And it must be adequately documented. In practice, PIs will spend most of their time on four of these tests: reasonableness, allocability, consistency and documentation. Reasonableness, defined at 200.404, asks whether the cost, in its nature and amount, does not exceed what a prudent person would incur in the circumstances prevailing when the decision was made. The test looks at whether the cost is of a type generally recognised as ordinary and necessary, whether it reflects sound business practice and arm's-length bargaining, whether market prices for comparable goods were considered, and whether the individuals concerned acted with prudence. A premium-priced centrifuge chosen for no stated reason when a comparable instrument costs half as much is a reasonableness problem. So is business-class travel when economy is available, or a catered working lunch for a lab of twelve that happens to cost as much as a restaurant dinner. Allocability, at 200.405, asks whether the cost is chargeable to this particular award in accordance with the relative benefits received. A cost is allocable to an award if it is incurred specifically for that award, if it benefits both that award and other work and can be distributed in reasonable proportion to the benefit, or if it is necessary to the overall operation of the institution and is assignable in part to the award. The first two are direct costs; the third is the domain of indirect costs. Allocability is the principle most often violated in the ordinary life of a lab, because it is so easy to charge a shared purchase to whichever grant has the most money left, rather than dividing it among the projects that actually use it. The Uniform Guidance anticipates this temptation and forbids it explicitly: a cost allocable to one award may not be charged to another to overcome funding deficiencies, avoid restrictions or for other reasons of convenience. Consistency, which appears in both 200.403 and in the rules on direct and indirect costs, asks whether the institution treats like costs alike. If a university's normal practice is to recover departmental administrative salaries through its indirect cost rate, it cannot also charge those salaries directly to a federal grant except in the specific circumstances the rules allow. Otherwise the government would pay for the same thing twice: once through the rate and once through the direct budget. This is why so many seemingly sensible direct charges, from office supplies to general-purpose laptops to the department's administrative assistant, draw scrutiny. They are the kinds of cost that are ordinarily in the indirect pool. Documentation is the least intellectual of the tests and the most frequently failed. A cost that is necessary, reasonable and allocable is still unallowable if the institution cannot show, with records created at or near the time, what was bought, why, for which project, and who approved it. Auditors routinely disallow costs not because they doubt the purchase was legitimate but because the file does not prove it. The practical lesson is to write the justification when the cost is incurred, not when the auditor asks. What is expressly unallowable Beyond the general tests, Subpart E contains some fifty sections addressing specific items of cost, from advertising to travel. Most confirm that normal research expenses are allowable under conditions. A smaller number identify costs that are unallowable regardless of how the general tests come out. These are worth knowing by heart, because charging them to a federal award is not an error of judgement but a clear violation. Alcoholic beverages are unallowable. Entertainment costs, including amusement, diversion, social activities and associated costs such as tickets and meals, are unallowable unless they have a programmatic purpose and are authorised in the approved budget or with prior written approval. Fundraising and investment management costs are unallowable. Lobbying is unallowable, as are most costs of organising or influencing legislation. Bad debts, contributions and donations made by the institution, fines and penalties resulting from violations of law, and goods or services for personal use are unallowable. Costs of alumni activities, commencement and convocation, and student activities are generally unallowable as direct or indirect costs, subject to limited exceptions. Membership dues in social or country clubs are unallowable; membership in professional and scientific societies generally is allowable. The pattern is instructive. Expressly unallowable costs are, for the most part, costs that benefit the institution or individuals rather than the research, or costs whose public funding would be politically indefensible. When a PI proposes a charge that falls near one of these categories, such as a laboratory end-of-year dinner, a gift for a departing technician, or the reception after a symposium, the right question is not whether it helped morale but whether it is on the list. If it is, other funds must pay. Why the 2024 revision matters to investigators The April 2024 revision of the Uniform Guidance was the most substantial since 2014, and several of its changes land directly on decisions PIs make. It applies to new awards and to funding amendments made on or after 1 October 2024, which means that for several years many investigators will hold awards under both the old and the new thresholds at the same time. Knowing which rule applies to which award is a task for the sponsored programs office, but knowing that the difference exists is the PI's responsibility. The threshold for equipment rose from $5,000 to $10,000. An item with a useful life of more than one year and a per-unit acquisition cost at or above the lesser of the institution's capitalisation level or $10,000 is now equipment; below that, it is a supply. Because most universities set their own capitalisation thresholds, the practical effect depends on whether your institution has raised its own figure. Many did so in 2024 and 2025; some did not. The de minimis indirect cost rate, available to organisations that do not have a negotiated rate, rose from 10 percent to up to 15 percent of modified total direct costs. This matters to PIs mainly through their subrecipients: a small company or community organisation receiving a subaward can now recover more of its overhead without negotiating a rate. The portion of each subaward included in the modified total direct cost base, on which indirect costs are calculated, rose from the first $25,000 to the first $50,000 over the life of the subaward. The result is that a prime institution now recovers indirect costs on more of each subaward it manages, which recognises the real work of subrecipient monitoring. The threshold at which a non-federal entity must undergo a single audit rose from $750,000 in annual federal expenditures to $1,000,000. Some small subrecipients that previously had single audits no longer do, which shifts more of the monitoring burden onto the prime institution. Several other thresholds moved in the same direction: the value of residual unused supplies at the end of an award above which the government must be compensated rose from $5,000 to $10,000, and the ceiling on fixed amount subawards without prior approval rose from $250,000 to $500,000. The revision also clarified language throughout, replaced "should" with "must" where requirements were intended to be mandatory, and addressed matters such as the treatment of data and information technology costs. The general direction of these changes is to reduce administrative burden by raising thresholds that had not been adjusted for inflation in a decade. None of them loosens the basic tests of allowability. Where the PI stands in all of this Institutions assign responsibilities differently, but some features are nearly universal. The PI is responsible for the scientific and technical direction of the project and for ensuring that expenditures are appropriate to it. The PI approves or initiates purchases, hires and assigns staff, certifies or confirms effort, reviews monthly financial statements, and is usually the first person asked to explain any questioned cost. The department administrator prepares budgets, processes transactions and reconciles accounts. The central sponsored programs office submits proposals, negotiates and accepts awards, handles prior approvals, and interprets policy. The post-award or research accounting office invoices sponsors, draws funds, prepares financial reports and manages closeout. The internal audit function and external auditors test whether all of this works. The risk in this division of labour is that each party assumes someone else is checking. A PI who approves charges without reading them because the department administrator is competent, and an administrator who processes whatever the PI approves because it is the PI's grant, together produce a system in which nobody actually applies the tests of allowability. Most institutions now make the PI's role explicit through training requirements, signed acknowledgements at proposal submission, and periodic certification of expenditures and effort. These are not formalities. They are the institution's way of establishing, for later reference, that the person closest to the costs accepted responsibility for them. A useful way to hold all of this in mind is to imagine the award's financial record being read, years later, by someone who knows nothing about your science and everything about the rules. Every charge should make sense to that reader on its face, or have a note beside it that makes it make sense. Every promise in the proposal should have a matching trace in the records. Everything that follows in this booklet is a way of making that imagined reading uneventful. Chapter 2. Direct Costs: Building a Budget You Can Defend Direct costs are the costs that can be identified specifically with a particular project and assigned to it with a high degree of accuracy. That is the definition in 2 CFR 200.413, and it sounds simple enough: the salary of the graduate student who runs the experiments, the antibodies she uses, the flight to the conference where she presents the results. The complications arise at the edges, where a cost benefits several projects, or where it is of a kind the institution normally treats as overhead, or where the proposal promised one thing and the project needs another. A budget that can be defended is one in which every line has a traceable reason, a basis for its amount, and a plan for how it will be documented when spent. This chapter works through the main categories in roughly the order they appear on a typical federal budget form, with attention to the points where investigators most often go wrong. People: salary, effort and the cap For most research awards, personnel costs are the largest single category, often between half and three quarters of direct costs. They are also the category most exposed to audit, because the link between a salary charge and the work it paid for depends entirely on the institution's system for recording how people spend their time. The basic principle, in 2 CFR 200.430, is that compensation for personal services is allowable to the extent that it is reasonable for the services rendered, conforms to the institution's established written policy applied consistently to federal and non-federal activities, and is supported by records that accurately reflect the work performed. The share of an individual's salary charged to an award must match the share of that individual's total professional effort devoted to the award. The key word is total. Effort is expressed as a percentage of everything a person does for the institution in return for their institutional base salary, including research, teaching, clinical work, administration, committee service and the writing of new proposals. It is not a percentage of a forty-hour week. A faculty member who works sixty hours a week and spends fifteen of them on a particular project is devoting 25 percent effort to it, not 37.5 percent. Effort is also not the same thing as time in the lab: a PI who spends the morning analysing data for the project at home and the afternoon teaching is dividing effort between the project and instruction, wherever the work happens. Institutional base salary is the annual compensation the institution pays for an individual's appointment, whether that time is spent on research, teaching or other activities. It excludes income earned outside the institution, such as consulting or honoraria, and for most universities excludes supplements that are not part of the base appointment. Institutional base salary cannot be increased because a grant is paying for it. The rate used to charge a federal award must be the same rate the institution pays for all the individual's other duties. For NIH awards, an additional limit applies. Since the 1990s, appropriations acts have limited the direct salary that NIH grants may pay to any individual to the rate of Level II of the federal Executive Schedule. For 2026 that figure is $228,000 a year, up from $225,700 in 2025. The cap is not a limit on what an investigator may earn. It is a limit on the rate at which NIH will reimburse that salary. A faculty member whose institutional base salary is above the cap can still devote any percentage of effort to an NIH project, but the salary charged to the grant is calculated as that percentage of the capped rate, and the institution must pay the difference from non-federal funds. The worked example in Table 1 shows how the arithmetic runs for a hypothetical investigator. Table 1. Applying the NIH salary cap to an investigator above the cap (illustrative). Item Amount Institutional base salary (12-month) $260,000 Committed effort on NIH award 25% Salary that effort represents $65,000 2026 NIH cap (Executive Level II) $228,000 Salary chargeable to the award (25% of cap) $57,000 Difference paid from institutional funds $8,000 The $8,000 difference in that example is not cost sharing in the formal sense unless the institution chooses to treat it as such, but it is a real cost the department must fund, and a PI above the cap who commits heavy effort to multiple NIH awards can create a significant liability for the department. For faculty on nine-month appointments the cap is prorated, which for 2026 gives $171,000 for the academic year. NSF takes a different approach. Its policy, set in the PAPPG, is that salary for senior personnel is normally limited to no more than two months of their regular salary in any one year, counting all NSF-funded grants together. Compensation above that amount must be disclosed in the proposal budget and justified, and specifically approved in the award. The two-month rule is the source of a common confusion: it limits the salary NSF pays, but it does not limit the effort a faculty member devotes during the academic year, which may be substantial and uncompensated by NSF. Fringe benefits follow salary. Most universities charge fringe benefits through a negotiated fringe benefit rate, applied as a percentage of salary to each salary charge, with different rates for different employee categories, such as faculty, staff, postdocs and students. Because the rate is negotiated with the federal government, the PI does not choose it. The budgeting task is simply to apply the correct rate to each category and to remember that when salary moves, fringe moves with it. Graduate students present particular issues. Their compensation typically consists of a stipend or salary, tuition remission, and sometimes health insurance. Tuition remission is allowable as part of compensation where it is consistent with institutional policy and the student is performing necessary work on the project. For NIH research grants, the combined compensation for a graduate student, including salary, fringe and tuition remission, is limited to the amount paid to a first-year postdoctoral scholar at the NIH training stipend level in effect when the grant is awarded. Tuition remission is excluded from the base on which indirect costs are calculated, which is one reason its budgeting treatment varies between institutions. Materials, services and travel Materials and supplies are allowable when they are necessary to carry out the project, and the costs should be net of any discounts or rebates. The principal budgeting question is specificity. A proposal line that reads "laboratory supplies, $40,000" is less defensible than one that explains the expected consumption of reagents, animals, sequencing kits and consumables per experiment and multiplies by the number of experiments planned. Agencies rarely challenge supply budgets at proposal time, but a well-reasoned budget justification becomes a useful document later, when the PI needs to show that spending was consistent with the approved plan. Supplies purchased for general use across a lab must be allocated among the projects that benefit. A case of pipette tips used by every project in the lab should not be charged entirely to one grant. Most institutions allow reasonable allocation methods, such as distribution by headcount on each project or by proportion of use, provided the method is documented and applied consistently. The failure mode is the lab in which each month's purchases go to whichever award has the most remaining balance, which is a textbook allocability violation even if every item was genuinely used for research. Computing devices merit a specific note. Under 2 CFR 200.453, computing devices, meaning machines used to acquire, store, analyse, process and publish data, are allowable as supplies when they cost less than the equipment threshold, provided they are essential and allocable to the award, even if not solely dedicated to it. A laptop used by a graduate student for data analysis on the project is therefore generally allowable as a direct cost. A laptop for the PI that will be used for teaching, email and administration as much as for research is harder to defend, and many institutions require a specific justification or a partial allocation. Services, including core facility charges, sequencing, animal care per diems and analytical services, are allowable when necessary and when the rates charged are consistent. University core facilities that bill internal users are service centres, and they must set their rates to recover no more than their actual costs over time and charge all users, federal and non-federal, the same rate. A PI does not need to understand the rate-setting method, but should know that core rates are themselves subject to federal rules and that it is not permissible to negotiate a special internal rate for a federally funded project. Travel is allowable for employees on official business related to the award, under the institution's written travel policy, provided costs are reasonable and consistent with what the institution normally allows. The Uniform Guidance requires that airfare be the lowest reasonable commercial fare, typically economy, and that higher classes be justified by specific circumstances such as medical need or the unavailability of economy seats that would require travel at unreasonable hours. Foreign travel on federal awards is subject to the Fly America Act, which requires the use of US flag air carriers, or foreign carriers operating under code-share arrangements with them, with limited exceptions. Many agencies require that foreign travel be identified in the proposal. Conference registration fees are allowable where attendance is related to the project, and meals included in registration should be deducted from any per diem claimed. Participant support and other special categories Participant support costs are direct costs for items such as stipends, subsistence allowances, travel allowances and registration fees paid to or on behalf of participants or trainees, but not employees, in connection with conferences, workshops or training projects. They are defined at 2 CFR 200.1 and carry two rules that surprise PIs. First, they are excluded from the base on which indirect costs are calculated. Second, NSF and several other agencies require prior approval to rebudget funds out of the participant support category into other categories. The logic is that participant support is money the sponsor intended to pass through to people outside the institution, and the sponsor wants to control whether it is diverted. A PI who plans a summer workshop, reduces it in scale, and spends the savings on a postdoc without approval has made an unallowable rebudget, however sensible the decision was scientifically. Consultant costs are allowable for professional services by individuals who are not employees of the institution, where the services are necessary and the fees are reasonable. The consultant must be distinguished from a subrecipient, a distinction considered in Chapter 6, and from an employee. A faculty member at the same institution is not normally eligible to be paid as a consultant on an award from the same institution, because the compensation belongs within their institutional base salary. Publication costs, including page charges and open-access fees, are allowable where they relate to the project. Following changes in federal public access policy, agencies increasingly expect investigators to budget for making publications and data available. NIH's public access policy, updated to require that accepted manuscripts be made publicly available on PubMed Central without an embargo, took effect in July 2025; data management and sharing costs have been an explicit allowable budget item at NIH since its Data Management and Sharing Policy took effect in January 2023. Budgeting these costs deliberately, rather than hoping residual funds will cover them, is now a normal part of a defensible budget. Animal and human subject costs, including per diems for animal care and payments to research participants, are allowable, with the latter often requiring particular documentation to protect participant privacy while demonstrating that payments were made. Costs that are usually indirect The trickiest direct-cost questions concern items that are ordinarily part of indirect costs. The general principle is consistency: a cost of a type the institution normally recovers through its F&A rate cannot be charged directly unless the circumstances are genuinely different. The Uniform Guidance addresses the most common case explicitly. Under 2 CFR 200.413(c), salaries of administrative and clerical staff should normally be treated as indirect. Direct charging may be appropriate only if all of four conditions are met: the services are integral to the project or activity; the individuals involved can be specifically identified with the project; the costs are explicitly included in the budget or have the prior written approval of the federal agency; and the costs are not also recovered as indirect costs. "Integral" means essential to the project's success, as in a large multi-site clinical trial that needs a dedicated coordinator, or a centre grant with substantial administrative workload specific to the award. A departmental administrator who spends part of her time processing purchases for the project does not meet the test; those services are what the administrative component of the F&A rate pays for. Similar reasoning applies to office supplies, postage, local telephone service, memberships, and general-purpose software. These are presumed to be indirect unless the project creates a need that is distinct from ordinary departmental operations. A project that runs a mailed survey to ten thousand households can charge postage directly. A lab that sends ordinary correspondence cannot. The practical test a PI can apply is simple. Ask whether the cost would exist if this project did not. If it would, because it is part of running a department or a laboratory in general, it is probably indirect. If it arises only because of the project and scales with the project's activities, it is probably direct. The test is not perfect, but it anticipates most of the questions the sponsored programs office will ask. Writing the budget justification Every federal proposal budget is accompanied by a justification: a narrative explaining each category and the basis for its amount. Investigators often treat it as a formality and fill it with generic text. This is a mistake, because the justification is part of the application the agency approves, and it becomes a reference point for what the institution represented it would do. A good justification names each person, their role, their effort and the basis for their salary. It explains supplies in terms of the experiments that will consume them. It identifies each trip, its purpose and its approximate cost. It names each piece of equipment, explains why it is needed and why existing institutional resources cannot serve. It explains every subaward and what the partner will do. It explains any cost that might look unusual to a reviewer who knows the rules, such as direct-charged administrative support, in terms that show why it meets the applicable test. For NIH modular budgets, which apply to many research grant applications requesting $250,000 or less in direct costs per year and are requested in $25,000 modules, the application does not include a detailed categorical budget, but the institution still expects the PI to have one internally, and the personnel justification remains a required element. A modular budget is a simplification of what is submitted, not of what is spent. A budget built this way does two things. It gives the reviewer confidence that the project has been planned seriously, and it gives the PI, years later, a document that explains why the money was spent as it was. The second function matters more than most investigators realise. When an auditor asks why a cost was incurred, the most persuasive answer is a sentence written before the cost existed, predicting that it would. Chapter 3. Indirect Costs: What the Rate Pays For No part of research finance generates more resentment among scientists than indirect costs. An investigator who wins a grant of, say, $1.5 million in total costs may discover that only $1 million or so is available to spend on the project, with the rest going to the university as "overhead". It is natural to see this as a tax, and natural to suspect that the money funds administrators rather than science. Both perceptions are understandable and both are largely mistaken, and a PI who understands why will make better budgeting decisions and argue more effectively with both the institution and the sponsor. Indirect costs, formally facilities and administrative costs or F&A, are real costs of doing research that cannot practically be assigned to individual projects. The laboratory building in which the research takes place has to be heated, cooled, lit, cleaned, secured and maintained. It has to be paid for, either through depreciation of its construction cost or through interest on the debt used to build it. Its fume hoods, vivarium, shared instruments and hazardous-waste systems have to operate. Someone has to process the purchase orders, run the payroll, maintain the accounting system, prepare the financial reports, ensure compliance with animal welfare and human subjects regulations, and handle export controls and research security. None of this can be charged to a single grant with any precision, but all of it is needed for any grant to be performed. The F&A rate is the mechanism by which each project pays its proportionate share. What goes into the pools Appendix III to the Uniform Guidance, which applies to institutions of higher education, divides indirect costs into two broad groups, each made up of several cost pools. The facilities group includes depreciation on buildings, building improvements and equipment; interest on debt associated with those buildings and equipment; operation and maintenance expenses such as utilities, custodial services, repairs, security, and environmental health and safety; and library expenses. These costs are driven by space and by the physical infrastructure of research. The administrative group includes general administration, meaning institution-wide functions such as the president's office, accounting, payroll, purchasing and legal services; departmental administration, meaning the administrative work performed in academic departments, including a portion of the time of faculty chairs and administrative staff; sponsored projects administration, meaning the offices that handle proposals, awards, compliance and post-award accounting; and student administration and services, to the limited extent those functions benefit research. Each pool is allocated among the institution's major functions, principally instruction, organised research, other sponsored activities and other institutional activities, using bases specified in Appendix III. Building costs are allocated using detailed space surveys that assign each room, by square footage, to the functions it serves. Administrative costs are allocated using expenditure bases. The result, after a long and detailed calculation, is the portion of each pool attributable to organised research. Dividing that total by the organised research base produces the F&A rate for research. Two features of the pool structure are worth noting. First, the administrative component of the rate for universities has been capped at 26 percentage points since 1991, when OMB revised Circular A-21 following controversies over university indirect cost recovery, most prominently at Stanford. Any administrative cost above that cap is unrecovered and must be paid by the institution from other sources. For many research-intensive universities, actual administrative costs substantially exceed the cap. Second, some institutions are entitled to a utility cost adjustment of 1.3 percentage points, added to reflect the higher energy consumption of research space, a provision now carried in Appendix III. The implication for PIs is that most of the rate is facilities. In typical negotiated rates at research universities, the facilities component is often similar to or larger than the administrative component, and the administrative portion is held down by the cap. When faculty imagine that indirect costs fund administrators, they are thinking of a portion of the rate that is, by rule, limited and in practice under-recovered. Rates, bases and the arithmetic of a budget An institution's F&A rate is negotiated with its cognizant agency for indirect costs, which for most universities is either the Department of Health and Human Services, through its Cost Allocation Services, or the Department of Defense, through the Office of Naval Research. The institution prepares a detailed proposal based on a recent fiscal year's actual costs; the cognizant agency reviews and negotiates it; and the result is a negotiated indirect cost rate agreement. The agreement is binding on all federal agencies, which must accept the negotiated rate except where a statute, regulation or agency-level policy approved by OMB provides otherwise, or where a program has a published limit on indirect costs. Rate agreements typically contain several rates: an on-campus research rate, an off-campus research rate, rates for instruction and other sponsored activities, and sometimes special rates for particular facilities. They also specify whether the rates are predetermined, meaning fixed for a period and not adjusted after the fact; provisional, meaning used until final rates are set; fixed with carry-forward, meaning adjusted in a later period for differences between estimated and actual costs; or final. Most research universities operate under predetermined rates negotiated for several years at a time. The rate is not applied to the total budget. It is applied to a base, which for universities is modified total direct costs, or MTDC. Under the definition in 2 CFR 200.1, MTDC consists of all direct salaries and wages, applicable fringe benefits, materials and supplies, services, travel, and up to the first $50,000 of each subaward, regardless of the period of performance of the subawards under the award. It excludes equipment, capital expenditures, charges for patient care, rental costs, tuition remission, scholarships and fellowships, participant support costs and the portion of each subaward in excess of $50,000. The excluded items are either capital costs, which are paid for in full as direct costs or recovered separately through depreciation, costs that pass through the institution without consuming much of its infrastructure, or costs, such as rent for off-campus space, that would duplicate what the facilities component pays for. Table 2 shows how this works for a hypothetical one-year budget at an institution with a 57 percent on-campus research rate. Table 2. Calculating indirect costs on a modified total direct cost base (illustrative, 57% rate). Budget item Direct cost In MTDC base? Amount in base Salaries and fringe benefits $210,000 Yes $210,000 Materials, supplies and services $45,000 Yes $45,000 Travel $8,000 Yes $8,000 Equipment (one instrument) $62,000 No $0 Tuition remission $18,000 No $0 Subaward (first year of $120,000 total) $40,000 First $50,000 only $40,000 Totals $383,000 $303,000 At a 57 percent rate, the indirect costs on that budget are about $172,700, and total costs are about $555,700. The effective rate as a proportion of total direct costs is about 45 percent, lower than the negotiated rate because of the excluded items. In the second year, if the subaward continues with $40,000 more, only $10,000 of it enters the base, because the first $50,000 is counted once across the subaward's life. Investigators sometimes treat MTDC exclusions as a budgeting strategy: shift money into equipment or participant support to reduce indirect costs and free more for direct spending within a fixed total. There is nothing wrong with budgeting for genuinely needed equipment, but a budget distorted to minimise the base invites the question of whether each item is actually necessary and correctly classified. More to the point, the institution's F&A recovery on a given award does not change what the research costs the institution to support. Every exclusion is paid for by someone. Where the PI touches the rate Although no PI negotiates the rate, investigators contribute to the data from which it is built, usually without realising it. The most important contribution is the space survey. Because the facilities component is allocated by square footage, the institution periodically asks departments to report how each room is used: for organised research, instruction, other sponsored activities, departmental administration or other institutional purposes. A laboratory used entirely for federally funded research is assigned to organised research; a room used half for teaching labs and half for research is split. The accuracy of those assignments determines how much of the building's depreciation, interest and operating costs flows into the research rate. A survey that overstates research use inflates the rate and is exactly the kind of error that cognizant agency reviewers and auditors look for. A survey that understates it leaves recoverable costs unclaimed. The PI's role is simply to answer the survey accurately, which usually means knowing which funding sources pay for the people working in each room. The second contribution is the organised research base itself. Salaries and other costs charged to sponsored research form the denominator of the rate calculation. When costs that belong to research are recorded as instruction or departmental activity, or the reverse, the rate is distorted. The same effort records that support salary charges on individual awards therefore also support the institution-wide calculation. This is one reason the 2001 clarification on voluntary uncommitted cost sharing, discussed in Chapter 5, mattered so much to universities: it confirmed that uncommitted faculty effort did not have to be added to the research base, where it would have diluted the rate. The third is timing. Rate agreements run for fixed periods, and proposals are budgeted using the rates in effect, or expected to be in effect, during each year of the project. Under the Uniform Guidance and long-standing federal policy, the rate in effect at the time of the initial award generally applies for the life of that competitive segment, even if the institution negotiates a new rate in the meantime. A PI planning a multi-year budget should use whatever rate the sponsored programs office provides for each year and should not assume that the rate on a colleague's older award applies to a new one. On-campus, off-campus and the Columbia lesson The distinction between on-campus and off-campus rates has produced one of the most instructive enforcement cases in research finance. The on-campus rate applies to research performed in facilities owned by the institution or for which it pays depreciation, use charges or rent that is included in its facilities costs. The off-campus rate, much lower because it omits most facilities costs, applies when the research is performed in space the institution does not own and whose costs are not in its pools, such as leased space whose rent is charged directly, or facilities owned by another entity. In July 2016, Columbia University agreed to pay $9.5 million to settle a False Claims Act case concerning this distinction. According to the US Attorney's Office for the Southern District of New York, Columbia had applied its on-campus rate of 61 percent, rather than its off-campus rate of 26 percent, to some 423 NIH grants for research performed between 2003 and 2015 in buildings owned by New York State and New York City, principally at the New York State Psychiatric Institute. The government alleged that Columbia did not own or operate those facilities, did not for most of the period pay the state for their use, and did not disclose the non-ownership to NIH. For a PI, the lesson is not that one should memorise the rate agreement. It is that the location of the research is a representation. Proposals ask where the work will be performed. If the answer changes during the project, because a lab moves into leased space, a clinical study is performed at a partner hospital, or fieldwork occupies most of the award period, the applicable rate may change, and the sponsored programs office needs to know. Many rate agreements also specify that when a project is performed off campus for more than a stated portion of its duration, the off-campus rate applies to the whole project or to the relevant portion. Getting this right is the institution's job, but the PI is the only person who knows where the research is actually happening. The fight over the rate In February 2025, NIH issued a notice, NOT-OD-25-068, announcing that it would cap indirect cost rates for institutions of higher education at 15 percent for new awards and for existing awards going forward, replacing negotiated rates. Had the policy taken effect, research universities whose negotiated rates were typically in the 50 to 65 percent range would have lost a large fraction of their indirect cost recovery on NIH awards almost overnight. The response was immediate. Associations of universities and medical colleges, joined by a coalition of state attorneys general, sued in the US District Court for the District of Massachusetts. The court issued a temporary restraining order within days, then a preliminary injunction, and in April 2025 a final judgment vacating the policy. The plaintiffs' central argument was that the policy violated both the Uniform Guidance's requirements for deviating from negotiated rates and the appropriations rider that has, since fiscal year 2018, prohibited HHS from modifying the indirect cost provisions in effect for NIH. In January 2026 the US Court of Appeals for the First Circuit affirmed, holding that Congress had deliberately prevented NIH from displacing negotiated rates in this way. Similar caps announced during 2025 by NSF, the Department of Energy and the Department of Defense were also challenged in court and blocked, and appropriations language for fiscal year 2026 extended protections to additional agencies. The litigation settled the question of whether an agency could impose a flat cap by notice. It did not settle the policy debate about whether the negotiated-rate system is the right one. Critics argue that it is opaque, that it rewards institutions with expensive facilities, and that the government cannot easily see what its indirect payments buy. Defenders argue that the rates reflect audited actual costs and that the alternative is for universities to subsidise federal research even more heavily than they already do through the administrative cap and cost sharing. In 2025 a coalition of higher education associations known as the Joint Associations Group, including the Council on Governmental Relations, the Association of American Universities and others, published an alternative called the Financial Accountability in Research, or FAIR, model. Its central proposal is to replace the single negotiated rate with a structure in which more research-specific support costs would be budgeted and justified project by project, alongside a fixed percentage for general institutional operations. The final version was released in the second half of 2025. Whether it, or some other reform, will be adopted by the government remained an open question as of 2026. For the working PI, the practical point is this. The indirect cost rate on your award is not a number the university chose arbitrarily, and it is not a number you can negotiate away to make your budget more competitive at a federal agency. It is the product of an audited, negotiated process, and it is legally binding. But the system around it is under more pressure than at any point in the past three decades, and a PI who understands what the rate pays for is better placed to take part in the argument about what should replace it. Foundations, industry and limited rates Not every sponsor pays the negotiated rate. Many private foundations cap indirect costs at a fixed percentage, often 10 to 15 percent of direct costs, and some pay none at all. Certain federal programs have statutory or regulatory limits, such as the 8 percent limit on indirect costs for NIH training grants and fellowships of certain kinds, and limited rates for some Department of Agriculture programs. Industry sponsors typically pay a full rate, often higher than the federal research rate, because universities treat industry work as outside the federal negotiation and price it accordingly. When a sponsor pays less than the negotiated rate, the difference is unrecovered indirect cost. The costs themselves do not disappear; the institution pays them from other resources, typically unrestricted funds, endowment income or tuition. Many institutions have policies requiring that proposals to sponsors with low indirect rates be approved at a senior level, precisely because every such award increases the institution's subsidy of research. In some cases, discussed in Chapter 5, unrecovered indirect costs can be counted as cost sharing on federal awards. In all cases, a PI who understands that indirect costs are real costs will see why the university cares so much about the rate on each award, and why a large foundation grant with no overhead is, from the institution's perspective, not free money but a purchase of research partly paid for by the university itself. Hashtags: #ScientificGrantBudgeting #FinancialCompliance #ResearchFinance #GrantManagement #PrincipalInvestigator #SponsoredResearch #UniformGuidance #FederalResearchGrants #DirectCosts #IndirectCosts #FacilitiesAndAdministrativeCosts #CostAllowability #CostAllocability #CostReasonableness #CostSharing #EffortReporting #NIHGrants #NSFGrants #Subawards #BudgetJustification #ResearchAccounting #AuditCompliance #GrantCloseout #ResearchAdministration #FutureOfResearchFinance

  • Research Software Engineering (Version Control, Unit Testing, and Modular Code)

    Download the Book (PDF): Introduction In 2006 the structural biologist Geoffrey Chang and his colleagues retracted five papers, three of them from Science. The papers described the three-dimensional structures of membrane transport proteins, and they had been cited hundreds of times. Other laboratories had built on them, designed experiments around them, and in some cases struggled to reconcile their own data with them. The cause of the collapse was not fraud, not contaminated samples, and not a flawed theory. It was a program written in the laboratory that swapped two columns of data, inverting the sign of a set of measurements before they went into the structure calculation. Years of careful crystallography had been passed through a few lines of code that nobody outside the group had examined and nobody inside it had tested against a case with a known answer. The story is often told as a cautionary tale about one unlucky lab. It is better understood as an ordinary event that happened to be visible. Most research now runs through software. When the Software Sustainability Institute surveyed 417 researchers at fifteen research-intensive UK universities in 2014, 92 per cent said they used research software, 69 per cent said their research would not be practical without it, and 56 per cent said they wrote their own. Of those who wrote their own, around one in five had received no training in software development at all. Surveys in the United States and elsewhere since then have found much the same pattern. The laboratory's most heavily used instrument is frequently the one its members were never taught to build. This booklet is about building that instrument well. Its controlling argument is simple to state and, in practice, surprisingly hard to act on: the engineering practices that make research code trustworthy are the same practices that make it cheap to change, and a researcher who adopts them in small, proportionate steps pays less for reliability than they are already paying, invisibly, for its absence. Version control, automated tests, modular design, continuous integration and containerised environments are usually presented to academics as a tax — a set of professional niceties that software companies can afford and that a PhD student with a deadline cannot. The chapters that follow argue the reverse. These practices are how you stop losing afternoons to the question "which version of the script produced the third plot in the paper?", how you make a reviewer's request for an extra analysis a morning's work rather than a fortnight's, and how you make it possible for a new student to pick up a project when its author leaves. Who this is for The reader I have in mind writes code as part of doing research rather than as the research itself. You might be a doctoral student in ecology whose thesis rests on a few thousand lines of R; a postdoc in physics maintaining a Fortran solver inherited from a predecessor; an economist whose replication package has to satisfy a journal's data editor; or a principal investigator who suspects the group's code is fragile but does not know what to ask for. You can write a loop and a function. You may have used Git, perhaps by copying commands from a colleague. You have probably never written a test on purpose. You do not need to become a professional software engineer. The research software engineering movement — which took its name at a workshop in Oxford in 2012, gained a UK society in 2019, and now has sister organisations in Germany, the Netherlands, the United States, Australia and elsewhere — exists partly so that universities can employ specialists for the hardest problems. But specialists cannot be everywhere, and most research code will always be written by researchers. The goal here is the level of craft that lets your code do what you think it does, lets you prove it, and lets someone else run it in five years. What the booklet covers, and what it leaves out The booklet follows the life of a piece of research code from its first commit to its eventual archiving. The first chapter sets out why research software fails and why those failures are hard to see. Chapters 2 and 3 cover version control with Git: first as a personal laboratory notebook, then as the basis for collaboration through branches, pull requests and review. Chapter 4 turns to the structure of the code itself — how to break a sprawling analysis script into functions and modules that can be understood and tested one piece at a time. Chapter 5 addresses unit testing, with particular attention to the problem that makes scientific testing distinctive: you often do not know the right answer in advance. Chapter 6 covers continuous integration, the practice of having a machine run your tests on every change. Chapter 7 deals with computational environments and containers, including Docker and the HPC-friendly Apptainer. Chapter 8 addresses the long view: documentation, licensing, citation, versioning, and what happens to code when its author moves on. Some things are deliberately absent. There is no tutorial in a specific language; the examples lean on Python and R because those dominate academic computing, but the principles transfer to Julia, MATLAB, C++ or Fortran. There is nothing on high-performance optimisation, GPU programming or parallel computing, each of which deserves its own treatment and none of which matters if the serial code is wrong. Workflow managers such as Snakemake and Nextflow are mentioned where they bear on reproducibility but not taught. And the booklet does not attempt to cover the governance of large community software projects with dozens of contributors; it is aimed at the individual and the small group, where most research code lives. A note on proportion The most common objection to everything in this booklet is that it is overkill. For a thirty-line script that will run once and be discarded, it is. Research code exists on a spectrum: exploratory notebooks, analysis pipelines behind a paper, tools shared within a lab, and libraries used by a community. Each position on that spectrum warrants a different level of engineering, and one of the aims of the chapters that follow is to help you recognise where your code sits and what it needs. A throwaway exploration needs version control and nothing else. An analysis that supports a published claim needs tests for its core calculations and a recorded environment. A tool used by other people needs continuous integration, documentation, versioned releases and a licence. The mistake is not to under-engineer throwaway code. It is to fail to notice when throwaway code has quietly become the foundation of a thesis. That transition happens without announcement, usually around the time the code has grown too large to be rewritten comfortably, and the practices in this booklet are cheapest to adopt just before it. It helps to think of the practices as layered, each one making the next possible. Version control gives you a history you can trust, which makes it safe to restructure code into modules. Modular code can be tested one piece at a time. Tests are what a continuous integration service runs. A recorded environment is what lets those tests, and your results, mean the same thing on another machine. And all of it together is what makes code maintainable by someone other than its author. You can start at the bottom of that stack this afternoon, and each layer you add pays for itself before you need the next. The chapters are written to be read in order for that reason, though each also stands on its own for the reader who arrives with a specific problem. Chapter 1. The Instrument Nobody Calibrated Every experimental scientist learns to calibrate instruments. A mass spectrometer is run against a reference standard before samples go in; a thermometer is checked in ice water; a new antibody is validated against a known positive and a known negative. Nobody regards this as bureaucratic overhead. It is simply what it means to take a measurement seriously. Yet the same researchers routinely pass their data through code that has never been checked against anything, written in a hurry, modified dozens of times, and run on a machine whose configuration nobody could reproduce. Software is an instrument too, and the argument of this chapter is that it fails in ways that are unusually hard to see — which is precisely why it needs the kind of deliberate checking that other instruments receive as a matter of course. How research code goes wrong The failures that make headlines are rarely exotic. They are mundane errors that survived because nothing was in place to catch them. Consider the case of the "Willoughby–Hoye" scripts. In 2014 Patrick Willoughby, Matthew Jansma and Thomas Hoye published a protocol in Nature Protocols for assigning the structures of small molecules by computing NMR chemical shifts, together with a set of Python scripts that automated part of the calculation. The protocol was widely used. In 2019 a graduate student at the University of Hawaii, Yuheng Luo, working with the chemist Rui Sun on the characterisation of compounds from a cyanobacterium, noticed that the scripts gave different answers on different computers. On one machine a computed value came out around 172.4; on another, around 173.2 — enough to change which candidate structure looked correct. The cause, reported in Organic Letters by Jayanti Bhandari Neupane and colleagues, was a single call to Python's glob function, which lists files matching a pattern. The Python documentation states that glob returns results in arbitrary order. On some operating systems the order happened to be alphabetical; on others it was not. The scripts silently assumed it was, and so matched the wrong output files to the wrong inputs. The fix was one line: sort the list. The authors estimated that well over a hundred published studies might have used the scripts. Nothing about this bug was unusual. The code was not badly written by the standards of academic software. It worked, repeatedly and convincingly, on the machines of the people who wrote it. It failed only when run somewhere else, and even then it did not crash — it produced plausible numbers that were wrong. That combination, plausible output from a quiet error, is the signature of most serious research software failures. The same pattern appears in the case with which this booklet opened. Geoffrey Chang's group did not publish nonsense; they published protein structures that were entirely credible to reviewers and to the field until other groups produced structures of related proteins that could not be reconciled with them. The in-house program that flipped two columns had presumably been run many times without complaint. There was no oracle — no case with a known answer — against which its output was ever compared. Spreadsheets deserve mention here, because much research computation still happens in them and they illustrate the same failure mode with particular clarity. In 2010 the economists Carmen Reinhart and Kenneth Rogoff published an influential paper associating high public debt with sharply lower growth. In 2013 Thomas Herndon, Michael Ash and Robert Pollin, working from the spreadsheet the authors had kindly shared, found that a formula averaging across countries omitted several rows, among other issues concerning data exclusion and weighting. The coding error was only one of several disputed choices, and its effect on the headline conclusion has been argued over ever since, but it was the error that captured public attention because it was so easy to understand. A range in a formula stopped a few cells short. Nothing warned anyone. Genomics offers a still more pervasive example. Microsoft Excel, by default, converts certain gene symbols into dates: SEPT2 becomes 2-Sep, MARCH1 becomes 1-Mar. In 2016 Mark Ziemann, Yotam Eren and Assam El-Osta screened supplementary files attached to papers in leading genomics journals and found that roughly a fifth of the papers with Excel gene lists contained such corrupted names. The problem was persistent enough that in 2020 the HUGO Gene Nomenclature Committee renamed the affected genes — SEPT1 became SEPTIN1, MARCH1 became MARCHF1 — partly to stop software mangling them. When a scientific naming authority changes the names of human genes to accommodate a spreadsheet's behaviour, the scale of the problem is hard to dispute. Why the errors stay hidden It would be comforting to think that these are the failures of careless people. They are not. They are the predictable result of three features of research software that distinguish it from most other code. The first is that research code often has no independent specification. A payroll system can be checked against the law and the contract: if an employee is paid the wrong amount, someone notices. Research code frequently computes something that nobody has computed before, which is the point of the research. If the output looks surprising, that may be a discovery or a bug, and the researcher has strong incentives to hope it is the former. If the output looks unsurprising, nobody checks it at all. The absence of a known right answer is what makes scientific testing distinctive, and Chapter 5 is largely about techniques for testing when the answer is unknown. The second is that research code evolves by accretion. A typical analysis starts as a short script to load some data and make a plot. It gains a cleaning step, then a special case for a malformed file, then an extra model, then a flag to switch between two versions of a preprocessing step because a reviewer asked. Each change is small and sensible. After two years the script is two thousand lines long, contains several abandoned approaches that are commented out or still run but whose output is ignored, and relies on global variables set at the top that are silently changed halfway down. Nobody designed it. Its structure is the fossil record of the project's history, and it is almost impossible to reason about any one part of it in isolation. Chapter 4 addresses how to prevent and undo this. The third is that research code runs in an environment nobody records. The result of an analysis depends not only on the code but on the versions of the language, the libraries, the operating system and occasionally the hardware. Default arguments change between library versions. Random number generators change their algorithms. A function deprecated in one release is removed in the next. The glob bug was an environmental failure in exactly this sense: the code was identical everywhere, and the environment determined the answer. Chapter 7 deals with capturing environments so that they can be reconstructed. These three features compound. Code with no specification is hard to test; code that grew by accretion is harder still; and code whose behaviour depends on an unrecorded environment may pass a test on one machine and fail it on another. The result is software whose correctness rests almost entirely on the confidence of its author. What the evidence says about reproducibility Anecdotes can mislead, so it is worth asking how common these problems are in the aggregate. Several large studies have tried to find out, and their results are sobering. In 2022 Ana Trisovic, Matthew Lau, Thomas Pasquier and Mercè Crosas reported in Scientific Data on an attempt to re-execute research code deposited in the Harvard Dataverse repository. They collected more than 9,000 R files from over 2,000 replication datasets published between 2010 and 2020 and ran them in a clean environment. Seventy-four per cent of the files failed to complete without error on the first attempt. After the authors applied automated cleaning — fixing common problems such as hard-coded absolute file paths and missing library calls — 56 per cent still failed. These were not random scripts found on the internet; they were code that researchers had deliberately packaged and deposited so that others could reproduce their work, often to satisfy a journal's policy. Earlier, Christian Collberg and Todd Proebsting at the University of Arizona attempted to build the software described in several hundred papers from computer systems conferences and journals — a field whose researchers are, if anyone is, professional programmers. They reported in Communications of the ACM in 2016 that for a substantial fraction of papers the code could not be obtained, and that of the code they did obtain, a large share could not be built without significant effort. Their project became a widely cited illustration that sharing code is not the same as sharing working code. Studies of Jupyter notebooks tell the same story. In 2019 João Felipe Pimentel and colleagues analysed over a million notebooks on GitHub and found that only a small minority could be re-executed to produce the same results, with failures caused by missing dependencies, cells run out of order, and references to files that did not exist in the repository. The notebook format encourages running cells in whatever order is convenient during exploration, which means that the state of the notebook at the moment its author saved it may be unreachable by running it top to bottom. Pimentel and his co-authors put numbers on this: of the notebooks they could attempt to run, only about a quarter executed without error, and only around four per cent reproduced the outputs stored in them. The common thread is that code which works on its author's machine, today, frequently does not work anywhere else or at any later date. This is not a moral failing. It is the natural state of software that has never been run anywhere but where it was written. The case for treating code as a method There is a useful reframing that runs through the rest of this booklet. When a paper describes an experimental procedure, the methods section is supposed to contain enough detail for a competent colleague to repeat it. Journals, funders and reviewers take this seriously, at least in principle. The code that transforms raw data into reported results is part of the method, often the most consequential part, and for most of the history of computational science it was described, if at all, in a sentence: "analyses were performed in R version 3.4 using custom scripts." That sentence is not a method. It is a promissory note. The practices in the chapters that follow are, collectively, how to turn it into something a colleague could actually use. Version control records exactly which code produced which result. Tests are the calibration record — the evidence that the instrument reads correctly against known standards. Modular structure makes the procedure legible, so that a reviewer can find and check the step that matters. Continuous integration shows that the tests pass on a machine other than the author's. A recorded environment specifies the conditions under which the procedure was carried out. Documentation and a licence make it possible for someone else to use it lawfully and correctly. The same reframing explains why these practices are not merely defensive. A laboratory with calibrated instruments runs experiments faster, not slower, because it does not have to rerun them when a result looks wrong. A researcher with tested, versioned, modular code can respond to a reviewer's request by changing one parameter and rerunning a pipeline, confident that nothing else has moved. The cost of engineering is paid once; the cost of its absence is paid every time the code must change. When the code is the public argument The cases above concern errors. A different kind of episode shows what happens when research code, even correct code, is exposed to scrutiny it was never built to withstand. In March 2020 the epidemic model developed by Neil Ferguson's group at Imperial College London informed the United Kingdom's decision to move towards a national lockdown. The model's code had been developed over many years, originally in C, as a single large file of many thousands of lines. When the group released a cleaned-up version on GitHub in the spring of 2020, working with software engineers from Microsoft and GitHub to restructure it, the code attracted intense public criticism. Commentators — some of them professional programmers, some with political motives — pointed to its size, its lack of tests, and reports that it could produce different results from the same inputs under some configurations. The debate was heated and often ill-informed about how stochastic simulations work; run-to-run variation is expected in such models, and the question is whether it is controlled and understood. The more careful response came from the reproducibility community. Stephen Eglen of the University of Cambridge, through the CODECHECK initiative, independently ran the released code and in June 2020 reported that he could reproduce the key results in the relevant Imperial report, within the variation expected from a stochastic model. The code, in other words, did what its authors said it did. But the episode showed how much credibility depends on things other than correctness. A model that has tests, a documented history, a declared environment and a structure that outsiders can follow is far easier to defend than one whose correctness must be established after the fact, under public pressure, by a volunteer. The Imperial team's code was vindicated; it would have been better for everyone had it never needed to be. The lesson for an ordinary researcher is not that their thesis code will be debated in newspapers. It is that code which cannot be inspected easily will be trusted, or distrusted, on grounds that have nothing to do with whether it is right. Engineering practices are, among other things, a way of making the evidence of correctness available before anyone asks for it. The costs researchers already pay That last claim deserves defending, because it is the heart of the booklet's argument and because it cuts against a strong intuition. Researchers tend to believe that engineering practices are expensive and their absence is free. In fact the absence has costs that are real but diffuse, and so go unaccounted. There is the cost of archaeology: the hours spent working out which of analysis_final.R, analysis_final_v2.R and analysis_final_REALLY.R produced the figures in the submitted manuscript, and whether the version on the laptop matches the version on the cluster. There is the cost of fear: the reluctance to improve code because any change might break something in a way that will not be noticed, which causes code to ossify around its earliest and worst decisions. There is the cost of revision: when a reviewer asks for an analysis with a different exclusion criterion, the researcher with a monolithic script must trace every place the criterion is used, while the researcher with a modular pipeline changes a configuration value. There is the cost of succession: when a doctoral student graduates, their code often becomes unusable within months, and the next student starts again from nothing, as if the previous years of work had been lost in a fire. And there is the tail risk: the small but real chance of an error that reaches publication, with consequences ranging from an embarrassing correction to a retraction and years of reputational damage. None of these costs appears on a budget line, which is why they are so easy to ignore. But anyone who has spent a week reconstructing an analysis to answer a reviewer, or inherited a predecessor's code and given up on it, has paid them. The chapters that follow are, in effect, an argument that the same effort spent up front buys a good deal more. Where to begin If this chapter has been persuasive, the natural temptation is to try to adopt everything at once. That rarely works. Engineering practices stick when each one solves a problem the researcher already feels, and when it is introduced at a moment when the cost of adoption is low. The order of the chapters reflects a sensible order of adoption. Version control comes first because it is cheap, immediately useful, and makes every subsequent change safe. Modular structure and tests come next, because they are what make the code trustworthy. Automation and environments follow, because they extend that trust beyond the author's own machine. Long-term maintainability comes last, because it matters most once the code has other users. The next chapter starts where every research project should: with a history of the code that you can trust. Chapter 2. Version Control as a Laboratory Notebook Every researcher has seen the folder. It contains model.py, model_old.py, model_new.py, model_new_fixed.py, model_thesis.py and model_thesis_JB_comments.py, along with a directory called backup_march whose relationship to the others is unclear. The folder is a version control system: an improvised, manual, unreliable one. Its owner knows that the history of the code matters and has tried to preserve it. What the folder cannot do is say what changed between any two files, why, or which one produced the result in the paper. Version control software does exactly those things, and the dominant tool, Git, is now so widely used that learning it is less a choice than a basic literacy. This chapter treats Git not as a tool for software teams but as what it can be for an individual researcher: a laboratory notebook for code, recording what was done, when and why, in a form that cannot be quietly altered and can always be returned to. What Git actually records Git was written by Linus Torvalds in 2005 to manage the development of the Linux kernel, after the project lost access to the proprietary tool it had been using. Its design reflects that origin: it is distributed, meaning every copy of a repository contains the full history, and it is built around content rather than files. Understanding a small amount of how it works makes the rest of it far less mysterious. A Git repository is a directory whose history Git tracks. The unit of history is the commit: a snapshot of the state of every tracked file at a moment in time, together with the name of the person who made it, a timestamp, a message describing it, and a pointer to the commit or commits that came before. Each commit is identified by a hash — a long string of hexadecimal characters computed from its contents and its ancestry. Because the hash depends on everything that came before, altering any past commit changes its hash and the hash of every commit after it. That property is what makes the history trustworthy: a commit hash cited in a paper identifies exactly one state of the code, and nobody can change that state without the change being detectable. Between the files on disk and the committed history sits the staging area, sometimes called the index. When you run git add on a file, you are not saving it to history; you are placing its current contents into the next commit you are assembling. git commit then records the staged snapshot. Newcomers often find the extra step irritating, but it has a real purpose for researchers: it lets you commit one logical change even when you have made several unrelated edits. If you fixed a bug in the data loader and also started experimenting with a new plotting style, you can stage and commit the bug fix alone, with a message that describes it, and leave the experiment for later. The last concept needed for solo work is the working tree — the files as they currently sit on your disk. git status compares the working tree and the staging area against the last commit and tells you what has changed. git diff shows the changes line by line. git log shows the history. With these five commands — status, add, commit, diff and log — a researcher already has most of the value of version control. Commits as notebook entries A laboratory notebook is useful to the extent that its entries are made at the right moments and say the right things. The same is true of commits. The right size for a commit is one coherent change: a bug fixed, a function added, a parameter changed, a figure restyled. A commit that does one thing can be understood, reviewed, and if necessary reversed on its own. A commit that bundles a week of work — "lots of changes" — is almost as unhelpful as no commit at all, because when something breaks it offers no way to find which of the many changes was responsible. Committing little and often is the single most important habit. Most experienced users commit several times an hour when actively working. The message is the notebook entry proper. Its first line should be a short summary, conventionally under about fifty characters, written as an instruction: "Fix off-by-one error in sliding window," "Add bootstrap confidence intervals to summary table." After a blank line, a longer body can explain why the change was made, which is the part future readers most need and the code itself cannot tell them. The code shows that the exclusion threshold changed from 0.05 to 0.01; only the message can record that this followed a discussion with a co-author about the reviewer's second comment. The diff is the what; the message is the why. There is a particular kind of commit that researchers should learn to make deliberately: the commit that corresponds to a result. When you generate the figures for a manuscript submission, commit first, and then record the commit hash alongside the output — in the figure's metadata, in a log file, or in the lab's own notes. Git's tags exist for this purpose. A tag is a human-readable name attached permanently to a commit: submitted-2026-03, revision-1, thesis-final. Six months later, when the reviewer's report arrives, git checkout submitted-2026-03 restores the exact code that produced the submitted results, and git diff submitted-2026-03 shows everything that has changed since. Making provenance automatic Recording hashes by hand works until the day it is forgotten, which is usually the day it matters. A more robust habit is to have the analysis record its own provenance. Both Python and R can ask Git for the current commit hash when a script starts — in Python by calling git rev-parse HEAD through the subprocess module, in R through system() or the gert package — and write it into every output: a line in a log file, a field in the metadata of a saved results table, or a small text label in the corner of a draft figure. The same call can check whether the working tree is clean. If there are uncommitted changes, the script can print a loud warning or refuse to run in "production" mode, because an output generated from uncommitted code cannot be traced to any state in the history. This small piece of machinery closes a gap that otherwise stays open for the life of a project. Every figure and table can be traced to exactly one commit, and every commit can be restored. When a co-author asks, eighteen months later, whether the numbers in Table 2 of the thesis were computed before or after the fix to the outlier filter, the answer is a lookup rather than an argument. A related discipline concerns rewriting history. Git allows local commits to be amended or reorganised before they are shared — git commit --amend fixes the message or contents of the most recent commit, and an interactive rebase can squash several work-in-progress commits into one. This is harmless and often tidy as long as the commits exist only on your machine. Once commits have been pushed to a shared remote, and especially once a hash has been cited anywhere, they should be treated as permanent. Rewriting published history breaks the guarantee that a hash identifies one state of the code, which is the whole point of recording it. What belongs in the repository A question that trips up many researchers is what to put under version control. The short answer is: everything a human writes, and nothing a machine generates or that is too large or sensitive to share. Source code belongs, obviously. So do configuration files, documentation, the text of the paper if it is written in LaTeX or Markdown, small hand-curated reference files, and the scripts that download or generate everything else. The file that records the software environment — requirements.txt, environment.yml, renv.lock or equivalent — belongs, for reasons Chapter 7 explains. Generated outputs generally do not belong. Figures, intermediate data, compiled binaries and cached results can be regenerated from the code, and committing them clutters the history and invites conflicts. There are sensible exceptions — a small table of final results kept under version control makes it easy to see when a code change alters a number — but the default should be to exclude them. Large raw data does not belong in a Git repository. Git stores every version of every file forever, and a repository containing a few gigabytes of binary data becomes slow to clone and painful to use. Data should live in an appropriate data repository or institutional store, with the repository containing a script that fetches it and, ideally, a checksum that verifies it. Tools such as Git LFS (Large File Storage), git-annex and DVC (Data Version Control) exist to bridge this gap for projects where data and code must be versioned together; they store large files elsewhere and keep lightweight pointers in Git. Secrets must never enter the repository: passwords, API keys, access tokens, and any personal or confidential data. Because Git preserves history, a secret committed once and deleted in the next commit is still present in the repository, and anyone with a copy can retrieve it. Removing it properly requires rewriting history, which is disruptive and, once the repository has been pushed to a shared service, may already be too late. Credentials should live in environment variables or local configuration files excluded from version control. The same caution applies with extra force to research data governed by ethics approvals or data protection law: participant data committed to a public repository is a reportable breach, not a technical inconvenience. The mechanism for exclusion is a .gitignore file in the repository's root, listing patterns for files Git should never track: data/raw/, .env, *.pdf in an output directory, the pycache folders Python creates, the .Rhistory file R leaves behind. Setting it up at the start of a project prevents a great deal of trouble later. Templates for common languages are widely available. Remotes, backups and sharing Git on a single laptop is a notebook that can be lost with the laptop. Its value multiplies when the repository is pushed to a remote — a copy hosted elsewhere, most often on GitHub, GitLab, Bitbucket, or an institution's own GitLab server. git push sends local commits to the remote; git pull fetches and merges commits from it. For an individual researcher, the remote is first of all a backup, and a better one than most, because it preserves not just the latest state but the entire history. It is also how code moves between machines: write on the laptop, push, pull on the cluster, run. This alone eliminates a class of errors in which the version on the cluster and the version on the laptop diverged without anyone noticing. The choice of host has practical consequences. GitHub is the largest and has the richest ecosystem of integrations, including the continuous integration service discussed in Chapter 6 and the Zenodo integration discussed in Chapter 8. GitLab can be self-hosted, which many universities do, keeping code on institutional infrastructure — sometimes a requirement for sensitive projects. Repositories can be private until publication and made public afterwards; there is no need to expose work in progress. What matters most is that the remote exists and that pushing to it becomes habitual. Using the history A history is only valuable if you use it, and Git offers tools that make the history an active aid to research rather than an archive. The simplest is reading it. git log with a file name shows every commit that touched that file; git log -p shows the changes themselves. When a number in the output changes unexpectedly, the log of the relevant files is the first place to look. git blame, despite its name, is not about assigning fault. It annotates each line of a file with the commit that last changed it. When you encounter a strange line — a magic constant, an unexplained special case — blame leads you to the commit that introduced it, and a well-written commit message explains why it is there. In inherited code this is often the only surviving explanation of a decision. The most powerful tool for researchers is git bisect. Suppose a result that was correct in January is wrong in June, and there are two hundred commits in between. bisect performs a binary search over the history: you mark a known good commit and a known bad one, Git checks out the commit halfway between, you test it and report good or bad, and Git halves the range again. After about eight steps — two to the eighth is 256 — it identifies the exact commit that introduced the problem. If the test can be scripted, git bisect run automates the whole search. This is only possible if commits are small and each one leaves the code in a runnable state, which is another reason to commit carefully. Finally, the history makes undoing things safe. git restore discards uncommitted changes to a file. git revert creates a new commit that undoes an earlier one, preserving the record that both happened. git checkout or git switch moves the working tree to any past commit or branch. The existence of these commands changes how people work: when every state can be recovered, experimentation stops being risky, and the collection of old and backup files becomes unnecessary. Notebooks, binary files and the limits of Git Git works best with plain text, because its diffs are line-based. Several formats common in research sit awkwardly with it. Jupyter notebooks are stored as JSON files that mix code with outputs, including embedded images and execution counters. A notebook whose code is unchanged but which has been rerun produces a large, meaningless diff. Merging two people's changes to the same notebook is often impractical. Several remedies exist. Tools such as nbstripout remove outputs before committing, so that only the code and text are tracked. Jupytext pairs each notebook with a plain-text script version — a Python file or an R Markdown file — that diffs cleanly. Increasingly, researchers use notebook formats that are plain text from the start, such as Quarto documents or R Markdown. More fundamentally, the analytical core of a project should not live in a notebook at all; notebooks are excellent for exploration and presentation, but the functions they call belong in modules that can be tested, as Chapter 4 argues. Word documents, Excel files, images and compiled binaries can be committed, but Git can only record that they changed, not how. For manuscripts written in Word, Git adds little beyond backup. For spreadsheets used as data entry, exporting to CSV alongside the original makes changes visible. Starting today Version control is the rare practice whose benefits begin immediately and whose cost is almost entirely a matter of habit. For an existing project, the steps are: initialise a repository in the project folder, write a .gitignore that excludes data, outputs and secrets, make a first commit of the current state, and push it to a private remote. From that moment, every change is recorded. The _old files can be deleted once they are safely committed, because the history now does their job better. The discipline that follows is small: commit whenever a coherent piece of work is done, write a message that says why, tag the commits that produce results, and push at the end of every session. Within a few weeks this becomes automatic, and it becomes difficult to remember how research was done without it. The next chapter builds on this foundation to address what happens when a project has more than one line of development — and more than one person. Chapter 3. Branches, Merges and the Shape of Collaboration A single line of history is enough for a researcher working alone on one thing at a time. Research is rarely like that. A doctoral student is halfway through restructuring the data pipeline when a supervisor asks for a quick variant of last month's figure. A postdoc wants to try a new estimation method without disturbing the version that the rest of the group depends on. Two students need to modify the same simulation code for different projects. A collaborator in another institution sends a fix. Each of these situations involves more than one line of development at once, and Git's answer to all of them is the branch. This chapter explains what branches are, how the main branching strategies used in industry differ, and which of them suits the scale and rhythm of academic work. Its argument is that most research groups should adopt a deliberately simple strategy — short-lived branches merged into a single main line through review — and resist the more elaborate models that were designed for problems researchers do not have. What a branch is In Git, a branch is nothing more than a movable label pointing at a commit. When you make a new commit on a branch, the label moves forward to point at it. The default branch in a new repository has historically been called master; since 2020, GitHub, GitLab and Git itself have moved to main as the default name, and this booklet uses main throughout. Creating a branch — git switch -c new-estimator — makes a new label pointing at the current commit. Commits made from then on advance the new label while leaving main where it was. You can switch back to main at any time, and the working tree returns to the state it was in before the experiment began. The two lines of development coexist, each with its own history, sharing everything up to the point where they diverged. Because branches are just labels, they are cheap. Creating one takes no time and no disk space. This matters because it changes behaviour: when branching is free, it becomes natural to start every piece of non-trivial work on its own branch, keeping main in a known, working state at all times. Merging, and what conflicts mean Eventually a branch's work either gets abandoned — in which case the branch is simply deleted — or brought back into main. Bringing it back is a merge. If main has not moved since the branch was created, Git can simply advance the main label to the branch's latest commit, which is called a fast-forward merge. If both have moved, Git creates a new merge commit with two parents, combining the changes from each. Git merges changes line by line. When the two branches modified different files, or different parts of the same file, the merge is automatic. When both changed the same lines in different ways, Git cannot know which version is right, and it reports a conflict: it marks the disputed region in the file with both versions and stops, asking a human to decide. Resolving a conflict means editing the file to the correct combined state, staging it, and completing the merge. Conflicts intimidate newcomers, but they are information, not failure. A conflict says that two lines of work made incompatible assumptions about the same piece of code, and a person needs to reconcile them. The practical lesson is that conflicts grow with the length of time branches live apart. A branch merged after a day rarely conflicts; a branch merged after three months almost always does, and the conflicts may be difficult to resolve because nobody remembers the reasoning behind either side. This observation, more than any other, shapes the choice of branching strategy. An alternative to merging is rebasing, which replays a branch's commits on top of the latest main as if the work had been started from there, producing a linear history without merge commits. Rebasing produces tidier histories and is popular in some teams. Its cost is that it rewrites the branch's commits, giving them new hashes, which causes trouble if anyone else has based work on them. For research groups the simple rule is: rebase your own unpublished branches if you like tidiness; never rebase anything others are using. Three strategies compared A branching strategy is a convention about which branches exist, how long they live, and how work flows between them. Three strategies dominate discussion, and they differ chiefly in how many long-lived branches they maintain and how quickly work returns to the main line, as Table 1 summarises. Table 1. Three common branching strategies compared. Strategy Long-lived branches Typical branch lifetime Designed for Fit for research groups Git Flow main, develop, plus release and hotfix branches Days to weeks Versioned releases supporting several versions at once Poor for most; useful only for large libraries with formal releases GitHub Flow main only Hours to days Continuously deployed services Good default for groups and shared tools Trunk-based development main only (the trunk) Hours, or direct commits Large teams with strong automated testing Good for solo work and mature, well-tested codebases Git Flow was described by Vincent Driessen in a 2010 blog post, "A successful Git branching model," which became one of the most widely copied pieces of writing about Git. It uses two permanent branches: main, which holds only released versions, and develop, where integration happens. Features are built on branches taken from develop; when a release is due, a release branch is cut, stabilised and merged into main with a version tag; urgent fixes to released versions go on hotfix branches taken directly from main. The model is thorough and handles the case of shipping numbered versions of software while maintaining older ones. It is also heavy. In 2020 Driessen added a note to the original post observing that for software delivered continuously, such as web applications, a simpler workflow like GitHub Flow was often more appropriate, and that Git Flow should not be treated as dogma. For a research group writing analysis code, the ceremony of release branches and a separate develop branch adds overhead with little benefit. GitHub Flow, described by Scott Chacon of GitHub in 2011, strips the model down to one rule: main is always in a working state. Any change, however small, is made on a descriptively named branch taken from main. When it is ready, the author opens a pull request, the change is discussed and reviewed, automated tests run, and the branch is merged into main and deleted. There are no develop or release branches. When a released version needs marking, it is tagged on main. Trunk-based development goes further still: developers commit to the main branch (the "trunk") directly or through branches that live only hours. It depends on strong automated testing, so that a broken commit is detected within minutes, and often on techniques such as feature flags that let incomplete work be merged without being switched on. Large software organisations use it at scale. For a single researcher working alone, committing directly to main is effectively trunk-based development, and it is perfectly reasonable as long as tests exist. A strategy fitted to research For most research groups, a lightly adapted GitHub Flow is the right default. It works like this. The main branch always contains code that runs and whose tests pass. Nobody commits broken work to it. Every piece of work — a new analysis, a bug fix, a refactoring — happens on a short-lived branch with a name that says what it is for: fix-date-parsing, add-mixed-model, reviewer2-sensitivity. When the work is done, it goes back into main through a pull request, ideally reviewed by someone else, with the automated tests described in Chapter 6 running first. The branch is then deleted. Results intended for publication are generated from main and tagged. Research does add one wrinkle that software teams rarely face: exploratory branches that may never be merged. A student may try three approaches to a modelling problem, each on its own branch, and adopt only one. This is a legitimate use of branches, and the abandoned ones need not be deleted immediately; they are a record of what was tried. But they should be recognisable as experiments, perhaps with a prefix such as explore/, and their conclusions should be written down somewhere more durable than a branch name before they are removed. A branch is not a lab notebook entry; a commit message or a short note in the project's documentation is. The other research-specific hazard is the branch that lives for a whole thesis. A student takes a copy of the group's shared model, makes changes for their own project over three years, and never merges anything back. By the time they graduate, their version and the group's have diverged so far that reconciling them is impossible, and the student's improvements are lost. The remedy is the same as for any long-lived branch: merge small, generally useful changes back to main promptly, and keep project-specific changes in project-specific code that uses the shared model rather than modifying it. Chapter 4's discussion of modularity is largely about making that separation possible. Pull requests and code review A pull request (GitLab calls it a merge request) is a proposal to merge one branch into another, presented on the hosting platform as a page showing every change, a thread for discussion, and the results of any automated checks. It is the natural point at which a second person looks at the code. Code review has a long pedigree in software engineering. Studies of practice at large companies, including Microsoft and Google, have found that its main benefits are not only catching defects but spreading knowledge of the codebase, maintaining consistency, and surfacing alternative approaches. For research groups, the knowledge-spreading benefit may be the most valuable. A group in which every change to the shared analysis code is seen by at least one other person is a group in which more than one person understands that code, which matters enormously when someone leaves. Review in a research setting should be proportionate and kind. It is not an examination. A reviewer might check that the change does what its description says, that there is a test for the new behaviour, that the code is readable, and that nothing obviously unsafe — a hard-coded path to someone's home directory, a password — has slipped in. Scientific review is also appropriate: does this preprocessing step make sense, is this the right statistical test? Two eyes on a statistical choice are worth a great deal. A reviewer should ask questions rather than issue verdicts, and authors should keep pull requests small enough to be reviewed in twenty minutes. A pull request of two thousand lines will not be reviewed; it will be approved. Small groups may worry that they lack anyone to review. Even a lone researcher benefits from the pull request as a discipline: opening one, reading your own diff on the web interface, and waiting for automated tests to pass before merging catches a surprising number of mistakes. And review need not come from a domain expert. A student in a neighbouring group, a research software engineer from a central team, or a co-author can review for clarity even if they cannot review the science. Protecting main Hosting platforms let you enforce the convention that main is always working. Branch protection rules can require that changes to main arrive only through pull requests, that automated tests pass before merging, and optionally that at least one other person approves. For a shared group repository these settings are worth turning on; they convert a convention that people forget under deadline pressure into a guard rail. For a solo project, requiring tests to pass is usually enough. Protection also reduces a specific academic hazard: the late-night fix pushed directly to the shared repository an hour before a co-author runs the final analysis for a submission. With protection in place, that fix goes through a pull request, its tests run, and the co-author sees it arrive. Collaboration across institutions Research collaborations frequently span groups that do not share infrastructure. Git's distributed design handles this well. The usual model on public platforms is the fork: an external collaborator makes their own copy of the repository under their account, works on branches there, and opens pull requests back to the original. The maintainers of the original retain control over what enters main, while outsiders can contribute without being granted write access. This is also how researchers contribute fixes to the open-source libraries they depend on — a practice worth encouraging, since a bug fixed upstream is fixed for everyone. Collaborations with sensitive code or data may need a private repository with invited members, or an institutional GitLab instance. The branching strategy remains the same; only the access controls change. Papers, revisions and frozen results One situation in research does call for something closer to a long-lived branch. A paper is submitted from a tagged commit on main. Work continues: the code is refactored, a dependency is upgraded, a new analysis is added for the next paper. Four months later the reviews arrive, asking for a sensitivity analysis on the submitted results. The researcher now needs to modify the code as it was at submission, not as it is today, because today's code may produce slightly different numbers for reasons unrelated to the review. The clean way to handle this is a revision branch taken from the submission tag: git switch -c revision-1 submitted-2026-03. The sensitivity analysis is added there, the revised results are generated and tagged, and any change that is generally useful — a bug fix discovered along the way, say — is merged or cherry-picked back into main. Git's cherry-pick command copies a single commit from one branch to another, which is exactly the operation needed when one fix should reach two lines of development. This is, in miniature, the problem that Git Flow's hotfix branches were designed to solve, and it is the one part of that model worth borrowing. The same logic applies to software released for others to use. A lab that publishes version 1.2 of a tool, then begins substantial work towards 2.0, may need to issue a bug-fix release 1.2.1 for users who cannot upgrade yet. A maintenance branch taken from the 1.2 tag handles this. Most research projects will never need it, and should not create such branches in anticipation; the point is that the simple strategy extends naturally when the need arises. A worked example Consider a group of four — a principal investigator, a postdoc and two doctoral students — sharing a repository that contains a library of preprocessing and modelling functions used by everyone's projects. Under the approach described here, the postdoc notices that a function for detrending time series mishandles missing values. She creates a branch named fix-detrend-missing, writes a test that reproduces the problem, fixes it, and opens a pull request. The automated tests run in a few minutes and pass. One of the students, who uses the function heavily, reviews it that afternoon, asks one question about how the fix treats a series that is entirely missing, and approves after a short exchange. The branch is merged and deleted. The next morning both students pull main and have the fix. Nobody emailed a zip file; nobody's thesis now depends on a private copy of the function; and the history records not only the fix but the discussion that shaped it. The whole episode took perhaps two hours of combined effort, most of it the fix itself. The underlying principle Underneath the details, every successful branching strategy rests on one idea: integrate often. Work that stays separate accumulates divergence, and divergence becomes conflict. GitHub Flow and trunk-based development are both, at heart, ways of shortening the time between writing code and merging it. Research groups should adopt whichever version of this idea they can sustain, and they should add process only when a specific problem demands it. The next chapter turns from the history of the code to its structure — because how easily code can be branched, merged and reviewed depends heavily on how it is organised. Hashtags: #ResearchSoftwareEngineering #ResearchSoftware #VersionControl #Git #GitHub #SoftwareReproducibility #UnitTesting #ModularCode #CodeQuality #ScientificSoftware #ResearchCode #ContinuousIntegration #SoftwareTesting #CodeReview #BranchingStrategies #PullRequests #SoftwareProvenance #ReproducibleResearch #ComputationalReproducibility #Containerization #Docker #Apptainer #SoftwareDocumentation #ResearchCodeManagement #FutureOfResearchSoftwareEngineering

  • Human Participant Research Ethics (Historical Scandals, IRBs, and Modern Informed Consent)

    Download the Book (PDF): Introduction In July 1972 the Associated Press reporter Jean Heller published a story that began with a sentence no official could explain away. For forty years, she wrote, the United States Public Health Service had been conducting a study in which Black men with syphilis went without treatment so that doctors could observe what the disease did to the human body. The study had begun in Macon County, Alabama, in 1932. Penicillin had become the standard cure in the late 1940s. The men had not been given it. Many had been told they were being treated for "bad blood." Some had died of the disease, some had infected their wives, and some of their children had been born with congenital syphilis. The study had been discussed at professional meetings and published in medical journals across its entire life. It was not hidden. It was simply not considered wrong by the people in a position to stop it. That last point is the one worth holding on to. The history of research ethics is often told as a story of villains exposed and rules adopted to stop them. There were villains, and there were rules. But the more troubling pattern is how often harmful research was carried out by respected investigators at respected institutions, reviewed by colleagues, funded by governments, and published without objection. The failures were not usually secret. They were ordinary. And the system of protections that followed was designed, above all, to interrupt ordinary professional judgement before it could go wrong again. This book traces how that system came to be, how it works now, and where it is straining. It begins with the Nazi medical experiments and the Nuremberg Code of 1947, a document that stated the principle of voluntary consent with a clarity that has never been bettered and was then largely ignored by the countries that wrote it. It moves through the scandals that finally forced change in the United States in the 1960s and 1970s, including Tuskegee, the hepatitis studies at the Willowbrook State School and the injection of live cancer cells into elderly patients at the Jewish Chronic Disease Hospital in Brooklyn. It follows the construction of the modern framework: the National Research Act of 1974, the Belmont Report of 1979, the federal regulations that became the Common Rule, the successive revisions of the World Medical Association's Declaration of Helsinki, and the international guidelines that now govern clinical trials run across many countries at once. The later chapters turn to how that framework operates in practice and where it falls short. They examine what institutional review boards actually do, and what they cannot do. They look closely at informed consent, the central ritual of the whole system, and at the stubborn evidence that participants often do not understand what they have agreed to. They consider how the rise of biobanks, genomic databases and digital health records has broken the assumption that consent is a single event with a clear beginning and end, and they assess the models, including broad, tiered and dynamic consent, that have been proposed to replace it. They examine the idea of the "vulnerable population," which began as a protective category and has sometimes turned into a reason to exclude whole groups of people from the research that might have helped them. And they follow clinical research as it has moved across borders, to countries where oversight may be weaker and where the gap between what participants need and what the trial offers may be wide. The argument of this book The argument running through these chapters is simple to state. The protections built after the great scandals were designed as gates. An investigator submits a protocol, a committee reviews it before anything happens, a participant signs a form before enrolment, and the ethical work is, in most practical respects, done. The gate model made sense for the abuses it was built to prevent. Tuskegee, Willowbrook and the Nazi experiments were failures of permission: people were enrolled without knowing what was being done to them, or without any real choice, in studies no independent body had ever questioned. Requiring prior review and prior consent was the right answer to those failures. But the risks of research have been moving. Increasingly they arise after the gate has closed. A blood sample given for a diabetes study is later used, without the donor's knowledge, to study schizophrenia and ancestral migration. A genome deposited in a research database under a promise of anonymity turns out to be re-identifiable from public genealogy records. A trial approved in one country is run in another, where the local committee has fewer resources and where participants have no access to the treatment once the trial ends. A participant consents in good faith at enrolment, and the study changes shape over the following decade in ways no one could have described to her at the start. None of these problems is solved by a better form or a stricter committee at the entrance. They are problems of stewardship over time. The claim of this book, then, is that human research protection is in the middle of a transition from gatekeeping to stewardship: from ethics as a checkpoint passed once, to ethics as a continuing relationship between researchers, institutions and the people whose bodies and data make research possible. The transition is incomplete and uneven. Some of the most important recent changes, such as the revised Common Rule that took effect in the United States in 2019, the European Union's Clinical Trials Regulation, the 2024 revision of the Declaration of Helsinki and the 2025 revision of the international Good Clinical Practice guideline, move in this direction without quite committing to it. Understanding why the old model was built the way it was is the best way to see what a better one would require. What the reader will find here The book is written for readers who need to understand research ethics without necessarily being specialists in it: students entering the health and social sciences, researchers preparing their first protocols, members of ethics committees, research coordinators, journalists, policy staff and participants who want to know what protections they are actually owed. It assumes no background in law or philosophy. Where it describes regulations it does so in plain terms, naming the actual provisions where that helps a reader find them. It concentrates on biomedical and clinical research, because that is where most of the formal machinery was built and where the stakes of failure are most visible, but it also draws on social and behavioural research where the lessons carry over. The famous obedience experiments of Stanley Milgram, the covert observation in Laud Humphreys's Tearoom Trade, and the 2014 Facebook experiment that altered the news feeds of nearly 700,000 users to study emotional contagion each show a different side of the same problem: what happens when people become the material of research without meaningfully agreeing to it. The United States features prominently, because its scandals and its regulatory responses shaped the international field and because its regulations are still the template many institutions around the world follow. But the book also describes the European framework, the Council for International Organizations of Medical Sciences guidelines, the ethics review systems of countries that now host large numbers of trials, and the international Good Clinical Practice standard that sponsors use to run studies across dozens of jurisdictions at once. A word about tone. Research ethics is sometimes written as a catalogue of horrors, and sometimes as a bureaucratic manual. Neither serves the reader well. The horrors matter because they reveal how decent people rationalise harm; the procedures matter because they are the mechanism by which a society tries not to repeat it. The aim here is to keep both in view, and to treat the people who were harmed not as illustrations but as the reason the subject exists at all. The men in Macon County were not asked. The families of the children at Willowbrook were offered a place in a crowded institution if they agreed to the research. The Havasupai who gave blood in the Grand Canyon believed they were helping to fight diabetes. Every rule described in the chapters that follow is, at bottom, an attempt to make sure that the next person in their position is asked, is told the truth, is free to say no, and is not forgotten once the form is signed. Whether the rules succeed at that last part is the question this book is most concerned to answer. Chapter 1. The Doctors' Trial and the Code Nobody Adopted On 9 December 1946, in the same Nuremberg courthouse where the leading officials of the Third Reich had just been tried, an American military tribunal opened proceedings against twenty-three defendants, twenty of them physicians. The case was formally United States of America v. Karl Brandt et al., and it became known as the Doctors' Trial. Brandt had been Hitler's personal physician and a senior administrator of the programme that killed disabled people in institutions across Germany. His co-defendants included professors, senior officers of the medical services and doctors who had worked in the concentration camps. The charges concerned both the so-called euthanasia killings and a long series of experiments carried out on prisoners without their consent. The trial ran until August 1947. Sixteen defendants were convicted, and seven were sentenced to death and hanged the following year. The verdict is remembered less for its sentences than for a passage in the judgment in which the judges set out ten principles for "permissible medical experiments." That passage, never given a formal title by the court, came to be called the Nuremberg Code. Its first sentence is the most quoted line in the history of research ethics: "The voluntary consent of the human subject is absolutely essential." This chapter asks two questions about the Code. What exactly did it respond to? And why, having stated the principle of consent so plainly, did it have so little effect for the next twenty years on the countries whose judges had written it? What the experiments were It is tempting to treat the Nazi experiments as simple sadism dressed in laboratory coats, and some of them were close to that. But most were designed to answer questions that the German military or health authorities regarded as urgent, and they were run by people with genuine scientific training. That is precisely what made them instructive to the tribunal. At Dachau, the Luftwaffe physician Sigmund Rascher conducted high-altitude experiments in which prisoners were placed in low-pressure chambers to simulate the conditions a pilot might face after bailing out at great height. Many died. Rascher also ran freezing experiments, immersing prisoners in tanks of ice water for hours to study hypothermia and methods of rewarming, because German pilots were ditching in the North Sea. At Ravensbrück, women prisoners were deliberately wounded, and the wounds contaminated with bacteria, wood shavings and glass, to test whether sulfonamide drugs could prevent the infections that were killing soldiers on the Eastern Front. Other Ravensbrück prisoners had bones, muscles and nerves removed in experiments on regeneration and transplantation. At Buchenwald and elsewhere, prisoners were infected with typhus to test vaccines. At Dachau, the tropical medicine specialist Claus Schilling infected more than a thousand prisoners with malaria. Experiments on mass sterilisation, by X-ray and by chemical injection, were carried out at Auschwitz and Ravensbrück as part of the regime's racial programme. Josef Mengele's experiments on twins at Auschwitz, perhaps the most infamous of all, did not feature in the trial because Mengele had escaped. Several features of these experiments recur, in milder forms, throughout the rest of this book. The subjects were drawn from populations already stripped of rights: prisoners, Jews, Roma, Poles, Soviet prisoners of war. Their availability was treated as an opportunity. The research questions were framed in terms of the needs of others, soldiers and airmen above all, so that the suffering of the subjects was weighed against a large and diffuse benefit to the nation. And the investigators were embedded in institutions that approved, funded and published their work. Some of the hypothermia data were later discussed at a German medical conference in 1942, to an audience of physicians who raised no public objection. The argument of the defence The defendants did not, for the most part, deny what they had done. Their lawyers argued instead that there was no settled international standard against which their conduct could be judged, and that doctors in the Allied countries had done similar things. The defence pointed to experiments conducted in the United States during the war, including the malaria studies on prisoners at the Stateville Penitentiary in Illinois, in which inmates were infected in the course of testing antimalarial drugs for troops in the Pacific. They cited older examples too: Walter Reed's yellow fever experiments in Cuba in 1900, in which volunteers were deliberately exposed to infected mosquitoes, and the work of researchers who had tested treatments on prisoners in the Philippines early in the century. The point was not that these experiments were identical to Dachau, but that the line between acceptable and unacceptable research was nowhere written down. The Reed example cut in more than one direction. The yellow fever commission had asked its volunteers, many of them recently arrived Spanish immigrants, to sign written agreements, in Spanish as well as English, which acknowledged that the disease could be fatal and offered payment in gold, with a further sum for those who fell ill. These documents are often described as among the earliest written consent forms in medical research. They show that the idea of telling people about the risk and securing their agreement in writing was familiar to at least some American investigators at the turn of the century. They also show, uncomfortably, how early the tension between payment and free choice appeared: a large sum offered to poor men for accepting the chance of a deadly infection raises exactly the question of undue inducement that ethics committees still debate today. Several volunteers did contract yellow fever, and Reed's colleague Jesse Lazear died of it after being bitten, whether deliberately or accidentally remains disputed. The prosecution relied heavily on two American medical experts. Andrew Ivy, a physiologist from Illinois, had been sent by the American Medical Association, and Leo Alexander, a Boston neuropsychiatrist of Austrian origin, served as a consultant to the prosecution. Both argued that there were in fact well-understood principles of medical ethics, including the requirement of consent, and that the defendants had violated them. Ivy's position was complicated by the Stateville studies, which had been conducted in his own state. While the trial was under way, a committee convened by the governor of Illinois examined the ethics of prison research and concluded, in terms that supported the prosecution's case, that prisoners could volunteer for research provided their consent was free. Historians have since noted how conveniently timed that conclusion was. The defence argument had one embarrassing weakness. Germany itself had issued detailed guidelines on human experimentation in 1931. The Reich Circular on new therapies and human experimentation required consent, prohibited experiments on dying patients and required special caution with children. By the standards of the time it was one of the most demanding research codes anywhere. It had never been repealed. Germany had not lacked rules; it had created conditions in which rules ceased to matter. This is an early version of a lesson the rest of this book repeats: written standards protect people only when institutions are willing to enforce them against their own members. There was an even older German precedent. In the late 1890s the Breslau dermatologist Albert Neisser injected serum from syphilis patients into women, several of them sex workers, without telling them, in an attempt to develop a vaccine. Some later developed syphilis. The resulting public controversy led the Prussian government in 1900 to issue a directive requiring unambiguous consent after proper explanation for any non-therapeutic intervention. Consent as a principle, in other words, was not invented at Nuremberg. It had been stated, and set aside, before. The ten points The Code the judges wrote contains ten principles. The first, on consent, is by far the longest. It requires that the person involved have legal capacity to consent, be able to exercise free power of choice without "any element of force, fraud, deceit, duress, over-reaching, or other ulterior form of constraint or coercion," and have sufficient knowledge and comprehension of the elements of the experiment to make an understanding and enlightened decision. It lists what the person must be told: the nature, duration and purpose of the experiment, the method and means, the inconveniences and hazards reasonably to be expected, and the effects on health that may follow. And it places the duty to ensure the quality of consent on each individual who initiates, directs or engages in the experiment, a duty that "may not be delegated to another with impunity." The remaining nine principles concern the design and conduct of the research. The experiment should be such as to yield fruitful results for the good of society, unprocurable by other means. It should be based on the results of animal experimentation and knowledge of the natural history of the disease. It should avoid all unnecessary physical and mental suffering. No experiment should be conducted where there is reason to believe that death or disabling injury will occur, except perhaps where the experimental physicians also serve as subjects. The degree of risk should never exceed the humanitarian importance of the problem. Proper preparations and facilities should protect the subject against even remote possibilities of injury. The experiment should be conducted only by scientifically qualified persons. The subject should be at liberty to bring the experiment to an end. And the scientist in charge must be prepared to terminate the experiment if continuation is likely to result in injury, disability or death. Read today, the Code is striking both for what it contains and for what it omits. It has no provision for anyone who cannot consent: children, people with serious mental illness, the unconscious. Taken literally, it forbids all research on them, including research designed to benefit them. It makes no distinction between research on healthy volunteers and research combined with medical care, a distinction that would later become central to the Declaration of Helsinki. It places the entire burden of protection on the individual investigator and says nothing about independent review. There is no committee in the Nuremberg Code. The judges assumed that a properly conscientious scientist, bound by these principles, would be enough. That assumption turned out to be the Code's deepest flaw, and the history of the next thirty years is largely the history of discovering it. A code for barbarians The Nuremberg Code had no formal legal force in any country. It was part of a judgment by a military tribunal applying international law to defendants accused of crimes against humanity. Outside that context, it was a statement of principles that anyone could choose to ignore, and in the United States, Britain and elsewhere, the medical profession largely did. The reasons were partly professional and partly psychological. American physicians could read the Code as a response to atrocities committed by people fundamentally unlike themselves. The legal scholar and psychiatrist Jay Katz, who spent much of his career studying the gap between the Code and practice, argued that many American researchers regarded it as a good code for barbarians but an unnecessary one for ordinary physicians. The ideal of the ethical physician, guided by conscience and professional tradition, seemed an adequate safeguard. Formal consent requirements were seen as a legalistic intrusion into the doctor's judgement about what was good for a patient, and the line between patient and research subject was, in the clinics of the 1950s, often not drawn at all. The scientific context also discouraged reflection. The postwar decades were a period of extraordinary expansion in medical research, driven in the United States by the growth of the National Institutes of Health and by the conviction, forged by penicillin, radar and the atomic bomb, that organised science could solve almost any problem. The federal research budget grew rapidly through the 1950s and 1960s. Clinical research became a career in its own right, with its own incentives to publish and to secure further funding. Subjects were needed in large numbers, and they were found in the places where they had always been found: charity wards, prisons, institutions for the disabled, and military bases. The attitude of governments was no better. In the early 1950s the United States Department of Defense adopted a version of the Nuremberg principles for its own research on atomic, biological and chemical warfare, in a memorandum issued by Secretary of Defense Charles Wilson in 1953. The memorandum was classified, and its requirements were unevenly followed. The later investigations of the Advisory Committee on Human Radiation Experiments, established by President Clinton in 1994, documented a long series of Cold War studies in which people were exposed to radiation without adequate information, including the injection of plutonium into hospital patients in the 1940s and the feeding of radioactive tracers to children at the Fernald State School in Massachusetts in the late 1940s and early 1950s. The committee found that many officials had been aware of the ethical principles and had not applied them. There is a further, less often discussed, case. Japan's Unit 731, a biological warfare research unit in occupied Manchuria led by the army physician Shiro Ishii, conducted lethal experiments on thousands of prisoners, mostly Chinese, during the 1930s and 1940s, including deliberate infection with plague and anthrax, vivisection and frostbite experiments. Its leaders were never tried by the Tokyo tribunal. Declassified documents have since shown that American officials, interested in the unit's biological warfare data, granted its members immunity in exchange for access to their findings. The Nuremberg Code was announced at almost the same moment that the Allied powers chose to treat another set of human experiments as a source of useful knowledge rather than a crime. What Nuremberg settled and what it left open It would be unfair to say that the Code achieved nothing. It gave the principle of voluntary consent a place in international law and a canonical wording that every later framework has had to address. It established that research ethics is not simply a private matter between doctor and patient but a subject on which courts and societies can pass judgement. And it provided a benchmark against which later abuses could be measured; the critics of the 1960s were able to say, with force, that American researchers were doing things the Code forbade. What Nuremberg did not settle was how such principles could be made to govern the everyday conduct of respectable science. Its model of protection rested on the conscience of the individual investigator. It assumed that the danger lay in bad people, and that good people, once reminded of the rules, would follow them. The following two decades demonstrated that this assumption was false. The researchers who ran the studies described in the next chapter were not Nazis. They were, by the standards of their profession, conscientious, accomplished and in many cases admired. Some of them believed sincerely that what they were doing was good for their subjects. It was precisely because conscience failed in such people that the answer, when it came, took the form of external review: a committee that would look at a protocol before any subject was enrolled and ask whether it should go ahead at all. That answer, the gate, is the institution whose strengths and limits run through the rest of this book. It was invented not because the Nuremberg principles were wrong, but because stating a principle had turned out to be the easy part. The difficulty was building a system that would apply the principle to people who were sure they did not need it. Chapter 2. Scandals at Home: Tuskegee, Willowbrook and the Beecher Exposé In June 1966 the New England Journal of Medicine published an article by Henry K. Beecher, a professor of anaesthesia research at Harvard Medical School, under the modest title "Ethics and Clinical Research." Beecher was not an outsider or a campaigner. He was one of the most eminent clinical investigators in the United States, and he had spent much of his career running studies on human beings, including research on pain and on the placebo effect. That was what gave his article its force. He described twenty-two studies, all published in reputable journals by researchers at reputable institutions, in which he believed subjects had been placed at serious risk without their informed consent. He did not name the investigators. He did not need to. Anyone in the field could look up the references. The examples included the withholding of known effective treatment from patients, among them servicemen with streptococcal infections who were denied penicillin in a study of rheumatic fever prevention, and charity patients who were given no treatment for typhoid in order to compare outcomes. They included studies in which liver damage was deliberately induced, in which cancer cells were injected into patients who did not know what they were receiving, and in which children with intellectual disabilities were deliberately infected with hepatitis. Beecher's central claim was that such research was not the work of a few rogues. It was, as he put it, a matter of widespread practice, and it arose mostly from thoughtlessness and carelessness rather than wilful disregard of the patient's rights. Beecher's article is the natural hinge of this chapter because it named the problem that Nuremberg had not. The trouble was not that American medicine contained monsters. The trouble was that the culture of clinical research allowed conscientious people to do things to their subjects that, looked at plainly, they could not defend. The Jewish Chronic Disease Hospital and Willowbrook Two of the cases Beecher alluded to became the subject of public controversy in their own right, and together they illustrate how professional prestige shielded questionable research. In July 1963 Chester Southam, a respected cancer researcher at the Sloan Kettering Institute in New York, arranged for live cultured cancer cells to be injected under the skin of twenty-two elderly, chronically ill patients at the Jewish Chronic Disease Hospital in Brooklyn. Southam was studying how the immune system rejected cancer cells, and his hypothesis was that debilitated patients would reject them more slowly than healthy people. He had previously injected cancer cells into prisoners at the Ohio State Penitentiary, who had volunteered. At the hospital, the patients were told they were receiving an injection to test their immune response. They were not told that the cells were cancerous. Southam's view was that the cells would be rejected and posed no real risk, and that telling patients the word "cancer" would cause them needless distress. Three young physicians at the hospital refused to take part and resigned. A member of the hospital's board, the lawyer William Hyman, went to court to obtain the patients' records. In 1965 the Board of Regents of the University of the State of New York found Southam and the hospital's medical director, Emanuel Mandel, guilty of fraud, deceit and unprofessional conduct. Their licences were suspended for a year, but the suspensions were stayed and they were placed on probation. The penalty was mild, and the professional consequences milder still: a few years later Southam was elected president of the American Association for Cancer Research. The case shows the profession's reluctance, even in the face of a formal finding of deceit, to treat such conduct as seriously wrong. The Willowbrook State School on Staten Island was an institution for children with intellectual disabilities, chronically overcrowded and understaffed. Hepatitis was endemic there. From the mid-1950s into the early 1970s, the paediatrician Saul Krugman of New York University led a research programme in which newly admitted children were deliberately infected with hepatitis, at first by being fed extracts made from the faeces of infected patients, in order to study the course of the disease and to test whether gamma globulin could protect against it. The research was scientifically important. It helped to distinguish the two forms of viral hepatitis now known as hepatitis A and hepatitis B and contributed to the path toward a hepatitis B vaccine. Krugman and his defenders offered a set of justifications that are worth examining because they recur in later debates. The children were likely to contract hepatitis anyway, given conditions in the institution. The infections were induced under controlled conditions in a special unit with better care than the rest of the school. The strain used was believed to produce a relatively mild illness in children. And the parents had given consent. Critics replied that the "they would have caught it anyway" argument made the institution's own failings a licence for research; that the proper response to endemic hepatitis was to improve conditions, not to study them; and that the consent was compromised. In the mid-1960s, when overcrowding led Willowbrook to close its general admissions, places remained available in the hepatitis research unit. Parents seeking a place for their child were, in effect, told that there was room if they agreed to the study. A choice offered to a family with no other option for care is not the free choice the Nuremberg Code had described. The Willowbrook case has a lasting place in teaching because the arguments on both sides are serious. Krugman was a humane and accomplished physician who went on to receive major honours, including a Lasker award in 1983. The research produced real knowledge. Yet it depended on a population that had been placed in an institution by a society that had largely stopped looking at it, and on a consent obtained under circumstances that made refusal costly. The case is the clearest early example of what later ethicists would call structural vulnerability: harm that arises less from any single decision than from the conditions under which decisions are made. Tuskegee No case shaped American research ethics more than the Tuskegee Study of Untreated Syphilis in the Negro Male, to give it the name used by the Public Health Service itself. It began in 1932 in Macon County, Alabama, one of the poorest counties in the country, in collaboration with the Tuskegee Institute. The original plan was a short observational study of the effects of untreated syphilis in Black men, justified in part by a belief that the disease behaved differently in Black and white patients. It lasted forty years. About six hundred men were enrolled, roughly four hundred with latent syphilis and two hundred without the disease as controls. They were not told they had syphilis. They were told they had "bad blood," a local term covering a range of ailments, and were offered free medical examinations, meals on examination days, transport and a burial stipend in exchange for allowing an autopsy after death. Some of the procedures presented to them as treatment, including painful spinal taps, were purely diagnostic. A Black public health nurse, Eunice Rivers, served as the men's point of contact throughout the study and was central to keeping them enrolled. The ethics of the study were questionable from the start, since treatments for syphilis in the 1930s, although toxic and imperfect, did exist. They became indefensible after penicillin was established as an effective cure in the 1940s. The men were not offered it. The Public Health Service took active steps to keep them from treatment elsewhere, including providing lists of study participants to local physicians and, during the Second World War, to draft boards, so that men called up for military service would not be treated as other syphilitic draftees were. The study continued through the Nuremberg trial, through the adoption of the Nuremberg Code, through the Declaration of Helsinki in 1964, and through the publication of Beecher's article. It ended because of one persistent employee and one reporter. Peter Buxtun, a venereal disease interviewer with the Public Health Service in San Francisco, learned of the study in the mid-1960s and began raising objections internally from 1966. In 1969 a panel convened by the Centers for Disease Control reviewed the study and recommended that it continue. Buxtun eventually took the story to the press. Jean Heller's Associated Press report appeared in July 1972, and the study was halted later that year. Hearings in the Senate, led by Edward Kennedy, followed in 1973. A class-action lawsuit brought on behalf of the men by the civil rights lawyer Fred Gray was settled in 1974 for about ten million dollars, and the government agreed to provide lifetime medical care to the survivors and, later, to their affected wives and children. In May 1997 President Clinton issued a formal apology at the White House, in the presence of several of the surviving men. The harm was not confined to the men and their families. Tuskegee became a symbol, in Black communities across the United States, of the willingness of medical institutions to treat Black lives as expendable. Its effects are measurable. In a study published in the Quarterly Journal of Economics in 2018, the economists Marcella Alsan and Marianne Wanamaker found that the 1972 disclosure was followed by increased medical mistrust and reduced use of health care among older Black men, and estimated that these effects were associated with a meaningful reduction in their life expectancy. Research abuse, in other words, damages not only its direct victims but the relationship between whole communities and the institutions meant to serve them. That damage makes it harder to recruit the very populations whose underrepresentation in research later chapters of this book will discuss. Guatemala: the scandal found later The same agency produced a scandal that remained hidden for more than sixty years. In 2010 the historian Susan Reverby, working in the papers of the Public Health Service physician John Cutler at the University of Pittsburgh, found records of experiments conducted in Guatemala between 1946 and 1948. Cutler, who later also worked on the Tuskegee study, had led a programme in which prisoners, soldiers, psychiatric patients and sex workers in Guatemala were deliberately exposed to syphilis, gonorrhoea and chancroid in order to test prophylactic methods. In some cases infected sex workers were brought into prisons; in others, bacteria were introduced through abrasions. The research was funded by the United States government and conducted with the cooperation of Guatemalan officials. Following Reverby's disclosure, the American government apologised to Guatemala, and President Obama asked the Presidential Commission for the Study of Bioethical Issues to investigate. Its 2011 report, titled "Ethically Impossible": STD Research in Guatemala from 1946 to 1948, found that more than 1,300 people had been deliberately exposed to these diseases, that far fewer were documented as receiving adequate treatment, and that the researchers had understood at the time that what they were doing would not have been accepted at home. The Commission noted that the experiments were being carried out in the same years that American judges were sitting in judgement at Nuremberg. Guatemala foreshadows a theme of the final chapter of this book: research that would not pass at home has repeatedly been conducted abroad, on populations with less power to object. Social and behavioural research The crisis was not limited to medicine. In the early 1960s the Yale psychologist Stanley Milgram ran his obedience experiments, in which participants were told to administer what they believed were increasingly severe electric shocks to another person, in fact an actor, as part of a supposed study of learning. Many continued to the highest levels on the instructions of the experimenter. The findings were profound, and the participants were debriefed, but many had experienced considerable distress, and the study relied on deception as its central method. In 1970 the sociologist Laud Humphreys published Tearoom Trade, a study of anonymous sexual encounters between men in public toilets. Humphreys had acted as a lookout to observe the encounters, noted the licence plates of the men involved, traced their addresses, and later interviewed them in their homes under a false pretext as part of a separate health survey. In 1971 the Stanford Prison Experiment, led by Philip Zimbardo, was stopped after six days when the simulated prison environment produced serious distress among participants. These studies raised different questions from the medical scandals. The risks were psychological and social rather than physical: distress, humiliation, loss of privacy, the threat of exposure in a period when homosexual conduct was criminal in most American states. The methods, including deception and covert observation, were thought by many researchers to be indispensable to answering important questions. They showed that any framework of protections would have to cover research in which harm was not a matter of needles and drugs, and that consent could be compromised by the design of a study as much as by the neglect of an investigator. From exposure to regulation By the late 1960s the pieces of a response were already forming, though slowly. The thalidomide disaster, in which a sedative prescribed to pregnant women in Europe and elsewhere caused severe birth defects in thousands of children, prompted the United States Congress to pass the Kefauver-Harris Amendments to the Food, Drug, and Cosmetic Act in 1962. Among their provisions was a requirement that investigators obtain the consent of people receiving experimental drugs, though with exceptions where the investigator judged consent not feasible or contrary to the patient's best interests. The Food and Drug Administration strengthened the requirement over the following years. More important for the long run was a policy issued by the Surgeon General, William Stewart, in February 1966, a few months before Beecher's article appeared. It required institutions receiving Public Health Service research funds to provide for prior review of research involving human subjects by a committee of the investigator's "institutional associates," which would assess the rights and welfare of the subjects, the appropriateness of the methods of obtaining consent, and the risks and potential benefits. This was the germ of the institutional review board. For the first time, the judgement of the individual investigator would be checked, before the research began, by someone else. The policy was an administrative condition of funding rather than a law, and its implementation was uneven. Some institutions set up serious committees; others treated review as a formality. What finally turned it into a national system was Tuskegee. The public outrage of 1972 and the Kennedy hearings that followed made it impossible to leave the matter to the discretion of the profession and its funders. The next chapter describes the statute that followed, the national commission it created, and the report that commission produced, which remains the moral foundation of human research regulation in the United States. Taken together, the cases in this chapter reveal a consistent pattern. The subjects were almost always people with limited power: the poor, the institutionalised, the imprisoned, the elderly, the colonised, members of racial minorities. The research was almost always conducted in the open and justified by appeal to scientific value or to the argument that subjects were no worse off than they would otherwise have been. And the people best placed to object, the investigators' own colleagues, usually did not. A system of protection built on the conscience of individual researchers had been tried for a generation after Nuremberg. It had failed, not dramatically but routinely, and that routine failure is what the gatekeeping model was designed to prevent. Chapter 3. From Outrage to Architecture: Belmont, the Common Rule and Helsinki Scandals produce anger, and anger produces demands that something be done. What gets done depends on who is in the room when the response is designed. In the United States after Tuskegee, the room contained legislators, lawyers, physicians, philosophers and theologians, and what they built was a combination of three things: a statute that required review, a set of regulations that specified how the review would work, and a short philosophical document that explained what the review was for. Outside the United States, the medical profession had already begun to build its own framework, the Declaration of Helsinki, which would evolve over sixty years into something quite different from where it began. And alongside both, the pharmaceutical industry and its regulators developed a technical standard, Good Clinical Practice, that now governs most drug trials in the world. This chapter describes those three layers, professional, governmental and industrial, and the way they fit together, often awkwardly. Understanding the architecture matters because nearly every practical question in research ethics, from what must go into a consent form to whether a study in another country is acceptable, is answered by reference to one or more of these documents. They do not always give the same answer. The National Research Act and the Belmont Report The National Research Act was signed into law in July 1974. It did two main things. It gave statutory backing to the requirement that institutions receiving federal research funds establish review boards to protect the rights of human subjects. And it created the National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research, an eleven-member body charged with identifying the basic ethical principles that should underlie research and with recommending how they should be applied. The Commission met from 1974 to 1978 and produced a series of reports on specific topics, including research on the fetus, on prisoners, on children, on people institutionalised as mentally infirm, and on institutional review boards. Its best-known work, however, was a short document that emerged from a period of intensive deliberation beginning with a meeting at the Smithsonian Institution's Belmont Conference Center in Maryland in 1976. The Belmont Report, formally titled Ethical Principles and Guidelines for the Protection of Human Subjects of Research, was published in the Federal Register in April 1979. The Report does three things in fewer than ten pages. First, it distinguishes research from practice. Medical practice consists of interventions designed solely to enhance the wellbeing of an individual patient with a reasonable expectation of success; research is an activity designed to test a hypothesis and contribute to generalisable knowledge. The distinction matters because research, by its nature, serves purposes beyond the individual, and therefore requires a different kind of scrutiny. Many of the scandals of the previous chapter had happened in the fog between the two, where physicians experimented on patients while telling themselves, and the patients, that they were providing care. Second, the Report identifies three basic ethical principles. Respect for persons incorporates two convictions: that individuals should be treated as autonomous agents, and that persons with diminished autonomy are entitled to protection. Beneficence is understood not as kindness but as an obligation: do not harm, and maximise possible benefits while minimising possible harms. Justice concerns the distribution of the burdens and benefits of research, and the Report makes the point with direct reference to the history: the burdens of research had fallen disproportionately on the poor and on ward patients, while the benefits of improved medical care flowed mainly to private patients, and the Tuskegee study had used disadvantaged rural Black men to study a disease that was by no means confined to that population. Third, the Report connects each principle to a practical requirement. Respect for persons is applied through informed consent, which the Report analyses as having three elements: information, comprehension and voluntariness. Beneficence is applied through the systematic assessment of risks and benefits. Justice is applied through the fair selection of subjects, at both the individual and the social level. The genius of the Belmont Report was its economy. Three principles, three applications, a clear distinction between research and practice. It gave review boards a vocabulary for their deliberations and gave the regulations a rationale. Its limitations follow from the same economy. It does not tell anyone what to do when the principles conflict, as they constantly do: when respecting a person's choice to enter a risky trial collides with the duty to protect them, or when a fair distribution of research burdens requires recruiting groups whom other rules treat as vulnerable. It was written with the individual subject in mind and has little to say about communities, about data, or about research that crosses borders. And its account of justice is almost entirely about burdens. It says much less about the injustice of being left out of research altogether, which later became a central concern. The regulations and the Common Rule The Department of Health, Education, and Welfare had issued its first regulations on the protection of human subjects in May 1974, codified at Title 45 of the Code of Federal Regulations, Part 46. Following the Commission's work, the Department of Health and Human Services substantially revised them in 1981, and the Food and Drug Administration issued parallel regulations on informed consent and institutional review boards for research on products it regulates, codified in Title 21, Parts 50 and 56. Additional subparts were added to 45 CFR 46 over the following years to govern research involving pregnant women, fetuses and neonates, prisoners, and children. The core of the regulations, known as Subpart A, was adopted in 1991 by more than a dozen federal departments and agencies as a single shared policy. This became the Federal Policy for the Protection of Human Subjects, universally called the Common Rule. Its reach is defined by funding: it applies to research conducted or supported by the signatory agencies. Research funded privately and not involving FDA-regulated products falls outside it, unless an institution chooses to apply it voluntarily, which many universities and hospitals historically did through their assurances with the federal government. The Common Rule translated the Belmont principles into procedures. It set out the membership requirements for review boards, the criteria a board must find satisfied before approving research, the elements that must be disclosed in informed consent, the circumstances in which consent may be waived or altered, the categories of research that are exempt from its requirements, and the requirement for continuing review of approved research. Chapter 4 describes how these procedures work in practice. The Rule remained substantially unchanged for a quarter of a century. By the 2010s it was widely regarded as out of date: too burdensome for low-risk social science, too permissive about the use of stored specimens and data, and poorly suited to multi-site research, in which the same protocol might be reviewed separately by dozens of boards, each demanding its own changes. After a lengthy consultation, a revised Common Rule was published in January 2017. After two delays, its main provisions took effect in January 2019, and a requirement that most federally funded cooperative research in the United States use a single review board took effect in January 2020. The revisions were significant. They introduced a requirement that consent forms begin with a concise presentation of the key information most likely to help a prospective participant decide, in response to evidence that participants were being overwhelmed by long forms. They created a mechanism for "broad consent" to the storage and future use of identifiable data and specimens. They eliminated routine continuing review for many minimal-risk studies. They expanded the categories of exempt research. And they required the posting of a consent form for each federally funded clinical trial on a public website. The revised rule also changed its language on vulnerability, replacing references to "mentally disabled persons" and removing pregnant women and "handicapped" persons from its list of examples of subjects vulnerable to coercion, in favour of "individuals with impaired decision-making capacity" and "economically or educationally disadvantaged persons." Some of the most ambitious proposals, including a plan to require consent for research on all non-identified biospecimens, were dropped after objections from researchers and institutions. The Declaration of Helsinki While the United States was building a regulatory system, the World Medical Association, an international federation of national medical associations, was writing its own. The Declaration of Helsinki was adopted at the Association's General Assembly in Helsinki in June 1964. It was, in part, the profession's answer to Nuremberg: a statement written by physicians, for physicians, that would govern medical research without importing the rigid legal framing of a war crimes tribunal. The 1964 Declaration differed from the Nuremberg Code in two important ways. It distinguished between clinical research combined with professional care, which it treated more permissively, and non-therapeutic research on healthy volunteers, where it applied stricter standards. And it provided for consent by a legal guardian where the subject lacked capacity, opening the door to research on children and incapacitated adults that the Nuremberg Code had implicitly closed. Critics have noted that the Declaration's first version was, in some respects, less protective than Nuremberg, and more accommodating of the physician's therapeutic judgement. The Declaration has since been revised many times, and the revisions trace the development of the field. The 1975 revision, adopted in Tokyo, introduced the requirement that research protocols be reviewed by an independent committee, bringing the Declaration into line with the emerging American model. The 2000 revision, adopted in Edinburgh, was a major rewrite that abandoned the therapeutic and non-therapeutic distinction as an organising principle and took positions on two issues that proved highly contentious: the use of placebo controls when a proven treatment exists, and the obligation to provide participants with access to beneficial interventions after a trial ends. Both positions were driven by controversies over trials in low-income countries, discussed in the final chapter of this book. Further revisions followed in 2008 and 2013. In October 2024 the Association adopted a further revision at its General Assembly, again in Helsinki, after a consultation of about two and a half years. The 2024 text consistently refers to research "participants" rather than "subjects," a change of vocabulary that signals a change of attitude. It broadens its intended audience beyond physicians to all those involved in medical research. It places new emphasis on meaningful engagement with participants and communities before, during and after research, on scientific integrity, and on fairness in the distribution of the burdens and benefits of research across populations. And it adopts a more contextual account of vulnerability, warning that excluding groups from research can itself entrench health disparities. The Declaration has no legal force of its own, but it is incorporated by reference into the laws and regulations of many countries, and journals commonly require authors to state that their research conformed to it. Global guidelines and Good Clinical Practice Two further frameworks complete the architecture. The Council for International Organizations of Medical Sciences, founded in 1949 under the auspices of the World Health Organization and UNESCO, has issued guidelines on biomedical research involving humans since the early 1980s, with the explicit aim of showing how the principles of the Declaration of Helsinki could be applied in low- and middle-income countries. Its most recent version, the International Ethical Guidelines for Health-related Research Involving Humans, was published in 2016 in collaboration with the World Health Organization. The 2016 guidelines are notable for grounding research ethics in both scientific and social value, for their detailed treatment of research in low-resource settings, and for an approach to vulnerability that looks at the specific characteristics that make a person or group susceptible to harm rather than simply listing categories. The second is Good Clinical Practice. The International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use, which brings together drug regulators and industry from Europe, Japan, the United States and, increasingly, many other regions, adopted its guideline E6 on Good Clinical Practice in 1996. It set a single standard for the design, conduct, monitoring, recording and reporting of clinical trials, so that data generated under it would be acceptable to regulators in all participating regions. Its ethical provisions are brief and largely derivative of the Declaration of Helsinki, but its practical influence is enormous, because it governs the day-to-day operations of the trials on which drug approvals depend. An addendum in 2016 introduced risk-based approaches to monitoring. A thorough revision, E6(R3), was finalised in January 2025. It restructures the guideline around principles of proportionality and quality by design, recognises decentralised trial designs and digital health technologies, and gives more attention to participants' perspectives and to informed consent. The European Medicines Agency applied the new principles and main annex from July 2025. The main features of these frameworks can be compared directly, as Table 1 sets out. The differences in authority and focus explain much of the tension between them: a physician running an industry-sponsored trial in a middle-income country may be bound simultaneously by the country's law, the sponsor's Good Clinical Practice obligations, the regulations of the country where the drug will be marketed, and a journal's requirement to comply with the Declaration of Helsinki. Table 1. The foundational frameworks of human research ethics compared. Framework and issuer Current version Legal status Distinctive contribution Nuremberg Code (US military tribunal) 1947 Part of a judgment, not binding law Absolute requirement of voluntary consent Declaration of Helsinki (World Medical Association) 2024 Professional code, adopted into many laws Independent review; placebo and post-trial rules Belmont Report (US National Commission) 1979 Ethical basis for US regulations Respect for persons, beneficence, justice Common Rule, 45 CFR 46 (US federal agencies) Revised 2018, in force 2019 Binding on federally funded research Procedures for review boards and consent CIOMS Guidelines (CIOMS with WHO) 2016 Guidance Low-resource settings; social value ICH E6 Good Clinical Practice (regulators and industry) E6(R3), 2025 Adopted by regulators for drug trials Operational standard for multinational trials An architecture built for the gate Looking at this architecture as a whole, its centre of gravity is clear. Every framework after Nuremberg places heavy emphasis on the moment before research begins: the protocol is reviewed, the risks weighed, the consent form approved, the participant informed and enrolled. These are the functions that the scandals of the previous chapter most obviously demanded, and they are the functions that the architecture performs best. What happens afterwards receives much less attention. Continuing review exists, but it has always been the weakest part of the system, a paperwork exercise at many institutions and now, for minimal-risk research in the United States, largely abolished. Monitoring of trial conduct is the domain of sponsors and regulators, whose concern is primarily data quality. The relationship between researchers and participants after enrolment, the use of data and samples after a study ends, the fate of participants after a trial closes, the effects of research on communities rather than individuals: these are the areas where the frameworks are thinnest and have been changing most. The 2024 Declaration of Helsinki, the revised Common Rule's provisions on broad consent, and ICH E6(R3)'s attention to participant perspectives are all efforts to strengthen them. None yet amounts to a fully developed account of what researchers owe the people they study across the life of a project. That gap is the subject of the second half of this book. The first task, however, is to look closely at the institution the architecture put at its centre: the committee at the gate. Hashtags: #HumanParticipantResearchEthics #ResearchEthics #HumanSubjectsResearch #InformedConsent #InstitutionalReviewBoards #IRB #NurembergCode #TuskegeeStudy #WillowbrookStudy #BeecherExposé #BelmontReport #CommonRule #DeclarationOfHelsinki #GoodClinicalPractice #ICHGCP #CIOMSGuidelines #ResearchGovernance #ParticipantProtection #VulnerablePopulations #ResearchMisconduct #ClinicalResearchEthics #ConsentModels #ResearchStewardship #GlobalResearchEthics #FutureOfResearchEthics

  • Community-Based Participatory Research (Collaborative Methodologies for Social Impact)

    Download the Book (PDF): Introduction In the late 1980s residents of West Harlem began organizing against the environmental burdens their neighborhood had been asked to carry: a sewage treatment plant on the Hudson waterfront, and a disproportionate share of Manhattan's diesel bus depots. Northern Manhattan's children were being hospitalized for asthma at rates far above the city average. Residents did not need a study to tell them the air was bad; they could smell it and they could see their children's inhalers. What they needed was evidence that would count in rooms where they had no seat. The organization that grew out of that fight, West Harlem Environmental Action, founded in 1988 and now called WE ACT for Environmental Justice, entered into a long partnership with researchers at Columbia University's school of public health. In one early pilot study, published by Patrick Kinney, Peggy Shepard and colleagues in 2000, community members and scientists together measured fine particles and diesel exhaust on Harlem sidewalks and found that concentrations tracked the traffic of trucks and buses street by street. Measurements of that kind became part of the case that pushed New York's transit authority toward cleaner buses and depot reforms. The story is often told as a success of community-based participatory research, and it was. But the most instructive thing about it is not the air monitors. It is the order of events. The community named the problem before any researcher arrived. The community decided that measurement was the tool it needed. The researchers were invited into a fight that was already under way, and the value of their contribution depended on whether they could bring the technical capacity of a university without taking the fight away from the people who had started it. That order of events is the subject of this book. Community-based participatory research, usually shortened to CBPR, is a family of approaches in which people affected by a problem take part as partners, not subjects, in studying it. In the definition that has shaped the field in public health, it is a collaborative approach that equitably involves community members, organizational representatives and researchers in all aspects of the research process, with each partner contributing unique strengths and shared responsibilities. The key words in that sentence are "equitably" and "all aspects." Many projects involve communities. Far fewer involve them equitably, and fewer still involve them in all aspects, from the choice of question to the ownership of what was learned. The argument The argument of this book is that participation is real only to the extent that it redistributes control over decisions that carry consequences. Research is a sequence of such decisions. Someone decides what question is worth asking and what counts as an answer. Someone decides who collects the data and on what terms. Someone decides what the data mean, and whose interpretation prevails when two readings conflict. Someone's name goes on the paper. Someone owns the samples, the data files and any invention that follows. Someone decides what happens to all of it when the grant ends. At each of these points, control can move toward the community or quietly return to the institution. A project is participatory at the points where it moved, and conventional at the points where it did not, whatever its funding application says. This way of seeing the work has a practical advantage. Much of the literature on participatory research describes principles: trust, respect, mutual benefit, colearning. The principles are sound, but principles are cheap to endorse and hard to audit. Decisions, by contrast, leave traces. One can ask who was in the room when the research question was fixed, who signed the data use agreement, who is listed as first author, and who holds the keys to the database. A partnership that cannot answer those questions clearly has not yet decided how power will be shared, and in the absence of a decision the default is the institution, because the institution holds the money, the ethics approval, the publication channels and the legal personality that contracts require. The argument also explains why CBPR is harder than its admirers sometimes admit. Redistributing control has costs. It takes time that grant cycles do not budget for. It gives community partners the ability to say no, including to questions a researcher cares about and to publications a researcher needs for promotion. It exposes academic credit systems, built around individual authorship and priority, to claims they were not designed to handle. And it forces uncomfortable conversations about ownership in a legal environment that generally assumes the institution owns what its employees produce. A partnership that has never felt any of these costs has probably not shared much control. What the book covers and what it leaves out The chapters follow the life of a research project, because the decisions that matter arrive in roughly that order. The first chapter traces where participatory research came from, since its competing traditions still shape what practitioners mean by it. The second examines power directly: the structures, agreements and habits through which partnerships allocate control, and the ways control leaks back to the institution. The next three chapters take the core research tasks in turn: framing the question, collecting data, and interpreting findings. The sixth and seventh chapters take up the two issues where participatory ideals most often collide with institutional rules: authorship and credit, and ownership of data and intellectual property. The eighth asks what a partnership leaves behind when the funding stops, which is where many communities judge whether the whole exercise was worth their time. The book draws its examples mainly from health and environmental research, because that is where CBPR has been most fully developed, evaluated and argued over, and from research with Indigenous peoples, because Indigenous nations have done more than anyone to turn the ethics of research into enforceable rules. The principles transfer to education, urban planning, criminal justice, disability research and development studies, and examples from those fields appear where they sharpen a point. The book does not attempt a manual of specific data collection techniques, and it does not survey the large literature on citizen science in ecology and astronomy, where volunteers typically contribute observations to projects designed by others. That arrangement has great value, but it answers a different question from the one pursued here. A note on words "Community" is an unstable word, and the instability matters. A community can be a neighborhood, a tribal nation with its own government, a group of people sharing a diagnosis, a workforce, or a population defined by a researcher's sampling frame. These are very different partners. A sovereign nation can pass laws governing research on its lands; a neighborhood association cannot. A patient advocacy group may have paid staff and a policy agenda; a loosely connected group of residents may have neither. When this book speaks of the community partner, it means whichever people and organizations have a legitimate claim to represent those most affected by the research, and it treats the question of who holds that claim as part of the work, not a preliminary to it. Similarly, "researcher" here usually means an academic or institutionally employed researcher, and "institution" means the university, hospital, agency or firm that employs them, holds the grant and signs the contracts. The division is a convenience. Many community members are researchers by training, and many academic researchers come from the communities they study. But the division tracks something real: the unequal distribution of the resources that research requires, and the unequal exposure to the harms it can cause. The book is written for researchers who want to do this work well, for community organizations deciding whether and how to partner with a university, for funders and ethics committees trying to judge whether a proposal's claims to participation are genuine, and for students entering fields where community engagement is now expected. Each will find that the questions are the same from different sides of the table. The answers, when they are good, are agreed between those sides and written down. Chapter 1. Two Lineages of Participation Community-based participatory research did not begin as a method. It began as a set of arguments about who is entitled to produce knowledge, and those arguments came from two directions that have never been fully reconciled. One tradition treats participation as a way to make research work better: more relevant questions, better recruitment, more valid measures, findings that are actually used. The other treats participation as a way to shift power: research as a tool with which oppressed people analyze and change their own situation. The first tradition asks how communities can help research. The second asks how research can help communities, and whether research controlled by outsiders can ever do so. Most contemporary CBPR sits somewhere between them, and many of the disputes inside partnerships are really disputes about which lineage the project belongs to. The northern tradition: action research and the usefulness of involvement The first lineage is usually traced to Kurt Lewin, the German-American social psychologist who coined the phrase "action research" in the 1940s. In a 1946 paper in the Journal of Social Issues, "Action Research and Minority Problems," Lewin argued that research which produced nothing but books would not suffice, and proposed a cycle of planning, action and fact-finding about the results of action. His concern was practical. He worked with community relations bodies trying to reduce intergroup prejudice, and he observed that the people charged with acting on a problem learned more, and acted more effectively, when they were involved in studying it. Lewin's framing proved durable because it made participation an instrument of effectiveness. In organizational development, education and later in health services research, action research came to mean practitioners studying their own practice in iterative cycles, often with an outside facilitator. Teachers investigated their own classrooms; nurses studied their own wards. The approach was participatory in that the people closest to the work did the inquiry, but it did not necessarily challenge who held power in the organization. A hospital could run an action research project that improved discharge planning without asking whether patients should have a say in what counted as a good discharge. This lineage has much to recommend it. It produced a large body of practical knowledge about how to run iterative inquiry, how to combine reflection with change, and how to make findings usable. It also produced one of the central claims of modern CBPR: that involving the people affected by a problem can improve the science itself, by surfacing variables outsiders miss, by designing instruments that people actually understand, and by recruiting participants who would not answer a stranger's knock. When funders today justify community engagement, they usually do so in these terms. The southern tradition: knowledge as power The second lineage emerged in Latin America, Africa and South Asia in the 1960s and 1970s, and its founding texts read very differently. Paulo Freire's Pedagogy of the Oppressed, first published in Portuguese in 1968 and in English in 1970, argued that education could either domesticate people or free them, and that liberating education began with learners naming their own world. Freire's literacy circles in northeastern Brazil started from "generative themes" drawn from peasants' own lives, not from textbooks written elsewhere. The method assumed that poor and illiterate people already possessed knowledge about their circumstances, and that the educator's job was to help them analyze it critically, not to deposit correct information into them. The Colombian sociologist Orlando Fals-Borda carried a similar conviction into research. Working with peasant movements on Colombia's Atlantic coast in the 1970s, he developed what came to be called participatory action research, in which the separation between researcher and researched was deliberately broken down and the results were returned to the communities in forms they could use, including popular publications and illustrated histories. Fals-Borda and Muhammad Anisur Rahman later collected accounts of this work from several continents in Action and Knowledge: Breaking the Monopoly with Participatory Action-Research (1991). The subtitle states the thesis. Academic research held a monopoly on legitimate knowledge, and that monopoly was one of the ways in which the powerful stayed powerful. In Tanzania, Budd Hall and colleagues working in adult education in the early 1970s used the term "participatory research" for similar work, and Hall went on to help build international networks around it. In India, Rajesh Tandon founded the Society for Participatory Research in Asia (PRIA) in 1982. Robert Chambers, at the Institute of Development Studies in Sussex, popularized participatory rural appraisal in the 1980s and 1990s, a family of techniques such as community mapping, seasonal calendars and wealth ranking through which villagers could analyze their own conditions. Chambers's book Whose Reality Counts? Putting the First Last (1997) asked the question in its title of development professionals who had been designing programs from capital cities. This tradition does not regard participation primarily as a route to better data. It regards the conventional research relationship, in which an outsider extracts information and takes it away to be analyzed and published elsewhere, as itself a form of domination. The remedy is not to involve communities more skillfully in the outsider's project but to make the project theirs. Other currents Two further currents fed into contemporary CBPR and gave it some of its sharpest tools. Feminist researchers from the 1970s onward challenged the claim that good research required detachment, argued that the researcher's position shaped what she could see, and developed methods such as collaborative interviewing that treated participants as knowers. Their emphasis on reflexivity, the discipline of examining how one's own identity and interests shape the research, is now standard in participatory work. Indigenous scholars made the most direct challenge. Linda Tuhiwai Smith's Decolonizing Methodologies: Research and Indigenous Peoples, first published in 1999, opened with the observation that "research" is probably one of the dirtiest words in the Indigenous world's vocabulary. Her book documented how research had served colonial projects: classifying, measuring and displaying Indigenous peoples, removing their remains and artifacts, and producing knowledge that justified dispossession. Smith, a Māori scholar, set out an agenda in which Indigenous communities would define research priorities themselves, grounded in their own values and protocols. In Aotearoa New Zealand this became Kaupapa Māori research; in Canada and the United States, tribal nations and First Nations organizations developed their own research codes and review boards. Indigenous research ethics now supplies much of the most concrete guidance available on the questions this book treats as central: ownership, control and benefit. Disability rights activism contributed the slogan that many participatory researchers now use, "Nothing about us without us," which James Charlton took as the title of his 1998 book on disability oppression. The slogan compresses the core claim of the power-sharing tradition into five words. The formation of CBPR in public health The term "community-based participatory research" took hold in North American public health in the 1990s. A key text was a 1998 review in the Annual Review of Public Health by Barbara Israel, Amy Schulz, Edith Parker and Adam Becker, "Review of Community-Based Research: Assessing Partnership Approaches to Improve Public Health." It set out a list of principles that has been quoted ever since: CBPR recognizes community as a unit of identity; builds on strengths and resources within the community; facilitates collaborative, equitable partnership in all phases of research; promotes colearning and capacity building among all partners; integrates knowledge and action for mutual benefit; addresses health from positive and ecological perspectives; disseminates findings and knowledge gained to all partners; and involves a long-term commitment by all partners. Later versions added attention to cyclical and iterative processes and to the social inequalities that shape health. Israel and colleagues were writing from experience. The Detroit Community-Academic Urban Research Center, established in 1995 with funding from the US Centers for Disease Control and Prevention, brought together the University of Michigan schools of public health, community organizations on Detroit's east side and the city's health department. Its board, which brought representatives of community-based organizations together with the university, the health department and a health system, reviewed research proposals, and the partners adopted written CBPR principles to govern their joint work. The center became one of the most studied examples of long-term CBPR infrastructure, and the principles it developed influenced many later partnerships. The field grew quickly. Meredith Minkler and Nina Wallerstein edited Community-Based Participatory Research for Health, first published in 2003 and now in its third edition (2018, with Bonnie Duran and John Oetzel as co-editors), which became the standard reference. Community-Campus Partnerships for Health, a network founded in 1996, promoted partnership principles across universities and community organizations. In 2004 the US Agency for Healthcare Research and Quality commissioned a systematic review, led by Meera Viswanathan, of the evidence on CBPR, which concluded that the approach was promising but that the evidence base was thin and inconsistently reported. The journal Progress in Community Health Partnerships was launched in 2007 to give partnerships a venue for publishing their work. By the 2010s US federal funders, including the National Institutes of Health through its Clinical and Translational Science Awards and the Patient-Centered Outcomes Research Institute, created by the 2010 Affordable Care Act, were requiring or rewarding community and patient engagement. Parallel streams: patient involvement and participatory design The same questions surfaced in fields that rarely cited Freire or Fals-Borda. In the United Kingdom, patient and public involvement in health research became an organized expectation from the mid-1990s. The Department of Health established a body in 1996 that later became INVOLVE, and the National Institute for Health Research, created in 2006, built involvement into its funding requirements, so that grant applications must explain how patients and the public shaped the proposal and will shape the study. The language is different from CBPR; the unit is usually a patient group or a panel of lay advisers rather than a geographically bounded community, and involvement often stops short of shared authority. But the underlying logic, that people with lived experience of a condition know things about it that researchers do not, is the same. In Scandinavia, participatory design emerged in the 1970s from collaborations between computer scientists and trade unions. Projects such as UTOPIA, a Swedish and Danish collaboration with graphic workers in the early 1980s, set out to design workplace technologies with the workers who would use them, on the explicit premise that technology choices were also choices about power in the workplace. The tradition survives in human-computer interaction and in the design of public services, and it has contributed practical techniques for involving non-specialists in technical decisions: prototypes people can handle, scenarios they can argue about, and workshops in which the specialist's role is to make options visible rather than to choose among them. These streams matter here for two reasons. First, they show that the questions CBPR raises are not peculiar to public health or to marginalized neighborhoods; they arise whenever the people who produce knowledge and the people who live with its consequences are different people. Second, they illustrate how the same vocabulary can cover very different allocations of power. A patient panel that comments on a lay summary and a patient panel that holds a veto over the study design are both described as "involvement." Only the second has shifted control. What the evidence shows Because participatory research is often defended by its effects, it is fair to ask what those effects are. The honest answer is that the evidence is substantial but uneven, and that it supports some claims more strongly than others. The 2004 review for the Agency for Healthcare Research and Quality found that community involvement was associated with better recruitment and retention and with interventions better matched to local conditions, but that few studies were designed to test whether participation itself caused these results. Margaret Cargo and Shawna Mercer, reviewing the field in the Annual Review of Public Health in 2008, argued that the value of participatory research lay along three dimensions: translating research into practice, advancing self-determination, and pursuing justice. They noted that the second and third were rarely measured at all. A more useful approach to the evidence came from a realist review led by Justin Jagosh and published in the Milbank Quarterly in 2012. Instead of asking whether participatory research "works," the team asked what mechanisms produce its benefits and under what conditions. Examining a set of long-running partnerships, they identified partnership synergy, the combination of perspectives, skills and resources that no partner could achieve alone, as the central mechanism, and they found that its effects were cumulative: early trust made later co-governance possible, which produced culturally appropriate interventions, which in turn deepened trust. They described "ripple effects" that extended beyond the original project, including new partnerships, community capacity and systemic changes. A follow-up realist evaluation published in BMC Public Health in 2015 elaborated how conflict, when handled well, could strengthen rather than break these partnerships. The Engage for Equity study, led by Nina Wallerstein and John Oetzel with colleagues at the University of New Mexico and elsewhere, took the quantitative route. Surveying academic and community partners in roughly two hundred federally funded community-engaged projects in the United States, the team examined associations between partnership practices, such as shared decision-making and community influence over resources, and outcomes ranging from the quality of the research to changes in policy and community capacity. Their findings, reported in a series of papers including a 2020 special section of Health Education & Behavior, supported the proposition that the practices through which power is shared are associated with better outcomes, not just with partner satisfaction. None of this proves that every participatory project produces better science. Partnerships can fail, and a study can be highly participatory and badly designed. But the weight of the evidence runs in a consistent direction: the benefits of participation come less from the presence of community members than from what they are able to decide. That finding is the empirical counterpart of the argument this book makes on ethical grounds. Why the lineages still matter Institutional adoption brought CBPR money, legitimacy and training programs. It also brought a drift toward the northern lineage. When participation is justified by what it does for research quality and recruitment, it becomes easy to measure success by research outcomes and to treat the community as a stakeholder to be consulted rather than a partner with authority. Nina Wallerstein and Bonnie Duran, writing in the American Journal of Public Health in 2010, described CBPR as sitting on a continuum from a utilitarian pole, concerned with making research more effective, to an emancipatory pole, concerned with social transformation, and argued that the field's contribution to health equity depended on not losing the second pole. The two lineages give different answers to almost every concrete question a partnership faces. Who should decide the research question? The utilitarian answer is that researchers should, with community input to make it relevant; the emancipatory answer is that the community should, with researchers helping to make it answerable. Who owns the data? The utilitarian answer defers to institutional policy; the emancipatory answer treats the community's claim as primary. What counts as success? A publication and a funded follow-up grant, or a change in the conditions people live under? Andrea Cornwall and Rachel Jewkes, in a much-cited 1995 paper in Social Science & Medicine, "What Is Participatory Research?", made the point that the key difference between participatory and conventional research lies not in methods but in the location of power in the research process. A focus group is not participatory because it is a focus group; a survey is not extractive because it is a survey. What matters is who decided to use it, who designed it, who holds the results and who decides what they mean. This is the thread the rest of this book follows. A warning came from the same tradition. In 2001 Bill Cooke and Uma Kothari edited a collection provocatively titled Participation: The New Tyranny?, arguing that participatory methods in development had become a ritual that legitimized decisions already made, gave outsiders access to local knowledge without transferring control, and could override local power dynamics in the name of "the community." The critique was not that participation was bad but that the word had become detached from any redistribution of power. It is still the best test to apply to any project that calls itself participatory: what, specifically, can the community decide that it could not decide before? Most partnerships will never be purely emancipatory. They operate inside universities and funding systems that impose deadlines, require principal investigators, demand ethics approval and reward publication. The realistic goal is not purity but honesty: knowing which decisions are genuinely shared, which are delegated to the community, and which remain with the institution, and saying so openly. The next chapter examines how partnerships make those allocations, and how they go wrong. Chapter 2. Power Is the Method Every research partnership has a governance structure, whether or not anyone designed it. If no one decides how decisions will be made, they will be made by whoever controls the budget, signs the ethics application and answers to the funder. In most projects that is the academic principal investigator. This is not usually a matter of bad faith. It is a matter of defaults: universities are built to route authority through principal investigators, and the path of least resistance runs straight to them. Sharing power therefore requires deliberate design, and the design has to be specific enough to survive the moment when partners disagree. The ladder and its limits The most widely cited account of participation remains Sherry Arnstein's "A Ladder of Citizen Participation," published in the Journal of the American Institute of Planners in 1969. Arnstein, who had worked on federal urban renewal and anti-poverty programs, described eight rungs grouped into three bands. At the bottom were manipulation and therapy, which she called nonparticipation: programs designed to educate or "cure" participants rather than let them influence anything. In the middle were informing, consultation and placation, which she called degrees of tokenism: citizens could hear and be heard but had no assurance their views would be acted on. At the top were partnership, delegated power and citizen control, the degrees of citizen power. Her central point was blunt. Participation without redistribution of power is an empty and frustrating process for the powerless, because it allows those in power to claim that all sides were considered while ensuring that only some benefit. Arnstein's ladder has been criticized as too linear. Real partnerships do not sit on a single rung; a community may control recruitment while the university controls analysis, or hold a veto over publication while having little say in the budget. The International Association for Public Participation's spectrum, which runs from inform through consult, involve and collaborate to empower, is a less judgmental version often used in public agencies. Other writers have argued that the top of the ladder is not always the right goal: some communities do not want to run a research project, only to ensure it is done on their terms. The value of the ladder is not that it tells a partnership where to stand, but that it forces the question of where it actually stands on each decision, as opposed to where it says it stands. A practical way to use it is to break a project into its main decisions and ask of each who proposed, who could object, and who had the final word. The pattern that emerges is usually more informative than any overall characterization. Table 1 sets out how several common forms of community involvement tend to allocate a few of the decisions that matter most. Table 1. How common forms of community involvement allocate key decisions. Form of involvement Research question Budget Data custody Publication Community advisory board Researcher sets; board comments Researcher controls Institution Researcher decides Community-engaged study Researcher sets with input Small subcontract to partner Institution Partner reviews drafts Equitable CBPR partnership Jointly negotiated Shared, written allocation Joint agreement Joint approval and authorship Community-led research Community sets Community holds grant Community Community decides The table simplifies, but it makes one pattern plain. The forms differ less in how often community members attend meetings than in who holds money, data and the right to decide what gets published. A community advisory board that meets monthly and is warmly thanked in the acknowledgments may exercise less power than a partner organization that meets rarely but holds a third of the budget and a veto over data release. Structures that share control Partnerships have developed a set of tools for making power-sharing concrete. None is sufficient alone, and each can be hollowed out, but together they turn principles into obligations. The first is a written partnership agreement, sometimes called a memorandum of understanding or a research agreement. At its best it states the project's goals, the roles and responsibilities of each partner, how decisions will be made and disputes resolved, how money will be allocated, who will own and have access to data, how findings will be reviewed before publication, how authorship will be decided, and what happens if a partner withdraws. Many partnerships derive theirs from published models; the Detroit Community-Academic Urban Research Center's principles and the research agreement templates developed by Indigenous health organizations are frequently adapted. The value of such an agreement lies less in its legal force, which is often limited, than in the conversations required to write it. A partnership that has argued about data ownership before any data exist has a much better chance of resolving the argument when the data arrive. The second is a governance body with real authority. The difference between an advisory board and a steering committee is whether its decisions bind. Some partnerships give a community-majority steering committee formal approval over research questions, instruments and publications; others require consensus of all partner organizations on major decisions. Voting rules matter. A committee where the university holds as many seats as all community organizations combined, and the principal investigator chairs, will reproduce institutional control even if every vote is unanimous, because the agenda and the information flow run through the chair. The third is money. Budgets are statements of priorities, and in CBPR they are also statements of power. The principle is simple: community partners who do research work should be paid for it, and community organizations should receive a share of the grant sufficient to cover their real costs, including staff time, overhead and the unglamorous work of convening residents. The practice is harder. Universities often apply their full indirect cost rate to the entire grant while passing through only direct costs to subcontracted community partners; payment systems can take months to process invoices, which a small nonprofit cannot absorb; and some funders restrict what can be paid to people without formal credentials. Partnerships that take equity seriously negotiate these details before submitting the grant, and some have persuaded funders to name a community organization as co-principal investigator or as the lead applicant. The fourth is time and process: meeting at times and places community partners can attend, providing food and childcare, translating materials, and allowing enough time for community partners to consult the people they represent before a decision is made. These are sometimes treated as courtesies. They are in fact conditions for shared decision-making. A decision taken at a weekday afternoon meeting on campus, on the basis of a document circulated the night before in academic English, has not been shared regardless of who was invited. Positionality and the invisible allocation of authority Formal structures do not capture everything. Power also moves through expertise, language and social position. When a researcher explains why a proposed survey question would compromise validity, community partners may defer even when their objection was sound, because the vocabulary of methodology carries authority. When a community partner describes what residents will and will not tolerate, researchers may treat it as anecdote rather than evidence. The literature calls attention to this through the concept of positionality: the recognition that each partner's race, class, gender, education, institutional role and relationship to the community shape what they see and how they are heard. Michael Muhammad and colleagues, in a 2015 article in Critical Sociology, drew on interviews with academic and community partners to describe how researchers' identities, including their institutional privilege and, for researchers of color, their complicated position as both insiders and outsiders, affected trust, decision-making and outcomes. Their recommendation was not that researchers should confess their privilege and move on, but that partnerships should build regular, structured reflection into their work so that these dynamics could be named and addressed. A practical form of this is the power analysis, conducted periodically, in which partners ask of recent decisions who raised the issue, whose information was decisive, whose preferences prevailed, and whether anyone felt unable to object. Such an exercise is uncomfortable, and it works best when a trusted facilitator not employed by the university runs it. It is also one of the few ways to detect the gradual drift of authority back toward the institution, which rarely happens through any single decision. Who speaks for the community? Power-sharing presupposes a partner with whom to share it, and identifying that partner is itself an exercise of power. Researchers often begin with the organizations they already know, or those with the capacity to manage a subcontract: established nonprofits, health clinics, churches with paid staff. These organizations may be deeply rooted, or they may represent the more organized and resourced part of a community while speaking for the rest. A partnership built entirely with service providers may miss the people those providers serve; a partnership with elected tribal leadership may not capture the views of urban tribal members or of dissenting factions. There is no formula for resolving this, but there are better and worse practices. Better partnerships involve several organizations with different constituencies, create channels for residents who are not affiliated with any organization, and revisit the question as the work evolves. They recognize that communities contain their own inequalities, of gender, age, class, caste, immigration status and disability, and that the most affected people may be those with least voice even within their own community. Cooke and Kothari's warning, that participation can reinforce local power structures by treating the loudest voices as the community's voice, applies with full force here. For sovereign Indigenous nations the question has a clearer answer: the nation's government, through whatever processes it has established, has authority to approve or refuse research involving its citizens and lands. Many tribal nations in the United States have established research review boards or research codes; the Navajo Nation's Human Research Review Board, created in the 1990s, is among the best known. In Canada, the Tri-Council Policy Statement on ethical conduct for research, the joint policy of the federal research agencies, devotes a chapter to research involving First Nations, Inuit and Métis peoples, requiring engagement with the relevant community where research is likely to affect its welfare. These arrangements do not eliminate internal debate, but they locate authority where a partnership can find it. A worked example: Kahnawake The Kahnawake Schools Diabetes Prevention Project, begun in 1994 in the Mohawk community of Kahnawake near Montreal, is one of the most thoroughly documented examples of how governance can be designed to keep authority with a community over decades. The project was initiated in response to high rates of type 2 diabetes in the community and was designed as a school- and community-based prevention program with an evaluation component led jointly by community and academic researchers. Several features of its design repay attention. First, the project adopted a written code of research ethics, first in the mid-1990s and revised later, which set out the obligations of community and academic researchers, the principles of the partnership and the procedures for approving publications and handling data. Second, it established a community advisory board, made up of community members, which was not a sounding board for researchers' plans but a body with authority over the project's direction, including review of research proposals and of manuscripts. Third, the research team included community researchers employed by the project alongside academic investigators, so that expertise in both research methods and community life was held inside the team rather than traded across its boundary. Fourth, the project was structured from the outset as long-term: the intervention and its evaluation continued for years, and the partnership outlived several funding cycles. Researchers who worked with the project, including Ann Macaulay and Margaret Cargo, have written about both its strengths and its tensions: the time required for community review, the occasional disagreements between the advisory board and academic researchers over what to measure and publish, and the effort required to keep a partnership vital as individuals moved on. What made the arrangement durable was not the absence of tension but the fact that the procedures for resolving it had been agreed in advance and were understood to bind everyone, including the academics whose careers depended on publication. The lesson generalizes. When the community's authority is written into the project's founding documents and exercised routinely, it becomes part of how the work is done rather than an exceptional intervention. When it exists only as goodwill, it tends to be exercised only when nothing important is at stake. The hidden costs of sharing It is worth being candid about what power-sharing costs, because partnerships that pretend it is free tend to abandon it when the costs arrive. For researchers, the costs are mainly time and control. Joint decision-making is slower. Community review of instruments and manuscripts adds months. A community partner's objection can remove a variable a researcher wanted, delay a paper past a promotion deadline, or end a line of inquiry altogether. Researchers also carry reputational risk inside their institutions, where colleagues may regard participatory work as less rigorous, and where tenure committees may not know how to credit a jointly authored report written for a city council. For community partners the costs are different and often larger. Participation consumes the time of people who are already stretched: staff of small organizations with no slack, residents with jobs and caring responsibilities. Community partners carry the relational risk of the project, because it is their credibility with neighbors that is spent if the research disappoints or harms. They are frequently asked to educate researchers about their community, a form of labor that is seldom paid and rarely acknowledged. And when a project ends, the researchers move to the next grant while community partners remain to answer for what happened. Recognizing these asymmetries is part of sharing power. A partnership that compensates the time of community partners, that protects their credibility by honoring agreements, and that accepts slower timelines as the price of legitimacy has done more to share power than one with an elaborate governance chart and no budget line for community staff. When partners disagree Genuine power-sharing means that the community can say no, and eventually it will. A community partner may object to a question that researchers consider scientifically essential, refuse to allow a finding to be published, or conclude that the partnership is no longer worth its time. How a partnership handles these moments reveals what it actually is. The agreement should anticipate them. Common provisions include a period of discussion before any partner can withdraw, a commitment to seek mediation from a trusted third party, and rules about what happens to data and publications if the partnership dissolves. Some agreements distinguish between the community's right to review and comment on all publications and its right to veto publications that would identify or harm the community; others give a broader veto. Researchers sometimes worry that such provisions threaten academic freedom. The counterargument is that academic freedom protects researchers from interference by their employers and the state; it does not entitle them to publish data that others contributed on condition of shared control. Where the conflict is real, the honest course is to negotiate the terms before the research begins, when both parties can walk away without loss. Disagreement is not failure. Jagosh and colleagues found that partnerships which worked through conflict openly often emerged stronger, because the experience demonstrated that community voices had real weight. What damages partnerships is not disagreement but the discovery that the community's agreement was never actually required. Chapter 3. Whose Question Is It? Of all the decisions in a research project, the choice of question is the one that most constrains every other. It determines what data will be collected and what will be ignored, which people will be counted and which variables will explain them, what kind of answer is possible and to whom it will be useful. A community brought in after the question has been fixed can improve the recruitment materials, suggest better wording for a survey and help interpret the results, but it cannot change what the study is about. For that reason the question is where participatory research most often fails without anyone noticing, because the failure takes the form of an absence: the question the community would have asked and never got to. How questions are usually made In conventional research, questions come from three sources: the researcher's discipline, which defines what is interesting and publishable; the funder, which defines what will be paid for; and the researcher's own career, which rewards questions that extend a line of work. None of these sources is illegitimate. Disciplines accumulate knowledge by building on previous findings; funders have mandates; researchers need to specialize. But none of them runs through the people whose lives the research is about. The consequences are visible in what gets studied. Funding in health research is organized largely by disease and organ system, while people experience their health through housing, work, income, neighborhood safety and the behavior of institutions. A community asked to partner on a diabetes study may care more about the absence of a grocery store, the cost of insulin, or the fact that the clinic closes before people get home from work. Those are researchable questions, but they fall between funding streams, and a researcher with a diabetes grant has limited room to pursue them. Framing also matters within a topic. Research on marginalized communities has a long tradition of what critics call deficit framing: asking what is wrong with the community, its behaviors, its knowledge, its compliance, rather than what is wrong with the conditions and institutions it faces. A study asking why residents fail to attend appointments will look for explanations in residents; a study asking what the clinic does that makes attendance difficult will look in the clinic. The data may overlap, but the findings, and the interventions they justify, will differ. Israel and colleagues' principle that CBPR builds on strengths and resources within the community is partly a corrective to this habit. It asks researchers to begin from what a community already does to protect its health, and from the conditions that undermine those efforts. Communities that arrive with questions Some of the most consequential participatory research began when a community already had a question and could not get anyone to take it seriously. The Flint water crisis is the best-known recent case. After the city of Flint, Michigan, switched its water source to the Flint River in April 2014 under state-appointed emergency management, residents complained of discolored, foul-smelling water, rashes and hair loss. Officials repeatedly assured them the water was safe. One resident, LeeAnne Walters, whose household water tested very high for lead, contacted Marc Edwards, a civil engineer at Virginia Tech who had previously exposed lead contamination in Washington, DC. In 2015 the Virginia Tech team and Flint residents organized a citywide sampling effort in which residents collected water samples from their own homes using kits the researchers supplied, and the results showed lead levels well above federal action levels in many homes. At about the same time the pediatrician Mona Hanna-Attisha and colleagues analyzed children's blood lead records and found that the proportion of children with elevated blood lead levels had increased after the water switch, most sharply in the neighborhoods with the highest water lead. The two lines of evidence together forced officials to acknowledge the problem. Flint is not a textbook CBPR partnership; the collaboration was assembled in an emergency and some residents later criticized how credit and attention were distributed. But it shows with unusual clarity what happens when research follows a community's question. The residents already knew something was wrong. Official monitoring had been designed, and in some respects conducted, in ways that failed to detect it. What the researchers contributed was not the question but the capacity to answer it in a form that institutions could not dismiss. The environmental justice movement has produced many such cases, because environmental hazards are often first detected by the people who live beside them. In California's San Joaquin Valley, residents of small rural communities, many of them Latino farmworker households, had long worried about nitrate and arsenic in their drinking water. Working with the Community Water Center, a local advocacy organization, researchers including Carolina Balazs and Rachel Morello-Frosch at the University of California, Berkeley, examined the distribution of contamination across community water systems and found that systems serving higher proportions of Latino residents and renters tended to have higher nitrate levels. Balazs and Morello-Frosch later used the experience to argue, in a 2013 paper in the journal Environmental Justice, that community participation strengthens what they called the "three Rs" of science: its rigor, relevance and reach. Community partners, they found, directed attention to the right questions: not only whether water exceeded regulatory limits, but who was exposed, who paid for the costs of treatment and replacement water, and which small systems lacked the capacity to fix their problems. Those questions made the research both more accurate about the real burden and more useful for policy. Methods for finding the question together When a partnership begins without a pre-formed question, it needs ways to generate one jointly. Several approaches have become established. Community assessment brings together existing data, from health departments, census records and service providers, with residents' own accounts, gathered through listening sessions, community forums or door-to-door conversations. The important feature is that residents help interpret the data rather than simply supplying stories to illustrate it. A map of asthma hospitalizations means one thing to an epidemiologist and something richer to residents who know which blocks have mold-infested housing, which schools sit next to truck routes and which landlords ignore complaints. Photovoice, developed by Caroline Wang and Mary Ann Burris and described in a 1997 article in Health Education & Behavior, gives participants cameras to document their own community's strengths and concerns, followed by group discussion of the photographs and, typically, an exhibition for policymakers. Wang and Burris drew explicitly on Freire's idea of critical consciousness and on feminist theory. Photovoice has been used thousands of times since, sometimes as a data collection method within a study defined elsewhere, but its original purpose was agenda-setting: letting people who are rarely asked show what they think matters. Priority-setting partnerships formalize the process at a larger scale. The James Lind Alliance, established in the United Kingdom in 2004 and now part of the National Institute for Health and Care Research, brings together patients, carers and clinicians to identify and rank the most important unanswered questions about a particular condition. Its method moves from an open survey of uncertainties, through checking which are genuinely unanswered in the existing literature, to a facilitated workshop at which participants agree a top ten. The resulting lists have repeatedly diverged from the questions researchers had been funding, with patients and carers placing more weight on quality of life, side effects and practical management than on new drugs. Funders, including the National Institute for Health and Care Research itself, have used these lists to commission studies. Concept mapping and deliberative workshops offer other structured ways to reach agreement. What these methods share is a sequence: open generation of concerns by those affected, joint sorting and prioritization, and a deliberate step in which researchers help translate priorities into questions that can be answered with available methods and resources, while the community checks that the translation has not changed the meaning. Translation without capture That last step is where control is most easily lost. Residents may name a concern such as "our kids can't breathe here," which must be turned into a research question: what exposures, measured how, compared with what, over what period, linked to what outcomes? Every one of these choices involves technical judgment, and the person making it holds power. If researchers make the choices alone, the resulting question may be answerable but no longer the community's. Good practice keeps the translation visible. Researchers lay out options and their consequences in plain terms: measuring particulate matter at fixed monitors is cheaper and comparable with official data, while personal monitors carried by children capture actual exposure but are costlier and burdensome; comparing hospitalization rates across neighborhoods is fast but will not show what happens inside homes. The community chooses among options with an understanding of the trade-offs, and it can insist on the more expensive or slower design if that is the one that answers its question. Where no feasible design can answer the community's question, researchers say so, and the partnership decides whether a narrower study is still worth doing. Translation also involves outcomes. What would count as an answer, and what would be done with it? Communities often want research that can be used in a specific venue, a zoning hearing, a school board decision or a funding application, and the design must produce evidence in a form that venue will accept. This is not a corruption of science. It is the same consideration that leads a pharmaceutical trial to use endpoints a regulator will recognize. Questions that harm Communities sometimes resist questions not because they lack interest in the topic but because they have learned what such questions can do. Research on stigmatized conditions, such as alcohol use, mental illness, HIV or violence, produces findings that attach to the whole group studied, including people who never took part. A community-level finding travels further than any individual's data, and it can be used by outsiders to justify prejudice, disinvestment or intervention. The Barrow Alcohol Study is a standard warning. In 1979 researchers from the University of Pennsylvania, working under contract to the North Slope Borough in Alaska, surveyed residents of the Iñupiat town of Barrow (now Utqiaġvik) about drinking. The results were released at a press conference in 1980 before the community had reviewed them, and national newspapers reported that alcoholism was threatening the survival of the Iñupiat people. The borough's bond rating reportedly suffered, residents felt publicly humiliated, and the episode became a case study in research ethics courses. Edward Foulks, one of the investigators, later wrote an account of what went wrong, describing a partnership in which the community had commissioned the research but had no say over how its findings were framed and released. The data might have been accurate; the harm lay in a framing that described the community through its pathology and in a release process that bypassed the people described. The lesson for question-setting is that communities have a legitimate interest in how they will be represented, and that this interest is engaged at the start, not only at publication. A community may agree to study alcohol use if the question is framed around the availability of treatment, the effect of alcohol outlet density or the success of community-led sobriety movements, and refuse if it is framed as the prevalence of drinking. Such preferences can look to researchers like spin. Often they are an accurate recognition that the choice of what to measure determines what story will be told, and that the community, not the researcher, will live with the story. Group harm is poorly handled by conventional research ethics, which focus on the risks to individual participants. An individual can consent to answer questions about her drinking; she cannot consent on behalf of her town to be described as a town of drinkers. Participatory research offers one of the few mechanisms through which a group can weigh these risks collectively, and it does so best at the moment the question is chosen. Questions researchers would rather not ask The reverse problem also arises. Communities sometimes want questions that researchers, or their institutions, would rather avoid. Residents near a university may want to study the university's own effects on housing costs and displacement. Patients may want to examine how a hospital treats them. Communities policed heavily may want research on police conduct rather than on crime. Workers may want to study their employer's safety practices, and the employer may be a research funder. These questions test whether a partnership's commitment to community priorities holds when the priorities become inconvenient. The institution may have legitimate concerns about conflicts of interest, and researchers may fear retaliation from colleagues or funders. But a partnership that systematically steers away from questions implicating the institutions involved has a structural bias worth naming. One sign of a mature partnership is that it has at least once pursued a question that made its academic partner uncomfortable, and survived. When the funder has already chosen A frequent difficulty is that the question is partly fixed before any partnership exists. Funding announcements define topics, and applications have deadlines that leave little time for community deliberation. Researchers then approach community organizations with a grant already half-written and ask them to sign a letter of support. There are honest ways to handle this. The first is to build partnerships before funding opportunities arise, so that when an announcement appears, the partnership already has a list of priorities and can decide together whether this opportunity matches any of them. Long-standing partnerships such as the Detroit center operated in exactly this way, reviewing proposed projects against the partnership's agreed priorities. The second is to be explicit about what is fixed and what is open: the topic is set by the funder, but within it the community will choose the specific question, the population and the outcomes. The third is to decline. A community organization that turns down an opportunity because it does not serve its priorities is exercising exactly the power participatory research claims to respect, and researchers who treat such refusals as obstacles have misunderstood the enterprise. Some funders have changed their practices in response. Planning grants that pay for partnership development before a full proposal, application processes that require evidence of community involvement in defining the question, and review panels that include community members all shift the point at which communities enter. The Patient-Centered Outcomes Research Institute in the United States, for example, has required applicants to describe how patients and other stakeholders were involved in developing the research question, and includes patients and stakeholders as reviewers alongside scientists. Such measures do not guarantee genuine partnership, but they make it harder to recruit a community after the fact. The question is the first allocation of power in a research project. Once it is made, the choice of data follows, and with it a new set of decisions about who collects that data, on what terms, and at whose risk. That is the subject of the next chapter. Hashtags: #CommunityBasedParticipatoryResearch #CBPR #ParticipatoryResearch #CommunityEngagement #CollaborativeResearch #ParticipatoryActionResearch #ActionResearch #CommunityAcademicPartnerships #PowerSharing #SharedDecisionMaking #CommunityGovernance #ResearchEquity #SocialImpact #CommunityEmpowerment #DecolonizingMethodologies #IndigenousResearch #Photovoice #CommunityAssessment #PrioritySetting #ResearchPartnerships #DataOwnership #CommunityAuthorship #CoLearning #HealthEquity #FutureOfParticipatoryResearch

Latest Book Releases:

WELCOME TO THE INTERNATIONAL STUDENTS LIBRARY

bottom of page