Welcome to the VBNN Digital Library
Unlock a Vast Knowledge Ecosystem
Featuring over 30,000 books, academic papers, illustrations, and expert insights—continuously updated to support your research and professional growth.
Welcome to our library!
Here, you will find an exclusive collection created 100% by our own faculty, meaning you will not find these resources anywhere else. Over the last 20 years, our team has written much more than what is currently online, and we are actively working to upload our complete back catalog. We update our platform regularly, so be sure to check back from time to time. If you ever need help finding a specific resource, you can always contact us!
Maximize Your Access
Log in to instantly view and download tailored resources directly aligned with your specific program and curriculum.
Ready to begin? Sign in above to explore your personalized dashboard.
Please note: Login is only possible using your institutional email address; otherwise, the system will not recognize your account.
VBNN Library AI
Introducing our fully integrated Library AI. Designed to support your research, you may submit inquiries in any language and receive precise, evidence-based responses drawn exclusively from our published scholarly articles and textbooks.
Search...
Latest Publications:
Search this site
Results found for empty search
- Mixed-Methods Instrument Design (Developing and Validating Psychometric Surveys)
Download the Book (PDF): Introduction Every questionnaire makes a quiet promise. When a researcher adds up a respondent's answers to twelve items about workplace burnout and reports a score of 38, the number carries an implied claim: that it reflects something real about that person, that it would come out much the same if we asked again next week, that it means the same thing for a nurse in Leeds as for a teacher in Lagos, and that a 38 signals more burnout than a 31. None of those claims is guaranteed by the arithmetic. They are earned, or not, by the way the instrument was built. Most researchers encounter measurement at the far end of this process. They pick up a published scale, cite its Cronbach's alpha, and move on to the substantive question they care about. That is often reasonable. But sooner or later many of them discover that no existing instrument captures what they need. The construct is new, or the population is one the existing measures were never tested on, or the available scales were written for a context that no longer exists. At that point they face a task that looks deceptively simple, writing some questions, and turns out to be one of the most methodologically demanding things a social, health, or organisational scientist can undertake. This book is a guide to that task. It follows the path that has become the standard route for serious instrument development: begin with qualitative work that listens to the people whose experience the instrument will measure, turn what they say into candidate items, test those items with experts and with respondents in cognitive interviews, pilot the draft in a field sample, and then gather the quantitative evidence that the scores behave as the construct says they should. In the vocabulary of mixed methods research this is an exploratory sequential design: a qualitative phase whose results build a quantitative phase. It is the design that most scale development in health outcomes, psychology, education, and management now follows, whether or not the authors use that name. Why the sequence matters The central argument of this book is that validity is not a property you test for at the end. It is an argument you build from the beginning, and every stage of the sequence contributes a different kind of evidence to it. The focus group transcript is evidence about what the construct means to the people who live it. The item-writing log is evidence that each question traces back to something respondents actually said or a theory actually requires. The expert ratings are evidence that the item set covers the domain without straying outside it. The cognitive interviews are evidence that respondents understand the questions in the way the designer intended. The pilot data are evidence about how the items behave as a set. The final validation study is evidence about how the scores relate to other things in the world. If any link in that chain is weak, the later links cannot repair it. A beautifully fitting confirmatory factor model tells you that a set of items share a common source of variance. It cannot tell you that the common source is the thing you meant to measure rather than, say, a shared tendency to agree with positively worded statements, or a shared misreading of a key term. Only the earlier stages can rule those alternatives out. This is why the qualitative phases are not a warm-up or a courtesy to stakeholders. They are where the meaning of the scores is established, and the statistics that follow are tests of whether that meaning has survived translation into numbers. The same logic runs in reverse. Qualitative work that is never tested quantitatively produces an instrument whose items sound right but whose scores may be unreliable, redundant, or dominated by one sub-theme. Mixed methods instrument design is valuable precisely because each strand checks the other. The psychometric data can reveal that two items respondents described in different words are in fact measuring the same thing, or that an item everyone in the focus groups endorsed does not discriminate between people at all. When that happens, the researcher goes back to the qualitative record to understand why. What the reader will need This book is written for researchers who are competent in their own field but not specialists in psychometrics or qualitative methodology: a public health researcher who needs a measure of vaccine trust in a particular community, an organisational psychologist building a measure of psychological safety for hybrid teams, a nurse scientist developing a patient-reported outcome for a condition that has none, an education researcher who wants to measure students' sense of belonging in online programmes. It assumes some familiarity with basic statistics, correlation, and regression, and a general acquaintance with what an interview or a focus group is. It does not assume that the reader has run a factor analysis or coded a transcript. The book does not try to be a statistics manual. Where analytic methods matter, such as exploratory and confirmatory factor analysis, item response theory, reliability coefficients, and tests of measurement invariance, the chapters explain what each method does, what question it answers, what its conventional decision rules are, and where those rules mislead. The Further Reading section at the end points to the texts that go deeper. The aim is judgement: to help the reader understand why each step exists, so that when a real project departs from the textbook, as every real project does, they can make a defensible choice. How the book is organised Chapter 1 sets out the measurement problem itself. It explains what a construct is, what validity means in its modern sense as an argument about the interpretation and use of scores, and how the exploratory sequential design maps onto the chain of evidence that argument requires. Chapter 2 deals with the work that should happen before anyone is recruited: defining the construct, reviewing existing measures, specifying the population and intended use, and making early decisions about the structure of the instrument. Many failed scales fail here, because a vague construct produces a vague item pool that no amount of later statistics can sharpen. Chapter 3 turns to the qualitative elicitation phase, and especially to focus groups. It covers who to recruit, how many groups to run, how to write a discussion guide that elicits the construct without planting it, when individual interviews serve better, and how to analyse transcripts with item generation in mind. Chapter 4 is about the craft of item writing: converting themes and quotations into candidate items, choosing response formats and recall periods, deciding how large an initial pool should be, and avoiding the well-documented ways in which wording, context, and scale labels change the answers people give. Chapter 5 covers expert review and the quantification of content validity, including the content validity index and the content validity ratio, how to choose and brief experts, and how to combine expert judgement with the judgement of the target population. Chapter 6 is devoted to cognitive interviewing, the method that lets a researcher watch a question being answered. It explains the cognitive model of survey response on which the method rests, the difference between think-aloud and verbal probing, how many interviews and rounds are needed, and how to analyse what respondents say. Chapter 7 covers the pilot study and item analysis: item distributions, item-total statistics, exploratory factor analysis, reliability, and the use of item response theory, with an emphasis on how quantitative and qualitative evidence together should drive decisions about which items to keep. Chapter 8 addresses construct validation in the full sense: confirmatory factor analysis, convergent and discriminant evidence, known-groups comparisons, measurement invariance, test-retest reliability, and responsiveness to change, along with the role of qualitative follow-up in explaining anomalous results. The Conclusion draws out what follows from treating instrument development as a single, cumulative argument, for how projects are planned, funded, reported, and judged. Throughout, the book uses examples from real, published instrument development programmes where the details are documented, and constructed illustrations, clearly presented as such, where a worked example makes a point more clearly than a published case can. It does not invent studies or statistics. Where the evidence on a methodological question is thin or contested, it says so. One final point about scope. The methods described here were developed largely for self-report instruments: questionnaires in which respondents describe their own attitudes, experiences, behaviours, or states. Much of what follows applies equally to observer-rated scales, interview schedules, and proxy reports, but the examples and the emphasis are on self-report, because that is where most researchers begin and where the gap between what a question says and what a respondent hears is widest. Chapter 1: The Measurement Problem In 1955 Lee Cronbach and Paul Meehl published a paper in the Psychological Bulletin that still frames how researchers think about questionnaires. Their problem was simple to state. Many of the things psychologists wanted to measure, such as anxiety, intelligence, or authoritarianism, could not be pointed at. There was no ruler to lay against them, no criterion so obviously correct that a new test could be judged by how closely it matched. What, then, could it mean to say that a test measured anxiety? Their answer was that the meaning of a test score comes from the network of theoretical relationships in which the construct is embedded. If anxiety, as theorised, should rise before examinations, correlate with physiological arousal, fall with effective treatment, and be distinguishable from depression, then a valid anxiety score should show those patterns. Validation becomes the process of checking whether it does. That idea, construct validity, has been refined, criticised, and extended for seventy years, but its core has held. It is the right place to begin a book about building instruments, because it explains why the work is so demanding. A researcher who builds a new scale is not just writing questions. They are making a claim about an unobservable attribute of people and committing to a body of evidence that can support or undermine that claim. What a construct is and what a scale claims A construct is a concept that a researcher has deliberately defined for scientific use. Burnout, health literacy, organisational commitment, fear of falling, and perceived neighbourhood safety are all constructs. They are abstractions, but they are not arbitrary. Each has a definition that specifies what it includes, what it excludes, and how it is expected to relate to other constructs. The word "construct" is a reminder that the concept was built, and that other researchers might build it differently. A psychometric scale is a set of observable indicators, usually questionnaire items, whose responses are combined to produce a score that represents a person's standing on the construct. The combination is typically a sum or an average, sometimes a weighted estimate from a statistical model. The scale therefore rests on an assumption about how the construct and the items are related. In the most common case, the reflective model, the construct is assumed to cause the item responses: people who are more burnt out tend to agree more strongly with "I feel emotionally drained by my work" and "I feel used up at the end of the workday", and the items correlate with each other because they share that cause. In the less common formative model, the items define the construct rather than reflect it: socioeconomic status, for example, is often treated as formed by income, education, and occupation, which need not correlate with each other at all. Chapter 2 returns to this distinction, because choosing the wrong model leads to the wrong statistics and, worse, to discarding the wrong items. Whatever model is chosen, a scale makes several distinct claims when it produces a score. It claims that the items are relevant to the construct and that together they represent it adequately. It claims that respondents understand the items as intended and answer them on the basis of the attribute being measured, not something else. It claims that the scores are sufficiently consistent to be useful. It claims that the internal structure of the responses matches the structure of the construct, for example that a scale said to have three subscales really does have three. And it claims that the scores relate to other variables as theory predicts. Each claim can fail independently of the others. Validity as an argument Early textbooks described validity as a trinity: content validity, criterion validity, and construct validity, each established by a different kind of study. That framing has been abandoned by the measurement profession. Samuel Messick, in a series of papers culminating in a 1995 article in the American Psychologist, argued that validity is a single, unified concept: the degree to which evidence and theory support the interpretations of test scores for their proposed uses. Content and criterion evidence are not separate kinds of validity. They are sources of evidence bearing on one question, whether a particular interpretation of the scores is justified. The Standards for Educational and Psychological Testing, published jointly by the American Educational Research Association, the American Psychological Association, and the National Council on Measurement in Education, adopted this view. The current edition, published in 2014, identifies five sources of validity evidence: evidence based on test content, on response processes, on internal structure, on relations to other variables, and on the consequences of testing. Chapter 8 uses this framework to organise the final stage of validation, but it is worth seeing now how neatly those five sources line up with the stages of mixed methods development. Michael Kane took the argument a step further. In a framework he developed over two decades and summarised in a long 2013 article in the Journal of Educational Measurement, Kane proposed that validation should begin by writing down the interpretation and use argument: the chain of inferences that leads from a respondent's answers to the conclusion someone will draw from the score. For a questionnaire used to screen employees for burnout, that chain might run from observed responses to a score (scoring inference), from the score to the person's typical level across occasions and item sets (generalisation inference), from that typical level to their standing on burnout as defined (extrapolation or explanation inference), and from their standing to a decision about referral (decision inference). Each inference rests on assumptions, and each assumption needs evidence. The more ambitious the use, the more evidence is needed. The practical consequence of Kane's approach is that validity is not a certificate a scale receives once. It belongs to a specific interpretation for a specific use in a specific population. A scale can have strong support for use as a research outcome in adults and weak support for use as a clinical screening tool in adolescents. Phrases such as "a validated scale" are shorthand at best and misleading at worst. When developers design a programme of work, they should know from the outset which interpretation and use they intend to support, because that determines which evidence they must collect. There is a dissenting view that the reader should know about. Denny Borsboom and colleagues argued in a 2004 paper in Psychological Review that validity should be defined more narrowly: a test is valid for measuring an attribute if the attribute exists and variations in it causally produce variations in the test scores. On this view, questions about consequences and uses belong to a different category of evaluation. The debate matters for philosophers of measurement, but for the working instrument developer the two positions converge on the same practical demand. Whether validity is defined as an argument about interpretations or as a causal claim about an attribute, the developer must show that the responses are produced by the construct rather than by something else. Establishing how responses are produced is exactly what the qualitative stages of the sequence are for. Why mixed methods It is possible to develop a scale without qualitative work. A researcher can read the literature, write items from theory, give them to a large sample, and let factor analysis sort them out. Many scales have been built this way, and some are good. But the approach has a recurring weakness: it can only find the construct that the researcher already had in mind. If the people being measured experience the phenomenon in ways the literature has not described, those experiences will never make it into the item pool, and no statistical analysis can detect what was never asked. Health outcomes research learned this lesson with particular force. Patient-reported outcome measures, questionnaires that ask patients to rate their own symptoms, function, and quality of life, were for many years written largely by clinicians. When developers began systematically asking patients what mattered to them, they found that clinician-written measures often missed dimensions patients considered central, such as fatigue in rheumatoid arthritis, and included items patients considered irrelevant. The United States Food and Drug Administration's 2009 guidance on patient-reported outcome measures used to support labelling claims made patient input a formal expectation: developers were expected to show, with documented qualitative research, that the concepts measured were important to patients and that patients understood the items. Two good research practices reports from the International Society for Pharmacoeconomics and Outcomes Research, published by Donald Patrick and colleagues in Value in Health in 2011, set out how concept elicitation and cognitive interviewing should be done to meet that expectation. The FDA's later patient-focused drug development guidance series has continued in the same direction. The same logic has spread well beyond health. In a 2004 paper in Psychological Assessment, Dawne Vogt, Daniel King, and Lynda King argued that focus groups with members of the target population enhance content validity by clarifying the domain, identifying its facets, and supplying the language respondents actually use. They illustrated this with their own work developing a measure of deployment stressors among military personnel, in which focus groups surfaced experiences that the existing literature had not captured. Anthony Onwuegbuzie and colleagues, writing in the Journal of Mixed Methods Research in 2010, formalised a ten-phase "instrument development and construct validation process" that interleaves qualitative and quantitative steps throughout, not just at the start. Brendon Luyt, in the same journal in 2012, proposed a framework for mixing methods across development, validation, and revision. Qualitative work contributes three things that are hard to get any other way. First, it defines the domain from the inside: what the phenomenon consists of, for the people who experience it, including facets the researcher did not anticipate. Second, it supplies language: the words and phrases respondents use, which make items easier to understand and less likely to be misread. Third, it provides evidence about response processes: how respondents interpret and answer questions, which is the only direct evidence that responses are produced by the construct rather than by misunderstanding, social desirability, or guesswork. Quantitative work contributes what qualitative work cannot. It shows how items behave across a large number of respondents, whether they vary enough to be informative, whether they cluster in the way the construct's structure implies, whether scores are consistent, and whether they relate to other variables as expected. It also shows when the qualitative picture was wrong or incomplete, for instance when two themes that seemed distinct in focus groups turn out to be empirically indistinguishable. The exploratory sequential design John Creswell and Vicki Plano Clark, whose textbook Designing and Conducting Mixed Methods Research has shaped the vocabulary of the field, describe three core mixed methods designs. In a convergent design, qualitative and quantitative data are collected in parallel and compared. In an explanatory sequential design, quantitative results come first and qualitative work follows to explain them. In an exploratory sequential design, qualitative exploration comes first and its results are used to build a quantitative phase. Instrument development is the textbook application of the exploratory sequential design: the qualitative phase explores the construct, an intermediate phase builds the instrument, and the quantitative phase tests it. In practice, however, a well-run instrument development project is not purely sequential. It contains smaller loops. Cognitive interviewing, a qualitative method, sits in the middle of a mostly quantitative development stream. The pilot study, a quantitative step, frequently produces results that send the researcher back to the transcripts. The final validation study may include a qualitative component to explain unexpected findings, which is an explanatory sequential loop nested within the larger exploratory design. Thinking of the whole project as a chain of evidence, rather than as two blocks of methods, helps keep these loops purposeful. Table 1 sets out the stages covered in this book, what each produces, and which of the Standards' sources of validity evidence each primarily supplies. Table 1. Stages of mixed methods instrument development and the validity evidence each supplies. Stage Main method Key output Primary evidence source Construct specification Literature and theory review Written construct definition and domain map Test content Qualitative elicitation Focus groups, interviews Themes, facets, respondent language Test content Item generation Structured writing from themes Candidate item pool with audit trail Test content Expert and target review Rating panels, CVI Relevance and coverage ratings Test content Cognitive interviewing Think-aloud, verbal probing Revised items, comprehension evidence Response processes Pilot and item analysis Field survey, EFA, IRT Reduced item set, provisional structure Internal structure Construct validation CFA, correlations, invariance, retest Structural, relational, and stability evidence Internal structure; relations to other variables The fifth source in the Standards, evidence based on consequences, does not map neatly onto a development stage, because consequences arise once an instrument is in use. It still deserves attention during development, particularly for instruments that will inform decisions about individuals, such as clinical screening or personnel selection. If a scale systematically produces lower scores for one group because of an item that group interprets differently, the consequences of using it may be unjust even if the scale is internally consistent. Several stages, notably the qualitative elicitation with diverse participants and the invariance testing in the final phase, bear on that risk. The audit trail A practical implication of treating validity as a cumulative argument is that every stage must leave a record that the next stage can use and that a reader can inspect. This book calls that record the audit trail, borrowing the term from qualitative research, where it refers to documentation that allows an outsider to follow the analyst's decisions. For instrument development, the audit trail links each final item back through every decision made about it. A reader should be able to take an item from the published scale and see which facet of the construct definition it represents, which qualitative themes and quotations it was derived from, what the experts said about it, what changed after cognitive interviews and why, how it performed in the pilot, and why it survived item reduction when others did not. In practice this is usually a spreadsheet or database, an item-tracking matrix, with one row per item and columns for each stage. It sounds bureaucratic. It is the single most useful document a development team can keep. The audit trail serves three purposes. It enforces discipline during development, because a team that must record why each item exists is less likely to include items for no reason. It makes decisions revisitable, because when pilot data show that an item behaves oddly, the team can go back to the cognitive interview notes and see whether respondents had flagged a problem. And it makes the finished instrument evaluable by others. Reviewers and later users can judge the content validity of a scale far better when they can see where its items came from. Consider a constructed illustration. A team developing a measure of financial stress among university students includes the item "I have skipped meals to save money." A reader of the final paper wants to know why that item is there and whether it measures stress or simply poverty. With an audit trail, the answer is available: the item was written from a focus group theme labelled "trade-offs with basic needs", which emerged in five of six groups; experts rated it highly relevant but two flagged the stress-versus-hardship distinction; cognitive interviews showed that students interpreted it as a behaviour, not a feeling, so the team kept it in a separate behavioural subscale; and in the pilot it loaded on that subscale rather than on the affective one. Without the audit trail, the reader has only the item and a factor loading, and can only guess. What can go wrong It is worth closing this chapter with the ways measurement fails, because each stage of the sequence exists to prevent one or more of them. The construct can be underrepresented. Messick called this construct underrepresentation: the instrument omits important facets of the construct, so scores reflect only part of it. A loneliness scale that asks only about lacking companionship, and never about feeling misunderstood, underrepresents loneliness if the latter is part of the construct as defined. Careful construct specification and qualitative elicitation are the main defences. The scores can contain construct-irrelevant variance, Messick's second threat: systematic influences other than the construct. Reading difficulty, acquiescence (the tendency to agree regardless of content), social desirability, and shared method effects all introduce such variance. So do items that different groups interpret differently. Cognitive interviewing and careful item writing are the primary defences, backed by statistical checks in the validation phase. The items can be unreliable, producing scores with so much random error that they cannot distinguish between people or detect change. The pilot and validation phases catch this. The structure can be misunderstood. A scale presented as measuring one thing may in fact measure several, or a scale presented as having five subscales may have two. Factor analysis in the pilot and validation phases tests structural claims, but it can only test the structure that the item pool makes possible. Finally, the whole enterprise can suffer from what Jessica Flake and Eiko Fried, in a 2020 paper in Advances in Methods and Practices in Psychological Science, called questionable measurement practices: decisions about measurement that are undisclosed or unjustified and that raise doubts about the validity of conclusions. These include creating ad hoc scales without reporting how, modifying existing scales without saying so, and dropping items without explanation. Flake, Jolynn Pek, and Eric Hehman had earlier reviewed a sample of articles in a leading social and personality psychology journal and found that validity evidence for the measures used was frequently thin or absent, with many scales supported by little beyond a reliability coefficient. The remedy is not a particular statistic. It is the transparent, staged process this book describes, documented well enough that others can see what was done. The chapters that follow take the stages in order. Before any focus group is convened, however, the construct must be defined well enough to know whom to ask and what to ask them about. That is the subject of Chapter 2. Chapter 2: Defining the Construct Before Talking to Anyone A common way to start a scale development project is to convene some focus groups and see what emerges. It feels open-minded and participant-centred. It is usually a mistake. Without a working definition of the construct, the researcher does not know whom to recruit, what to ask, when a discussion has drifted off topic, or how to decide which of the many things participants say belong in the instrument. The qualitative phase then produces a large, shapeless body of material, and the item pool that follows inherits its shapelessness. Lee Anna Clark and David Watson, in an influential 1995 paper in Psychological Assessment titled "Constructing validity", put the point bluntly: the first step in scale development is to develop a precise and detailed conception of the target construct and its theoretical context. They revisited the topic in 2019, in the same journal, and found the advice still widely ignored. Robert DeVellis, whose textbook Scale Development: Theory and Applications is now in a fifth edition co-authored with Carolyn Thorpe, makes "determine clearly what it is you want to measure" the first of his guidelines. Scott MacKenzie, Philip Podsakoff, and Nathan Podsakoff, in a 2011 paper in MIS Quarterly that set out a full procedure for construct measurement in management and information systems research, spend much of their first step on conceptual definition. The agreement across disciplines is striking. This chapter describes what a working construct definition contains, how to review what already exists, and which structural decisions should be made, at least provisionally, before fieldwork begins. The definition written now is not final. Qualitative work will refine it, sometimes substantially. But a draft definition is what makes the qualitative work productive. Writing a working definition A usable construct definition does more than name the concept. It states what kind of thing the construct is: an attitude, a belief, an emotional state, a behaviour, a capacity, a perception of the environment. It says whose attribute it is: an individual, a dyad, a team, an organisation. It specifies the domain: the range of content the construct covers. It draws boundaries by stating what the construct excludes, especially neighbouring constructs with which it might be confused. And it indicates the expected stability: whether the construct is a trait that should change little over years, a state that fluctuates from day to day, or something in between. Consider how much turns on these decisions for a construct such as psychological safety, which Amy Edmondson defined in a 1999 Administrative Science Quarterly paper as a shared belief held by members of a team that the team is safe for interpersonal risk taking. That definition fixes the construct as a belief, not a behaviour. It locates it at the team level, as a shared belief, which has consequences for how responses are aggregated and analysed. It limits the domain to interpersonal risk, so a question about whether employees feel physically safe at work falls outside it. And it implies a degree of stability tied to the team's climate rather than an individual's personality. A researcher who wants to measure psychological safety in hybrid teams starts with these boundaries and asks whether the qualitative work confirms, extends, or challenges them. Writing the definition forces decisions that are easy to postpone. Is health literacy the capacity to obtain, process, and understand health information, or does it include the capacity to act on it? Is fear of falling a specific anxiety about falling, or does it include the loss of confidence in performing everyday activities without falling? Different research groups have answered these questions differently, and their instruments differ accordingly. In the fear-of-falling literature, for instance, the Falls Efficacy Scale developed by Mary Tinetti and colleagues in 1990 measured confidence in performing activities without falling, while later instruments measured concern about falling, and researchers came to treat these as related but distinct constructs. The point is not that one definition is right. It is that the definition must be chosen, and the choice shapes everything after. It helps to write the definition in two forms: a one-sentence summary and a longer specification of a page or so. The longer version should list the facets or components the researcher currently expects, with a brief description of each, and should list the neighbouring constructs from which the target must be distinguished. This longer document becomes the first entry in the audit trail. When the qualitative phase begins, it serves as the sensitising framework: not a set of hypotheses to be confirmed, but a map that shows where the researcher expects to find things, so that surprises are recognisable as surprises. Jingle, jangle, and the review of existing measures Two old warnings from the history of psychology are useful here. The jingle fallacy, a term associated with Edward Thorndike, is the assumption that two measures with the same name measure the same thing. The jangle fallacy, named by Truman Kelley in 1927, is the assumption that two measures with different names measure different things. Both are common. There are many scales called "engagement" that measure substantially different constructs, and many scales with different names whose scores correlate so highly that they are practically interchangeable. A careful review of existing instruments protects against both fallacies. The review has several aims. It establishes whether a new instrument is needed at all: if an adequate measure exists for the intended population and use, adapting or simply using it will usually be better than starting again. It identifies how others have defined and operationalised the construct, which helps sharpen the working definition. It maps neighbouring constructs and their measures, which will later be needed for discriminant validity testing. And it provides a source of item ideas and a record of what has and has not worked. In health outcomes research, this review has been formalised. The COSMIN initiative (COnsensus-based Standards for the selection of health Measurement INstruments) has published a methodology for systematic reviews of patient-reported outcome measures, described by Cecilia Prinsen and colleagues in Quality of Life Research in 2018, along with a risk-of-bias checklist for the studies that report measurement properties. A COSMIN-style review rates each instrument's evidence on content validity, structural validity, internal consistency, reliability, measurement error, hypothesis testing, cross-cultural validity, criterion validity, and responsiveness, and grades the overall quality of that evidence. Outside health, few fields have such a formal apparatus, but the same logic applies. A reviewer should look not just at whether an instrument exists but at how strong the evidence behind it is and for which populations. The review often reveals that the problem is not the absence of a measure but the absence of evidence for a specific population or use. A well-validated depression scale developed in North American adults may have little evidence of content validity for older adults in rural South Asia. In that case the right project may be cross-cultural adaptation rather than new development, a process with its own methods (translation, back-translation or team translation, cognitive testing, and invariance testing) that overlaps substantially with the stages in this book. Dorcas Beaton and colleagues set out widely used guidelines for cross-cultural adaptation of self-report measures in Spine in 2000, and the ISPOR task force led by Diane Wild published principles for translation and cultural adaptation in Value in Health in 2005. Purpose, population, and context of use The intended use of an instrument shapes its design at least as much as the construct does. In a 1985 paper in the Journal of Chronic Diseases, Bart Kirshner and Gordon Guyatt distinguished three purposes for health measures: discriminative instruments, which distinguish between people at a point in time; predictive instruments, which classify people according to some future outcome; and evaluative instruments, which measure change within people over time. The distinction has practical force. A discriminative instrument benefits from items that vary widely across people and may include stable characteristics. An evaluative instrument needs items that can change when the underlying condition changes; an item about something that never changes, however well it discriminates between people, is useless for detecting treatment effects. A predictive instrument may be judged mostly by its accuracy against the future criterion, with less concern for internal coherence. The population must be specified with similar care. Age range, literacy, language, cultural background, clinical characteristics, and the settings in which people will complete the instrument all matter. They determine who should be recruited for focus groups and cognitive interviews, what reading level the items must meet, which response formats are feasible, and which subgroups will later need to be compared in invariance testing. A measure intended for adolescents and adults needs adolescents and adults in every qualitative phase. A measure intended for use across several countries needs participants from those countries early, not only at the translation stage. Context of use includes practical constraints. How long can the instrument be? A research questionnaire completed once by motivated volunteers can tolerate forty items; a measure given to patients at every clinic visit may need to be under ten. What mode of administration will be used: paper, web, smartphone, telephone, or interviewer? Mode affects visual design, the feasibility of long response scales, and the likelihood of socially desirable answers. Will the instrument stand alone or be embedded in a longer survey, where context effects from preceding questions may shape responses? These constraints do not need final answers before fieldwork, but they should be written down, because they affect how many items the pool must contain and how aggressive later item reduction must be. Early structural decisions Several decisions about the instrument's structure should be made provisionally at this stage, because they affect what the qualitative phase must explore and how the item pool is built. The first is whether the construct is reflective or formative, a distinction introduced in Chapter 1. The question to ask is about the direction of causality between the construct and its indicators. Would a change in the construct be expected to produce changes in all the indicators together? If a person becomes more burnt out, one would expect exhaustion, cynicism, and reduced efficacy items all to shift, which is the reflective pattern. Or would a change in a single indicator change the construct, even if the other indicators stayed the same? If a person's income rises while their education stays fixed, their socioeconomic status has changed, which is the formative pattern. Cheryl Jarvis, Scott MacKenzie, and Philip Podsakoff reviewed marketing research in a 2003 paper in the Journal of Consumer Research and argued that many constructs treated as reflective were better conceived as formative, with consequences for the conclusions drawn. The distinction matters because the statistical tools in later chapters, such as internal consistency, factor analysis, and item-total correlations, assume a reflective model. Applied to a formative construct, they will recommend dropping exactly the items that make distinct contributions, destroying the construct's content. Adamantios Diamantopoulos and Heidi Winklhofer described the different procedures appropriate to formative index construction in the Journal of Marketing Research in 2001. Most constructs measured by psychometric surveys are treated as reflective, and most of this book assumes that model. But the developer should confirm the choice rather than default to it, and the qualitative work can help. If focus group participants describe a set of experiences that tend to occur together and seem to express a single underlying state, a reflective model is plausible. If they describe a set of distinct contributors that combine to produce an overall condition, with no reason to expect them to co-occur, a formative model may fit better. The second decision concerns dimensionality. Is the construct expected to be unidimensional, or to consist of several related facets? If several, will the instrument report a total score, separate subscale scores, or both? A multidimensional construct requires an item pool with enough items for each facet to survive item reduction as a reliable subscale, usually several times the final number per facet. Deciding this in advance does not preclude discovering a different structure later, but it ensures the pool is built to test the structure that theory proposes. The third decision concerns the level of analysis. Constructs such as team climate or neighbourhood cohesion are attributes of collectives measured through individuals. For these, the wording of items matters in a specific way. An item can refer to the respondent ("I feel able to raise problems in this team") or to the collective ("People in this team are able to raise problems"). Organisational researchers call the latter a referent-shift item, and it is generally preferred when the aim is to measure a shared property of the group. The choice should be made before item writing, and ideally before the discussion guide for focus groups is written, because it affects how questions are posed to participants. Specifying the construct in practice Pulling these elements together, a construct specification document at the end of this stage typically includes: the one-sentence definition; an expanded definition with provisional facets; the list of neighbouring constructs and how the target differs from each; a summary of existing instruments and why none is adequate for the intended use; the target population and its relevant subgroups; the intended purpose (discriminative, evaluative, or predictive) and the decisions scores will inform; the context of use, including mode and length constraints; and the provisional structural decisions about measurement model, dimensionality, and level of analysis. It is also worth writing, at this stage, a first draft of the interpretation and use argument described in Chapter 1: the chain of inferences from item responses to the intended conclusion, with a note of the evidence each will need. This draft usually reveals gaps immediately. A team that intends its scale to detect change after an intervention, for example, will see that it needs test-retest data in a stable sample and responsiveness data in a changing one, which has implications for the design of the final validation study that are far cheaper to address now than later. A constructed example shows what this looks like for a real kind of project. Suppose a team intends to measure digital exclusion stress among older adults: the distress arising from being unable to use digital services that are increasingly required for everyday tasks such as banking, booking medical appointments, or claiming benefits. The one-sentence definition might read: the distress older adults experience when they need to use digital services and feel unable to do so. The expanded definition might propose facets of practical frustration, dependence on others, fear of making mistakes, and feelings of being left behind. Neighbouring constructs include technology anxiety, computer self-efficacy, loneliness, and general psychological distress, and the specification must say how digital exclusion stress differs from each: from technology anxiety, for example, in being tied to specific needs and consequences rather than to technology as such. The population might be adults aged 65 and over living in the community, including those who do and do not use the internet, with attention to differences by age band, education, and living situation. The purpose might be evaluative, to assess the effect of community digital support programmes, which immediately implies a need for items that can change. A length limit of fifteen items might be set because the measure will sit in a longer evaluation survey often administered by telephone, which rules out long or visually complex response formats. None of these decisions is final. The focus groups may show that "fear of making mistakes" is really about fear of fraud, a distinct and much more specific concern. They may reveal a facet the team did not anticipate, such as shame. They may show that the construct cannot sensibly be separated from loneliness in this population. But because the team wrote down what it expected, it will recognise each of these as a finding rather than simply absorbing it. The limits of the armchair It would be possible to take this chapter's advice too far, and some development projects do. A team that spends months refining a construct definition and then treats the qualitative phase as a formality to confirm it has missed the point of mixed methods design. The definition written at this stage is a hypothesis, and the qualitative phase is its first test. The appropriate stance is that of an investigator who has a clear question and genuinely does not know the answer. The balance is easier to strike when the construct specification is explicit about its uncertainties. A good specification says not only what the team believes but where it is least sure: which facets are borrowed from literature on a different population, which boundaries with neighbouring constructs are contested, which aspects of the population's experience are simply unknown. Those uncertainties then become the priorities of the discussion guide. The next chapter describes how to design and run the qualitative elicitation phase so that it addresses them. Chapter 3: Listening First: Focus Groups and Qualitative Elicitation The qualitative phase of instrument development has a specific job. It is not a general exploration of a topic, and it is not a study whose findings will be published for their own sake, although they often are. Its job is to find out what the construct consists of for the people who will be measured, and to capture the language they use to describe it, in a form that can be turned into items. In health outcomes research this phase is called concept elicitation, a term that captures its purpose well: the aim is to elicit, from the target population, the concepts the instrument should measure. This chapter concentrates on focus groups, because they are the method most often associated with the early stage of scale development and because their strengths and pitfalls are distinctive. But it also considers individual interviews, which are often the better choice, and a set of structured elicitation techniques that can be used within either format. Why groups, and when not A focus group is a facilitated discussion among a small number of people who share some relevant characteristic, typically lasting between one and two hours. David Morgan, whose writing did much to establish focus groups as a research method in the social sciences, argued that their distinctive contribution is the interaction among participants. People compare experiences, disagree, qualify each other's accounts, and build on each other's language. In the course of a discussion, a group can reveal the range of ways a phenomenon is experienced, which experiences are common and which idiosyncratic, and which words people reach for when they try to describe it. For item generation, these are exactly the right outputs. Focus groups have particular value when the construct is shared or social. A construct such as team psychological safety, neighbourhood cohesion, or perceived stigma is partly constituted by how people talk about it together, and a group discussion reproduces something of that social reality. They are also efficient: a single session can generate material from eight people. But groups have costs. Participants may conform to an emerging consensus or defer to a dominant speaker. They may be reluctant to describe experiences that are embarrassing, stigmatised, or dangerous to disclose in front of peers. And groups are poor at capturing the detailed sequence of an individual's experience, such as the way symptoms develop over a day or the steps of a decision. Individual interviews are therefore often preferable, or a necessary complement. They are better for sensitive topics such as sexual health, substance use, experiences of violence, or financial hardship. They are better when participants are hard to assemble, such as people with severe illness, shift workers, or senior executives. And they are better when the construct concerns individual inner experience that does not benefit from comparison, such as the phenomenology of pain or intrusive thoughts. Many successful development programmes use both: groups to map the breadth of the domain and establish shared language, and interviews to probe sensitive or complex facets in depth. The patient-reported outcome literature offers useful guidance on combining them. Meryl Brod, Laura Tesler, and Torsten Christensen, in a 2009 paper in Quality of Life Research on best practices for qualitative research in content validity, and Kathryn Lasch and colleagues, in the same journal in 2010, both describe concept elicitation programmes in which the choice of method is justified by the nature of the construct and the population, and in which the rationale is documented. That documentation is part of the audit trail. Who to recruit Sampling for concept elicitation is purposive. The aim is not to estimate how common each experience is in the population, which the quantitative phase will do far better, but to capture the full range of relevant experience. That means deliberately recruiting people who differ on the characteristics most likely to shape how they experience the construct. The construct specification from Chapter 2 identifies these characteristics. For a measure of digital exclusion stress among older adults, they might include age band, whether the person uses the internet at all, living alone or with others, level of education, and disability. For a measure of burnout among health workers, they might include profession, setting, seniority, and shift pattern. A sampling frame that crosses the most important of these characteristics, and ensures each is represented, is usually enough. Trying to cross every characteristic with every other produces an unmanageable number of cells. Groups are usually more productive when they are relatively homogeneous on the characteristics most likely to create status differences or inhibit disclosure. Mixing managers and their direct reports in a discussion of psychological safety would be self-defeating. Mixing people with and without a stigmatised condition in a discussion of that condition may silence those who have it. The standard approach, often called segmentation, runs separate groups for each segment and compares across them. Homogeneity within groups encourages open discussion; heterogeneity across groups ensures range. Group size is usually between five and ten participants. Richard Krueger and Mary Anne Casey, whose practical guide to focus groups has gone through five editions, have in recent editions favoured the smaller end of that range for complex or emotionally demanding topics, because larger groups leave each participant little time to speak. Over-recruiting by one or two people is standard practice, since some will not attend. How many groups The question of how many groups are enough has attracted more empirical attention than it once did. The traditional answer was to continue until saturation: the point at which new groups produce no new themes. The concept is intuitive but slippery, because whether a theme is new depends on how finely themes are defined, and a researcher can always find something new by looking closely enough. Virginia Braun and Victoria Clarke, who developed the widely used reflexive approach to thematic analysis, argued in a 2021 paper that saturation is a poor fit for interpretive qualitative research, where meaning is generated by the analyst rather than discovered. For concept elicitation, however, saturation has a more concrete meaning than in interpretive research, and it can be tracked. The question is whether additional groups are producing new concepts relevant to the instrument: new facets, new experiences, new consequences, or substantially new language. A practical technique is the saturation grid: a table with concepts as rows and groups or interviews as columns, filled in as analysis proceeds, so that the team can see when new columns stop adding new rows. The ISPOR good practice reports recommend presenting such a grid as part of the evidence for content validity. Empirical work gives some guidance on what to expect. Greg Guest, Emily Namey, and Kevin McKenna, in a 2017 study published in Field Methods, analysed data from a larger study that used focus groups and examined how quickly themes accumulated. They found that two to three groups were likely to capture about eighty per cent of the themes identified in the full data set, including the most prevalent ones, and three to six groups about ninety per cent. Those figures come from one study with a particular topic and population and should not be treated as a rule. But they support a common-sense conclusion: for a well-defined construct in a fairly homogeneous population, a handful of groups per segment will capture most of what matters, while a construct with many facets, or a population with many relevant segments, will require more. The saturation grid tells the team when it has reached that point for its own project. The discussion guide A discussion guide is the moderator's plan for the session. For concept elicitation it must balance two aims that pull against each other: covering the facets the construct specification anticipates, and leaving room for participants to reveal what it does not. Krueger and Casey describe a sequence of question types that remains the most useful template. Opening questions are quick, factual, and answered by everyone, to get people speaking. Introductory questions introduce the topic in general terms and invite participants to describe their experience of it. Transition questions move the discussion toward the key areas. Key questions address the central issues and take most of the session. Ending questions ask participants to reflect on what has been said, identify what matters most, and add anything missed. For concept elicitation, the key questions should begin broad and become more specific, a pattern often called the funnel. The broad questions come first because they allow participants to raise concepts spontaneously, which is the strongest evidence that a concept belongs to the construct. Only after that should the moderator probe for concepts from the construct specification that have not arisen. A concept raised spontaneously by participants in several groups carries more weight than one that participants agreed with only when prompted, and the analysis should record which was which. The most common error in discussion guides is to plant the construct: to ask questions that presuppose the researcher's framing. A question such as "How does technology make you feel stressed?" presupposes both that technology causes stress and that "stress" is the right word. A better opening is "Tell me about the last time you needed to do something online, like booking an appointment or checking your bank account. What happened?" Participants will describe their experience in their own terms, and the moderator can follow up on the emotions and consequences they mention. If they never mention anything resembling stress, that is important information about the construct. Guides for concept elicitation often include specific prompts about dimensions that item writing will later need: frequency ("How often does that happen?"), intensity ("How much does that bother you when it does?"), duration, triggers, and impact on daily life. They may also ask directly about language: "What words would you use to describe that feeling?" or "If someone else was going through this, how would they describe it?" Structured elicitation techniques Open discussion can be supplemented by structured exercises that produce data in forms closer to what item writing needs. Free listing asks each participant to write down, individually and before discussion, every word or experience that comes to mind in connection with a prompt. The lists capture individual perspectives before group influence operates, and the frequency with which terms appear across participants is a rough indicator of salience. Card sorting presents participants with cards describing experiences, drawn from earlier groups or from the literature, and asks them to sort the cards into piles of things that go together and to name the piles. The resulting groupings are evidence about how participants structure the domain, which bears on the dimensionality decisions from Chapter 2. Card sorts can also be analysed quantitatively, for example by counting how often each pair of cards is placed together. Ranking and rating exercises ask participants to judge which experiences are most important, most common, or most distressing. These should be treated cautiously, since small purposive samples do not estimate population prevalence, but they help prioritise concepts for item writing. Vignettes describe a hypothetical person in a situation relevant to the construct and ask participants to discuss how that person would feel or respond. They are particularly useful for sensitive topics, because participants can speak about the character rather than themselves, and for exploring boundaries between the target construct and its neighbours. Moderating for elicitation Moderating a focus group well is a skill that improves with practice and supervision. For concept elicitation, a few principles matter most. The moderator should speak little, listen closely, and follow up on participants' own words rather than paraphrasing them into the researcher's vocabulary. If a participant says "it just makes me feel stupid", the follow-up should use their word: "Can you say more about feeling stupid?" not "So you feel a loss of self-efficacy?" The moderator should actively manage dominant speakers and invite quieter participants, and should explicitly welcome disagreement ("Has anyone had a different experience?"), because variation is the point. A second person, often called the assistant moderator or note-taker, should record who says what, non-verbal reactions, and the moderator's own observations. Sessions should be audio-recorded, with permission, and transcribed verbatim. Summaries or notes are not a substitute for transcripts in concept elicitation, because the exact words participants use are part of the data. Online focus groups, conducted by video conference or as asynchronous text-based discussion boards, became much more common after 2020 and are now a standard option. They widen geographic reach and can make participation easier for people with mobility limitations or caring responsibilities. They also change the dynamics: interaction is more stilted on video, turn-taking is harder, and participants without good internet access or digital skills are excluded, which would be a fatal flaw in a study of digital exclusion. The mode should be chosen with the population in mind and reported. Analysing transcripts with items in mind Analysis of concept elicitation data is a form of thematic or content analysis, but its aim is specific. It must produce a structured account of the construct's domain, organised into concepts that can each be represented by one or more items, with supporting quotations and a record of how widely each concept was expressed. Several analytic approaches serve this aim. Framework analysis, developed by Jane Ritchie and Liz Spencer at the National Centre for Social Research in Britain and widely used in applied health research, is particularly well suited, because it produces a matrix of participants or groups by themes, with summarised data in each cell, making comparison across segments straightforward. Codebook approaches to thematic analysis, in which a coding frame is developed from the construct specification and refined as new codes emerge, serve the same purpose. Braun and Clarke's reflexive thematic analysis, first described in 2006 and substantially elaborated since, emphasises the analyst's interpretive role and is less concerned with counting; it can be used, but the analyst should be clear that the output needed is a map of concepts suitable for item generation. Whatever the approach, the analysis should produce several things. A concept list, organised hierarchically if the construct has facets, with a clear definition of each concept. For each concept, a record of whether it arose spontaneously or when prompted, in how many groups or interviews, and in which segments. For each concept, a set of illustrative verbatim quotations, especially those that capture how participants phrase the idea. A note of concepts that participants raised but that fall outside the construct as currently defined, with a decision about whether the definition should be revised or the concept excluded. And an updated construct specification that reflects what was learned. Two analysts coding independently and then reconciling differences is good practice. Some teams calculate an agreement statistic, but the more important output is the discussion of disagreements, which frequently reveals ambiguities in concept definitions that would otherwise carry through into ambiguous items. The WHOQOL programme, run by the World Health Organization in the 1990s to develop a cross-culturally applicable measure of quality of life, illustrates the value of this stage at scale. Focus groups were conducted in field centres across many countries, involving people with and without illness and health professionals, to identify the facets of quality of life and generate items in each culture rather than simply translating items written elsewhere. The resulting instrument, the WHOQOL-100, and its short form, the WHOQOL-BREF, described by the WHOQOL Group in Psychological Medicine in 1998, owe their cross-cultural coverage to that qualitative foundation. Revising the construct The qualitative phase almost always changes the construct specification, and it should. Facets may be split, merged, added, or dropped. Boundaries with neighbouring constructs may be redrawn. The dimensionality hypothesis may change. Occasionally the phase shows that the construct as conceived does not exist in the population's experience at all, or exists only for a subgroup, which is an uncomfortable but valuable finding. These changes must be recorded in the audit trail, with the evidence that prompted them. The revised specification then becomes the blueprint for item generation. By the end of this phase, a team should have a concept list with definitions, a body of verbatim language for each concept, an indication of each concept's salience across the population, and a clear sense of which concepts the instrument must cover. Chapter 4 turns to the craft of converting that material into items. Hashtags: #MixedMethodsInstrumentDesign #PsychometricSurveys #ScaleDevelopment #InstrumentDevelopment #ExploratorySequentialDesign #ConstructValidity #ContentValidity #ResponseProcesses #ConstructSpecification #ConceptElicitation #FocusGroups #CognitiveInterviewing #ItemGeneration #ContentValidityIndex #PilotTesting #ItemAnalysis #ExploratoryFactorAnalysis #ConfirmatoryFactorAnalysis #ItemResponseTheory #ReliabilityAnalysis #MeasurementInvariance #TestRetestReliability #CrossCulturalAdaptation #ValidityEvidence #FutureOfPsychometricMeasurement
- Mixed Methods in Public Health (Bridging Epidemiological Data with Human Experience)
Download the Book (PDF): Introduction We know the size of the gaps. That is the strange position public health finds itself in after four decades of serious disparities research. Pick almost any outcome — infant mortality, hypertension control, cancer stage at diagnosis, years of life lost — and a competent analyst can tell you within an afternoon how far apart the groups are, whether the distance is widening or narrowing, and roughly how much of it survives adjustment for income, education, and insurance status. The surveillance systems work. The cohorts are large and well characterised. The statistical machinery for decomposing a difference into its measured components has become sophisticated enough that the limiting factor is no longer the method but the data going into it. And yet the gaps are mostly still there. In the United States, pregnancy-related mortality among Black women has run roughly three times the rate among white women for as long as the federal government has published the comparison, and the ratio has been remarkably insensitive to the economic expansions, insurance reforms, and clinical guideline revisions that occurred across that period. It does not narrow with education: a Black woman with a college degree carries a higher risk than a white woman who did not finish high school. That single fact, reported repeatedly by the Centers for Disease Control and Prevention and by state maternal mortality review committees, is a quiet indictment of the explanatory models most of the field was using. Whatever is producing the difference is not well captured by the variables in the model. This is the point at which description stops paying. Another cycle of the same analysis on a larger sample will estimate the same gap with a tighter confidence interval. It will not tell you what happens in a prenatal visit when a woman reports pain and is not believed. It will not tell you why a family that qualifies for a programme does not enrol in it, or what a patient means when she says the clinic "isn't for people like me," or how a neighbourhood's history of displacement shapes whether a health department's outreach worker is let through the door. Those are questions about meaning, sequence, and mechanism as they are lived, and no amount of variance decomposition will answer them, because the answers were never encoded in the data. The obvious response — the one that has been urged in editorials and funding announcements for twenty years — is to add qualitative work. Interview people. Listen. Understand context. The advice is correct and almost entirely useless as stated, because it stops at the point where the difficulty begins. Adding qualitative work is easy. Most large studies now have some. What is hard, and what is rarely done well, is making the two kinds of evidence actually bear on each other: getting a regression coefficient and a phenomenological account of illness to sit in the same argument, so that each constrains, complicates, or corrects what the other would have concluded alone. That is the subject of this book, and it is narrower than "mixed methods" as the phrase is usually used. The claim I want to defend is this: integration is the whole of mixed methods, and it is a design decision made before data collection, not an interpretive gesture made after it. A study that runs a cohort analysis and a set of interviews side by side, and then writes a discussion section observing that the themes "resonate with" the quantitative findings, has not done mixed-methods research. It has done two studies and stapled them together. The staple is the problem. Everything that makes the combination worth more than the sum — the ability to explain an anomaly, to discover that a validated instrument is measuring something different in one population than another, to find that a protective factor in the model is experienced as a burden by the people who have it — depends on decisions about sampling, timing, instrument design, and analytic procedure that have to be made at the protocol stage. Miss them, and no amount of careful writing at the end will recover the lost purchase. The health-disparities case sharpens this. In a study of, say, a new surgical technique, the quantitative and qualitative strands are usually asking about the same object from two angles, and a loose integration still yields something useful. In disparities research the two strands are frequently in tension by construction. The quantitative strand works with categories — race, ethnicity, socioeconomic position, rurality — that are administrative summaries of vastly heterogeneous experience, and it is obliged to treat them as fixed attributes of individuals so that they can be entered into a model. The qualitative strand tends to find that those categories are not attributes at all but relations, produced and enforced in specific encounters, and that the thing doing the causal work is often something the quantitative variable is a poor proxy for. When you interview people about why a category predicts an outcome, you frequently learn that the category is the wrong object. That tension is not a nuisance to be smoothed over. It is the most productive thing a mixed-methods disparities study can generate, and a design that cannot surface it has wasted the qualitative strand. Much of what follows is about building studies that can register disagreement between their own components and treat it as a finding rather than a failure. I should be equally clear about what this book is not. It is not a general textbook on mixed-methods research; the standard references do that job well, and I point to them at the end. It is not an introduction to statistical methods for epidemiologists, who have better sources, nor a manual for conducting interviews, which cannot be learned from a book in any case. And it is not a plea for qualitative research on the grounds that numbers are cold and stories are warm. That argument has done real damage. It licenses the use of interviews as illustration — the quotation deployed to make a table more palatable — which is the single most common way the qualitative strand of a health study is wasted. Phenomenological inquiry, done properly, is a disciplined method for establishing the structure of an experience, with its own standards of adequacy and its own characteristic failures. It earns its place in a study because it produces knowledge nothing else produces, not because it is humane. The book proceeds in nine chapters. The first two establish what each tradition can and cannot deliver: the ceiling that epidemiological measurement runs into in disparities work, and what phenomenological inquiry actually is, as distinct from the loose "qualitative component" that appears in many protocols. The third takes up the paradigm question — whether these two ways of knowing can coherently be combined at all — and disposes of it quickly, because the practical question that matters is not philosophical compatibility but the location of the joints where the strands touch. Chapters four through six are about design. Four sets out the three core configurations — convergent, explanatory sequential, exploratory sequential — and the four decisions that distinguish them. Five and six build a convergent and a sequential design respectively, in enough detail to write a protocol from, including the failure modes each is prone to. Chapters seven and eight are about execution, and they are where most studies are actually lost. Seven concerns sampling: how you choose whom to interview when your sampling frame is a cohort of thirty thousand people, and why the answer depends on which design you chose. Eight concerns analysis: the specific procedures — joint displays, data transformation, following a thread, mixed-methods matrices — by which two datasets are brought into contact, and what to do when they disagree. The ninth chapter addresses rigour, reporting, ethics, and money: how integrated work is judged, how to write it so that reviewers can see the integration, what the reporting standards require, and the particular ethical weight of studying populations that have been studied before without benefit. The conclusion takes up what follows from all of it, and what remains genuinely unsettled. A word on the intended reader. I have written for someone who is competent in one of these traditions and a beginner in the other — the epidemiologist who has been told to add a qualitative aim to a resubmission, or the qualitative researcher invited onto a cohort study and unsure what is being asked of them. That describes most people doing this work. The greatest practical obstacle to good integration is not methodological but social: two people who cannot evaluate each other's evidence, deferring politely, producing two reports. Reading this will not make an epidemiologist into an interviewer or an ethnographer into a biostatistician. It should make each able to tell the difference between good and bad work in the other's domain, to ask the questions that matter at the design stage, and to recognise when a collaborator is proposing something that will not integrate. One more thing, said once and not repeated. Mixed-methods work is expensive, slow, and harder to publish than either strand alone. It should not be the default. Many important questions in public health are purely quantitative, and many are purely qualitative, and a study that mixes when it did not need to has spent a great deal of money buying complexity. The test is simple and is worth applying honestly before anything else: is there a question in this study that neither strand could answer alone? If you cannot state that question in a sentence, do not mix. If you can, the rest of this book is about building a study that can actually answer it. Chapter 1: The Ceiling of Description Epidemiology is a machine for detecting and sizing differences in the distribution of health. It is extraordinarily good at this. The discipline's great achievements — the link between smoking and lung cancer, the identification of cholera's waterborne transmission, the recognition that social position predicts mortality in a graded fashion across the whole social range rather than only at the bottom — were all fundamentally acts of comparison, executed with enough care that the comparison could bear weight. The Whitehall studies illustrate both the power and the limit. Michael Marmot and colleagues followed British civil servants, a population with universal healthcare access, stable employment, and no one in destitution, and found a clean stepwise gradient in coronary mortality running from the top of the employment hierarchy to the bottom. Each grade did worse than the one above it. The finding demolished a comfortable assumption — that health inequalities were about material deprivation at the bottom of society — and it could not have been produced by any other method. You need the cohort, the follow-up, the classification, and the statistics. But notice what the gradient does not contain. It does not say what it is about occupying a lower grade that damages the heart. Whitehall II went looking, and the leading candidate that emerged was control over work: low decision latitude, demands without authority. That is a real advance. It is also a construct measured by questionnaire items, and the step from "scores lower on a job control scale" to "experiences their working life in a way that produces chronic physiological stress" is a step the questionnaire cannot take on its own. The instrument was built from a theory of what mattered. If the theory was incomplete — if something else in the experience of subordination was doing part of the work — the instrument would not have detected it, and the analysis would have reported a residual. This is the ceiling. Epidemiological methods can establish, with great precision, that a difference exists, how large it is, how it is patterned across time and place, and how much of it is statistically accounted for by the variables you thought to measure. They cannot tell you what you failed to measure, and in disparities research what you failed to measure is usually the point. Residual confounding is not a technical problem Every epidemiologist learns that adjustment is imperfect. Measured covariates are noisy, categories are coarse, and the true confounder is often a construct for which the available variable is a proxy. Adjusting for "education, in four categories" does not adjust for education; it adjusts for a crude marker of a bundle of things — credential, cognitive exposure, network, signalling value in a labour market — whose composition differs by group. The standard treatment of this is as a measurement problem, to be reduced by better instruments and more categories. In disparities work it is not a measurement problem, or not only. Consider the repeated finding that adjustment for socioeconomic position reduces but does not eliminate racial differences in a wide range of outcomes. Two readings are available. The first is that the socioeconomic measures are imperfect and that with better ones the gap would close; this motivates ever-richer adjustment, wealth instead of income, neighbourhood composition, intergenerational measures. The second is that the residual is not error at all but the signal — that there is an exposure, operating independently of material position, which the model does not contain because nobody has built a good variable for it. The second reading has better support. Its most developed statement is Nancy Krieger's ecosocial theory, which asks how people literally embody their social conditions, including discriminatory ones, across the life course. On that account, the residual in a race-adjusted model is partly the accumulated physiological cost of racism, operating through pathways — vigilance, anticipatory stress, interrupted care, chronic dysregulation — that no existing item bank measures well. Bruce Link and Jo Phelan's fundamental cause theory makes a parallel argument about socioeconomic position: it persists as a predictor across centuries in which the proximate causes of death changed completely, because what it really indexes is flexible access to resources that can be deployed against whatever the current threats happen to be. Both theories predict that adjustment will underperform, and both explain why. Neither theory, however, tells you what the pathway looks like in a given population, in a given city, for a given condition. That is an empirical question about how people actually live, and answering it requires a method that can find exposures nobody has yet named. This is the first and most important thing the qualitative strand contributes: not colour, not illustration, but discovery of unmeasured constructs. The interview is one of the few instruments in the public health toolkit capable of returning a variable you did not know to look for. The category problem The deeper issue concerns what the grouping variables in a disparities analysis actually denote. When a model contains a coefficient for race, that coefficient is doing something specific and easily misread. It is not estimating the effect of a biological attribute; the field has, with some exceptions, absorbed that point. It is estimating an association between a socially assigned classification and an outcome, in a particular population at a particular time, net of whatever else is in the model. Whether that association is informative depends entirely on what one thinks the classification stands in for — and researchers frequently do not say. Lisa Bowleg made the sharpest version of this critique in her 2012 argument that the conventional handling of "women and minorities" is additive where lived experience is multiplicative: a study that enters sex and race as separate main effects models a world in which being a Black woman is being Black plus being a woman. Intersectionality theory, originating in Kimberlé Crenshaw's legal scholarship, holds that it is not; the position is qualitatively distinct, not the sum of two disadvantages. The methodological implication is uncomfortable, because the natural quantitative translation — interaction terms — is both underpowered in most samples and conceptually thin. An interaction term tests whether the effect of one variable differs by level of another. It does not tell you what the distinct position consists of. Chandra Ford and Collins Airhihenbuwa's Public Health Critical Race praxis pushes further, arguing that the standard research process itself encodes assumptions — about which comparisons are natural, which group is the implicit reference, whose outcomes constitute the norm — that reproduce the hierarchy under study. The habitual choice of white outcomes as the reference category is not neutral; it frames every finding as a deficit in the comparison group, and it directs attention to what those groups do rather than to what is done to them. This is a place where the two traditions have genuinely different instincts, and it is worth sitting with the disagreement rather than resolving it prematurely. The quantitative instinct is that categories must be fixed for the duration of an analysis, because a variable whose meaning shifts is not a variable. The qualitative instinct is that the categories are constituted in interaction, that a person may be racialised differently in a clinic than in a workplace, and that treating the classification as a stable attribute of a person imports the very assumption under investigation. Both are right within their frames. A mixed-methods design cannot dissolve the tension, but it can exploit it: use the quantitative strand to establish the pattern the classification predicts, and the qualitative strand to investigate what the classification is standing in for in this specific context — which may turn out to be different in a rural county than in an urban one, a finding with immediate implications for how you would intervene. Missing the moment of action A second limitation is temporal. Most epidemiological data are records of states and events at intervals — a blood pressure at a study visit, a prescription fill, a hospitalisation, a death. Between those points is process, and process is where most intervenable mechanisms live. Take the familiar cascade framework, applied to hypertension: a population is screened, some are diagnosed, some of those are treated, some of those achieve control. Each transition shows loss, and each loss is patterned by group. The cascade is a good analytic device precisely because it localises where a disparity is generated — a gap concentrated at diagnosis implies a different intervention than one concentrated at control. But the cascade describes the losses; it does not contain them. Knowing that a population's attrition is concentrated between prescription and control tells you that something happens in the months after a prescription is written. It does not tell you whether the pills produced side effects that made work impossible, whether the pharmacy is on a bus route that runs every ninety minutes, whether the patient stopped because their pressure normalised and nobody explained that this was the drug working, or whether the clinician's manner at the follow-up visit made returning unappealing. Those four explanations imply four different interventions with four different costs. Distinguishing them requires asking people what happened, in an order that follows their account rather than the protocol's, and being willing to hear an answer nobody anticipated. Administrative data can sometimes adjudicate between such hypotheses once they exist — you can test whether pharmacy travel time predicts discontinuation — but the hypotheses have to come from somewhere, and the usual somewhere is a literature written about a different population a decade ago. The null result that means something Consider the position of a trialist whose intervention did not work. A well-powered trial of a community health worker programme reports no significant improvement in the primary outcome. The result is published, the programme is deemed ineffective, and the field moves on. But "did not work" conflates at least four distinct situations. The programme may have been delivered as designed and been genuinely ineffective. It may not have been delivered as designed — workers improvising under caseload pressure, sessions shortened, the core component quietly dropped. It may have been delivered and received, but interpreted by participants as surveillance rather than support, producing avoidance. Or it may have worked well for a subgroup and been irrelevant or harmful for others, netting to zero. These are not equivalent, and the distinction has enormous practical consequence: three of the four situations are fixable. A trial with only quantitative outcome data cannot tell them apart. Process evaluation exists for exactly this reason, and the UK Medical Research Council's guidance on process evaluation of complex interventions — the framework most often cited in this space — treats implementation, mechanisms of impact, and context as things that must be studied alongside outcomes rather than inferred from them. The point generalises well beyond trials. Any quantitative null, and any unexpected positive, is a signal that the model of how the world works is wrong somewhere, and locating the error is qualitative work. Averages over people who are not alike A third limit follows from what a regression coefficient is. It is an average, and averages are informative in proportion to the homogeneity of what they average over. Disparities research routinely takes averages over populations whose heterogeneity is the substance of the question. The point is easiest to see in intervention research. A text-message reminder system for cervical screening raises uptake by a few percentage points on average. That average may be composed of a large effect in women who intended to attend and simply forgot, no effect in women who have decided screening is not for them, and a small negative effect in women for whom unsolicited health messages from an institution they distrust are an intrusion. The mean is real, and it is also a description of nobody. Any plan to scale the intervention needs to know which of those subpopulations dominates in the new setting, and that is a question about the composition of reasons, not about the point estimate. The standard quantitative response is subgroup analysis or effect-modification testing, and it helps, but only when the relevant subgroups correspond to variables already in the dataset. If the operative distinction is between women who have had a prior experience of being dismissed in a clinical encounter and women who have not, no administrative dataset contains it. The moderator has to be discovered before it can be measured, and discovery of that kind is what open-ended inquiry is for. There is a related and underappreciated problem in how disparities are quantified at all. A gap can be expressed as a difference or a ratio, on an absolute or relative scale, against a best-group reference, a population mean, or an external standard. These choices are not cosmetic: a programme can narrow the relative gap while widening the absolute one, and the two framings will support opposite policy conclusions from identical data. The choice among them is a value judgement dressed as a technical decision, and it is one of the places where consultation with the affected population — asking what a meaningful improvement would consist of — changes the analysis rather than merely commenting on it. An illustration: asthma in a single city Consider a concrete case of the kind that arises constantly. A city health department observes that paediatric asthma emergency department visits are four to five times higher in two adjacent neighbourhoods than in the city as a whole. The pattern is stable across five years of administrative data, survives adjustment for insurance status and age structure, and is not explained by differential prevalence of diagnosis in primary care records. A well-resourced quantitative team can take this a considerable distance. They can link addresses to air monitoring data and to traffic volumes, model exposure to particulate matter, extract housing code violation records, and estimate associations between violation density and visit rates. Suppose they find that housing violations are strongly associated with visits and that modelled outdoor air quality is not. That is a substantive result and it narrows the field. It also leaves the decisive question open. "Housing violations predict asthma visits" is compatible with mould in specific units, with pest infestation and the associated allergens, with cold-induced avoidance of ventilation in winter, with landlord retaliation making tenants unwilling to report problems, and with the sheer administrative burden of a complaint process that consumes days a working parent does not have. Each implies a different lever: remediation funding, pest control contracts, weatherisation, tenant legal protections, or a redesigned complaints system. The correlation cannot distinguish among them because the variable "violation" is an artefact of an inspection process, not a measure of exposure. Twenty carefully sampled interviews with parents who have taken a child to the emergency department in the past year can distinguish among them, and will typically add something nobody listed — in cases like this one, often a detail about the pathway into the emergency department itself, such as the absence of any after-hours option that does not involve an ambulance, or a prior experience in which presenting to primary care produced a referral back to the emergency department anyway. That is not context for the finding. It is the finding, and the administrative data can now be re-interrogated to see how far it generalises. What the ceiling is not It is worth being careful here, because critiques of quantitative disparities research have a tendency to overreach into a general scepticism about measurement, and that would be a serious error. The numbers are not optional. Without surveillance data nobody would know that the maternal mortality ratio differs by race, that it is not explained by education, or that most pregnancy-related deaths are judged preventable by the review committees that examine them. Those facts came from counting, carefully, at scale, over decades, and they are the reason the qualitative questions in this field are being asked at all. A phenomenological study of thirty women's birth experiences cannot establish that a population-level disparity exists; it does not have and cannot have the sampling structure to support such a claim, and a researcher who implies otherwise has misunderstood their own method. Nor is the qualitative strand a solution to bias in the quantitative one. Interviews have their own systematic distortions: people reconstruct, rationalise, tell researchers what seems expected, and have limited access to the causes of their own behaviour. A woman who stopped her medication may sincerely report a reason that is not the reason. Combining methods does not cancel error; it changes its structure, which is useful only if the design is built to exploit the change. The accurate statement of the ceiling is narrow. Epidemiological methods answer questions of the form how much, in whom, compared with what, changing how? They do not answer by what mechanism, experienced how, and why does the mechanism operate here and not there? The first set of questions is necessary for the second to be well posed. The second set is necessary for the first to lead anywhere. A field that has spent forty years producing high-quality answers to the first while treating the second as commentary has produced, predictably, excellent descriptions of gaps it cannot close. The rest of this book is about the second set of questions, beginning with the method best suited to them — and with what that method is actually for, which is less obvious than it looks. Chapter 2: What Phenomenological Inquiry Actually Is There is a sentence that appears in a great many protocols: "We will conduct semi-structured interviews to explore patients' experiences and perceptions, which will be analysed thematically." It is not wrong, exactly. It is also not a method. It is a placeholder where a method should be, and it is the reason so many qualitative strands in health research produce findings that a competent reader could have guessed in advance. The problem is partly that the word "experience" is doing unexamined work. In ordinary usage it means something like "what happened to someone, as they would tell it." In phenomenological research it means something considerably more specific, and the difference determines what a study can claim. The tradition, briefly and usefully Phenomenology began as a philosophical programme. Edmund Husserl's project, at the start of the twentieth century, was to describe the structures of consciousness as they present themselves, setting aside questions about whether the objects of consciousness exist independently. His central methodological move was the epoché or bracketing: the deliberate suspension of the natural attitude — our default assumption that the world is simply there, as we take it to be — in order to attend to how a thing is given to awareness. Martin Heidegger, Husserl's student, redirected the enterprise, arguing that human existence is always already embedded in a world of practical involvements and that description must start there rather than with a detached consciousness. Maurice Merleau-Ponty added the body as the site of perception rather than an object perceived. Alfred Schütz carried the apparatus into social science. None of this is directly a research method, and it would be dishonest to suggest that a health researcher needs to resolve the differences between Husserl and Heidegger before conducting an interview. What matters for our purposes is the operational inheritance, which is real and specific. Several distinct research methods descend from this lineage. Amedeo Giorgi's descriptive phenomenological method, developed at Duquesne University, is the most procedurally explicit: the researcher reads the whole account, divides it into meaning units, transforms each into disciplinary language while staying faithful to the participant's sense, and synthesises a general structure of the phenomenon. Clark Moustakas's transcendental approach, widely used in North American health research, emphasises horizonalisation — treating every statement as initially of equal weight — followed by clustering into themes and the construction of textural and structural descriptions. Max van Manen's hermeneutic phenomenology of practice, developed in nursing and education, treats writing itself as the method and organises inquiry around lived body, lived space, lived time, and lived relation. Jonathan Smith's interpretative phenomenological analysis (IPA) is the dominant approach in British health psychology; it works idiographically, case by case, with small samples, and is explicit that the analyst's interpretation is part of the product rather than a contamination of it. The differences among these matter methodologically — whether you bracket your preconceptions or foreground them, whether you seek a single essential structure or a set of convergent and divergent accounts, how many participants make sense — and a study should name which one it is using and follow it. But they share a common ambition that distinguishes them from most other qualitative work, and that ambition is what a mixed-methods design should be buying. What the method is for: structure, not opinion The ambition is this: to produce an account of the structure of an experience — what it consists of, what is essential to it, how its parts relate — rather than a catalogue of what people said about it. The distinction is easy to miss and it is the difference between a useful qualitative strand and a useless one. Consider a study of living with type 2 diabetes in a low-income population. A survey-like qualitative approach asks what patients find difficult and reports the frequencies: cost of testing supplies, dietary restrictions, forgetting medication, fear of complications. This is opinion collection with quotations. It yields a list that could have been generated from the existing literature, and its practical implication is a list of barriers to address. A phenomenological approach asks a different question: what is it like to live a life organised around a condition that has no symptoms until it has permanent consequences? Done well, it might return something like this — that the illness is experienced as a demand for continuous self-surveillance in the absence of any felt confirmation that the surveillance is working; that the body, which is normally the transparent medium through which one engages the world, becomes an object under suspicion; that this produces a particular temporal structure in which the present is always mortgaged to a future catastrophe that may not arrive; and that the moral language of "compliance" is experienced as a verdict on character delivered by someone who has not had to live inside that structure. That is not a list of barriers. It is an account of a form of life, and it does explanatory work the list cannot. It predicts, for instance, that an intervention adding more self-monitoring will intensify exactly the burden that drives disengagement, and that a clinician's reassurance that "your numbers look good" may land as a reprieve rather than as information — which in turn predicts a pattern of post-reassurance discontinuation that is testable in pharmacy refill data. The structural account generates quantitative hypotheses. The barrier list mostly generates programme components. This is the first thing to understand about integrating phenomenological work with epidemiological data: the integration is only worth doing if the qualitative strand is producing structural accounts. A list of perceived barriers merged with a regression model produces a joint display in which one column says "cost" and the other says "income is a significant predictor," and the reader learns nothing that either strand did not already say. Bracketing, and the honest version of it The most misunderstood element of the tradition is bracketing. In protocols it often appears as a claim to have set aside preconceptions so as to encounter the data freshly — a claim that, taken at face value, is not credible and is not what serious practitioners mean. The defensible version is procedural rather than psychological. It consists of writing down, before analysis, what you expect to find and why; conducting interviews in a way that does not lead the participant toward those expectations; attending deliberately to material that contradicts them; and documenting, as analysis proceeds, the points at which your prior commitments shaped an interpretive choice. In the descriptive traditions the aim is to keep the prior commitments from silently determining the result. In the hermeneutic and interpretative traditions — van Manen, IPA — the aim is different: the researcher's position is treated as the instrument through which understanding occurs, so it is made explicit rather than suspended. Both are respectable. What is not respectable is claiming the first while doing the second, or making the claim and producing no evidence of the work. For a mixed-methods study there is a specific and awkward version of this problem. The qualitative interviewer usually knows the quantitative findings. They know the hypertension control gap, the adherence figures, the neighbourhood pattern. That knowledge shapes what they hear as relevant and what they follow up. In a sequential explanatory design this is deliberate and appropriate — the whole point is to explain a specific quantitative result — but it must be handled explicitly, because an interviewer hunting for confirmation of a known coefficient will find it. The practical safeguards are ordinary: an interview guide whose opening questions are genuinely open and do not presuppose the finding, a deliberate discipline of asking about cases that would disconfirm, a second analyst without exposure to the quantitative results coding a subset, and an audit trail recording where the quantitative prior entered. Sample size, and the argument about saturation Epidemiologists reviewing a qualitative aim almost always ask the same question first: how do you know twenty-five is enough? The traditional answer, data saturation — sampling until new cases stop yielding new codes — has come under sustained and warranted criticism. As commonly used it is unfalsifiable: authors declare saturation without showing the evidence, and the declaration is made after recruitment stopped for practical reasons. It also imports an assumption from the wrong paradigm, treating themes as items in a finite population to be enumerated. The better available framework is Kirsti Malterud and colleagues' concept of information power, which holds that the number of participants needed depends on how much relevant information the sample holds, and offers five dimensions to reason about it: the breadth of the study aim, the specificity of the sample, whether the study is supported by established theory, the quality of the dialogue in the interviews, and whether the analysis is a cross-case thematic analysis or an in-depth analysis of few cases. A narrow aim, a highly specific sample, strong theoretical grounding, and rich interviews all reduce the number required. This is arguable in a protocol and assessable by a reviewer, which saturation as practised is not. The numbers this produces will still look small to a quantitative collaborator. IPA studies commonly involve six to twelve participants; Giorgi's method can work with three to five; a cross-case thematic study within a mixed-methods project might use twenty to forty. The size is not a weakness to be apologised for, because the sample is not doing the work a quantitative sample does. It is not estimating a population parameter and it carries no claim about prevalence. It is characterising a structure, and structures can be characterised from few well-chosen cases in the same way that the mechanism of an engine can be established from one engine. The corollary matters for integration and is frequently violated: a phenomenological strand must never be reported with counts implying prevalence. "Twelve of twenty-five participants described distrust of the clinic" is a sentence that does damage. It invites the reader to treat forty-eight per cent as an estimate, which it is not — the sample was chosen for information richness, not representativeness, so the proportion is an artefact of recruitment. If prevalence of a construct is the question, the right move is to develop a measure from the qualitative work and administer it to a probability sample, which is precisely the exploratory sequential design of Chapter 6. The interview as an instrument A phenomenological interview does not resemble a survey administered aloud, and the difference is not one of formality. A survey asks a fixed set of questions in a fixed order so that answers are comparable across respondents; comparability is the whole design principle, and deviation is error. A phenomenological interview asks the participant to return to concrete episodes and describe them in detail, and it follows wherever the description leads, because the structure being sought is the participant's, not the interviewer's. The practical craft turns on a small number of moves. The opening question should invite a narrative of a specific occasion rather than a generalisation — "tell me about the last time you went to the clinic" rather than "what do you think of the clinic." Generalisations are cheap and largely retrieved from a stock of socially available opinions; episodes are expensive and contain detail the participant did not curate. When a participant does generalise, the interviewer's job is to ask for the instance behind it: what happened, who was there, what was said, what did you do next. The second move is to pursue the taken-for-granted. Phenomenological data live in what participants assume needs no explanation, so the productive follow-up is often a request to explain the obvious — "you said you just knew it wasn't worth going in; how did you know?" This can feel rude, and interviewers new to the method tend to skip it. It is where most of the yield is. The third is tolerating silence and resisting the urge to supply a candidate answer. An interviewer who offers examples — "was it the cost, or the travel?" — has converted an open question into a multiple-choice item and closed off whatever the participant was about to say. In mixed-methods work this failure is endemic, because the interviewer is frequently a member of a team that already has hypotheses and finds them hard not to voice. None of this is neutral or frictionless. Who conducts the interview matters, particularly in disparities research, where the participant's reading of the interviewer's race, class, institutional affiliation, and apparent purpose will shape what is sayable. This is not a bias to be eliminated — there is no view from nowhere in an interview — but it is a condition to be understood and reported. A study in which university-affiliated interviewers asked residents of a historically surveilled neighbourhood about their distrust of institutions has produced data about what those residents will say to a university-affiliated interviewer, which is a real and interpretable thing, and not the same as what they would say to a neighbour. From transcript to structure The step that reviewers understand least is how a pile of transcripts becomes a claim about structure. It is worth describing concretely, because the credibility of the whole strand rests on it and because vagueness here is what invites the suspicion that qualitative analysis is impressionistic. Begin with a single transcript read whole, without coding, to get the shape of the account. Then work through it in meaning units — segments bounded by a shift in what is being talked about, not by sentence or paragraph. For each, write what is being described in the participant's own terms. This first pass is deliberately close to the surface and produces a great deal of material that will not survive. The second pass transforms these into disciplinary language: not the participant's words, and not a theory's words either, but a description of what the participant's account shows about the phenomenon. If a woman describes checking her blood sugar four times a day and says "I'm looking for something to be wrong," the meaning unit is not "monitoring behaviour" and it is not "health anxiety." It is closer to: the monitoring practice is oriented toward detection of a fault, not confirmation of wellbeing, so a normal result provides no positive information and the search continues. That formulation is answerable to the transcript and it is a claim about structure. The third step is cross-case. Working across transcripts, the analyst asks which of these transformed units recur, which are variants of a common structure, and which are genuinely idiosyncratic. What emerges is not a frequency table but a description of a general structure together with the range of its variations — including, importantly, cases that do not fit, which in the descriptive traditions are treated as boundaries of the structure rather than as outliers to be dropped. Throughout, the discipline that makes the result checkable is the audit trail: the preserved chain from transcript segment through transformation to claim, such that another analyst can ask where a particular assertion came from and receive an answer. This is the qualitative equivalent of a reproducible analysis script, and mixed-methods teams should treat it with the same seriousness. When an epidemiologist asks "how do you know that?", the correct response is not a statement about immersion in the data; it is to show the chain. What phenomenology cannot do Three limits should be stated plainly, because overclaiming is the most common way a qualitative strand loses the confidence of quantitative collaborators. It cannot establish prevalence, incidence, or effect size. Nothing about the method supports inference from the sample to a population's distribution. It cannot establish causation in the counterfactual sense. Participants' accounts of why they did something are data about how they understand their action, which is a genuine object of knowledge, and also a poor guide to what would have happened otherwise. People confabulate, and they do so most confidently about socially loaded behaviour. A design that treats participants' causal attributions as causal evidence has confused two different claims. And it cannot, on its own, tell you whether what it found is specific to the group studied or general to the human condition. A structural account of illness among low-income Black women in a southern US city may describe something about that position, or something about chronic illness, or something about poverty, or something about being a patient anywhere. Distinguishing these requires comparison — either a comparative qualitative design or, more efficiently, integration with quantitative data that can establish whether the associated pattern is population-specific. That last limit is the one that makes the case for mixing most directly. The phenomenological strand finds a structure. The epidemiological strand can say where that structure has purchase and where it does not. Neither can do the other's part, and the decision about how to join them is the subject of the chapters that follow. Chapter 3: The Paradigm Question, and Why It Is Not Your Problem Every introduction to mixed-methods research contains a chapter on paradigms, and most of them are a trap. The argument runs that quantitative research rests on a postpositivist worldview — a mind-independent reality, knowable approximately, through objective procedures — while qualitative research rests on constructivism, in which realities are multiple and locally constructed, and that these are incompatible at the level of ontology and epistemology. From this some writers concluded, in the 1980s, that mixing is incoherent: the incompatibility thesis. A researcher approaching the field for the first time can spend a great deal of time in this literature and emerge with nothing usable. The dispute deserves a short hearing, because it is not empty, and then it deserves to be set down, because in twenty years of practice it has almost never been the thing that determined whether a study worked. What the argument was actually about The incompatibility thesis was never mainly about philosophy. It was about power, and specifically about the standing of qualitative research in fields — evaluation, health services, education — where quantitative work held the money and the journals. Qualitative researchers who had spent a decade arguing that their work was not a soft preliminary to real research were understandably wary of an integration that in practice meant supplying quotations for someone else's report. Insisting on paradigmatic incommensurability was a way of defending the independence of a tradition that was being asked to serve. That was a reasonable defence of a real interest, and the interest has not gone away. Anyone who has watched a qualitative aim get cut from a resubmission for space, or seen a phenomenological analysis reduced to two illustrative quotations in a results section, knows that the concern was well founded. But as an argument about whether the two kinds of evidence can bear on each other, it does not survive contact with practice. The clearest response came from David Morgan, who argued that the paradigm framing imported an unnecessarily metaphysical picture of what a research community is. A paradigm, in Kuhn's original sense, is a shared set of exemplars and practices, not a metaphysical commitment; researchers learn how to work by apprenticeship to good examples, and the philosophical positions attributed to them are mostly reconstructions after the fact. Most working epidemiologists have never committed to postpositivism and would not recognise the description. Most interview researchers do not in fact believe that there is no such thing as a tumour. The pragmatist alternative, developed for this field by Abbas Tashakkori and Charles Teddlie and by John Creswell and Vicki Plano Clark, takes the research question as primary and treats methods as tools selected for their fitness to it. Truth, on this view, is what works in the sense of what withstands inquiry, and the dictates of any single method are not binding on what may be asked. A more demanding alternative, Burke Johnson's dialectical pluralism, does not dissolve the tension but institutionalises it: the differences between paradigms are kept live within a team, so that the disagreement between a constructivist and a realist reading of the same finding becomes a resource rather than a problem to be managed away. Both positions are defensible. Both give you permission to proceed. Where the tension is real Setting down the metaphysics does not mean the traditions get along automatically. There are practical frictions that repeat in almost every mixed-methods team, and they are worth naming because they masquerade as philosophical disagreements when they are usually disagreements about craft. The first concerns generalisation. The quantitative strand is built to support inference from sample to population, and everything about its sampling and reporting serves that end. The qualitative strand generalises differently — to theory, to the structure of a phenomenon, or by supplying enough contextual detail that a reader can judge transferability to their own setting. Neither of these is weaker; they are answers to different questions. The friction arises when a team writes a single abstract and has to decide what the study claims. The second concerns the fixity of constructs. A survey instrument must mean the same thing to every respondent or the responses cannot be pooled; equivalence of measurement across groups is testable and is a genuine achievement when it holds. Qualitative inquiry frequently finds that it does not hold — that "feeling depressed," "trusting your doctor," or "having enough food" are constituted differently across communities in ways that make the pooled score a composite of unlike things. This is not a philosophical objection to measurement. It is an empirical finding about a specific instrument, and it is one of the most valuable things a qualitative strand can produce, because it is directly actionable. The third concerns the role of the researcher. In one tradition, the investigator's characteristics are a source of bias to be minimised by procedure. In the other, they are part of the instrument and are addressed by being made explicit. A team that has not discussed this will produce a methods section that claims both. The fourth is about evidence hierarchies and is felt most acutely when the work meets policy. Guideline panels, health technology assessment bodies, and funders operate with ranking schemes in which experimental designs sit at the top and qualitative evidence, if it appears at all, appears as context. A team producing an integrated finding will discover that the machinery for appraising it is thin: systematic review methods for mixed-methods evidence exist but are young, and a reviewer trained on risk-of-bias tools has no established way to weigh a phenomenological account against a hazard ratio. This is a real constraint on what integrated work can accomplish downstream, and it is a reason to be deliberate about which strand carries the claim a decision-maker will act on. None of these requires a philosophical settlement. They require a team that can name the difference and decide, per study, what the study will claim and on what basis. The question that actually matters Here is the substitution I would recommend. Instead of asking what paradigm the study occupies, ask: at what points do the two strands touch, and what happens at each of those points? This is a design question with concrete answers, and it turns out to be the question that distinguishes competent mixed-methods work from the stapled-together kind. Michael Fetters, Leslie Curry, and John Creswell set out the useful version of it, identifying integration as occurring at three levels — at the level of design, where the strands are sequenced and prioritised; at the level of methods, through connecting, building, merging, or embedding; and at the level of interpretation and reporting, through narrative, joint displays, or data transformation. A study should be able to say, of each level, exactly what it does. Connecting means that one strand's results determine the sample for the other — the cohort analysis identifies who gets interviewed. Building means that one strand's results shape the other's instrument — the interviews generate items for a questionnaire, or an unexpected qualitative finding prompts a new covariate to be extracted from the records. Merging means the two datasets are brought together for joint analysis after both are complete. Embedding means one strand is nested inside a larger design, as a process evaluation inside a trial. These are not alternatives to be chosen once; a single study may build early, connect in the middle, and merge at the end. But each must be planned, because each imposes requirements on what is collected and when. Why you are mixing Before the design question comes a prior one, and it is the one most often skipped: what is the mixing supposed to accomplish? The most durable answer remains the typology Jennifer Greene, Valerie Caracelli, and Wendy Graham proposed in 1989 for mixed-method evaluation designs, which has held up because it is specific about what each purpose requires of a design, as Table 1 sets out. Table 1. Purposes for mixing, after Greene, Caracelli and Graham (1989), with what each demands of a design. Purpose What it seeks Design requirement Failure signature Triangulation Convergence on the same construct from independent methods Strands must be genuinely independent and address the same construct Strands share a data source, so agreement is circular Complementarity Elaboration and clarification of results from one method by the other Overlapping but distinct facets; comparable domains Strands address unrelated questions; nothing to elaborate Development One method informs the instrument or sample of the other Strict sequencing; time between phases Phases run in parallel, so the informing cannot happen Initiation Paradox and contradiction; new frames of reference Capacity to surface and report discordance Disagreement is smoothed into a consensus narrative Expansion Extending breadth by using different methods for different components Clear scope for each strand Presented as integration when it is parallel reporting Two entries in that table deserve emphasis for disparities research. Triangulation is the purpose most often claimed and least often achieved. The classical idea — that convergence from independent methods increases confidence — requires that the methods really be independent in their sources of error. In practice, interviews and surveys conducted by the same team, with the same framing, recruited from the same clinic, share a great deal of error structure. When they agree, it may be because they are picking up the same artefact. Janice Morse's early formulation of methodological triangulation was careful about this; the casual usage that followed was not. Initiation is the purpose least often claimed and most often valuable. Greene and colleagues meant by it the deliberate seeking of paradox — designing so that contradiction between strands can appear and be taken seriously, on the grounds that a contradiction is a signal that some assumption shared by both analyses is wrong. In disparities research, where the quantitative categories are contested and the measures are frequently non-equivalent across groups, contradiction is likely, and a design that can register it is worth a great deal. Most designs cannot, because they are built to produce an integrated narrative and treat disagreement as a problem of interpretation. Chapter 8 returns to this at length. A case where the difference bites The abstract claim that constructs may not be equivalent across groups becomes concrete in screening instruments, and depression screening is the standard example. A brief depression screen asks about low mood, anhedonia, sleep, appetite, energy, self-worth, concentration, psychomotor change, and suicidal ideation. It was developed and validated in populations where distress is habitually reported in psychological vocabulary, and its scoring assumes that the items are indicators of a single underlying dimension with stable loadings. Applied across populations, it frequently produces group differences in prevalence that are then reported as disparities in depression. Qualitative work in cross-cultural psychiatry has documented for decades that idioms of distress differ: in many communities the same underlying state is expressed and experienced through bodily complaints — pain, fatigue, heaviness, disturbed sleep — rather than through statements about mood, and the psychological vocabulary the instrument requires may be unavailable, stigmatised, or simply not the way the experience presents itself. If that is so, the instrument is not measuring the same thing in both groups, and the observed prevalence difference is a mixture of a real difference and a measurement artefact in unknown proportion. Notice what this is and is not. It is not a philosophical objection to the possibility of measurement. It is a testable psychometric claim — differential item functioning, measurement non-invariance — for which quantitative methods exist and which is checked far less often than it should be. What the qualitative work supplied was not a refutation but a hypothesis about where to look, and a specification of which items to suspect. That is the characteristic shape of a productive interaction between the strands: the interview generates a structured account, the account implies something testable about the instrument, and the quantitative strand tests it in a sample large enough to support the test. The disparities implication is severe and worth stating baldly. If a measure functions differently across the groups being compared, every comparison built on it — prevalence, trend, treatment effect, disparity magnitude — inherits the distortion, and no amount of adjustment for confounders touches it. A field that compares groups on instruments never examined for invariance in those groups is reporting a quantity that is part substance and part artefact. Mixed-methods work is one of the few routes to finding out which. Who holds the integration There is a final practical point that belongs in this chapter because it is where the paradigm question actually shows up in working life: integration is done by people, and it fails for organisational reasons more often than intellectual ones. Alan Bryman's interview study of British researchers who had conducted mixed-methods projects found that a striking number reported not having integrated their data, and gave reasons that were largely structural: the strands were run by different people with different timetables, journals wanted one kind of paper, audiences expected one kind of finding, and individual skills ran in one direction. Nobody in these projects had decided against integration. It simply did not happen, because no one's job was to make it happen and nothing in the project's structure forced it. The remedies are unglamorous and effective. Name a person responsible for integration, with time allocated to it, rather than assuming it will emerge from a collaboration. Hold joint analysis sessions at which both datasets are on the table simultaneously, scheduled before either strand's analysis is finalised — if the qualitative analysis is complete and written up before the quantitative team sees it, the only available integration is commentary. Write the integrative paper first, or at least plan it first, rather than treating it as what is left after the two strand papers are placed. And build in enough shared vocabulary that each side can interrogate the other: an epidemiologist who cannot ask a pointed question about how a theme was derived will defer to it, and deference is not integration. Teams sometimes attempt to solve this by hiring a single person trained in both. That helps, and it is not sufficient, because the volume of work in a serious study exceeds what one person can hold. What actually works is a team in which at least two people can read both datasets competently and are scheduled to do so together. The test, restated Set against all this, the useful discipline at the design stage is a short list of questions that a protocol should answer in plain sentences, before any philosophy is invoked: · What is the question that neither strand could answer alone? · Which of the five purposes above is the mixing serving? If more than one, which is primary? · At what points do the strands touch — connecting, building, merging, embedding — and in what order? · What would disagreement between the strands look like, and what would the study do if it occurred? · Which strand takes priority if resources are cut, and what does the study still claim if the other is reduced? A protocol that answers these is doing mixed-methods research. A protocol that discusses pragmatism for two paragraphs and then describes two parallel workplans is not, however sound its epistemology. The last question on that list deserves particular attention, because the situation it describes is not hypothetical. Budgets are cut, recruitment runs late, and one strand almost always ends up compressed. Teams that have decided in advance which strand is load-bearing make that cut coherently. Teams that have not tend to preserve the quantitative strand by default, because it has the primary outcome and the power calculation attached to it, and discover at the end that the design they described no longer exists. With the philosophical question set aside and the practical one in its place, the next task is to look at the configurations available — what the three core designs actually are, what each buys, and what each costs. Hashtags: #MixedMethodsInPublicHealth #MixedMethodsResearch #PublicHealthResearch #HealthDisparities #EpidemiologicalData #QualitativeResearch #PhenomenologicalInquiry #EvidenceIntegration #ConvergentDesign #ExplanatorySequentialDesign #ExploratorySequentialDesign #JointDisplays #MixedMethodsMatrices #DataTransformation #ConnectingMethods #BuildingMethods #MergingMethods #EmbeddedDesign #ProcessEvaluation #Intersectionality #MeasurementInvariance #CommunityExperience #MechanismDiscovery #IntegratedEvidence #FutureOfMixedMethodsPublicHealth
- Missing Data Mechanics: Multiple Imputation by Chained Equations (MICE) and FIML
Download the Book (PDF): Introduction Every data set that matters has holes in it. A survey respondent skips the income question. A patient in a twelve-month trial stops attending after month four. A sensor drops its readings during a power cut. A school declines to release test scores for one cohort. A laboratory assay fails for samples below a detection threshold. The holes are rarely the point of the study, and so they are treated as a nuisance: something to be cleaned up before the real analysis begins. That framing is the root of most bad practice in the field. Missing data are not a cleaning problem. They are a modelling problem, and the decisions made about them are substantive claims about the world that deserve the same scrutiny as any other assumption in an analysis. The claim this book defends is simple to state and demanding to live up to. Every method for handling incomplete data, including the methods that appear to make no assumptions at all, rests on an assumption about why the data are missing. Multiple imputation by chained equations (MICE) and full information maximum likelihood (FIML) are two principled ways of acting on the most useful of those assumptions, the assumption that data are missing at random. Both work when, and only when, the model they use contains the information that explains the missingness and respects the structure of the analysis that follows. Neither can verify its own central assumption from the observed data. The competent analyst therefore does three things: builds an imputation or estimation model rich enough to make the missing-at-random assumption plausible, checks that the machinery has actually done what it claims, and then asks how far the conclusions would move if the assumption were wrong. The software handles the arithmetic. The judgement cannot be delegated to it. Why the default is wrong Most statistical software, faced with a row that has a missing value in any variable used in a model, silently drops that row. This is listwise deletion, or complete-case analysis, and it is the default in R's lm(), in most regression procedures in SAS and Stata, and in countless spreadsheets. Its appeal is that it seems neutral: the analyst has not invented any data. But dropping cases is itself a model of the missing values. It implicitly asserts that the people who were dropped are, in the respects that matter to the analysis, the same as the people who were kept. When that assertion is false, and it frequently is, complete-case estimates are biased, sometimes badly. Even when it is true, the analysis throws away information. A regression with ten predictors, each missing independently in five per cent of cases, retains only about sixty per cent of its rows, because the probability that a row is complete on all ten is 0.95 raised to the tenth power. The analyst who began with a thousand participants now reports on six hundred without having decided to. The other common improvisations fare no better. Replacing a missing value with the variable's mean shrinks its variance and weakens its correlations with everything else. Carrying the last observation forward in a longitudinal study asserts that nothing changes after a participant leaves, which is seldom what anyone believes. Adding a missing-value indicator to a regression, a trick once widely recommended, produces biased coefficients even when the data are missing completely at random, as several methodologists have shown. Each of these methods is easy to run, and each embeds an assumption that is usually stronger and less credible than the one principled methods require. Two principled routes The modern alternatives descend from a single intellectual source. In 1976 Donald Rubin published a short paper in Biometrika titled "Inference and Missing Data" which introduced a formal way of describing the process that generates missingness, and showed under what conditions that process can be ignored when drawing inferences. The vocabulary that paper introduced, of data missing completely at random, missing at random, and missing not at random, has organised the field ever since, though, as later chapters show, the words are often used more loosely than the mathematics allows. From that foundation grew two families of method. The first is likelihood-based. If the data are missing at random and the analyst can write down a model for the joint distribution of the variables, the likelihood of the observed data alone is sufficient for valid inference. Maximising that likelihood directly, case by case, using whatever each case has observed, is what structural equation modellers call full information maximum likelihood. The expectation-maximisation algorithm of Dempster, Laird and Rubin (1977) is a closely related route to the same estimates. The second family is multiple imputation, which Rubin developed in the 1970s and set out in book form in 1987. Instead of maximising a likelihood, the analyst fills in each missing value several times with plausible draws from its predictive distribution, analyses each completed data set in the ordinary way, and combines the results with a set of simple formulas now called Rubin's rules. The variation between the imputed data sets carries the uncertainty that the missing values introduce. Multiple imputation by chained equations is the most widely used practical implementation of that second idea. Rather than specifying one joint distribution for all variables at once, the analyst specifies a separate conditional model for each incomplete variable, a regression of that variable on the others, and cycles through them repeatedly. The approach was developed in parallel by Stef van Buuren and colleagues in the Netherlands, who wrote the mice package for R, and by Trivellore Raghunathan and colleagues at the University of Michigan under the name sequential regression multivariate imputation. Its flexibility is its great strength, since each variable can be imputed with a model suited to its type, and its great danger, since that flexibility allows an analyst to build a set of conditional models that do not fit together or do not match the analysis to come. What this book covers The chapters move from concepts to machinery to judgement. Chapter 1 describes the anatomy of missingness: how to represent and summarise patterns of missing values, and why deletion and single imputation fail even in simple cases. Chapter 2 states Rubin's three mechanisms precisely, using the missingness indicator and the factorisation of the joint distribution, and explains the conditions under which the missingness process can be ignored. Chapter 3 turns to diagnostics: what the observed data can reveal about the mechanism, including Little's test of missing completely at random and the modelling of missingness indicators, and what they fundamentally cannot reveal. Chapter 4 develops the theory of multiple imputation: what makes an imputation "proper", why a single imputation understates uncertainty, how Rubin's rules work, and how the fraction of missing information guides the number of imputations. Chapter 5 turns that theory into practice with MICE, covering the chained equations algorithm, the choice of imputation methods including predictive mean matching, and the construction of the predictor matrix. Chapter 6 deals with the hard cases that break naive imputation: interactions and nonlinear terms, derived variables and scales, multilevel data, and the principle of congeniality between imputation and analysis. Chapter 7 presents full information maximum likelihood, its casewise likelihood, its implementation in lavaan, the use of auxiliary variables through Graham's saturated correlates approach, and how it compares with multiple imputation. Chapter 8 addresses validation: checking convergence, comparing observed and imputed distributions, overimputation and simulation by amputation. Chapter 9 goes beyond missing at random to sensitivity analysis under missing-not-at-random mechanisms: delta adjustment, pattern-mixture and selection models, reference-based imputation in clinical trials, and tipping-point analysis. The conclusion draws out what follows for practice and reporting. The examples use R, principally the mice and lavaan packages, because they are free, widely used and well documented, with occasional Python where the comparison is instructive. The code is short and meant to teach the logic rather than serve as a template to paste. Equations are written inline in plain text: P(R | Y_obs, Y_mis) is the probability of the observed pattern of missingness given the observed and missing parts of the data, and so on. Readers need a working knowledge of regression and a willingness to think about probability distributions; no knowledge of Bayesian computation is assumed, though the ideas are introduced where they matter. A note on attitude Missing data methods attract two kinds of misplaced confidence. The first is the belief that because a principled method has been used, the missing data problem has been solved. Multiple imputation run with default settings, on a model that omits the variables that predict dropout, is not principled in any useful sense; it is complete-case analysis with extra steps and narrower confidence intervals. The second is the belief that, because the key assumption cannot be tested, all methods are equally arbitrary and the choice between them does not matter. That is also wrong. Some assumptions are more plausible than others, and plausibility can be increased by design, by measurement, and by modelling. A study that records the reasons participants drop out, measures predictors of non-response, and plans its imputation model before seeing the outcome data is in a far stronger position than one that does none of these things, even though neither can prove its assumption holds. The goal throughout is to make the assumptions visible. An analysis of incomplete data is persuasive not when it uses the most sophisticated algorithm but when a sceptical reader can see what was assumed about the missing values, why that assumption was reasonable, what was done to check the machinery, and how the conclusions would change under plausible alternatives. The rest of this book is about how to earn that persuasiveness. Chapter 1: The Anatomy of Missingness Before any question of mechanism or method, there is a question of description. Which values are missing, in which variables, for which cases, and in what combinations? Analysts who skip this step and go straight to an imputation routine often discover, too late, that a third of their sample is missing an entire block of variables, or that two variables are never observed together, or that the outcome is missing for exactly the people the study was designed to understand. Description is cheap and it frequently changes the plan. The data matrix and its shadow Think of the data as a matrix Y with n rows, one per unit of observation, and p columns, one per variable. Alongside it sits a second matrix of the same shape, R, the response indicator, whose entry r_ij equals 1 if y_ij is observed and 0 if it is missing. (Some authors code it the other way round; the logic is the same.) The matrix R is the shadow of the data. It is always fully observed, and it is itself data: the pattern of zeros and ones can be tabulated, modelled, and correlated with the observed values of Y. Much of the theory in the chapters that follow is a theory about the joint distribution of Y and R. Partition each row of Y into its observed part, Y_obs, and its missing part, Y_mis. Which values fall into each part differs from row to row, so it helps to think of a missing data pattern as a distinct row of R. A data set with five variables has at most 32 possible patterns, but real data usually show far fewer. Grouping cases by pattern is the first descriptive step. Several structural patterns recur often enough to have names. In a univariate pattern, only one variable has missing values, and every other variable is complete. This is the textbook case, and it is surprisingly common: a trial in which baseline covariates are fully recorded but the primary outcome is missing for dropouts fits it exactly. In a monotone pattern, the variables can be ordered so that if a case is missing variable j, it is also missing every variable after j. Longitudinal studies with attrition, where people who leave never return, produce monotone patterns. They matter because monotone data admit simpler, sequential estimation and imputation: each variable can be modelled from the ones before it, and no iteration is needed. In a general or arbitrary pattern, missing values are scattered with no such ordering. Most observational data sets look like this. Two further patterns deserve mention because they create special difficulties. In file matching, two variables are never observed together in the same case, as when two surveys sharing some common questions are merged. Their association is not identified by the data at all, and any imputation will simply reproduce whatever the model assumes about it, usually conditional independence given the common variables. In planned missingness designs, the analyst deliberately withholds some items from some respondents, for instance by randomly assigning each respondent three of four questionnaire blocks, to reduce burden. Because the missingness is created by the researcher's randomisation, its mechanism is known, and such designs are among the few situations in which a strong assumption about missingness can be guaranteed rather than hoped for. Summarising what is missing The mice package provides a quick way to see patterns. The function md.pattern() lists each distinct pattern, the number of cases with that pattern, and the number of missing values in each variable. Using the small nhanes data set that ships with mice, which has 25 cases and four variables (age group, body mass index, hypertension and cholesterol): library(mice) md.pattern(nhanes, plot = FALSE) colMeans(is.na(nhanes)) # proportion missing per variable mean(complete.cases(nhanes)) # proportion of complete rows The output shows that age is always observed, that 13 of the 25 rows are complete, and that the remaining rows fall into a handful of patterns, including several in which body mass index, hypertension and cholesterol are all missing together. That last detail matters. When variables go missing in blocks, the observed data carry little information about the missing block for the affected cases, and imputations for them rely heavily on the model. Three further summaries repay the effort. The first is pairwise coverage: for each pair of variables, the proportion of cases in which both are observed. The function md.pairs() returns the counts. Coverage is the raw material from which any estimate of a correlation must be built; if two variables are jointly observed in only thirty cases, no method can estimate their association precisely, however sophisticated. The second is influx and outflux, a pair of statistics proposed by van Buuren. Influx measures how well a variable's missing entries are connected to observed data in other variables; outflux measures how well a variable's observed entries are connected to missing data elsewhere. A variable with high outflux is a useful predictor for imputing others; one with low outflux contributes little and can often be left out of the predictor set. The function flux() computes both. The third summary is the distribution of the number of missing values per case. A small number of cases missing almost everything behaves very differently from a large number of cases each missing one item, even when the overall proportion of missing cells is the same. The former are nearly empty rows whose imputation is almost entirely model-driven; the latter carry most of their information and are easy to impute well. A common and harmful habit is to judge a data set by a single figure, "twelve per cent missing", and to choose a method on that basis. Rules of thumb such as "multiple imputation is unnecessary below five per cent" or "imputation is unreliable above forty per cent" have no sound basis. Madley-Dowd and colleagues, writing in the Journal of Clinical Epidemiology in 2019, showed by simulation that the proportion of missing data should not by itself guide the decision to impute; what matters is how much information the observed data and auxiliary variables carry about the missing values, and whether the assumptions of the method hold. Ten per cent missing on an outcome that is strongly predicted by recorded variables is a mild problem. Ten per cent missing on an outcome that is missing because of its own value, with no good predictors, can be severe. Not every blank is a missing value Before counting missing values, it is worth asking whether every blank cell is really missing in the sense this book means: a value that exists, or would exist, but was not recorded. Several kinds of blank are not. A question about the age of a respondent's youngest child is not applicable to a respondent with no children; there is no value to recover, and imputing one creates a fictional child. Skip patterns in questionnaires generate many such structural blanks, and they should be coded as a distinct category or handled by conditioning the analysis on eligibility, not imputed. A laboratory value reported as "below the limit of detection" is not missing either; it is censored, and the analyst knows it lies in a specific interval. Treating it as missing throws away that knowledge, and imputing it without respecting the bound can produce values above the limit. Models for censored data, or imputation constrained to the known interval, are the right tools. A "don't know" response to an attitude question may be a genuine answer that deserves its own category rather than a missing value to be guessed. Longitudinal studies raise a sharper version of the question. When a participant dies before the one-year visit, is their one-year quality-of-life score missing? For some purposes the answer is no: the value does not exist, and an estimand that includes it is describing a hypothetical world in which the person survived. The estimand framework of the ICH E9(R1) addendum, adopted in 2019 for clinical trials, forces this distinction by asking analysts to specify how intercurrent events such as death, treatment discontinuation or use of rescue medication are to be handled before deciding what counts as missing. The same discipline serves observational research well. Every blank should be classified as missing, not applicable, censored, or a substantive response before a single imputation model is written. Why deletion fails Complete-case analysis discards every row with any missing value on the variables in the model. Its failures come in two forms: bias and waste. Consider the bias first, through a hypothetical example. Suppose a study measures systolic blood pressure at baseline and at one year in 1,000 adults, and the one-year value is missing for 300 of them. Suppose further that people with higher baseline blood pressure are more likely to miss the follow-up visit, perhaps because they are older and less mobile. The complete cases are then disproportionately people with lower baseline pressure, and because baseline and follow-up pressures are strongly correlated, they are also disproportionately people with lower follow-up pressure. The mean of the observed follow-up values underestimates the population mean. No amount of additional sample size fixes this; complete-case analysis converges confidently to the wrong answer. Now change the target. Instead of the mean of follow-up pressure, suppose the analyst wants the regression of follow-up pressure on baseline pressure. Here complete-case analysis is unbiased, because the probability of missingness depends only on the covariate, baseline pressure, and the regression of follow-up on baseline is the same among complete cases as in the whole population. This is a general and useful result: in a regression, if the probability that a case is complete depends only on the covariates and not on the outcome given the covariates, complete-case estimates of the regression coefficients are unbiased. White and Carlin examined the implications in a 2010 paper in Statistics in Medicine comparing complete-case analysis with multiple imputation for missing covariates, and found that in some settings complete-case analysis is the better choice. The lesson is not that deletion is always wrong. It is that whether deletion is biased depends jointly on the mechanism and on the estimand, and the analyst has to think about both. The waste is less subtle. Every case dropped for a missing value in one variable also throws away its observed values in all the others. In a model with many covariates, each with modest missingness, the complete-case sample can shrink dramatically. Standard errors inflate, power drops, and the analyst may be tempted to remove variables from the model simply to recover cases, which introduces a different kind of bias. Pairwise deletion, which computes each correlation or covariance from all cases observed on that pair of variables, recovers some of the waste but introduces its own problems. Different entries of the covariance matrix are based on different subsets of cases, so the matrix need not be positive definite, and regressions built from it can produce impossible results such as R-squared values above one. There is also no single sample size to use for standard errors. Pairwise deletion is unbiased only under the strictest mechanism, missing completely at random, and has no principled advantage over the methods described later. Why single imputation fails Filling each missing value once and then analysing the completed data as if it were fully observed avoids discarding cases, but every variant of single imputation misrepresents either the distribution of the data or the uncertainty of the estimates, and usually both. Mean imputation replaces each missing value with the variable's observed mean. The mean of the completed variable is unchanged, but its variance is understated, because a block of identical values has been inserted, and its covariances with other variables are attenuated, because the inserted values do not vary with anything. Correlations shrink toward zero. Even under the most benign mechanism, mean imputation biases every analysis that involves variability or association. Regression imputation replaces each missing value with its prediction from a regression on the observed variables. This preserves means and, under missing at random, preserves regression relationships of the imputed variable on its predictors, but it overstates them: the imputed values lie exactly on the regression surface, with no residual scatter. Correlations are inflated and variances understated. Adding a random residual to each prediction, known as stochastic regression imputation, fixes the distributional problem and produces unbiased estimates of means, variances and correlations under missing at random. It still fails on uncertainty: the analysis treats the imputed values as if they were real observations, so standard errors are too small and confidence intervals too narrow. The completed data look like a larger sample than was actually collected. Last observation carried forward, long used in clinical trials, replaces a missing follow-up value with the participant's most recent observed value. It assumes that the outcome freezes at the moment of dropout. In a trial of a progressive disease, this makes dropouts look better than they would have been; in a trial where patients improve over time, it makes them look worse. The direction of bias depends on the trajectory and on differential dropout between arms, and the method also understates variance. Regulatory guidance moved away from it after the National Research Council's 2010 report on missing data in clinical trials, which recommended that single imputation methods like it not be used as the primary approach unless their assumptions are scientifically justified. The missing-indicator method for covariates sets missing values to a constant, often zero, and adds a dummy variable flagging which values were missing. It has intuitive appeal and is easy to implement, but Jones showed in a 1996 paper in the Journal of the American Statistical Association that it gives biased regression coefficients even when data are missing completely at random, except in special cases such as randomised trials where the incomplete covariate is independent of treatment. It remains reasonable for baseline covariates in randomised trials, where its bias does not affect the treatment effect, and for prediction models where the missingness itself is predictive and will be present at deployment. It is not a general solution. Hot-deck imputation, which replaces each missing value with an observed value from a similar donor case, has a long history in government surveys and preserves the distribution of the variable well, because every imputed value is a real observed value. Performed once, it shares the uncertainty problem of all single imputation. Performed several times with appropriate randomness, it becomes a form of multiple imputation, and the predictive mean matching method discussed in Chapter 5 is its principled descendant. The comparison is summarised in Table 1, which describes each ad hoc approach by what it preserves and what it distorts under the most favourable mechanism, missing completely at random. Under less favourable mechanisms, the distortions typically grow. Table 1. Ad hoc approaches to missing data and their typical consequences. Method Means Variances and correlations Standard errors Listwise deletion Unbiased only if MCAR (or estimand-specific) Unbiased only if MCAR Valid for reduced sample; inefficient Pairwise deletion Unbiased if MCAR Matrix may be non-positive definite No clear sample size Mean imputation Preserved Variances shrunk; correlations attenuated Too small Regression imputation Preserved under MAR Variances shrunk; correlations inflated Too small Stochastic regression Preserved under MAR Preserved under MAR Too small Last observation carried forward Biased unless outcome truly stable Variances distorted Too small From description to assumption The pattern of failures in Table 1 has a common cause. Each method either discards information or treats estimated values as known, and each embeds an assumption about why the data are missing that is usually left unstated. Deletion assumes the dropped cases resemble the kept ones in the relevant respects. Mean imputation assumes the missing values are unrelated to anything else in the data. Carrying forward assumes nothing changes after dropout. Principled methods do not escape the need for assumptions; no method can, because the missing values are, by definition, unobserved. What distinguishes them is that the assumption is explicit, is usually weaker, and is made in a form that the analyst can reason about and, to a limited extent, strengthen by collecting and using the right information. The descriptive work of this chapter feeds directly into that reasoning. A table of patterns shows where the information is thin. Coverage shows which associations are poorly supported. Outflux identifies useful predictors. The distribution of missing values per case identifies near-empty rows that may deserve separate treatment. All of this should be done, and reported, before any imputation is run. The next step is to make precise what "why the data are missing" means, which is the subject of Chapter 2. Chapter 2: Rubin's Taxonomy Stated Precisely The terms missing completely at random, missing at random and missing not at random appear in almost every applied paper that handles incomplete data, usually in a sentence asserting that the data were assumed to be missing at random. Few of those papers say what the assumption means, and many use it in ways its author would not recognise. The phrase "missing at random" sounds like a claim that missingness is haphazard. It is not. It is a precise, conditional statement about the relationship between the missingness indicator and the data, and it can hold when missingness is strongly systematic. Getting the definitions right is not pedantry; it determines what information the analyst must include in the model for principled methods to work. The missingness model Recall the data matrix Y and the response indicator R from Chapter 1. Rubin's insight was to treat R as a random variable with its own distribution, and to write the joint distribution of data and missingness as a product of two parts: f(Y, R | theta, psi) = f(Y | theta) * f(R | Y, psi) The first factor, f(Y | theta), is the model for the complete data, the thing the analyst actually cares about, governed by parameters theta such as means, regression coefficients and variances. The second factor, f(R | Y, psi), is the missingness mechanism: the probability of each pattern of observed and missing values, given the values of the data, governed by parameters psi. This way of writing the joint distribution is called the selection model factorisation, because the second factor describes how units are selected into being observed. Splitting Y into Y_obs and Y_mis, the three mechanisms are defined by what the mechanism is allowed to depend on. The data are missing completely at random (MCAR) if f(R | Y, psi) = f(R | psi): the probability of any pattern of missingness does not depend on the data at all, observed or missing. A blood sample dropped on the floor, a questionnaire page lost in the post, or a random subsample chosen for an expensive measurement all generate MCAR data. Under MCAR, the complete cases are a simple random subsample of the full sample, which is why complete-case analysis is unbiased under this mechanism, though inefficient. The data are missing at random (MAR) if f(R | Y, psi) = f(R | Y_obs, psi): the probability of the missingness pattern may depend on observed values but, once those are taken into account, does not depend on the values that are missing. Suppose, hypothetically, that older respondents in a health survey are less likely to report their weight, and age is recorded for everyone. If, among respondents of the same age, the probability of reporting weight does not depend on the weight itself, the weights are MAR given age. The missingness is strongly systematic, since older people are far more likely to be missing, yet it is still MAR. The data are missing not at random (MNAR) if the probability of missingness depends on the missing values even after conditioning on everything observed. If, among respondents of the same age, heavier people are less likely to report their weight, the weights are MNAR. The observed weights at each age are then a biased sample of all the weights at that age, and no method that relies only on the observed data can correct the bias without additional assumptions. These definitions have a hierarchical structure. MCAR is a special case of MAR, since a mechanism that depends on nothing trivially depends only on observed values. MNAR is everything else. Ignorability and why it matters The payoff of the MAR assumption becomes clear when we write the likelihood of what was actually observed. The analyst sees Y_obs and R, and their joint density is obtained by integrating over the missing values: f(Y_obs, R | theta, psi) = integral of f(Y_obs, Y_mis | theta) * f(R | Y_obs, Y_mis, psi) dY_mis Under MAR, the mechanism does not depend on Y_mis, so it can be pulled outside the integral: f(Y_obs, R | theta, psi) = f(R | Y_obs, psi) integral of f(Y_obs, Y_mis | theta) dY_mis = f(R | Y_obs, psi) f(Y_obs | theta) The likelihood now factorises into a piece that depends only on psi and a piece that depends only on theta. If, in addition, the parameters theta and psi are distinct, meaning that knowing one tells you nothing about the other (in Bayesian terms, their prior distributions are independent), then for inference about theta the first factor is a constant and can be ignored. The analyst can base inference on f(Y_obs | theta), the likelihood of the observed data under the complete-data model, without modelling the missingness mechanism at all. This combination, MAR plus distinctness, is what Rubin called ignorability. It is the formal licence behind both full information maximum likelihood, which maximises f(Y_obs | theta) directly, and multiple imputation, which draws the missing values from their distribution given the observed data under the complete-data model. The word "ignorable" is itself a trap. It does not mean the missing data can be ignored. It means the mechanism can be left unmodelled, provided the analysis uses a correct model for the complete data and uses all the observed data. Complete-case analysis does not satisfy this condition, because it discards observed values. Under MAR, complete-case analysis is generally biased while likelihood and multiple imputation methods are not. Distinctness is almost always taken for granted and rarely causes practical difficulty. The MAR condition does the heavy lifting. Refinements that matter in practice Rubin's 1976 definition was stated for the particular data set at hand: the mechanism evaluated at the realised observed values. Later work showed that this version is enough for likelihood and Bayesian inference but not for frequentist statements about repeated sampling, such as the claim that a confidence interval has 95 per cent coverage. For those, the MAR condition must hold for all possible data sets that could have arisen, not just the one observed. Seaman, Galati, Jackson and Carlin, in a 2013 paper in Statistical Science titled "What is meant by 'missing at random'?", distinguished "realised MAR" from "everywhere MAR" and showed that confusion between them runs through the applied literature. Mealli and Rubin revisited the definitions in Biometrika in 2015, introducing the term "missing always at random" for the stronger, everywhere version. For practical purposes, when an applied analyst assumes MAR, the stronger version is what is being assumed, and it is what the simulation studies that justify MICE and FIML actually build in. A second refinement is that the mechanisms are defined relative to a particular set of variables. Whether data are MAR depends on what is included in Y. Return to the weight example. If older people are less likely to report weight and weight itself does not affect reporting once age is known, the weights are MAR given age. But if the analyst's model excludes age, the missingness depends on a variable outside the model that is correlated with weight, and from the model's point of view the weights behave as MNAR. The mechanism has not changed; the analyst's information has. This is the single most important practical consequence of the taxonomy: MAR is something the analyst helps to make true by including the right variables. Variables that predict missingness, or predict the missing values themselves, are called auxiliary variables when they are not part of the substantive analysis, and including them in the imputation or estimation model is how an MNAR problem is moved toward MAR. Collins, Schafer and Kam demonstrated the value of this "inclusive" strategy in a 2001 simulation study in Psychological Methods, showing that including auxiliary variables rarely hurts and can remove substantial bias. A third refinement concerns which variables the mechanism depends on in regression settings. As noted in Chapter 1, when the analysis is a regression of an outcome on covariates, complete-case analysis is unbiased whenever missingness depends only on the covariates, whether the missing values are in the outcome or the covariates themselves. That includes some MNAR mechanisms: if the probability that a covariate is missing depends on the value of that covariate, but not on the outcome, the covariate is MNAR, yet complete-case regression remains unbiased. Conversely, MI and FIML, which assume MAR, would be biased in that situation. The taxonomy tells you which general-purpose methods are valid; it does not rank them for every estimand. Careful analysts ask the specific question, "does missingness depend on the outcome given the covariates?", alongside the general one. A small simulation to fix ideas The differences between mechanisms are easiest to feel by simulating them. The following R code generates a covariate x and outcome y with correlation about 0.6, then deletes half the values of y under each mechanism and compares the complete-case mean of y with the truth. set.seed(2026) n <- 10000 x <- rnorm(n) y <- 0.6 * x + rnorm(n, sd = 0.8) mcar <- runif(n) < 0.5 # depends on nothing mar <- runif(n) < plogis(2 * x) # depends on observed x mnar <- runif(n) < plogis(2 * y) # depends on y itself c(truth = mean(y), MCAR = mean(y[!mcar]), MAR = mean(y[!mar]), MNAR = mean(y[!mnar])) Under MCAR, the complete-case mean is close to zero, the true value. Under MAR and MNAR, the complete cases over-represent low values of y, and the complete-case mean is noticeably negative. Now fit the regression of y on x among the complete cases under each mechanism. Under MCAR and MAR, the slope is close to 0.6, because missingness depends only on the covariate. Under MNAR, the slope is attenuated, because selection on the outcome distorts the conditional distribution. Finally, impute y from x under MAR, for instance with mice, and the mean of y is recovered, because the imputation model contains x, which explains the missingness. Under MNAR, imputation from x alone does not recover it: the imputation model believes that people missing y have the same conditional distribution of y given x as those observed, which is false. That last point bears restating. Imputation and likelihood methods are not magic; they recover the truth under MAR because the observed data, conditional on the right variables, are representative of the missing data. When that conditional representativeness fails, they fail too. Pattern-mixture factorisation The selection model is not the only way to factorise the joint distribution. The alternative is to condition in the opposite direction: f(Y, R) = f(Y | R) * f(R) This is the pattern-mixture factorisation. It describes the data distribution separately within each missingness pattern and then mixes the patterns according to their frequencies. In a trial with dropout, it asks: what is the distribution of outcomes among completers, and what is the distribution among dropouts? The first is estimable from data; the second is not, because the dropouts' outcomes are missing. The pattern-mixture view makes the identification problem vivid. Something must be assumed about f(Y_mis | Y_obs, R = dropout), and MAR is one such assumption, namely that the conditional distribution of the missing values given the observed ones is the same in the dropouts as in the completers. Other assumptions, such as "dropouts' outcomes are on average five units worse than those of comparable completers", are equally expressible and are the basis of the sensitivity analyses in Chapter 9. The two factorisations are complementary. The selection model is natural for thinking about why data go missing and for stating MAR. The pattern-mixture model is natural for stating departures from MAR in terms clinicians and subject experts can evaluate, because it specifies how the unobserved values differ rather than how the probability of being observed depends on values nobody has seen. Reasoning about mechanisms in real data Because the mechanism cannot be read off the data, it has to be reasoned about from knowledge of how the data came to be collected. The most productive question is not "which of the three categories applies?" but "what process decided, for each case, whether this value was recorded, and what did that process know?" If the process knew only things that are also in the data set, the data are MAR with respect to those things. If it knew something that is not in the data set and that is related to the missing value, they are not. Consider attrition in a hypothetical twelve-month trial of an antidepressant, with depression scores recorded monthly. Participants may leave because they feel better and see no reason to continue, because they feel worse and lose faith in the treatment, because of side effects, or because they move house. If the decision to leave at month five is driven by how the participant felt at months one to four, and those scores are recorded, dropout is MAR given the observed history. If it is driven by a sudden deterioration in the weeks after the month-four visit, which was never measured, it is MNAR. The truth is usually a mixture. What makes MAR more plausible here is frequent measurement: the closer the last observation to the moment of dropout, the more of the information the dropout decision used is in the data. What makes it less plausible is a long gap between visits and reasons for dropout that are tied to the outcome itself. Electronic health records present a harder case. A laboratory test appears in a patient's record because a clinician ordered it, and clinicians order tests when they suspect an abnormality. The absence of a potassium measurement is therefore informative: it suggests the clinician saw no reason to check. Some of the information behind that decision is recorded, in diagnoses, vital signs and prior results, but much of it, the clinician's impression at the bedside, is not. Missingness in routinely collected data is often driven by exactly the kind of unrecorded judgement that makes MAR doubtful, and analyses of such data need to be especially careful about both auxiliary information and sensitivity analysis. Measurement devices produce yet another kind of process. A sensor that saturates at a maximum reading, or an assay with a lower limit of quantification, produces missing values determined entirely by the unobserved value. These are MNAR in the strictest sense, but they are also censored, and the censoring point is known, which turns an intractable problem into a tractable one: the value is known to lie beyond a threshold, and models for censored data can use that. Survey income questions blend several processes: some respondents refuse because the amount is high and they are cautious, some because it is low and they are embarrassed, some because they genuinely do not know a household total. Each process has a different relation to the missing value. Two general lessons follow. First, a single variable can have missing values generated by several mechanisms at once, and the overall behaviour is dominated by whichever component is related to the missing value. Second, the analyst's best tool for moving toward MAR is not statistical but practical: recording the reasons for missingness, measuring predictors of non-response, and capturing information as close as possible to the moment data went missing. A study designed with missing data in mind can make MAR far more credible than any amount of post hoc modelling. The taxonomy at a glance Table 2 summarises the three mechanisms, the formal condition each imposes, what it implies for the methods discussed in this book, and whether the observed data can test it. Table 2. Rubin's missingness mechanisms compared. Mechanism Formal condition Example Complete-case analysis MI and FIML Testable from data? MCAR P(R given Y) = P(R) Samples lost in transit Unbiased, inefficient Valid Partly: can be rejected MAR P(R given Y) = P(R given Y_obs) Older people skip weight item, age recorded Often biased Valid if model includes predictors of missingness No MNAR Depends on Y_mis given Y_obs Heavier people skip weight item Often biased Biased without extra assumptions No The last column is the uncomfortable one, and it is the subject of Chapter 3. The observed data can sometimes show that MCAR is false, by showing that missingness is related to observed variables. They can never show that MAR is true, because the evidence needed, the relationship between missingness and the missing values, is precisely what is missing. Everything the analyst does to handle incomplete data rests, ultimately, on a judgement about the plausibility of an untestable assumption. The discipline lies in making that judgement well-informed, stating it clearly, and testing how much depends on it. Chapter 3: Diagnosing the Mechanism Having defined the mechanisms, the natural next question is how to tell which one applies. The honest answer is that the observed data can tell you some things and not others, and it is essential to know which is which. They can reveal that missingness is related to observed variables, which rules out MCAR and identifies variables that must be in the model. They can reveal which observed variables predict the incomplete ones, which identifies useful auxiliaries. They cannot reveal whether, after conditioning on everything observed, missingness is related to the missing values themselves. That boundary is not a limitation of current methods that better statistics will one day overcome. It is a logical consequence of the values being missing. Comparing the observed and the missing The simplest diagnostic compares cases with and without a missing value on some variable in terms of their other, observed characteristics. If respondents missing income are older, less educated and more likely to live alone than respondents who reported it, the income data are not MCAR, and age, education and household composition are candidates for the imputation model. Such comparisons are best reported as standardised mean differences rather than significance tests, for two reasons. In large samples trivial differences are significant; in small samples important differences are not. And the purpose is not to test a hypothesis but to identify variables that carry information about missingness, for which the magnitude of the difference is what matters. A more complete version models the missingness indicator directly. For each incomplete variable, fit a logistic regression of its response indicator on the other variables, using cases where those variables are observed: dat$r_income <- as.integer(!is.na(dat$income)) fit_r <- glm(r_income ~ age + educ + sex + health + region, family = binomial, data = dat) summary(fit_r) Predictors with substantial coefficients are associated with missingness and belong in the imputation or estimation model, whether or not they appear in the substantive analysis. The discrimination of the missingness model, measured by the area under the receiver operating characteristic curve, gives a rough sense of how systematic missingness is with respect to observed data. An area near 0.5 suggests missingness is weakly related to what was observed. An area of 0.8 suggests strong selection on observed variables, which rules out MCAR decisively and means that complete-case analysis is likely to be biased for any estimand related to those predictors. Flexible methods such as classification trees or random forests can be used in place of logistic regression to detect interactions and nonlinear relationships, since the goal here is description, not inference. Two cautions apply. First, the predictors in a missingness model are themselves often incomplete, so the model is fitted on a subset. That is acceptable for an exploratory diagnostic, but its results should be read as suggestive. Second, a variable that predicts missingness is useful for the imputation model only if it is also related to the incomplete variable itself. A variable that predicts whether income is reported but is unrelated to income cannot reduce bias in estimates involving income. The most valuable auxiliary variables predict both missingness and the missing values. Collins, Schafer and Kam's simulation study found that auxiliaries correlated with the incomplete variable are useful even when they do not predict missingness, because they add precision, while auxiliaries that predict only missingness contribute little. Little's test and its relatives In 1988 Roderick Little published in the Journal of the American Statistical Association a test of the null hypothesis that multivariate data are MCAR. The idea is intuitive. If data are MCAR, the mean of each variable should be the same across all missingness patterns, apart from sampling error. The test groups cases by pattern and compares the observed means within each pattern with the overall means estimated by maximum likelihood. With J patterns, where pattern j has n_j cases and observes a subset of p_j variables, the statistic is d2 = sum over j of n_j (ybar_j - mu_j)' inverse(Sigma_j) * (ybar_j - mu_j) where ybar_j is the vector of observed means in pattern j, and mu_j and Sigma_j are the corresponding sub-vector and sub-matrix of the maximum likelihood estimates of the overall mean and covariance, usually obtained by the EM algorithm. Under MCAR and multivariate normality, d2 has approximately a chi-square distribution with sum(p_j) - p degrees of freedom. In R the test is available, for example, as mcar_test() in the naniar package. The test has real uses and serious limitations. It can reject MCAR, and a rejection is informative: it tells the analyst that the complete cases differ systematically from the others, and that deletion will probably bias some estimates. A failure to reject, however, is weak evidence for MCAR. The test has low power when patterns are small, it examines only means and would miss mechanisms that affect variances or relationships, it relies on multivariate normality, and, most fundamentally, it can only detect relationships between missingness and observed values. An MNAR mechanism in which missingness depends only on the missing value itself, and the missing value is unrelated to the observed variables, is invisible to it. Jamshidian and Jalal proposed in 2010, in Psychometrika, a complementary test that examines the homogeneity of covariances across patterns, implemented in the MissMech package, which catches some mechanisms Little's test misses, but it faces the same fundamental limit. In practice, the decision whether to use principled methods should almost never hinge on an MCAR test. Even under MCAR, multiple imputation and FIML are more efficient than deletion because they use the partially observed cases. The test is best treated as one descriptive tool among several, not as a gate. Why MAR cannot be tested It is worth understanding precisely why MAR cannot be verified, because the point is routinely misunderstood. Suppose a variable Y is missing for some cases and fully observed covariates X are available. The observed data consist of the joint distribution of X and R, and the distribution of Y given X among cases with R = 1. They contain no information whatever about the distribution of Y given X among cases with R = 0. MAR asserts that this unobserved conditional distribution is the same as the observed one. An MNAR model asserts that it differs in some specified way. Both are consistent with every observable feature of the data, because they agree on everything that can be observed and differ only on what cannot. Molenberghs, Beunckens, Sotto and Kenward made this formal in a 2008 paper in the Journal of the Royal Statistical Society, Series B, titled "Every missingness not at random model has a missingness at random counterpart with equal fit". For any MNAR model fitted to incomplete data, there is an MAR model that produces exactly the same fit to the observed data but different predictions for the unobserved data. Goodness of fit to observed data therefore cannot adjudicate between MAR and MNAR. When an MNAR selection model appears to "estimate" the dependence of missingness on the missing value, as the classic Heckman selection model does, the estimate is driven not by information about the missing values but by the distributional assumptions of the model, such as bivariate normality of the errors, or by exclusion restrictions that the analyst has imposed. Change those assumptions and the estimate changes, often dramatically. The practical consequence is that MAR is a working assumption justified by argument, not a hypothesis confirmed by data. The argument draws on knowledge of the data collection process, as in the examples of Chapter 2, and on the richness of the observed information. When the variables that plausibly drive missingness have been measured and are included in the model, MAR is a reasonable approximation. When the obvious drivers are unmeasured, it is not, and the analysis should be framed as conditional on an assumption known to be doubtful, with sensitivity analysis carrying more of the weight. Should the missingness indicator itself be a predictor? A question that follows naturally from modelling the response indicator is whether that indicator should enter the imputation or analysis model. If missingness on one variable predicts the values of another, perhaps because people who decline the income question also tend to have different health, then the indicator carries information and could help impute the other variable. Including indicators of missingness on other variables as predictors in an imputation model is legitimate and occasionally useful: they are fully observed, and if they are related to the target variable they sharpen its imputations. What is not legitimate is to use a variable's own missingness indicator to impute that variable under MAR, since within the cases being imputed it is constant and carries no information, and any attempt to make it do so is in effect an MNAR model in disguise. Chapter 9 shows how the indicator can be used deliberately, with a stated sensitivity parameter, for exactly that purpose. Including missingness indicators in the substantive analysis model is a different matter, and generally a mistake for inference, for the reasons Chapter 1 gave in discussing the missing-indicator method. The diagnostic use of indicators, to discover what predicts missingness and to understand the structure of the problem, is where their value lies. Evidence from outside the data set Although the incomplete data cannot test MAR, other sources of information sometimes can shed light on it. Several are worth pursuing when the stakes justify the effort. Follow-up of non-respondents. In a double-sampling design, a random subsample of the people who initially failed to respond is pursued intensively, perhaps with home visits or incentives, until most of them provide data. Because the subsample is random, its values are representative of the non-respondents, and they can be compared directly with what an MAR model predicted. This is expensive but decisive, and it is used in some national surveys for exactly this reason. Continuum of resistance. Survey methodologists sometimes compare early respondents with late respondents who needed several reminders, on the theory that late respondents resemble non-respondents. The theory is not guaranteed, but systematic differences between early and late responders on the incomplete variable, after conditioning on covariates, are a warning sign of MNAR. Linkage to external records. If some participants can be linked to administrative data, such as hospital records, tax records or death registrations, the linked values may reveal how people with missing study data differ from those without. In longitudinal cohorts, linkage of dropouts to health records has been used to check whether dropout is related to health outcomes beyond what the study's own measurements predict. Recorded reasons for missingness. Trials that record why each participant discontinued, whether adverse event, lack of efficacy, relocation or withdrawal of consent, allow the analyst to argue that some missingness is plausibly MAR and some is not, and to tailor sensitivity analyses to each reason. Benchmarks. Where population totals or distributions are known from a census or register, comparing imputed or weighted estimates with the benchmark can reveal residual bias, though this confounds non-response with other sources of error such as coverage and measurement. None of these sources proves MAR for the whole data set, but each can move the argument from bare assertion to a position supported by evidence. A worked diagnostic example Suppose a hypothetical school-based cohort measures adolescent wellbeing on a validated scale at three annual waves, alongside sex, family socioeconomic status, prior academic attainment, a bullying questionnaire and school attendance records drawn from administrative data. The wave-three wellbeing score, the primary outcome, is missing for about a quarter of students, mostly because they were absent on the day of data collection or had moved school. The pattern table shows that missingness is nearly monotone: almost every student missing wave two is also missing wave three, and a handful are missing wave three only. Coverage between wave-one wellbeing and wave-three wellbeing is therefore about three quarters of the sample, which is adequate for estimating their correlation. Attendance records, being administrative, are complete for everyone still enrolled in the region. Comparing students with and without a wave-three score on their observed characteristics shows sizeable standardised differences on attendance, prior attainment and wave-one wellbeing, and smaller ones on socioeconomic status. The missingness model, a logistic regression of the wave-three indicator on these variables, confirms that low attendance is the strongest predictor, which is unsurprising since absence on the survey day is the main route to missingness. MCAR is plainly false, and Little's test, run for completeness, rejects it decisively. The substantive argument then runs as follows. Missingness is driven mainly by absence and school moves. Attendance is recorded for everyone, and wave-one and wave-two wellbeing are recorded for most. If, among students with the same attendance, prior attainment and earlier wellbeing, those who were absent on the survey day have the same distribution of wellbeing as those who were present, the data are MAR given these variables. That is plausible for school moves, which are mostly driven by family circumstances largely captured by socioeconomic status. It is less plausible for absence, because students with worsening wellbeing may be more likely to be absent, and the decline since wave two is not observed. The direction of the likely departure is clear: missing students are probably somewhat worse off than MAR would predict. This reasoning dictates the plan. Attendance must be in the imputation model even though it is not in the substantive analysis, because it is the main driver of missingness and is correlated with wellbeing. Earlier waves of wellbeing must be included for the same reason. The primary analysis will assume MAR given these variables. The sensitivity analysis will ask how much worse the missing students' wave-three wellbeing would need to be, relative to MAR predictions, to change the study's conclusions, and whether such a difference is plausible given what teachers and pastoral staff know about persistently absent students. None of this required a test of MAR. It required knowing why the data were missing, measuring the drivers, and being explicit about the likely direction of error. A diagnostic workflow The tools of this chapter serve different purposes, and it helps to be explicit about what each can and cannot establish. Table 3 summarises them. Table 3. Diagnostic tools for the missingness mechanism. Tool What it examines What it can show What it cannot show Pattern tables and coverage Structure of missingness Where information is thin Anything about mechanism Comparisons of observed characteristics Differences by missingness status That data are not MCAR Whether data are MAR Logistic model for missingness indicator Predictors of missingness Variables needed in the model Dependence on missing values Little's MCAR test Mean differences across patterns Evidence against MCAR That data are MCAR or MAR Non-respondent follow-up or linkage External values for non-respondents Direct evidence on MNAR Coverage beyond the linked subset A sensible sequence for most projects runs as follows. Describe the patterns, proportions, coverage and distribution of missing values per case, and classify blanks as missing, not applicable, censored or substantive. For each important incomplete variable, compare the observed characteristics of cases with and without the value, and fit a missingness model to identify predictors of missingness. Identify variables correlated with each incomplete variable, using the observed data, as candidate auxiliaries. Collect whatever external evidence is available on the reasons for missingness and the characteristics of non-respondents. Write down, in plain language, the argument for why MAR is or is not a reasonable working assumption given the variables available, and which departures from MAR are plausible and in which direction. Carry that written argument forward into the imputation model, the analysis, and the sensitivity analysis. The fifth step is the one most often omitted, and it is the most important. A clearly stated argument about the mechanism, even an informal one, disciplines every later decision. It tells the analyst which auxiliaries are indispensable, which parts of the analysis are most exposed to MNAR bias, and what a sensitivity analysis should explore. It also tells readers what they are being asked to believe. With that argument in hand, the analyst is ready to choose a method, and the first of the two principled routes, multiple imputation, is the subject of the next three chapters. Hashtags: #MissingDataMechanics #MissingData #MultipleImputation #MICE #FullInformationMaximumLikelihood #FIML #MissingCompletelyAtRandom #MissingAtRandom #MissingNotAtRandom #RubinRules #ChainedEquations #PredictiveMeanMatching #ImputationModels #AuxiliaryVariables #Congeniality #CompleteCaseAnalysis #LittleMCARTest #MissingnessDiagnostics #PatternMixtureModels #SelectionModels #DeltaAdjustment #SensitivityAnalysis #TippingPointAnalysis #Lavaan #FutureOfMissingDataAnalysis
- Microbiome Informatics (Analyzing 16S rRNA and Shotgun Metagenomic Data)
Download the Book (PDF): Introduction A gram of human stool contains somewhere in the region of a hundred billion bacterial cells. A teaspoon of garden soil holds thousands of species, most of which have never been grown in a laboratory and never will be. For most of the history of microbiology, these communities were effectively invisible. You could culture what would grow on a plate, stain what you could see under a microscope, and guess at the rest. The guess was badly wrong. When researchers began amplifying ribosomal RNA genes directly from environmental samples in the 1980s and 1990s, they discovered that the organisms which grew in culture were a small and unrepresentative fraction of what was actually there. Today the invisible has become measurable, at least in a particular sense. A single run on a modern sequencer can generate tens of millions of reads from hundreds of samples, and a laptop can turn those reads into a table: samples down one side, microbial features across the top, numbers in every cell. From that table come the results that fill journals and newspapers — the gut bacteria that differ between people with and without a disease, the diversity that falls after a course of antibiotics, the metabolic pathways enriched in one soil and depleted in another. This booklet is about the journey from reads to that table, and from that table to claims about biology. Its controlling idea is simple to state and surprisingly easy to forget: a microbiome dataset is not a census of a community but a distorted, relative, and heavily processed sample of it, and every number in the final analysis inherits the choices made along the way. Good microbiome informatics is the discipline of knowing what those choices are, what each one does to the data, and which conclusions survive them. Why the table deceives Consider what happens between a person's gut and a spreadsheet. A sample is collected, stored at some temperature for some time, and has its DNA extracted by a kit that breaks open some cells more readily than others. A short stretch of one gene is amplified by the polymerase chain reaction using primers that match some taxa better than others. The amplified fragments are sequenced by an instrument that makes errors at a characteristic rate. The machine's output is filtered, merged, denoised or clustered, and compared with a reference database that knows some lineages intimately and others not at all. The resulting counts are then normalized, transformed, summarized into diversity indices, and fed into statistical tests. At every one of those steps, the relationship between what was in the sample and what appears in the table is bent a little. Some of the bending is random and averages out. Much of it is systematic: the same organism is under-counted in every sample processed with the same protocol. And one distortion is structural and universal: because the sequencer produces a roughly fixed number of reads per sample regardless of how much biomass went in, the counts carry information only about proportions. If one taxon doubles in absolute abundance, every other taxon appears to shrink. The table cannot, on its own, tell you which happened. None of this makes microbiome data useless. It makes it a particular kind of data, with particular rules, and much of the confusion in the field's literature comes from analyzing it as though it were something else — a set of independent counts, for instance, or a direct measurement of cell numbers. Two ways of reading a community There are two dominant strategies for sequencing a microbial community, and this booklet treats both. The first is amplicon sequencing, most commonly of the 16S ribosomal RNA gene in bacteria and archaea. Every bacterium carries this gene. It contains stretches that are nearly identical across the whole domain, which make convenient landing sites for PCR primers, interleaved with variable regions whose sequence differs between lineages. Amplify a variable region, sequence the products, and the differences between sequences become a proxy for the differences between organisms. Amplicon sequencing is cheap, it works on samples with little DNA or a great deal of host contamination, and decades of accumulated reference sequences make it possible to put names to most of what comes out. Its limits are equally clear: it reads one gene, it resolves organisms only as far as that gene's variation allows, and it says nothing directly about what the organisms can do. The second is shotgun metagenomics, in which all the DNA in a sample is fragmented and sequenced without targeted amplification. The reads come from every genome present — bacterial, archaeal, viral, fungal, and host. They can be matched to reference genomes to estimate which organisms are present, assembled into longer contigs and sometimes into near-complete genomes, and mapped to catalogues of genes and pathways to estimate what the community is capable of. Shotgun data cost more per sample, need more computation, and are hampered by host DNA in some sample types, but they avoid PCR amplification bias and open questions that a single marker gene cannot answer. The two approaches share more than it first appears. Both produce a sparse, compositional table of features by samples. Both depend heavily on reference databases. Both feed the same downstream ecology: the same diversity indices, the same distance measures, the same statistical frameworks. So the booklet treats them as two front ends to a shared analytical core. What this booklet covers The chapters follow the data. The first examines what happens before any computation — sample handling, DNA extraction, primer choice, sequencing platforms, and the contamination and controls that decide whether a dataset can be trusted at all. The next two cover how amplicon reads become features: first the older approach of clustering sequences into operational taxonomic units at a similarity threshold, then the newer approach of inferring exact amplicon sequence variants by modelling sequencing error. A chapter on reference databases and taxonomic classification follows, since naming organisms is where many analyses quietly go wrong. Two chapters then turn to shotgun data: one on estimating which organisms are present, by read classification, marker genes, and genome assembly; one on functional profiling, from gene families to metabolic pathways, including the inference of function from 16S data alone. The last four chapters form the analytical core common to both data types. One confronts the compositional nature of sequencing counts and the long argument over how to normalize them. One covers alpha diversity, the richness and evenness of individual samples. One covers beta diversity, the dissimilarity between samples, with the ordinations and permutation tests that usually accompany it. The final chapter addresses differential abundance testing — the attempt to say which specific organisms differ between groups — and closes with what reproducible reporting now requires. Who this is for The reader this booklet has in mind is someone who has, or soon will have, microbiome data in hand: a graduate student in microbiology or medicine, a bioinformatician moving into the field from genomics or statistics, a clinician-scientist whose trial includes stool samples, an ecologist whose soil cores are about to come back from the sequencing core. It assumes comfort with the idea of DNA sequencing and basic statistics, but not prior experience with microbiome software. It is not a software manual. Tools change quickly, command-line options change faster, and any tutorial written today will be partly stale within two years. What changes slowly is the reasoning: why a denoising algorithm needs an error model, why a closed-reference pipeline discards novel organisms, why Bray–Curtis and UniFrac can disagree about the same pair of samples, why a significant PERMANOVA result may say nothing about the difference in community centroids. Tools are named where naming them helps, because readers will meet them — DADA2, QIIME 2, Kraken, MetaPhlAn, HUMAnN, and the rest — but the aim is that a reader who understands the reasoning can pick up whatever tool is current and use it well. A word on humility Microbiome science has had an uneven decade. Early, striking associations — between particular bacterial ratios and obesity, for example — failed to replicate as cohorts grew and methods improved. Some of that was simply the ordinary correction of small, noisy early studies. But some of it was methodological: differences in extraction kits, primer sets, and bioinformatics pipelines were large enough to generate apparent biological differences that were in fact technical. Large coordinated efforts such as the Microbiome Quality Control project showed that the same samples, sent to different laboratories, could come back looking substantially different. The lesson is not that microbiome results are unreliable. Many are robust and important: the profound disruption of gut communities by broad-spectrum antibiotics, the effectiveness of faecal microbiota transplantation against recurrent Clostridioides difficile infection, the reproducible gradients of soil communities along pH. The lesson is that reliable results come from analysts who know where the distortions enter and who design, process, and interpret with them in mind. That knowledge is what the following chapters try to supply. Chapter 1: From Sample to Reads The most consequential decisions in a microbiome study are usually made before anyone opens a terminal. By the time the sequencing files arrive, the choice of sample type, storage method, extraction kit, primer pair, and sequencing platform has already fixed what the data can and cannot show. No algorithm can recover a taxon whose cells were never lysed, correct a primer that does not bind to a whole phylum, or distinguish a genuine low-abundance organism from a contaminant that entered through a reagent. This chapter follows a sample from collection to raw reads and identifies where the systematic distortions enter, because an analyst who does not know where they came from cannot reason about what they have done. Collection and storage Microbial communities do not stop living when they leave their host. A stool sample left at room temperature continues to change: facultative anaerobes and fast-growing organisms can bloom within hours, while others decline. The practical standard for many human studies has been immediate freezing at −80 °C, but that is impractical for home collection or field sampling, so a range of stabilizing buffers and preservation cards has been developed. Comparisons of these methods generally find that the differences they introduce are smaller than the differences between individuals, but not negligible, and — this is the key point — they are systematic. If cases are collected in a clinic and frozen immediately while controls mail samples from home, storage method is confounded with disease status, and no amount of downstream statistics will separate them. The same logic applies to every other pre-analytical variable. Time of day, time since the last meal, the part of a stool sample that was scooped, whether a swab touched skin before the mucosa: each introduces variation. The analyst's defence is design rather than computation. Record every such variable as metadata, keep it balanced across the comparison of interest, and if it cannot be balanced, at least make sure it can be modelled. DNA extraction Extraction is where the first large, systematic bias enters. Microbial cell walls differ enormously in how easily they break. Gram-negative bacteria, with their thin peptidoglycan layer and outer membrane, lyse readily. Gram-positive bacteria, with thick peptidoglycan, and some archaea and spores, are much tougher. An extraction protocol that relies on chemical or enzymatic lysis alone will under-represent the tough organisms; one that includes vigorous mechanical disruption — bead-beating — releases them more completely but may shear DNA from the easy ones. The size of this effect is not trivial. The International Human Microbiome Standards project, reported by Costea and colleagues in 2017, compared many extraction protocols on the same faecal material and found substantial protocol-dependent differences in the apparent composition, notably in the Gram-positive phylum then called Firmicutes (now Bacillota). Their recommendation of a standardized protocol with bead-beating reflects a broad consensus that mechanical lysis is necessary for representative results from stool. The companion Microbiome Quality Control project, reported by Sinha and colleagues in the same year, sent identical samples to many laboratories and found that the laboratory — its extraction, amplification, and bioinformatics — was a major source of variation in the resulting profiles. Two practical consequences follow. First, samples that will be compared must be extracted with the same protocol, ideally in randomized batches so that extraction run is not confounded with any variable of interest. Second, cross-study comparisons of relative abundance should be approached with great caution: a taxon that appears at 20 per cent in one study and 8 per cent in another may simply have been extracted differently. Choosing a marker and its primers For amplicon studies, the next decision is what to amplify. The 16S rRNA gene of bacteria and archaea is roughly 1,500 base pairs long and contains nine hypervariable regions, labelled V1 to V9, separated by conserved stretches. Short-read sequencers cannot read the whole gene in a single fragment, so most studies target one or two adjacent variable regions — V4 alone, V3–V4, or V1–V3 are common choices. No region is neutral. Different regions resolve different lineages with different efficiency: some genera are distinguishable in V1–V3 but nearly identical in V4, and vice versa. More importantly, the primers themselves have mismatches with some lineages. A primer that has one or two mismatches with a group's 16S sequence near its 3′ end will amplify that group poorly, and the group will be under-counted or missed. The widely used 515F/806R primer pair for V4, promoted by the Earth Microbiome Project, was revised in the mid-2010s precisely because the original version under-amplified certain marine and freshwater lineages, including the abundant SAR11 clade. Any published dataset carries the fingerprints of its primers. For fungi, the internal transcribed spacer (ITS) region is the usual marker; for microbial eukaryotes, the 18S rRNA gene. ITS in particular varies greatly in length between taxa, which interacts with the length limits of sequencing and with PCR, favouring organisms with shorter amplicons. The general lesson travels across markers: the choice of target is a choice about which organisms will be visible. A further complication is that the 16S gene is present in multiple copies in many bacterial genomes. Copy number ranges from one to more than a dozen, and the copies within a genome are not always identical. An organism with ten copies contributes roughly ten times as many amplicons per cell as one with a single copy, so 16S relative abundances are biased towards high-copy taxa. Correction methods based on predicted copy number exist, but their accuracy depends on how well the organism's copy number can be inferred from its relatives, and evaluations have been mixed. Many analysts leave counts uncorrected and simply remember that amplicon abundance is not cell abundance. PCR and its artefacts Amplification multiplies the target fragments over twenty-five to thirty-five cycles, and it does not multiply them evenly. Differences in primer binding, GC content, and amplicon length mean some templates are copied more efficiently than others, and those differences compound with each cycle. More cycles generally mean more distortion, which is why protocols try to use as few as the input DNA allows. PCR also creates sequences that were never in the sample. The most important are chimeras: when an extension is incomplete, a partially copied fragment can anneal to a different template in the next cycle and be extended into a hybrid whose first half comes from one organism and second half from another. Chimeras can make up a substantial fraction of raw amplicon reads in some protocols, and because they look like plausible novel 16S sequences, they inflate apparent diversity unless detected and removed. Polymerase errors add point mutations as well, generating a halo of rare variants around each true sequence. Both artefacts are central to the processing choices in the next two chapters. Shotgun library preparation Shotgun metagenomics avoids targeted PCR of a marker gene, but it is not free of bias. DNA is fragmented, by sonication or enzymes, and adapters are ligated or inserted, often by a transposase-based method. Transposases have mild sequence preferences, and many library protocols include a short amplification step that favours fragments of moderate GC content. Very high or very low GC genomes can be under-represented as a result. The dominant practical problem in shotgun work, though, is host DNA. Stool is mostly microbial, so a human stool metagenome typically contains only a small percentage of human reads. A skin swab, a biopsy, a saliva sample, or a blood sample can be overwhelmingly host: in tissue biopsies, host reads can exceed 99 per cent, leaving too few microbial reads for meaningful profiling at an ordinary sequencing depth. Host-depletion methods, which selectively lyse host cells or remove methylated DNA before sequencing, can help, though they introduce biases of their own. For such low-microbial-biomass samples, 16S amplicon sequencing, which amplifies only bacterial targets, is often the only affordable option. Sequencing platforms Most microbiome data to date have come from Illumina short-read instruments. For amplicons, the MiSeq has been a workhorse because it produces paired-end reads of up to 300 bases each, long enough that the forward and reverse reads of a V3–V4 amplicon overlap and can be merged into one sequence. Illumina errors are predominantly substitutions, with quality declining towards the ends of reads — especially the reverse read. Low-diversity libraries, which amplicon libraries are because every molecule begins with the same primer sequence, can cause problems for Illumina's cluster calibration; the usual remedy is to spike in a proportion of the PhiX control library or to stagger the primers. Long-read platforms — Pacific Biosciences circular consensus sequencing and Oxford Nanopore — can read the entire 16S gene, or even the full ribosomal operon including the 23S gene and the spacer between them. Full-length 16S improves taxonomic resolution considerably; an analysis by Johnson and colleagues in 2019 showed that full-length sequences can often resolve organisms to species or even strain level where single variable regions cannot, partly because intragenomic differences between 16S copies become visible. PacBio's circular consensus reads now reach very high accuracy; Nanopore accuracy has improved rapidly with new chemistries and basecallers, though its error profile, historically rich in insertions and deletions in homopolymer runs, still needs careful treatment. For shotgun metagenomics, long reads greatly improve assembly, producing complete circular genomes from complex communities that short reads leave fragmented. Sequencing depth — the number of reads per sample — should be chosen for the question. For 16S profiling of stool, tens of thousands of reads per sample usually capture the community's dominant structure, with diminishing returns beyond that. For shotgun profiling of species composition, a few million reads per sample is often adequate; for detecting low-abundance organisms, recovering genomes by assembly, or profiling function in detail, tens of millions may be needed. Depth also varies, sometimes by an order of magnitude, between samples in the same run, which is one reason normalization gets a chapter of its own. Contamination and why controls are not optional Every extraction kit, every reagent, and every laboratory surface carries a small amount of bacterial DNA. In a stool sample with abundant microbial DNA, this background is swamped and usually harmless. In a sample with very little microbial DNA — placenta, blood, lung, some tissue biopsies, cleanroom surfaces — the contaminants can make up most or all of what is sequenced. This is not a hypothetical. Salter and colleagues showed in 2014 that serial dilutions of a pure bacterial culture, processed with common kits, produced sequencing profiles increasingly dominated by contaminant taxa as the input decreased, and that the contaminant profile depended on the kit lot. The so-called "kitome" includes genera such as Ralstonia, Bradyrhizobium, Burkholderia, Pseudomonas, and Sphingomonas, among others — organisms that have appeared in many published reports of microbiomes from supposedly sterile or low-biomass sites. The long debate over whether the healthy human placenta has its own microbiome was driven largely by this problem; later studies with rigorous controls found that the signals in most placental samples were indistinguishable from contamination. The response is to sequence controls alongside the samples and to use them. · Extraction blanks — reagent-only samples carried through extraction, amplification, and sequencing — reveal the kit and laboratory background. · PCR no-template controls reveal contamination introduced at amplification. · Mock communities — mixtures of known organisms in known proportions, available commercially as cells or DNA — reveal how the entire pipeline distorts a known truth: which organisms it under-counts, which spurious features it creates, how far observed proportions stray from expected ones. · Positive controls and replicates of real samples reveal run-to-run consistency. Controls are useful only if the analysis consults them. Statistical tools such as the `decontam` package, described by Davis and colleagues in 2018, identify likely contaminants by two signatures: contaminant features tend to have higher relative abundance in samples with lower DNA concentration, because the fixed background makes up a larger share of low-input libraries, and they tend to be more prevalent in negative controls than in true samples. Neither signature is infallible, and simply subtracting everything seen in blanks is too crude — cross-contamination from real samples into blanks is common, so blanks often contain genuine sample organisms. But a study of a low-biomass environment that does not report its negative controls, and show how they were used, should not be believed. Cross-talk and index hopping A subtler form of contamination arises within the sequencing run. Samples are multiplexed by attaching short index sequences, and the instrument assigns each read to a sample by reading its index. Errors in index reading, and the phenomenon of index hopping on some patterned flow cells, can assign a small fraction of reads to the wrong sample. The effect is usually well below one per cent of reads, but it means that an abundant organism in one sample can appear at low levels in every other sample on the run. Unique dual indexing, in which each sample receives a unique combination of two indices, reduces the problem sharply. Where it cannot be eliminated, it argues for scepticism about very low-abundance detections, especially of taxa that are abundant elsewhere on the same run. Metadata as data The final pre-analytical element is the least glamorous and among the most important: the sample metadata. Every downstream analysis — every diversity comparison, every model, every test — depends on a sample table that records who or what each sample came from and under what conditions. Errors in this table, such as swapped sample labels, are among the most common sources of irreproducible results, and they are invisible in the sequence data. Community standards help. The Genomic Standards Consortium's MIxS checklists specify the minimum information to record about a sequenced sample from a given environment, and public repositories such as the Sequence Read Archive and the European Nucleotide Archive expect such metadata on deposit. For human studies, the variables most often found to associate with gut community composition — among them medication use (especially antibiotics, proton-pump inhibitors, and metformin), stool consistency, diet, age, and body mass index — should be recorded as a matter of course, because they are the confounders that a later reviewer will ask about. Designing for the comparison All of these distortions share one property that makes them manageable: they are largely consistent within a batch. An extraction kit that under-counts a tough-walled genus does so in every sample it touches, and a primer that misses a lineage misses it everywhere. For comparisons within a study, a consistent bias partly cancels, because both groups are distorted in the same way. The danger comes when the bias is not shared — when one group's samples were extracted in March with one kit lot and the other group's in September with another, or when cases and controls were sequenced on different runs. The remedy is old-fashioned experimental design. Samples should be allocated to extraction batches, PCR plates, and sequencing runs at random with respect to the variables of interest, or in a balanced block design, and the batch assignment should be recorded so that it can be modelled or checked later. Where a study accrues over years, as clinical cohorts do, it is often better to store samples and process them together than to sequence them as they arrive. Technical replicates of a few samples across batches let the analyst estimate how large batch effects actually are. Sample size deserves the same forethought. Microbiome data are highly variable between individuals — two healthy adults typically share only a fraction of their gut taxa, and the dominant genera can differ several-fold — so modest effects need large samples to detect. Early studies with a dozen subjects per group were badly underpowered for anything but the largest differences, which is one reason so many of their findings failed to replicate. Power calculations for microbiome endpoints are awkward, because the endpoints themselves are multivariate, but pilot data or published datasets from comparable populations can be used to simulate the power of a planned diversity comparison. The uncomfortable conclusion of such exercises is often that a study needs several times as many samples as its budget first allowed. What arrives at the terminal At the end of this process, the analyst receives files of reads — usually one or two FASTQ files per sample, containing sequences and per-base quality scores — together with a metadata table. It is tempting to treat these files as the beginning of the analysis. They are better thought of as the middle. Everything upstream has already shaped them, and the most useful thing an analyst can do before running a single command is to find out, sample by sample, how they were made: which kit, which primers, which run, which controls. That information decides which of the steps in the following chapters matter most, and which comparisons the data can honestly support. Chapter 2: Operational Taxonomic Units and the 97 Per Cent Convention For roughly the first fifteen years of high-throughput amplicon sequencing, the standard way to turn reads into a community table was to group similar sequences into clusters and count the reads in each cluster. The clusters were called operational taxonomic units, or OTUs, a term borrowed from numerical taxonomy, where it had long meant any grouping of organisms defined by an explicit operational rule rather than by a formal species concept. The rule in microbiome work was almost always the same: sequences that were at least 97 per cent identical belonged to the same OTU. The field has largely moved on to exact sequence variants, the subject of the next chapter, but OTUs are worth understanding in detail for three reasons. An enormous body of published work, including landmark projects such as the Human Microbiome Project and the early Earth Microbiome Project, was built on them. Many long-running studies and some specialized applications still use them. And the reasons the field moved away from them illuminate what a good feature definition needs to do. The preprocessing every pipeline shares Before clustering or denoising, amplicon reads pass through a sequence of steps that are broadly common to all pipelines. Demultiplexing assigns each read to its sample by its index sequence. Sequencing facilities usually do this before delivering data, so most analysts receive one pair of files per sample. Primer and adapter removal strips the primer sequences from the start of each read. Primers are not biological information — they were added by the experimenter, and variation within them reflects synthesis errors or degenerate positions rather than organisms — so they should be removed, usually with a tool such as cutadapt. Reads in which the primer cannot be found are often discarded, which usefully removes off-target products. Quality inspection examines the per-base quality scores that the sequencer assigns. These Phred scores encode the estimated probability of a base-calling error: a score of 20 corresponds to one error in a hundred, 30 to one in a thousand, 40 to one in ten thousand. Illumina reads typically start with high quality that declines along the read; the reverse read usually declines faster. Summary plots of quality by position across all reads guide where to truncate. Merging joins overlapping forward and reverse reads into one sequence. Where the two reads overlap, the merger compares them base by base and, where they disagree, keeps the base with higher quality and recalculates the quality score. Merging requires sufficient overlap — commonly at least twenty bases — so the amplicon length and read length together determine whether it can succeed. A V4 amplicon of about 250 bases sequenced with 2 × 250 reads overlaps almost completely; a V3–V4 amplicon of about 460 bases with 2 × 300 reads overlaps by a margin that aggressive truncation can destroy. Truncate too much to remove poor-quality tails and the pairs no longer meet. Quality filtering discards reads likely to contain errors. The most principled approach, introduced by Robert Edgar and Henrik Flyvbjerg in 2015, is to compute each read's expected number of errors — simply the sum of the per-base error probabilities implied by its quality scores — and to discard reads whose expected errors exceed a threshold such as one. Unlike an average quality score, which can hide a few very bad bases in an otherwise good read, expected errors directly estimate what matters: how many mistakes the read probably contains. Dereplication collapses identical sequences into a single representative, recording how many times each was seen. Because amplicon data are highly redundant — the same true sequence appears thousands of times — dereplication shrinks the computational problem dramatically. Why cluster at all After these steps, a typical dataset still contains a vast number of distinct sequences. Some are genuine biological variants. Many are errors: a single true sequence of 250 bases, read ten thousand times at a per-base error rate of a few tenths of a per cent, produces thousands of reads carrying at least one error, scattered across a cloud of single-mismatch variants. PCR adds its own errors and its chimeras. If each distinct sequence were counted as an organism, apparent diversity would be inflated by orders of magnitude. Clustering was the pragmatic solution. If errors are mostly one or two bases away from their true parent, then grouping everything within a few per cent of a central sequence will absorb the errors into the cluster of the organism that generated them. The price is resolution: genuinely distinct organisms whose sequences differ by less than the threshold are also merged. Where 97 per cent came from The 97 per cent figure has a specific history. In 1994, Erko Stackebrandt and Barbara Goebel compared 16S rRNA sequence similarity with DNA–DNA hybridization, which was then the gold-standard molecular criterion for bacterial species: strains sharing at least 70 per cent hybridization were considered the same species. They observed that strains with less than about 97 per cent 16S similarity essentially never reached the 70 per cent hybridization threshold. The 97 per cent figure was therefore a lower bound for species membership — below it, two strains are almost certainly different species — not a definition of a species. Many distinct species share more than 97 per cent 16S identity. Later work refined the number. A 2014 analysis by Yarza and colleagues, using full-length 16S sequences from type strains, estimated that 98.7 per cent identity better corresponds to the species boundary, with around 94.5 per cent for genera. And these figures concern full-length genes. A short variable region such as V4 may be more or less variable than the gene as a whole, so 97 per cent over 250 bases of V4 does not correspond to 97 per cent over 1,500 bases. Robert Edgar returned to the question in 2018 and concluded that for full-length sequences the optimal species-level threshold was around 99 per cent, and that for the V4 region, 100 per cent identity — that is, exact sequences — gave the best correspondence with species. The convention persisted mostly because it was the convention. So an OTU at 97 per cent is best understood as a coarse grouping, somewhere between species and genus depending on the lineage and the region, and not as a stand-in for a species. Reports that describe "species-level OTUs" from short-read 16S data overstate what the method delivers. Three ways to form clusters Once a threshold is chosen, there remains the question of how to form clusters, and the answer turns out to matter a great deal. There are three broad strategies, distinguished by whether they use a reference database. De novo clustering groups the sequences in the dataset with one another, without reference to any external database. Its virtue is that it captures every organism in the sample, including those absent from any database — a real advantage in poorly studied environments such as deep sediments or novel host species. Its weakness is that the resulting OTUs are specific to the dataset. Run the same algorithm on a different set of samples and you get different OTUs, which cannot be directly compared with the first set. De novo clustering is also the most computationally demanding approach. Closed-reference clustering compares each read with a reference database of sequences that have already been clustered at the chosen threshold — for many years, the Greengenes 97 per cent OTU collection was the standard — and assigns each read to the reference OTU it matches. Reads that do not match anything within the threshold are discarded. The advantages are speed, since the work is a search rather than an all-against-all comparison, and comparability: two studies that use the same reference and threshold produce OTUs with the same identifiers, which can be merged directly even if they sequenced different variable regions. The cost is that novel organisms are thrown away, and in environments that are poorly represented in databases, that can be most of the data. Open-reference clustering combines the two. Reads are first matched against a reference, and those that fail to match are then clustered de novo. The result is a table in which well-characterized organisms carry stable reference identifiers while novel ones are still retained. For some years, open-reference clustering was the recommended default in the QIIME pipeline, the most widely used analysis platform of its era. The trade-offs are summarized in Table 1. Table 1. Strategies for forming operational taxonomic units. Strategy Uses reference? Retains novel organisms? Comparable across studies? Main risk De novo No Yes No Dataset-specific units; heavy computation Closed-reference Yes, only No Yes, with same reference Discards unmatched reads Open-reference Yes, then de novo Yes Partly Mixed unit definitions in one table Algorithms and their quirks Within de novo clustering, several algorithms compete, and they produce noticeably different results from the same data. Hierarchical clustering, used by the mothur pipeline developed in Patrick Schloss's laboratory, computes pairwise distances between all unique sequences and builds clusters by linkage rules. Single linkage joins a sequence to a cluster if it is within the threshold of any member, which tends to chain distinct organisms together through intermediate errors. Complete linkage requires a sequence to be within the threshold of every member, producing tight but numerous clusters. Average linkage, which uses the mean distance, was found in benchmarking to give a reasonable balance. Later, mothur adopted the OptiClust algorithm, which iteratively reassigns sequences to maximize a measure of agreement between the clustering and the pairwise distance matrix. Greedy clustering, used by UCLUST, CD-HIT, and the open-source VSEARCH, processes sequences one at a time in a chosen order. Each sequence either joins an existing cluster whose centroid is within the threshold or becomes the centroid of a new one. Greedy algorithms are fast, but their results depend on the order in which sequences are processed. Sorting by abundance, so that the most common sequences become centroids first, is the usual choice and has a sensible logic: an abundant sequence is more likely to be a true biological sequence than an error. UPARSE, published by Edgar in 2013, took this logic further. It processes sequences in order of decreasing abundance, but a sequence becomes a new centroid only if it is not within the threshold of an existing centroid and is not a chimera of two existing, more abundant centroids. This integrated chimera filtering, together with the discarding of singletons, made UPARSE results much closer to the true number of organisms in mock communities than earlier approaches, which had typically produced many times too many OTUs. Swarm, developed by Frédéric Mahé and colleagues, abandoned the global threshold altogether. It grows clusters by linking sequences that differ by a small number of differences — often a single one — and then uses abundance patterns to decide where one cluster ends and the next begins. Because it has no fixed radius, it adapts to the natural structure of the data and avoids the arbitrary splitting of a continuous cloud of variants that a threshold imposes. The existence of this many algorithms, each giving different answers, is itself one of the lessons. The number of OTUs in a dataset is not a property of the community; it is a property of the community together with the algorithm, the threshold, and the preprocessing. Chimeras Chimera detection deserves its own attention because chimeras, left in the data, masquerade as novel organisms. Detection relies on the observation that a chimera can be explained as a combination of two parent sequences: its left portion closely matches one sequence, its right portion another, and it matches neither as well as it matches the combination. Reference-based detection compares each query against a curated database of chimera-free sequences. De novo detection uses the dataset itself, exploiting the fact that a chimera's parents were present in the same PCR and, having been amplified from the start, should be more abundant than the chimera. Edgar's UCHIME algorithm, published in 2011, implemented both modes and became a standard. De novo detection generally performs better for amplicon data, because the parents of a chimera are by definition in the sample, whereas a reference database may lack them. Even so, chimeras between closely related parents are hard to detect, because there are few positions at which the two halves are distinguishable, and some always escape. Singletons and the rare biosphere A decision that looks minor has large consequences: what to do about sequences observed only once. In a 2006 study of deep-sea communities, Sogin and colleagues described a vast "rare biosphere" of low-abundance lineages, drawing on early high-throughput sequence data. Subsequent work showed that much of this apparent rarity was sequencing and PCR error, which produces exactly the long tail of rare variants that a rare biosphere would. Real rare organisms certainly exist, but distinguishing them from error at the single-read level is close to impossible. The common response was to discard singletons, or all OTUs below some small count, before or after clustering. This greatly improves the accuracy of OTU counts in mock communities, at the cost of losing some genuine rare organisms. It also, as later chapters discuss, has consequences for richness estimators that rely on the number of singletons to estimate unseen diversity. Why the field moved on By the mid-2010s, three lines of criticism had accumulated against OTU clustering. The first was arbitrariness. The threshold did not correspond to any biological unit, the results depended on the algorithm and even the order of input, and de novo OTUs could not be compared across studies without reprocessing all the raw data together. The second was lost resolution. Clustering at 97 per cent merges organisms with genuinely different sequences, and some of those organisms differ ecologically or clinically. Pathogens and commensals can share OTUs. Once merged, they cannot be separated. The third was reference dependence. Closed-reference clustering throws away novelty; open-reference clustering produces a mixed table in which some units are defined by a curated reference and others by the dataset. What made it possible to answer all three criticisms at once was a change in how errors were handled. Clustering treated errors indirectly, absorbing them into clusters wide enough to contain them. If instead one could model the error process directly — estimating, from the data themselves, how likely it is that a given rare sequence was produced by error from a given abundant one — then the true sequences could be recovered exactly, with no threshold at all. That is the idea behind amplicon sequence variants, and it is the subject of the next chapter. When OTUs still make sense OTU clustering has not disappeared, and there are defensible reasons to use it. Some long-running monitoring programmes keep OTU pipelines for continuity with their historical data. In markers with high intragenomic variation, such as the fungal ITS region, where a single organism can carry several distinct copies, exact variants may split one organism into many features and some clustering can be useful. For long-read data with residual error rates higher than Illumina's, clustering at a high threshold such as 99 per cent may absorb errors that denoising does not model well. And some analysts cluster exact variants at a threshold afterwards, to obtain coarser units for particular ecological questions. In each of these cases, the defensible approach is to treat clustering as an explicit, reported choice, applied for a stated reason, and to recognize that the resulting units are operational in the fullest sense of the word: defined by the procedure, not by nature. Chapter 3: Amplicon Sequence Variants and the Denoising Turn Around 2016 the standard unit of amplicon analysis changed. Instead of grouping reads into clusters defined by a similarity threshold, the major pipelines began to infer the exact biological sequences present in each sample, distinguishing them from sequencing errors by a statistical model. The inferred sequences became known as amplicon sequence variants, or ASVs; some tools call them exact sequence variants or zero-radius OTUs, but the idea is the same. Within a few years, ASVs had become the default in QIIME 2, the successor to the original QIIME, and in most published amplicon work. The shift was not just technical. It changed what a feature in a microbiome table is. An OTU is a cluster, defined relative to the other sequences in a dataset and a chosen threshold. An ASV is a DNA sequence, defined by itself. That difference has consequences for resolution, for comparability, and for the kinds of mistakes an analyst can make. The core idea: model the errors, not the organisms Clustering treats sequencing errors indirectly, by drawing a boundary wide enough that most errors fall inside the cluster of their parent sequence. Denoising treats them directly. It asks, of every distinct sequence in the data: is this sequence more abundant than it would be if it had been produced by errors from some other, more abundant sequence? If yes, it is inferred to be real. If no, its reads are attributed to the parent. To answer that question, a denoiser needs to know how often errors occur and what kinds of errors they are. The three widely used denoisers differ mainly in how they obtain that knowledge. DADA2 DADA2, published by Benjamin Callahan and colleagues in 2016 and distributed as an R package, is the most widely used denoiser and the most statistically elaborate. Its key move is to learn an error model from the data themselves. The model estimates, for each possible substitution — an A read as a C, a G read as a T, and so on — the probability of that substitution as a function of the quality score the sequencer assigned. It begins by assuming that the most abundant sequences are true and that everything else is error, estimates error rates from the differences, uses those rates to re-infer the true sequences, and alternates between the two steps until they stabilize. The fitted model describes how error rates fall as quality scores rise. Analysts are advised to inspect this relationship: if the fitted rates track the rates implied by the nominal quality scores and decline sensibly with increasing quality, the model is plausible. With error rates in hand, DADA2 partitions the reads. It starts with all unique sequences in a single partition centred on the most abundant one. For each other sequence, it computes the probability of observing at least as many copies as were actually seen, given the sequence's distance from the partition centre, the quality scores of its reads, and the error model. If that probability — an abundance p-value — is vanishingly small, the sequence is too abundant to be explained by error, and it becomes the centre of a new partition. Reads are reassigned to the partition that best explains them, and the process repeats until no further sequence can be split off. Each final partition centre is an ASV. A simple illustration makes the logic concrete. Suppose a sample contains a true sequence read 20,000 times, and suppose the error model says that, at a particular position with typical quality scores, the probability of a specific substitution is about one in two thousand. Around ten reads carrying exactly that substitution are then expected purely from error. If the variant with that substitution is observed twelve times, error explains it comfortably and it is folded back into its parent. If it is observed four hundred times, the chance of error producing that many copies is astronomically small, and the variant is inferred to be a real sequence — perhaps a closely related strain, perhaps a second 16S copy in the same genome. The numbers here are invented for illustration, but the reasoning is exactly what the algorithm formalizes, with the twist that it does so for every sequence against every candidate parent, and uses the quality scores of the actual reads rather than a typical value. This procedure has an attractive property: it can distinguish sequences that differ by a single nucleotide, provided the less abundant one is observed far more often than error could produce. It can also discard a relatively abundant sequence that differs from a very abundant parent at a position where errors are common. Several practical choices shape DADA2's output. Reads are filtered and truncated before denoising, usually by trimming at positions where quality collapses and discarding reads with more than a set number of expected errors. Forward and reverse reads are denoised separately and then merged, which means the truncation lengths must leave enough overlap. Chimeras are removed after denoising, by identifying ASVs that can be reconstructed exactly as a left segment of one more abundant ASV joined to a right segment of another. Denoising can be run sample by sample, on all samples pooled, or in an intermediate "pseudo-pooling" mode. Sample-by-sample inference is fastest and scales to very large studies, but it may miss a variant that is rare in every sample yet present in many, because the evidence in any one sample is too thin. Pooling gathers that evidence across samples at greater computational cost. Pseudo-pooling approximates pooling by using variants found in a first pass as priors for a second. Studies interested in rare organisms should consider the pooled modes. Deblur Deblur, published by Amnon Amir and colleagues in 2017 and integrated into QIIME 2, takes a simpler and faster approach. Instead of learning an error model from each dataset, it uses a fixed upper-bound error profile derived from Illumina data. It sorts sequences by abundance and works downward: for each sequence, it subtracts from its less abundant neighbours the number of reads that the error profile predicts it would have produced as errors at each Hamming distance. Whatever remains above zero after all subtractions is kept. Deblur requires all reads to be trimmed to the same length, which simplifies comparison but discards information from longer reads. It also includes a positive filter that removes sequences not resembling any known 16S sequence, which helps to remove off-target amplification products — host mitochondrial or chloroplast sequences, for instance — but could in principle discard genuinely novel lineages. Because Deblur works one sample at a time with a fixed model, it is fast, parallelizes easily, and produces the same result for a given sample regardless of what other samples are processed alongside it. UNOISE Robert Edgar's UNOISE algorithm, in its UNOISE2 and UNOISE3 versions, took the abundance-based logic of his UPARSE clustering method to its limit. A less abundant sequence is judged to be an error of a more abundant neighbour if the ratio of their abundances, often called the abundance skew, is large relative to the number of differences between them. The permitted skew shrinks as the number of differences grows, via a parameter called alpha: a variant one mismatch away from a parent must be much less abundant than the parent to be counted as its error, while a variant several mismatches away can be quite rare and still count as real. A minimum abundance, by default eight reads across the dataset, excludes very rare sequences outright. UNOISE3 is implemented in USEARCH, and an open-source reimplementation exists in VSEARCH. Its outputs are called zero-radius OTUs. The three denoisers are compared in Table 2. Table 2. The three widely used amplicon denoisers. Denoiser Error model Read length handling Distinctive feature Typical environment DADA2 Learned from each dataset, by quality score Truncated per read direction, then merged Abundance p-value; optional pooling R package; QIIME 2 plugin Deblur Fixed upper-bound profile All reads trimmed to one length Positive filter for 16S-like sequences QIIME 2 plugin UNOISE3 Abundance skew vs. distance Merged reads of any length Minimum-abundance cutoff USEARCH or VSEARCH How much does the choice matter? Benchmarks with mock communities show that all three denoisers recover the known sequences well, and much more accurately than traditional OTU clustering did. They differ in the margins. An independent comparison by Nearing and colleagues in 2018 found that DADA2 tended to recover the most true variants, including more rare ones, at the cost of occasionally retaining more spurious ones, while Deblur and UNOISE3 were more conservative. On real datasets the three produced broadly similar community-level conclusions: the same patterns of diversity, the same separation of groups in ordination. The differences concentrated in the rare tail. This is the recurring pattern in microbiome methods comparisons. Large, robust biological signals survive most reasonable pipelines. Small effects and conclusions about rare features are sensitive to processing, and should be checked against alternatives. The case for ASVs In a widely cited 2017 perspective, Callahan, Paul McMurdie, and Susan Holmes argued that exact sequence variants should replace OTUs in marker-gene analysis, and their argument turned on a property that is easy to overlook: consistent labelling. An ASV is a DNA sequence, and a DNA sequence means the same thing in every study. If one laboratory finds a particular 253-base V4 sequence associated with a disease, another laboratory using the same primers can look for exactly that sequence in its own data, with no need to re-cluster anything. Tables from different studies that used the same region can be merged by sequence. Features can be tracked across time and across cohorts. A de novo OTU, by contrast, is defined only in relation to the other sequences it was clustered with; its identity is local to the dataset. ASVs also bring resolution. Because they can differ by a single nucleotide, they can distinguish organisms that a 97 per cent cluster would merge. In practice this sometimes separates ecologically distinct strains — for example, closely related members of a genus with different associations with host health. And ASVs are reference-independent: like de novo OTUs, they retain novel organisms, yet like closed-reference OTUs, they can be compared across studies. The costs and caveats ASVs are not without problems, and several of them catch analysts unawares. One genome, several variants. Many bacteria carry multiple copies of the 16S gene, and those copies are not always identical. When they differ within the amplified region, a denoiser will correctly infer several ASVs from what is, biologically, one organism. Analyses of complete genomes have found that intragenomic variation is common, and that exact variants therefore sometimes split single genomes — while, conversely, identical short variants can be shared by different species. An ASV is a sequence, not an organism, and counting ASVs is not counting strains. Region-specific identity. An ASV's identity depends on the exact primers and trimming. A V4 ASV cannot be matched to a V3–V4 ASV, and even the same region trimmed to different lengths yields different sequences. The promise of cross-study comparability holds only among studies with matching amplicons, processed to the same length. Sensitivity to error-model assumptions. DADA2's error model assumes that quality scores carry information about error rates. Some newer Illumina instruments, including the NovaSeq, report binned quality scores with only a few distinct values, which weakens the relationship the model depends on and can produce poorly fitted error rates. Workarounds, such as constraining the fitted error rates to decline monotonically with quality, are now commonly used, but analysts working with such data should inspect the error model rather than trust the defaults. Rare-variant uncertainty. Denoisers are, by design, conservative about sequences seen only a few times, and some legitimately rare organisms are discarded. Studies whose question concerns the rare tail — the detection of a low-abundance pathogen, for example — need to consider the pooling mode, the sequencing depth, and possibly a targeted assay. Loss of singletons for estimation. Denoised data contain few or no singleton features, because a sequence seen once can almost never be distinguished from error. This undermines richness estimators, discussed in a later chapter, that extrapolate unseen diversity from the frequency of singletons. Long reads and full-length variants Denoising extends naturally to long-read data if the error model is adapted. Callahan and colleagues showed in 2019 that DADA2 could infer exact full-length 16S variants from PacBio circular consensus reads, resolving organisms at single-nucleotide resolution across the whole gene. Full-length ASVs make intragenomic variation visible: the several copies of a genome's 16S gene appear as a set of variants in fixed ratios, which is both a complication and, in some analyses, a way to recognize strains. Nanopore data, with a different error profile rich in insertions and deletions, have required different approaches, and denoising at full single-nucleotide resolution from Nanopore amplicons has been an active area of development as chemistry and basecalling have improved. Tracking what the pipeline discarded Every step before denoising removes reads, and the proportion removed is itself diagnostic. A well-behaved run on stool with a common V4 protocol might retain most of its raw reads through primer removal, filtering, merging, and chimera removal. When retention collapses at a particular step, the cause is usually identifiable. Heavy losses at merging usually mean the truncation lengths left too little overlap, or that the amplicon was longer than expected. Heavy losses at chimera removal — say, a third or more of the reads — suggest either that the primers were not removed, so that ambiguous primer bases are confusing the chimera check, or that the PCR was over-cycled. Losses at quality filtering that differ sharply between samples may point to a problem with a particular run or lane. For this reason, a table of read counts retained at each step, per sample, should be produced for every amplicon analysis and examined before any biology is attempted. It takes minutes and catches a large share of processing errors. It also belongs in the supplementary material of the eventual paper, because a reviewer or a later reader cannot otherwise judge whether the feature table reflects the samples or the choices made in handling them. Platforms such as QIIME 2 help here by recording provenance automatically: each output file carries a record of the commands, parameters, and software versions that produced it, all the way back to the raw data. Whatever software is used, the principle is the same. The processing decisions — truncation lengths, expected-error thresholds, pooling mode, chimera method, reference versions — are part of the result, and a result reported without them cannot be reproduced. From variants to a table Whatever the denoiser, the output is the same kind of object: a feature table, with samples in one dimension and ASVs in the other, each cell holding the number of reads assigned to that ASV in that sample, together with the sequence of each ASV. The table is typically very sparse — most ASVs appear in only a few samples — and its row totals vary between samples for reasons that have nothing to do with biology. Two further objects are usually built from the sequences. The first is a taxonomic assignment for each ASV, discussed in the next chapter. The second is a phylogenetic tree, built by aligning the ASV sequences and inferring their relationships, often by inserting them into a reference tree of full-length sequences. The tree is needed for phylogenetically aware measures such as Faith's phylogenetic diversity and UniFrac distances, which later chapters treat in detail. A final, sensible step is to review the table against the controls sequenced with the samples. Mock communities should yield their expected sequences, with only a handful of spurious extras. Extraction blanks should be sparse and should share their features mainly with the low-biomass samples. ASVs that match mitochondria or chloroplasts, which amplify with many 16S primers because those organelles descend from bacteria, should be identified and usually removed. Only then does the table deserve to be treated as data about the community. Hashtags: #MicrobiomeInformatics #MicrobiomeAnalysis #16SrRNASequencing #ShotgunMetagenomics #AmpliconSequencing #MetagenomicSequencing #OperationalTaxonomicUnits #AmpliconSequenceVariants #DADA2 #QIIME2 #Deblur #UNOISE #TaxonomicClassification #MicrobialCommunityProfiling #MetagenomeAssembly #FunctionalProfiling #CompositionalData #AlphaDiversity #BetaDiversity #DifferentialAbundance #ContaminationControl #MicrobiomeQualityControl #ReferenceDatabases #MicrobialEcology #FutureOfMicrobiomeInformatics
- Metascience (The Empirical Study of Science Itself)
Download the Book (PDF): Introduction Somewhere in a university office this morning, a researcher is deciding what to work on next. The decision looks like a scientific judgement. In practice it is shaped by a dense lattice of non-scientific facts: which questions a funding agency's current call will reward, which results a high-prestige journal is likely to accept, how many months remain on a contract, what a hiring committee three years from now will be able to read off a curriculum vitae in ninety seconds. The researcher may well believe the choice was driven by curiosity. It was also driven by an allocation system, and that system has properties that can be measured. Measuring them is what this book is about. For most of its history, science has studied everything except itself. Chemistry has a chemistry of chemistry only in the trivial sense; physics does not run controlled trials on physicists. The institutions that decide which research gets done — peer review, grant competitions, journals, tenure committees, prizes — were built by practitioners using intuition, precedent and professional judgement, and then largely left alone. When they were defended, they were defended on the grounds that they had produced antibiotics and general relativity and the transistor, which is a claim about the output of the whole enterprise rather than about the specific mechanisms inside it. Nobody asked whether a different arrangement would have produced the same antibiotics faster, or three more of them, or the same number with less waste. Over the past two decades that has changed. A field has assembled itself under several names — metascience, the science of science, research on research — with a single defining commitment: that the machinery of science is an empirical object. How reliably does peer review identify good work? What predicts whether a finding will hold up when someone repeats the experiment? Does a larger grant buy proportionally more discovery? Who gets funded, and does the pattern track quality or something else? These are not philosophical questions about the nature of knowledge. They are questions with numbers in the answers, and over the past twenty years a great many of those numbers have been produced. The numbers are not reassuring. When two independent committees at a major machine-learning conference reviewed the same set of papers without knowing it, roughly half the papers accepted by one committee were rejected by the other. When forty-three experienced oncology researchers were recruited to replicate the National Institutes of Health grant review process on real applications, the agreement between reviewers evaluating the same proposal was statistically indistinguishable from zero. When a consortium of psychologists attempted careful, high-powered replications of a hundred published findings, thirty-six per cent produced a statistically significant result in the same direction as the original, and the average replication effect was half the size of the published one. When a comparable effort was made in preclinical cancer biology, the median effect size in the replications was eighty-five per cent smaller than in the original experiments — and the project managed to complete only a quarter of the experiments it had planned, because the published methods sections did not contain enough information to repeat the work. Each of these findings arrived with caveats, and most were contested. Some of the contest was productive. But the accumulated weight of two decades of this kind of work supports a claim that would once have seemed intemperate and is now close to consensus among people who study the question: the allocation machinery of science is far noisier, far more path-dependent, and far more responsive to incentives unrelated to truth than the people inside it generally believe. That is the first half of the argument. The second half is the reason this book exists rather than a long complaint. Noise, path dependence and perverse incentives are not moral failings. They are design properties. A system in which reviewers disagree is not a system staffed by bad reviewers; it is a system asking reviewers to make fine-grained distinctions that the available evidence cannot support, and then treating the resulting numbers as though they were measurements. A literature in which ninety-six per cent of reported results are positive is not a literature written by liars; it is the predictable output of a filter that publishes confirmations and discards disconfirmations, operating on researchers who respond rationally to what the filter rewards. Cumulative advantage — the process by which an early win compounds into a career and an early loss compounds into an exit — does not require anyone to be biased. It requires only that past success be used as evidence of future quality, which is a reasonable heuristic that becomes a self-fulfilling one when applied at scale. Design properties can be redesigned. That is the practical promise of metascience, and it is a promise that has begun, in a few places, to be kept. Journals that adopted a format in which studies are reviewed and accepted before the results exist publish positive findings forty-four per cent of the time, against ninety-six per cent in the conventional literature — a difference so large that it is hard to read as anything other than a measurement of how much the ordinary publication process distorts the record. A psychology journal that offered authors a small badge for sharing their data saw the data-sharing rate rise from under three per cent to thirty-nine per cent in eighteen months, while comparison journals did not move. Large cardiovascular trials began reporting null results at a rate that jumped from roughly two in five to more than nine in ten after prospective registration became mandatory. These are interventions, with before-and-after data, on real scientific systems. They are the beginnings of an engineering discipline. The argument of this book, stated once and then developed, is this: the reliability and the fairness of science are properties of its allocation machinery rather than of the virtue of the people inside it; that machinery is measurable; and measuring it reveals problems that are tractable by design rather than by exhortation. Several things follow from taking that seriously, and they run through what comes next. The first is that reliability and equity turn out to be the same problem, seen from two angles. It is tempting to treat "is the science any good?" and "is the system fair?" as separate concerns, the first technical and the second political. The evidence does not cooperate. The same cumulative-advantage dynamics that concentrate citations in a small elite also concentrate faculty jobs in a handful of departments; the same reviewer noise that makes funding decisions arbitrary at the margin is what allows small biases to determine outcomes; the same incentive to publish positive results quickly is what makes careers depend on luck. A system that allocates on the basis of past allocation is both less accurate and less fair than one that does not, and for exactly the same reason. The second is that the units of analysis are institutions, not individuals. Almost every popular treatment of these issues reaches, sooner or later, for a rogue. Fabricators exist and they matter, and a handful of them have done enormous damage. But the deepest finding of metascience is that you can get a badly distorted literature out of a population of entirely honest researchers, provided you reward them for the right things in the wrong proportions. A well-known model of this shows that if publication rates determine survival and lower methodological standards produce more publications, then standards will decline across the community over generations even though no individual ever makes a dishonest choice — and that replication, by itself, is too weak a corrective to stop it. Blame is an unproductive framework for a problem whose mechanism is selection. The third is a caution about the evidence itself. Metascience is science, which means it is subject to everything this book documents about science. Its studies are often observational, its samples are frequently one discipline or one funder, its own incentives reward striking findings about crisis, and several of its most-cited results have been vigorously disputed on methodological grounds. Where that is the case — the claim that science has become less disruptive over time is the sharpest example — I will say so and lay out the dispute rather than reporting the headline. A book arguing that scientific claims should be held to their evidence cannot exempt its own. A word on scope. This is not a survey of everything that has been written about how science works. It leaves out the sociology of laboratory life, the philosophy of confirmation, the history of particular discoveries, and the long argument about whether scientific knowledge is socially constructed — all of which are serious subjects with their own literatures. It concentrates on the quantitative study of science's allocation systems: citation, evaluation, funding, publication, replication, and careers. It draws most heavily on evidence from biomedicine, psychology, economics and computer science, because those are the fields where the work has been done, and it tries to be honest about how far the findings travel. And it is written for someone who is intelligent, interested, and not a specialist — a researcher in another field, a programme officer, a science journalist, a graduate student wondering what they have signed up for. The stakes are not academic. Global spending on research and development runs to more than two trillion dollars a year. Public funders allocate tens of billions through mechanisms whose reliability has been tested perhaps a dozen times in total. Clinical practice, technology policy, and a great deal of what we collectively believe about the world rest on a published literature whose error rate is unknown but demonstrably not small. If the machinery that decides which questions get asked is even modestly miscalibrated, the compounding loss over decades is very large — and it is invisible, because we never see the science that was not done. We do, however, now have the tools to look. What follows is an account of what has been found. Chapter 1: Science as an Object of Study The idea that science could be counted arrived, as such ideas often do, from someone with a filing problem. In the 1950s Eugene Garfield was working on a project to index medical literature and confronting a difficulty familiar to anyone who has tried to find their way into an unfamiliar field: subject headings are unreliable. Papers get filed under terms their authors did not use and readers do not think of. Garfield's insight, published in Science in 1955, was that scientific papers already contain a rich and standardised set of pointers to the literature they depend on, in the form of references — and that these pointers can be inverted. If paper A cites paper B, then from B's perspective A is a forward link, a later document that found B relevant. Build the index of forward links and you have a map of the literature drawn by the literature itself. Garfield's motivation was retrieval. What he actually built, in the Science Citation Index that launched in 1964, was the first large-scale instrument for observing science as a system. Every citation is a small, dated, attributed act of attention. Aggregate enough of them and you can see structure: which papers are being used, which fields are connecting to which, how quickly attention decays, where new specialties are forming. The index was designed to help researchers find things. It became the raw material for an entire empirical discipline, and — in a development Garfield later regarded with ambivalence — the basis for the metrics by which researchers came to be judged. A parallel line of thought was developing at the same time. Derek de Solla Price, a physicist turned historian of science, noticed that the quantity of scientific output was growing exponentially and had been doing so for centuries. His 1963 book Little Science, Big Science treated the scientific enterprise the way a demographer treats a population: as an aggregate with growth curves, size distributions, and saturation limits. Price observed that the number of scientists had been doubling roughly every fifteen years, that this could not continue indefinitely, and that the transition from exponential growth to something flatter would change the character of scientific life. He also noticed that scientific productivity was extraordinarily unequal — a small fraction of researchers produce a large fraction of the papers — and connected this to an earlier observation by Alfred Lotka, who had found in 1926 that the number of authors producing n papers falls off roughly as one over n squared. Price's later work gave the field its central dynamic principle. In a 1976 paper he described what he called cumulative advantage: a process in which the probability of receiving a new citation is proportional to the number already received. A paper that has been noticed is more likely to be noticed again; success compounds. This was a formalisation of something the sociologist Robert K. Merton had described, more memorably, in 1968. Merton had interviewed Nobel laureates and found that when a laureate and a lesser-known scientist produced similar work, the laureate received the lion's share of the credit — an effect he named after the verse in the Gospel of Matthew about those who have being given more. The Matthew effect in science, Merton argued, was not corruption. It was the ordinary operation of a system that uses reputation as a signal of quality, and it had the consequence that the distribution of recognition would be far more unequal than any plausible distribution of underlying ability. Merton's broader programme mattered too. In 1942 he had set out four norms he took to be constitutive of scientific practice — communalism, universalism, disinterestedness, organised scepticism. These have been criticised at length as idealised, and Merton himself later documented the gap between them and behaviour. But they served an important function: they made the social structure of science a legitimate object of study rather than a background condition. If science has norms, the norms can be violated, and whether they are violated is an empirical question. For four decades, that question was answered with small samples and heroic effort. The bibliometricians worked with printed indexes and punch cards. The sociologists of science did interviews and archival work. There were landmark studies — Jonathan and Stephen Cole's 1981 experiment on the National Science Foundation's peer review system remains one of the most-cited pieces of evidence in the field, and it was done by hand. But the scale of the available data was tiny compared to the scale of the thing being studied, and most of the interesting questions were out of reach. What changed Three things converged in the 2000s and 2010s to turn a specialty into a field. The first was data. Digital publishing made the full text and the full reference list of essentially every paper machine-readable. Crossref, established in 2000 by publishers to manage the digital object identifier system, incidentally created a shared registry of scholarly outputs and, eventually, an open citation graph. Commercial databases — Web of Science, Scopus — grew to cover tens of millions of records. Open alternatives followed: Microsoft Academic Graph, and after its retirement, OpenAlex, which indexes well over two hundred million works and makes the citation network freely available. Grant agencies digitised their application and award records; some made them public. ORCID gave researchers persistent identifiers, which made it possible to track careers rather than just papers. By the mid-2010s it was possible for a single researcher with a laptop to analyse a data set covering most of the published record of modern science. The second was computation, which is the boring half of the same story. A citation network with two hundred million nodes is not an object you can reason about with summary statistics. Network analysis, natural language processing over abstracts and full texts, and — more recently — large language models capable of extracting structured claims from prose have all made it possible to ask questions that would previously have required an army of coders. The third, and least appreciated, was the arrival of credible causal inference. Most early bibliometrics was descriptive, and description is easy to misread. That highly cited scientists receive more grants is a fact about correlation; it tells you nothing about whether the grants caused the citations, the citations caused the grants, or both track something else. What transformed the field was the systematic import of the identification toolkit from empirical economics: regression discontinuity around funding thresholds, difference-in-differences around policy changes, matched controls, instrumental variables, and where possible actual randomisation. The best metascience of the past fifteen years is not a better-powered version of the old bibliometrics. It is a different kind of claim. What counts as metascience The field has no agreed boundary, but a workable division sorts the research into five questions. How is research done? Studies of methods: what statistical practices are used, what designs are chosen, how large the samples are, whether the analyses match what was planned. This is where the evidence on statistical power lives. How is it reported? Studies of the gap between what was done and what appears in print: selective outcome reporting, missing methods detail, publication bias, spin in abstracts. How is it verified? Replication, reproduction of analyses from shared data, and the meta-research on both. How is it evaluated? Peer review, funding decisions, journal selection, hiring and promotion — the allocation machinery. How is it rewarded and incentivised? Career structures, metrics, prizes, institutional policy, and the behavioural responses they produce. A sixth category, increasingly, is what works to fix it — the study of interventions. This is the part of the field that most resembles a normal applied science, with treatment groups and outcome measures, and it is the part that has grown fastest. The size of the thing being measured Any account of modern science has to start from its scale, which is genuinely difficult to hold in mind. The best available estimates put long-run growth in the number of scientific publications at around four per cent a year since the start of the modern scientific era — an annual rate that implies a doubling roughly every seventeen years. The growth has not been uniform. A careful segmented analysis of four large bibliographic databases identifies distinct regimes: about 2.9 per cent a year in the pre-industrial period, rising to 5.6 per cent during the industrial revolution, falling back to 3.8 per cent through the era of economic crises and world wars, and settling at roughly 5.1 per cent a year since 1952 — a doubling time of about fourteen years. In the life sciences and the physical and technical sciences the long-run rates are near 5.1 and 5.5 per cent respectively. Sustained for seventy years, five per cent compounds into something remarkable. A field that produced a thousand papers a year in 1952 produces something like thirty thousand today. The consequence is that no individual can read their own field. Specialties fragment; reading is replaced by search; and the question of what to pay attention to becomes, itself, a resource allocation problem that the system must solve somehow. Citation counts, journal prestige, and institutional affiliation are the heuristics it reached for, and much of what follows in this book is an account of what those heuristics do when they are load-bearing. Growth also changes what the average scientific career looks like. Price anticipated this: exponential growth in output implies exponential growth in producers, and exponential growth in producers cannot continue once the population of potential scientists is exhausted or the funding stops expanding. What happens instead is that the number of people trained for research grows faster than the number of positions that let them do it — a mismatch that shows up in postdoctoral queues, in the average age at first independent grant, and in the intensity of competition for the signals that determine who continues. The instruments and their blind spots Every empirical field is shaped by what its instruments can see, and metascience is no exception. Its primary instrument is the bibliographic database, and the major databases share a set of biases that propagate into findings. Coverage is the first. Web of Science and Scopus are selective: journals are included on the basis of editorial criteria, and the selection has historically favoured English-language, North American and Western European publications, and journal articles over books, conference proceedings, and the grey literature. This matters enormously for any claim about national productivity, about the humanities and much of the social sciences, or about fields such as computer science where conferences rather than journals carry the important work. The open databases have broader coverage but looser quality control; OpenAlex indexes more than two hundred million works, including a long tail of material that no selective index would admit, which is an advantage for questions about the whole system and a liability for questions about its core. Identity is the second. Attributing papers to people is harder than it sounds. Common surnames, inconsistent initials, name changes, and institutional moves all corrupt author disambiguation, and the algorithms that resolve them make systematic errors — they are worse, for example, on Chinese and Korean names, which means studies of career trajectories are least reliable exactly where the workforce has grown fastest. ORCID has improved this for researchers who use it, which is a self-selected and increasingly but not universally representative group. The third and deepest problem is that the citation itself is a poor unit. A reference in a paper can mean any of a dozen things: this is the method I used, this is the claim I am refuting, this is the obligatory nod to the field's founder, this is my own earlier paper, this is a friend's paper, this is what the reviewer told me to add. Counting them treats all of these as the same event. Attempts to classify citations by function — and more recently to do so automatically from the surrounding sentences — show that a substantial fraction are perfunctory, and that negative citations exist but are rare and often coded in language too polite for a classifier to catch. Any measurement built on citation counts inherits this ambiguity, and the measurements do not become less ambiguous by being aggregated. None of this makes bibliometric evidence worthless. It makes it evidence of a particular kind: good for describing large-scale structure and change, weak for evaluating individuals, and always in need of a second, independent line of attack. The strongest studies in this book combine the bibliometric record with something else — an administrative data set of grant applications, a randomised assignment, a survey, a hand-coded sample. A field acquires institutions Fields become real when they get money, venues and jobs, and metascience acquired all three quickly. The Center for Open Science, founded in 2013 by Brian Nosek and Jeffrey Spies, built infrastructure — a preprint and project repository, standardised guidelines for journals, a registry for preregistered studies — and ran the large replication projects that gave the field its public profile. The Meta-Research Innovation Center at Stanford, founded the same year by John Ioannidis and Steven Goodman, concentrated on the methodological side. In the United Kingdom, the Research on Research Institute was established in 2019 as a consortium of funders and universities explicitly to conduct experiments on funding and evaluation practice. Philanthropic funders, notably the Alfred P. Sloan Foundation, Arnold Ventures and the Wellcome Trust, put substantial money into work that no disciplinary funder had an obvious reason to support. In 2016 the first Metascience conference was held at Stanford; by the 2020s there were dedicated journals, a growing number of academic posts with "meta-research" in the title, and — a decisive sign — national funders running randomised trials on their own procedures. Government interest followed. The US Defense Advanced Research Projects Agency funded SCORE, a programme that attempted to build automated systems for estimating the replicability of social and behavioural science claims, which produced both a large hand-replicated data set and a surprising finding: crowds of forecasters, and eventually algorithms, could predict replication outcomes considerably better than chance from features of the original paper. If replicability is predictable in advance, it is not random noise; it is a property that the publication system is currently failing to select on. The reflexivity problem There is a difficulty at the centre of this enterprise that deserves to be stated early rather than discovered late. Metascience is a scientific field that studies scientific fields, and it is not exempt from its own findings. Its practitioners publish in journals, apply for grants, and accumulate citations, and they face the same reward structures they are documenting. A metascientific paper reporting that a widely used practice is fine is much less publishable than one reporting that it is broken. The field's most famous title — John Ioannidis's 2005 essay "Why Most Published Research Findings Are False" — is a modelling exercise whose conclusion depends on assumptions about prior probabilities and bias that reasonable people dispute, and which has been cited many thousands of times, very often by people summarising its title rather than its argument. The honest position is that metascience has produced findings of very different evidential grades, and that they should not be treated alike. At one end sit results from designs that would be considered strong in any field: the randomised assignment of identical papers to independent review committees, the regression discontinuity around a funding cutoff, the before-and-after comparison at a journal that introduced a policy while comparable journals did not. At the other end sit correlational claims about large bibliometric data sets, where the measurement instrument is a citation count whose meaning is contested and where the number of analytic choices available to the researcher is enormous. Several of the field's most publicised results are of the second kind, and at least one of them — the claim that scientific papers have become steadily less disruptive over sixty years — has been substantially challenged on grounds of measurement artefact. None of this argues against the project. It argues for holding it to the standard it recommends. The useful test to apply to any metascientific claim is the same one it applies to the science it studies: what exactly was measured, on what population, compared to what counterfactual, and how many other ways could the analysis have gone? With that caution in place, we can start with the most basic instrument in the field's toolkit, and the one most widely misused: the citation. Chapter 2: The Arithmetic of Citation Most quantities in everyday life are distributed in a way that makes the average informative. Human height has a mean of about 1.7 metres and almost everyone is within twenty-five centimetres of it; knowing the mean tells you a great deal about any individual you might meet. Our statistical intuitions are calibrated on distributions like this, and they fail badly on distributions that are not. Citations are not. The number of times a scientific paper is cited follows a distribution in which the mean is dominated by a small tail and describes almost nobody. In a typical field, the median paper accumulates a handful of citations over its lifetime, a large minority accumulate none at all, and a tiny fraction accumulate thousands. The mean sits somewhere in the empty space between the bulk and the tail, describing a paper that does not exist. This is not a curiosity. It is the single most consequential fact about scientific evaluation, because nearly every metric in use is either an average of this distribution or a rank order derived from it, and averages and ranks both assume a shape the data does not have. The shape of the distribution The pattern was noticed before anyone had the data to characterise it properly. Alfred Lotka, working in 1926 on the productivity of chemists and physicists, found that the number of authors publishing n papers declined approximately as the inverse square of n: for every hundred authors with one paper, about twenty-five have two, about eleven have three, and so on. Derek de Solla Price found similar inequality in citations and in the distribution of papers across journals, and proposed a generative mechanism: cumulative advantage, in which the probability that a paper acquires its next citation rises with the number it already has. The mechanism is easy to state and hard to overestimate. Suppose that scientists find papers partly through search and partly through other papers' reference lists. A paper that has been cited a hundred times appears in a hundred reference lists and is therefore a hundred times more likely to be encountered than a paper cited once. Even if the two papers are identical in quality, the first will pull further ahead, and the gap will widen over time. The same logic, under the name preferential attachment, generates the heavy-tailed degree distributions observed in web links, film box-office receipts, city sizes and the number of sexual partners. There is nothing specifically scientific about it. What is specifically scientific is that we then use the outcome as a quality measurement. If the distribution of citations were generated purely by quality, citation counts would be excellent evaluation instruments. If it were generated purely by cumulative advantage acting on random initial fluctuations, they would be worthless. The truth is in between, and where in between turns out to be the central empirical question — one we will take up properly in the next chapter, where the evidence comes from experiments rather than observation. For now, take the shape of the distribution at face value and consider what it does to the metrics built on top of it. The impact factor and the fallacy of the mean The journal impact factor is a two-year mean: the number of citations in a given year to items published in the previous two, divided by the number of citable items. Garfield devised it in the 1960s as a tool for librarians deciding which journals to buy. It became, within thirty years, the dominant currency of scientific prestige, used to evaluate papers, people, departments and national research systems. The problem is arithmetic. A mean is a reasonable summary of a symmetric distribution and a misleading one of a skewed distribution, and citation distributions within a journal are extremely skewed — skewed enough that most papers in a high-impact-factor journal receive fewer citations than the impact factor suggests, while a handful receive vastly more. Nature has documented this about itself. In a 2005 editorial the journal reported that of roughly 1,800 citable items it had published in 2002 and 2003, only around fifty received more than a hundred citations in 2004, while the majority received fewer than twenty — and that eighty-nine per cent of that year's impact factor was generated by just a quarter of its papers. The number that conferred prestige on every paper in the journal was, in the main, earned by a small subset of them. The consequence for evaluation is severe. Using a journal's impact factor to estimate the likely citation impact of a specific paper in it is a prediction with enormous error bars, and the distributions of different journals overlap so extensively that a paper in a journal with an impact factor of four is quite often cited more than a paper in a journal with an impact factor of twenty. Yet hiring and promotion committees across much of the world have used journal placement as the primary signal, partly because reading the papers is expensive and partly because the metric has the appearance of objectivity. A second problem is that the impact factor is a negotiated quantity rather than a measured one. The denominator counts "citable items", a category determined in discussion between publishers and the index provider; editorials, news pieces and correspondence may be excluded from the denominator while the citations they attract remain in the numerator. Journals have been documented encouraging authors to cite the journal, publishing large numbers of reviews (which are cited more than primary research), and timing publication to maximise the two-year window. These are not accusations of fraud. They are the predictable responses of organisations to a metric that determines their standing. The h-index and what it hides In 2005 the physicist Jorge Hirsch proposed an individual-level measure intended to be robust to the problems of both totals and means: a researcher has an h-index of h if h of their papers have at least h citations each. It is a clever construction. It cannot be inflated by one enormously cited paper, nor by a long list of uncited ones, and it correlates reasonably well with expert judgement of seniority within a field. Its defects are equally structural. The h-index cannot decrease, which makes it a measure of accumulated career length as much as of contribution — a thirty-year-old and a sixty-year-old of identical talent will have very different values. It is bounded by total publications, so it penalises fields, methods and career paths that produce fewer, larger outputs. It varies by an order of magnitude between disciplines, because citation densities do, which makes cross-field comparison meaningless without normalisation that is rarely applied. It ignores author position and contribution entirely, so a name appended to a thousand-author consortium paper counts the same as sole authorship. And because it counts papers above a moving threshold, it can be raised efficiently by producing a steady stream of moderately cited work and not at all by producing one transformative result — precisely the incentive gradient one would not choose. The h-index is worth dwelling on because it illustrates a general pattern. Every attempt to fix a metric's known defect introduces a new one, because the underlying problem is not any particular formula. It is that a single number is being asked to stand in for a multidimensional judgement, under conditions where the people being measured can see the formula and adjust. Table 1 sets out what the commonest indicators actually measure and where each of them breaks. Table 1. What common research indicators measure, and where they fail. Indicator What it is What it can support Where it fails Raw citation count Forward links to one work Comparing similar-aged work within one subfield Heavy skew; field and age dependence; treats all citation types alike Journal impact factor Two-year mean citations per citable item Rough comparison of journals within a field Mean of a skewed distribution; negotiated denominator; says little about any single paper h-index Largest h with h papers of h+ citations Coarse indication of established output Never falls; confounded with career length and field; blind to author role Field-normalised citation score Citations relative to field and year baseline Cross-field comparison at aggregate level Depends heavily on field definition; still skewed; unstable for small samples Altmetric attention Mentions in news, policy documents, social media Detecting public and policy reach Measures attention, not quality; easily gamed; strongly topic-dependent None of these is useless. Each is useful for something narrow and misleading for the thing it is generally used for, which is ranking individuals. Uncitedness, and the myth of the unread paper A statistic that circulates widely in commentary about scientific waste holds that half of all published papers are never read by anyone other than their authors, editors and reviewers, and that ninety per cent are never cited. Both figures are wrong, and the way they went wrong is instructive. The real uncitedness rate depends on three choices: the field, the citation window, and the database. Measured in Web of Science over a five-year window, the proportion of research articles with zero citations is substantial in some humanities fields and small in the biomedical sciences, where it is typically in the single digits to low teens. Lengthen the window and the figure falls further, because citation accrual is slow in some fields and papers continue to be discovered for decades. Broaden the database to include the long tail of journals that selective indexes exclude and the figure rises sharply, because much of that tail genuinely goes unread. The often-quoted ninety per cent appears to have originated in a conference presentation and to have been repeated without a source ever since — a small demonstration that scientists are no more careful with statistics about themselves than anyone else is. The corrected figures matter because the rhetorical use of the inflated one is to argue that most science is worthless. What the real distribution shows is different and less dramatic: most papers contribute modestly, a few contribute enormously, and the modest contributions are the substrate on which the large ones are built. A field in which every paper was highly cited would be a field that had stopped exploring. There is also a category of work that the short-window metrics are structurally incapable of seeing. Bibliometricians call them sleeping beauties: papers that attract almost no attention for a decade or more and then are suddenly and heavily cited when a technique matures or an adjacent field catches up. Systematic searches of the citation record find thousands of them, distributed across all disciplines; some of the most extreme cases slept for more than half a century. Any evaluation system operating on a two-to-five-year horizon will classify a sleeping beauty as a failure, and will have done so at exactly the moment when continued support mattered most. What citations do measure It would be a mistake to move from "citation metrics are misused" to "citations are meaningless". They are not. At aggregate level, citation counts correlate respectably with expert judgement. Studies comparing bibliometric indicators with peer assessment of the same research units — departments, programmes, national evaluation exercises — generally find correlations strong enough that the two approaches produce similar rank orders at the top and the bottom, and diverge in the crowded middle. Work that later wins major prizes is, on average, heavily cited beforehand. Papers describing methods and reagents that many laboratories adopt accumulate citations for the good reason that many laboratories use them. The reasonable conclusion is about resolution. Citation data can distinguish a body of work that the field has used extensively from one it has largely ignored. It cannot reliably distinguish the fortieth-ranked candidate from the fifty-fifth, which is exactly the distinction that hiring and funding committees are asked to make. Using a coarse instrument for a fine-grained decision produces decisions that look precise and are in fact close to arbitrary — a theme that recurs when we turn to peer review. There is one further limitation worth naming. Citations measure uptake within the scientific literature, and a great deal of what science is for happens outside it. A paper that changed a clinical guideline, or a regulatory standard, or an engineering practice, may be cited rarely because the people acting on it do not write papers. Efforts to capture this — tracking citations in policy documents, patents, clinical guidelines and news coverage — have produced a family of alternative indicators, but these measure attention of a different kind rather than impact of a better kind, and they are at least as gameable. Goodhart's arrival The economist Charles Goodhart observed in 1975 that a statistical regularity tends to collapse once pressure is placed on it for control purposes. The anthropologist Marilyn Strathern gave the sharper form: when a measure becomes a target, it ceases to be a good measure. Scientific metrics have provided an unusually clean demonstration. Once citation counts became consequential, behaviours appeared that increase them without increasing the contribution they were meant to track. Self-citation, which is legitimate in moderation, becomes strategic at the margins; a handful of researchers have self-citation rates above fifty per cent. Citation cartels — informal arrangements in which groups of authors or journals systematically cite one another — have been detected by network analysis and have led to journals being delisted from indexes. Reviewers have been caught requiring authors to cite the reviewer's own work as a condition of acceptance, a practice known as coercive citation that surveys suggest a substantial minority of researchers have experienced. Salami publication — dividing one study's findings into the maximum number of separately publishable units — raises publication counts and, because each unit cites the others, citation counts too. The deeper effect is not cheating but reorientation. If what is rewarded is citations within a two-to-five-year window, the rational strategy is to work on topics with large, active, fast-citing communities, using methods that many others use, on questions where a result of some kind is likely. That is a description of useful normal science. It is also a description of a field that will systematically underproduce work that is slow to pay off, addressed to small communities, methodologically unusual, or likely to produce a null result. Concentration, and where it is heading If cumulative advantage operates and metrics reward it, we should expect the concentration of citations to increase over time. It has. An analysis of nearly twenty-six million papers and four million authors found that the share of all citations going to the top one per cent of authors rose from fourteen per cent in 2000 to twenty-one per cent in 2015. More than one citation in five now goes to one author in a hundred, and the trend over that fifteen-year window was steadily upward. The authors of that study attributed the rise partly to the growth of large collaborations, which multiply the number of papers an individual can be an author of, and partly to the increasing use of metrics in evaluation, which rewards exactly the accumulation strategies that produce concentration. A related finding concerns the literature rather than the people. A study of ninety million papers and 1.8 billion citations across 241 subject areas found that as a field grows larger, its list of most-cited works becomes more rather than less entrenched: citations concentrate on already-prominent papers, turnover in the canon slows, and the probability that a newly published paper ever reaches high-citation status falls. The intuitive expectation — that more papers means more chances for a good idea to be noticed — turns out to be backwards. In a field producing a hundred thousand papers a year, the only tractable way for a reader to decide what to read is to follow existing prominence, which reproduces it. This is the first place where the book's two themes meet. Concentration of citations is simultaneously an efficiency problem and an equity problem. It means that the system's attention is allocated increasingly by prior attention rather than by content, which degrades its accuracy as a filter; and it means that the returns to being early, well-placed and well-connected compound, which shapes who can build a career. The two are not separate pathologies. They are the same mechanism described in different vocabularies. Whether that mechanism is really cumulative advantage — whether status genuinely causes further status, rather than merely tracking quality that was there all along — is the question that observational citation data cannot settle. For that we need experiments. Chapter 3: Success Breeds Success The central claim of the previous chapter was that scientific attention is distributed in a way consistent with cumulative advantage. Consistency is not proof. A distribution in which a few papers get most of the citations is exactly what you would see if a few papers were genuinely much better than the rest, and quality in creative work is plausibly heavy-tailed. To distinguish the two stories you need a situation in which something that is not quality varies, and you can observe what happens next. Over the past fifteen years several such situations have been found or created. The evidence they produce is the strongest in metascience, and it points one way: status has independent causal force. Being recognised makes future recognition more likely, holding the underlying work constant. A threshold with nothing behind it The cleanest study of this kind exploits a feature of competitive grant schemes: the funding line. Applications are scored, ranked, and funded down to the point where the money runs out. Just above and just below that line sit proposals whose scores differ by an amount smaller than the measurement error of the scoring process — applications that are, for practical purposes, indistinguishable. One group receives money and the label of a winner. The other receives nothing. Thijs Bol, Mathijs de Vaan and Arnout van de Rijt applied this design to early-career grants from the Netherlands Organisation for Scientific Research, tracking applicants for eight years after the decision. Because the comparison is restricted to a narrow band around the threshold, the two groups can be treated as effectively randomised with respect to everything the review process could detect. The result was stark. Over the following eight years, the narrow winners accumulated more than twice as much research funding as the narrow losers. Since the initial grant itself was only a fraction of that difference, most of the gap was generated afterwards, by subsequent competitions that the winners entered with a funded grant on their record and the losers entered without one. The mechanism turned out to have two parts. Part of the advantage came from how later reviewers treated a prior award — prestige acting as evidence. The other part was self-inflicted: the near-miss applicants were substantially less likely to apply again at all. Losing narrowly did not merely fail to help; it discouraged. A system that allocates on the basis of past allocation does not need biased reviewers to compound inequality. It only needs discouraged applicants, and it produces them reliably. The design has limits worth naming. It identifies the effect of winning at the margin, which may not generalise to applicants far from the line. It comes from one funder in one country, at one career stage. And the outcome measured is subsequent funding, not subsequent discovery. But the logic is hard to escape: two populations that the evaluation system could not tell apart ended up, eight years later, in very different places, and the difference was manufactured by the evaluation system itself. Manufacturing status directly A second line of evidence comes from experiments in which status is assigned at random and the consequences are observed. Arnout van de Rijt and colleagues ran a set of field experiments on this logic outside science, in settings where success is public and cumulative: they gave arbitrary early boosts — donations to fundraising campaigns, positive ratings to product reviews, awards to contributors on a collaborative platform, endorsements to petition signatories — to randomly selected recipients and left matched controls alone. Across the domains, the arbitrary boost produced durable advantage: recipients went on to attract more support than controls, and in several cases the gap widened. A related experiment, run earlier on an artificial music market, gave participants access to the same set of unknown songs. Some participants saw download counts of others in their group; some did not. In the independent condition, song popularity tracked an underlying quality signal reasonably well. In the social-influence conditions, the same songs achieved wildly different outcomes across parallel worlds — a song that came near the top in one world came near the bottom in another — and the inequality of outcomes was much greater. Quality mattered, but it mattered less than history, and history was arbitrary. These are not studies of science. They are studies of the mechanism that operates in science, in settings where randomisation is ethically and practically possible. Their relevance is that they rule out the most reassuring interpretation of concentration in the citation record. Concentration does not require that the concentrated few be better; it arises spontaneously whenever visible past success feeds into present choice. Status inside the scientific record Within science itself, several natural experiments exploit sudden changes in a researcher's status with no corresponding change in their work. The most elegant use prizes. When a scientist is elected to a national academy, or wins a major award, nothing changes about the papers they have already published. Studies tracking citation trajectories across such events find that the already-published work is cited more afterwards — a pure Matthew effect, since the only thing that changed is the label attached to the author's name. The effect is not enormous, but it is consistently detectable, and it is largest for work that was previously obscure, which is what the theory predicts: reputation provides the most lift where prior attention was lowest. A darker natural experiment runs in the opposite direction. Pierre Azoulay and colleagues studied what happens to a scientific subfield when a superstar researcher dies unexpectedly in mid-career. The collaborators of the deceased suffer a lasting decline in output, which is unsurprising. What is surprising is what happens to everyone else: the rate of publication by scientists who had not collaborated with the star rises, and the newcomers who enter the field bring in different ideas and different methods. The presence of a dominant figure appears to have suppressed entry. That finding has a straightforward reading in terms of this chapter's argument. If evaluation weights the judgement of established figures heavily, and established figures have views about what counts as a promising direction, then the field's gatekeeping will be correlated with one person's taste. Remove the person and the correlation loosens. This is cumulative advantage operating not on individual citations but on intellectual agendas. The complication: what doesn't kill you Honest reporting requires the counter-evidence, and there is a significant piece of it. Yang Wang, Benjamin Jones and Dashun Wang applied the same near-threshold logic to early-career applicants for National Institutes of Health R01 grants, comparing 623 junior scientists whose proposals fell just below the funding cutoff with 561 whose proposals fell just above it. In the short run the result matched the Dutch study: the near-misses were more likely to leave the system entirely, showing about eleven per cent higher attrition in the following year and roughly a thirteen per cent chance of disappearing from NIH records altogether over the next decade. But among those who stayed, the pattern inverted. Over the following five years the near-miss group produced about twenty-one per cent more papers in the top five per cent of citations than the narrow winners, and the advantage was still around nineteen per cent in years six to ten. Early failure, for those who survived it, was associated with better subsequent work. Several explanations compete. Survivorship is the obvious one: the near-misses who persisted may have been a selected group, more determined or with better outside options, so the comparison is no longer between equivalent populations. The authors tested a range of such explanations and found the effect robust to the ones they could measure, but they could not rule out unobserved screening. A second possibility is a genuine treatment effect — that a setback prompts reappraisal, harder work, or a change of direction. A third is that the narrow winners were slowed down by what they won: a funded grant commits a young researcher to a specific programme of work for five years, which is not obviously optimal at the beginning of a career. The two studies are not in contradiction. They measure different outcomes — subsequent funding in one case, subsequent high-impact publication in the other — on different populations, in different systems. Together they suggest a more precise claim than either alone: the funding system's chief cumulative-advantage effect is on who remains in the system, not on how well those who remain perform. The Matthew effect operates most powerfully as a filter on participation. That is a more disturbing finding than a pure quality-distortion story, because attrition is invisible in the published record. We can count the papers of those who stayed. We cannot count the papers of those who left. Luck inside a career A different angle on the same question asks not how careers compare but how a single career unfolds. Roberta Sinatra and colleagues assembled the complete publication and citation records of thousands of scientists with long careers and asked where in the sequence a scientist's single highest-impact paper falls. The intuitive answers — early, in the burst of youthful creativity, or in mid-career at the peak of skill and resources — are both wrong. The highest-impact paper is located at random within the sequence. A scientist who publishes forty papers is about as likely to have their biggest hit at paper thirty-seven as at paper four. The model the authors fit to this pattern separates two components. One is productivity: how many attempts a scientist makes. The other is an individual parameter capturing a scientist's ability to take a given idea and make something of it, which they found to be roughly constant across a career and to vary substantially between individuals. Impact, in this account, is the product of a stable personal factor and a draw from a common luck distribution, repeated once per paper. If that is even approximately right, it has an uncomfortable implication for evaluation. Predicting which of a scientist's future papers will be the important one is not merely difficult; it is, within a career, impossible in principle. What can be predicted is the personal factor and the rate of attempts — which is to say, that the way to get more high-impact work out of a research system is to let capable people take more shots, not to identify in advance which shot will land. That is close to the opposite of how competitive grant systems, with their detailed prospective evaluation of specific proposed experiments, actually operate. What the mechanism does not require It is worth being precise about what is and is not being claimed, because the Matthew effect is frequently described as if it were a synonym for corruption. It requires no dishonesty. A reviewer who treats a prior grant as evidence of competence is making a defensible inference: prior grants do carry information. A reader who chooses to read the heavily cited paper first is allocating scarce attention sensibly. A department that hires from the most prestigious programmes is using a signal that is, on average, informative. Every individual step is reasonable. It requires no conspiracy. There is no coordination, no shared intent, and no identifiable decision at which the system chose to work this way. What it requires is only that (a) past outcomes are visible, (b) past outcomes are used as evidence about future quality, and (c) the evidentiary value of past outcomes is overestimated relative to the noise in how they were produced. Condition (c) is the one that turns a reasonable heuristic into a pathology, and it is the one that the evidence in this book keeps establishing. If grant decisions at the margin are near-arbitrary, then treating a grant as strong evidence of quality is a mistake — and the mistake is what converts an arbitrary event into a durable career difference. Merton's second thought Merton, who named the effect, did not think it was purely destructive, and his reasoning deserves a hearing. In a world of overwhelming information, he argued, attaching credit to recognised names performs a genuine function: it directs attention to work that is more likely to reward it, and it provides a stable currency of reputation that lets scientists judge whom to trust in fields they cannot evaluate directly. A system with no reputational shortcuts would not be more accurate; it would be paralysed. Merton also noted that eminent scientists are frequently aware of the effect and uncomfortable with it — his interviews with Nobel laureates are full of them describing how their names attracted credit they knew belonged elsewhere. The useful question is therefore not whether reputation should count but how much, and on what time scale. A reputational signal that decays — where a grant from ten years ago counts for less than one from two years ago, and where being a promising thirty-year-old confers less than being a productive forty-year-old — compounds far less than one that persists. Systems differ in this. A citation count never decays; an h-index never falls; institutional prestige is close to permanent. These are choices, not laws, and they are among the few levers that are straightforwardly adjustable. A note on eponymy There is a small, well-documented corner of this territory that makes the point with unusual economy. Scientific laws, effects and distributions are routinely named after someone other than their discoverer. The statistician Stephen Stigler proposed this as a law in its own right — that no scientific discovery is named after its original discoverer — and, in a joke that is also an argument, attributed the law to Merton, who had described the underlying dynamic first. Eponymy is credit at its most concentrated: a single name attached permanently to an idea, displacing everyone else who contributed. That it so reliably attaches to the more famous of two claimants, rather than the earlier, is a natural experiment the history of science has been running for centuries. It tells us that when the scientific community allocates a lump of indivisible credit, it does so in the direction of existing prominence — not occasionally, but as a rule strong enough to be stated as one. The cost of a compounding system There are three distinct costs, and they are often conflated. The first is unfairness to individuals, which is the one most often discussed. Careers are ended by margins of measurement error. This matters morally, and it matters practically because people observe it happening and draw conclusions about whether to enter the profession. The second is a loss of accuracy. A system that uses past allocation as evidence progressively contaminates its own evidence base. After several rounds, the signal a committee is reading is substantially composed of the outputs of earlier committees, and the correlation between the signal and the underlying quality it is meant to track degrades. This is a measurement problem, not an ethical one, and it is the reason that even a committee with no interest in fairness should care. The third and largest cost is a loss of diversity in what gets tried. If resources flow toward those who already have them, and those who already have them got them by doing work that a previous committee found promising, the portfolio of research the system supports narrows over time toward whatever the system already liked. Concentration of funding is therefore not only an equity question; it is a question about the variance of the bet the system is placing on the future. We will return to this in the chapter on money, where the evidence on whether large concentrated grants buy more discovery than small distributed ones turns out to be surprisingly clear and surprisingly ignored. For now, the point to carry forward is that the evaluation mechanisms this book examines do not operate on a blank slate. Every peer review, every funding panel, every hiring decision takes place inside a system where the inputs have already been shaped by previous decisions of the same kind. The question is how much signal those inputs actually carry — which means asking how reliable the evaluation mechanisms are in the first place. Hashtags: #Metascience #ScienceOfScience #ResearchOnResearch #ScientificEvaluation #ResearchReliability #Reproducibility #Replication #PeerReview #ResearchFunding #PublicationBias #SelectiveReporting #RegisteredReports #OpenScience #Bibliometrics #CitationAnalysis #CumulativeAdvantage #MatthewEffect #ResearchMetrics #ImpactFactor #HIndex #ResearchIncentives #ScientificCareers #CausalInferenceInMetascience #SciencePolicy #FutureOfMetascience
- Meta-Regression (Investigating Heterogeneity and Subgroup Nuances in Large Syntheses)
Download the Book (PDF): Introduction Every meta-analyst eventually meets the forest plot that refuses to behave. The confidence intervals do not overlap. A handful of studies report large benefits, a handful report none, and one or two point the other way. The pooled estimate sits in the middle like a compromise nobody signed. The heterogeneity statistics confirm what the eye already sees: the between-study variance is large, the I-squared value is high, and the prediction interval stretches across zero. At that point the analyst faces a choice that shapes everything that follows. One option is to report the average and move on. The other is to ask why the studies differ, and to reach for meta-regression. This booklet is about the second option, and about doing it honestly. Meta-regression is the natural tool for the question "what explains the spread?" It extends the familiar weighted average of meta-analysis into a weighted regression, in which each study's effect estimate is modeled as a function of study-level characteristics such as dose, duration, population age, setting, year of publication, or risk of bias. When it works, it turns an unexplained nuisance into a finding: the intervention works better at higher doses, the association is stronger in older populations, the early trials exaggerated the benefit. Those are the findings that guideline panels, policy makers and trial designers actually want. The trouble is that meta-regression is also one of the easiest places in evidence synthesis to manufacture a false discovery. The number of observations is the number of studies, not the number of participants, and in most syntheses that number is small. A review with thirty trials and fifty thousand patients has thirty data points for the purposes of meta-regression. The candidate explanatory variables are usually numerous, often correlated with one another, and rarely chosen before the data were seen. The outcome variable, an effect estimate, arrives with its own sampling error, which must be accounted for carefully. The covariates are summaries of groups rather than measurements of individuals, which opens the door to ecological bias. And the standard software output reports p-values that, under the default settings of many programs, are too small when the number of studies is modest. Each of these problems is well documented. Together they explain why so many published moderator findings fail to replicate when later syntheses add more studies. The controlling idea The argument of this booklet can be put in a sentence. Meta-regression is a legitimate and powerful way to investigate heterogeneity, but only when it is treated as a small-sample, observational analysis of study-level data, which means allowing for residual heterogeneity, using inference methods calibrated for few studies, limiting and pre-specifying the questions asked, and interpreting any association as a hypothesis about why studies differ rather than a proof of why patients respond differently. Each clause in that sentence corresponds to a technical decision. "Allowing for residual heterogeneity" means preferring random-effects (strictly, mixed-effects) meta-regression over fixed-effect meta-regression in almost every realistic situation, and understanding exactly why the fixed-effect version over-states certainty. "Inference methods calibrated for few studies" means the Knapp-Hartung adjustment, permutation tests and, where effect sizes are dependent, small-sample corrected robust variance estimation. "Limiting and pre-specifying" means taking seriously the multiplicity problem that Higgins and Thompson quantified two decades ago. "Study-level data" means respecting the difference between an across-study association and a within-study effect modifier. None of these ideas is new. What is still common is to see them applied piecemeal, or to see them described in methods sections and then ignored in the interpretation. Who this is for The intended reader is a practising meta-analyst: an epidemiologist, clinical researcher, psychologist, education researcher, ecologist or health technology assessor who already knows how to run a basic random-effects meta-analysis and has now hit a wall of heterogeneity. You do not need to be a statistician, but you should be comfortable with the idea of a weighted average, a regression coefficient and a standard error. Where software is shown, it is R with Wolfgang Viechtbauer's metafor package, because metafor implements essentially every method discussed here in a consistent way and is widely used across disciplines. The code snippets are short and meant to illustrate the decision being made, not to replace the package documentation. Readers who work in Stata, where the meta regress suite and the community-contributed metareg command by Harbord and Higgins provide similar functionality, or in other environments, will find the logic transfers directly. Throughout, the booklet uses a small number of hypothetical worked examples. The most frequent is an imagined synthesis of randomized trials of a supervised exercise program for depressive symptoms, reported as standardized mean differences, with candidate moderators such as weekly exercise minutes, mean participant age, whether sessions were supervised, and risk of bias. The numbers in these examples are invented for illustration and are labelled as such. They are chosen to be realistic in scale, so that the arithmetic of heterogeneity and the behavior of different tests can be seen concretely, but they do not describe any real body of evidence. How the book is organized The first chapter sets out what heterogeneity is and how it is measured, because a surprising amount of confusion in meta-regression flows from misreading the heterogeneity statistics that motivate it. Cochran's Q, I-squared, the between-study variance tau-squared and the prediction interval answer different questions, and only some of them tell you whether there is anything worth explaining. The second chapter builds the meta-regression model from the ground up, starting with subgroup analysis, which is simply meta-regression with a categorical covariate. It covers how to code moderators, how to read the coefficients, what residual heterogeneity means, and what the pseudo-R-squared statistic can and cannot tell you. The third chapter addresses the central modeling decision named in this book's subtitle: fixed-effect versus random-effects meta-regression. It explains why a fixed-effect meta-regression, which assumes that the covariates explain all between-study variation, produces standard errors that are too small whenever that assumption fails, which is nearly always. It also covers the estimation of residual between-study variance and the choice among estimators. The fourth chapter deals with inference when studies are few. The standard Wald-type z-test in a random-effects meta-regression ignores the uncertainty in the estimated between-study variance and consequently rejects the null hypothesis too often. The Knapp-Hartung method corrects this, with some subtleties that matter in practice. The fifth chapter is devoted to permutation tests, following the approach Higgins and Thompson proposed in 2004 for controlling spurious findings, including the use of permutation to adjust for the testing of multiple candidate covariates. The sixth chapter concerns diagnostics, and in particular the bubble plot: the scatter plot of effect estimates against a covariate, with each study drawn as a circle sized by its weight. It explains, in words rather than pictures, how to read one well, what patterns should alarm you, and how influence and outlier diagnostics complement it. The seventh chapter confronts the interpretive problems that no statistical adjustment can fix: ecological bias, confounding between study characteristics, and the special trap of regressing treatment effects on baseline risk. The eighth chapter extends the model to the settings that large syntheses actually produce: multiple effect sizes per study, several moderators at once, continuous covariates with non-linear relationships, and the risk of overfitting. The ninth chapter pulls these threads together into a protocol for credible moderator analysis, drawing on the Cochrane Handbook's guidance, published criteria for judging subgroup claims, and reporting standards. A short conclusion argues for what follows from all of this for the way meta-analysts plan, conduct and report their work. A note on terminology Terminology in this area is inconsistent across disciplines, and the inconsistency causes real errors. In medicine, "fixed-effect" (singular) usually refers to the model that assumes a single common true effect, while "random-effects" allows true effects to vary. In the social sciences, "fixed-effects" (plural) sometimes denotes a model that conditions on the included studies without assuming a common effect. The metafor package now distinguishes an "equal-effects" model from a "fixed-effects" model for this reason. A meta-regression that includes covariates and a between-study variance term is properly called a mixed-effects model, because the covariate coefficients are fixed parameters while the residual study effects are random; many authors, including the Cochrane Handbook, simply call it random-effects meta-regression. This booklet uses "fixed-effect meta-regression" for the model with no residual between-study variance and "random-effects meta-regression" or "mixed-effects meta-regression" for the model that includes it. The words "moderator", "covariate", "effect modifier" and "explanatory variable" are used for the study-level characteristics entered into the model, with the caution, developed in the seventh chapter, that a study-level moderator is not necessarily a patient-level effect modifier. The phrase "heterogeneity" is used in its statistical sense, meaning variation in true effects across studies beyond what sampling error would produce. Clinical and methodological diversity, meaning differences in populations, interventions, outcomes and designs, is the raw material that may or may not produce statistical heterogeneity. Meta-regression is the bridge between the two: it asks whether measured diversity accounts for observed heterogeneity. The rest of this booklet is about building that bridge so that it bears weight. Chapter 1: What Heterogeneity Is and When It Deserves Explanation Meta-regression begins with a diagnosis. Before asking what explains the variation among studies, the analyst has to establish that there is variation to explain, estimate how much there is, and decide whether it matters for the conclusions. Those three tasks are often collapsed into a glance at a single number, usually I-squared, and the collapse is the source of a surprising number of downstream mistakes. This chapter separates them. Two sources of spread in a forest plot Imagine a set of k studies, each reporting an estimate of the same kind of effect: a log odds ratio, a standardized mean difference, a correlation. Write the estimate from study i as y_i and its within-study sampling variance as v_i, which the analyst treats as known. The simplest model says that every study is estimating the same true effect, theta, and that the estimates differ only because each study sampled a finite number of participants. Under that model, y_i is theta plus a sampling error with variance v_i. The spread of points in the forest plot is then entirely explained by the width of the individual confidence intervals. The random-effects model adds a second layer. Each study has its own true effect, theta_i, and those true effects are drawn from a distribution with mean mu and variance tau-squared. The observed estimate is theta_i plus sampling error. Now the spread in the forest plot has two components: the within-study sampling variance v_i, which differs from study to study and shrinks as sample size grows, and the between-study variance tau-squared, which is common to all studies and does not shrink no matter how large each study becomes. This distinction is the foundation of everything in this booklet. Heterogeneity, in the statistical sense, is tau-squared: the variance of the true effects. It is a property of the population of studies, not of the precision with which they were conducted. Meta-regression is an attempt to replace some of that unexplained variance with a systematic account, by letting the mean of the distribution of true effects depend on study characteristics. If tau-squared is zero, there is nothing for meta-regression to explain. If tau-squared is large, there may be a great deal, although whether the available covariates can explain it is a separate question. The four statistics and what each one answers Four quantities are routinely reported to describe heterogeneity. They are frequently treated as interchangeable, and they are not. Cochran's Q is the weighted sum of squared deviations of each study estimate from the fixed-effect pooled estimate, with weights equal to the inverse of the within-study variances. Under the hypothesis that all studies share one true effect, Q follows approximately a chi-squared distribution with k minus 1 degrees of freedom, so it yields a test of homogeneity. The test has low power when studies are few or small, which is why some analysts use a significance level of 0.10 rather than 0.05, and very high power when studies are many or large, in which case it will detect trivial heterogeneity. A non-significant Q is therefore weak evidence of homogeneity, and a significant Q says nothing about whether the heterogeneity is large enough to matter. I-squared, introduced by Higgins and Thompson in 2002 and popularized in a widely read BMJ paper the following year, re-expresses Q as the proportion of total variation in the estimates that is attributable to between-study heterogeneity rather than sampling error. Computed as (Q minus its degrees of freedom) divided by Q, truncated at zero, it has the appealing property of not depending directly on the number of studies. It is, however, a relative measure. It compares tau-squared to the typical within-study variance. Two meta-analyses with identical between-study variance can have very different I-squared values if one consists of large, precise trials and the other of small, imprecise ones. As trials grow larger, I-squared rises toward 100 percent even when the absolute variation in true effects is unchanged. Rücker and colleagues made this point forcefully in 2008 under the title "Undue reliance on I2 in assessing heterogeneity may mislead." I-squared tells you how much of what you see is signal rather than noise; it does not tell you how much the true effects vary. Tau-squared, and its square root tau, is the direct estimate of how much the true effects vary, expressed on the scale of the effect measure. A tau of 0.30 for standardized mean differences means that the true effects have a standard deviation of about 0.30 around their mean. That is an absolute, interpretable quantity. Its weakness is that it is estimated imprecisely when k is small, and confidence intervals for tau-squared (available through the Q-profile method among others) are typically very wide in syntheses of fewer than twenty or so studies. The prediction interval, advocated by Riley, Higgins and Deeks in the BMJ in 2011 and now recommended in the Cochrane Handbook, combines the uncertainty about the mean effect with the estimated between-study variance to give a range within which the true effect in a new study, similar to those included, would be expected to lie. It is the most clinically meaningful summary of heterogeneity because it answers the question a decision maker actually asks: if this intervention is implemented somewhere new, what effect should we expect? Table 1 summarizes the differences. Table 1. What the common heterogeneity statistics measure. Statistic Question it answers Scale Main limitation Cochran's Q Is there evidence of any heterogeneity? Test statistic Power depends on number and size of studies I-squared What share of observed spread is not sampling error? Percentage Rises with study precision; relative, not absolute Tau-squared (tau) How much do true effects vary? Effect measure (squared) Imprecise when studies are few Prediction interval Where would a new study's true effect lie? Effect measure Assumes normal true effects; wide with few studies A worked example of reading the statistics together Suppose a hypothetical synthesis includes 26 randomized trials of a supervised exercise program for adults with depressive symptoms, each reporting a standardized mean difference in symptom scores at the end of treatment, with negative values favoring exercise. The random-effects pooled estimate is minus 0.45, with a standard error of 0.07. Cochran's Q is 78 on 25 degrees of freedom. I-squared is therefore (78 minus 25) divided by 78, or about 68 percent. The restricted maximum likelihood estimate of tau-squared is 0.09, so tau is 0.30. The pooled estimate and its confidence interval, roughly minus 0.59 to minus 0.31, suggest a moderate, reliably non-zero average benefit. The prediction interval tells a different story. Using the approximation recommended by Higgins, Thompson and Spiegelhalter, with a t distribution on k minus 2 degrees of freedom, the interval is minus 0.45 plus or minus 2.064 times the square root of (0.09 plus 0.0049), which is minus 0.45 plus or minus 0.64, giving roughly minus 1.09 to plus 0.19. The true effect in a new setting could be large and beneficial, or it could be close to zero or even slightly unfavorable. That gap between the confidence interval and the prediction interval is the practical definition of heterogeneity that matters. The average effect is well estimated; the effect you would get in any particular place is not. The question that motivates meta-regression is whether some of that spread can be tied to features of the trials. If the high-benefit trials are the ones with more supervised minutes per week, for example, that would be both scientifically interesting and practically useful. In metafor, the basic model and its heterogeneity summaries come from a few lines. library(metafor) dat <- escalc(measure = "SMD", m1i = m_ex, sd1i = sd_ex, n1i = n_ex, m2i = m_ctl, sd2i = sd_ctl, n2i = n_ctl, data = trials) res0 <- rma(yi, vi, data = dat, method = "REML") res0 # Q, I^2, tau^2, H^2 confint(res0) # Q-profile intervals for tau^2 and I^2 predict(res0) # includes the prediction interval The confint call deserves a moment of attention. In the hypothetical example, a 95 percent interval for tau-squared might run from something like 0.04 to 0.22. That width is typical. It means the analyst does not know with any precision how heterogeneous the trials are, only that they are heterogeneous. This uncertainty will reappear in the fourth chapter, because failure to account for it is the main reason standard meta-regression tests are too liberal. Heterogeneity as a finding, not a defect There is a tendency to treat heterogeneity as a flaw in a meta-analysis, something to be apologized for in the limitations section or, worse, reduced by excluding the studies that disagree. This is a mistake of framing. Heterogeneity is information. It tells you that the effect of an intervention or the strength of an association depends on something. The Cochrane Handbook, in its chapter on analysing data and undertaking meta-analyses (Chapter 10 in the version 6 series, authored by Deeks, Higgins and Altman), lists among the options for dealing with heterogeneity: checking the data for errors, not doing a meta-analysis at all, exploring heterogeneity, ignoring it (by fitting a fixed-effect model, which it does not recommend as a response to heterogeneity), incorporating it through a random-effects model, changing the effect measure, and excluding studies. Of these, exploring heterogeneity through subgroup analysis and meta-regression is the one that can turn the problem into knowledge. It is worth dwelling on two of the other options, because they bear on meta-regression. The first is changing the effect measure. Heterogeneity is scale-dependent. Trials of a treatment that produces a roughly constant relative risk reduction will show heterogeneous risk differences if their baseline risks differ, and vice versa. Before building a meta-regression to explain heterogeneity in risk differences, it is worth asking whether the heterogeneity largely disappears on the ratio scale. Relative measures are generally more consistent across studies than absolute ones, which is one reason they are the default for binary outcomes. A meta-regression on the wrong scale can find a spurious moderator whose only role is to proxy for baseline risk. The second is checking the data. A study that appears to be a wild outlier in a forest plot is, surprisingly often, the product of a data extraction error: a standard error entered as a standard deviation, a change score confused with a final score, an outcome coded in the wrong direction. A standardized mean difference computed from a standard error rather than a standard deviation will be inflated by a factor equal to the square root of the group size, which can easily turn a modest effect into an implausibly large one. Any meta-regression is only as good as the effect estimates it regresses, and extreme estimates exert disproportionate influence on regression slopes. A worked illustration of scale dependence The point about scale is easy to accept in the abstract and easy to forget in practice, so it is worth making concrete. Suppose, hypothetically, that a preventive treatment reduces the risk of an event by exactly one quarter in every population, a constant risk ratio of 0.75. Three trials are run in populations with control-group risks of 4 percent, 12 percent and 30 percent. On the risk ratio scale, the three trials estimate the same true effect, and any spread among them is sampling error. On the risk difference scale, the true effects are 1, 3 and 7.5 percentage points: a sevenfold range. A meta-analysis of risk differences would find substantial heterogeneity, and a meta-regression of risk differences on baseline risk, or on any characteristic correlated with it, such as mean age or disease severity, would find a strong moderator. That moderator would not be wrong, exactly. The absolute benefit really is larger in higher-risk populations, and that fact matters for decisions, since the number needed to treat depends on it. But the analyst who reported that "the treatment works better in older patients" on the strength of such an analysis would be describing a consequence of arithmetic rather than a biological interaction. The appropriate approach is usually to pool on the relative scale, where effects are more stable, and then to apply the pooled relative effect to the baseline risks of the populations of interest to derive absolute effects. The Cochrane Handbook recommends this route for presenting absolute effects in summary tables. The reverse can also happen. If a treatment produces a roughly constant absolute benefit across populations, relative measures will appear heterogeneous. Which scale is more stable is an empirical question, and for binary outcomes the evidence across many meta-analyses generally favors relative measures. The practical rule is that before modeling heterogeneity on one scale, the analyst should check whether it largely disappears on another, and should choose the scale for the meta-regression with that in mind. When heterogeneity deserves a meta-regression Not every heterogeneous meta-analysis warrants meta-regression. Four conditions make it worthwhile. The first is that heterogeneity is substantively important, which is best judged from tau and the prediction interval rather than from I-squared or the Q test. A tau of 0.05 on the standardized mean difference scale is unlikely to change any decision, even if I-squared is 70 percent because the trials are enormous. A tau of 0.30 around a mean of 0.45 plainly matters. The second is that there are enough studies. The Cochrane Handbook advises that meta-regression should generally not be considered when there are fewer than ten studies in a meta-analysis, and that guidance is widely interpreted as a floor of roughly ten studies per characteristic examined. Ten is not a magic number, and later chapters will show that even with twenty or thirty studies, inference requires care, but below ten the analysis is almost always too underpowered to detect realistic moderator effects and too unstable to produce credible estimates. The third is that there are candidate explanations with a plausible rationale, specified before the data were examined. Thompson and Higgins, in their 2002 paper on how meta-regression analyses should be undertaken and interpreted, put this first among their recommendations. The value of a moderator finding depends heavily on whether it was hypothesized in advance. A protocol that lists three pre-specified moderators with reasons, and reports all three results whatever they show, produces far more credible evidence than a search through fifteen extracted variables for the one that reaches significance. The fourth is that the covariates are measured with reasonable consistency across studies. Dose, duration, and design features are often reported well. Participant characteristics such as mean age or proportion female are reported as summaries and carry the ecological problems discussed in the seventh chapter. Some characteristics, such as the "intensity" of a behavioral intervention, are coded by reviewers with considerable judgment, and coding reliability should be established before the variable enters any model. When these conditions are not met, the honest course is usually to report the heterogeneity, present the prediction interval, describe the clinical and methodological diversity narratively, and refrain from a formal moderator analysis, or to present one explicitly labelled as exploratory with the limitations spelled out. The limits of what can be explained One further point frames the rest of the booklet. Even when meta-regression is well conducted, it can explain heterogeneity only in terms of variables that vary across studies and are reported by them. Many important sources of heterogeneity are invisible at the study level: differences in how an outcome scale was administered, in the skill of therapists, in co-interventions, in the local population's underlying biology. Some of the variation in true effects will remain unexplained in almost every synthesis, and that residual heterogeneity is not a failure of the analysis. It is an honest acknowledgment that the studies differ in ways the meta-analyst cannot see. This matters because the most consequential error in meta-regression is behaving as if the covariates have explained everything. That assumption is precisely what fixed-effect meta-regression makes, and the third chapter shows what it costs. Before turning to that choice, though, the second chapter sets out the model itself, beginning with its simplest and most familiar form: the subgroup analysis. A short checklist before modeling In practice, the steps before fitting any meta-regression can be written as a short sequence. It is included here because skipping any one of them is a common source of later embarrassment. Verify the extracted effect estimates and variances against the source papers for any study that appears extreme. Confirm that the effect measure is appropriate and consider whether heterogeneity is driven by the choice of scale. Estimate tau-squared with a recommended estimator and report a confidence interval for it. Compute and report the prediction interval alongside the pooled estimate. Judge the practical importance of heterogeneity from tau and the prediction interval, not from I-squared alone. Count the studies that report each candidate covariate, since missing covariate data reduce the effective k. Return to the protocol and list the moderators that were specified in advance, with their rationale. Only after these steps does the modeling itself begin. Chapter 2: From Subgroups to Meta-Regression Most meta-analysts encounter heterogeneity analysis first as subgroup analysis: split the studies by some characteristic, pool each group separately, and compare. Meta-regression can look like a more advanced, separate technique. It is not. A subgroup analysis is a meta-regression with a categorical covariate, and seeing the connection clearly makes both easier to do well. This chapter builds the model from that starting point, explains how to read its output, and sets out what residual heterogeneity and the pseudo-R-squared statistic mean. Subgroup analysis done properly Take the hypothetical exercise synthesis from the previous chapter and suppose that 14 of the 26 trials delivered sessions under direct supervision and 12 provided an unsupervised home program. The analyst pools each group. The supervised trials give a random-effects estimate of minus 0.58; the unsupervised trials give minus 0.31. The first estimate has a confidence interval excluding zero; so, narrowly, does the second. The tempting conclusion is that supervised exercise works better. The correct question is whether the difference between the two subgroup estimates, minus 0.27, is larger than would be expected by chance. That requires a test of the difference, not a comparison of the two separate significance tests. The most common fallacy in subgroup analysis is to observe that the effect is statistically significant in one subgroup and not in another and to conclude that the subgroups differ. That inference is invalid. A subgroup with fewer or smaller studies can fail to reach significance while having exactly the same true effect as its counterpart. The only valid evidence of a difference is a direct test of the difference, sometimes called a test for subgroup differences or a test for interaction. The standard test for subgroup differences partitions heterogeneity. The total Q statistic across all studies is split into the sum of the within-subgroup Q statistics and a between-subgroup Q statistic, which, under the null hypothesis of no difference, follows a chi-squared distribution with degrees of freedom equal to the number of subgroups minus one. Review software such as Cochrane's RevMan reports this test for random-effects subgroup analyses by comparing the subgroup pooled estimates using their random-effects standard errors. A subtle but consequential choice in random-effects subgroup analysis is whether to assume a common tau-squared across subgroups or to estimate a separate tau-squared within each. Separate estimates are the default in RevMan and in many applications, and they allow the supervised trials to be more or less heterogeneous than the unsupervised ones. A common tau-squared is what a meta-regression with a categorical covariate assumes, and Borenstein and colleagues, in their textbook "Introduction to Meta-Analysis," recommend pooling the within-subgroup estimates of tau-squared unless there is reason to expect different amounts of heterogeneity and enough studies in each subgroup to estimate them. Their reasoning is practical: with a dozen studies per subgroup, each separate estimate of tau-squared is very imprecise, and the subgroup with the smaller estimate, often by chance, receives an artificially narrow confidence interval. When the subgroups are small, a pooled tau-squared produces more stable comparisons. The same analysis as a regression Now write the subgroup analysis as a regression. Create an indicator variable, supervised, equal to 1 for supervised trials and 0 otherwise. The mixed-effects meta-regression model says that the observed effect in study i is: y_i = b0 + b1 * supervised_i + u_i + e_i where e_i is the sampling error with known variance v_i, and u_i is the study's residual random effect with variance tau-squared, now interpreted as the between-study variance not explained by the covariate. The coefficient b0 is the expected effect in unsupervised trials; b1 is the difference in expected effect between supervised and unsupervised trials. Fitting this model with a common residual tau-squared gives exactly the random-effects subgroup comparison with pooled tau-squared, and the test of b1 equal to zero is the test for subgroup differences. res_sub <- rma(yi, vi, mods = ~ supervised, data = dat, method = "REML") res_sub The same model can be fitted without an intercept, which yields the two subgroup means directly rather than one mean and a difference. The fit is identical; only the parameterization changes. res_cells <- rma(yi, vi, mods = ~ factor(supervised) - 1, data = dat) Once subgroup analysis is seen as regression, three generalizations follow naturally. The covariate can have more than two categories, represented by several indicator variables, with an omnibus test of whether any category differs. The covariate can be continuous, such as weekly minutes of exercise. And several covariates can enter the model at once, which allows the analyst to ask whether supervision still matters after accounting for dose. These are the capabilities that give meta-regression its value, and also, as later chapters will make clear, the capabilities that multiply its risks. Continuous covariates and why categorizing them is costly Suppose each trial reports its prescribed weekly minutes of supervised exercise, ranging from 60 to 240. A common approach is to dichotomize at some cut-point, perhaps 150 minutes, and run a subgroup analysis. This throws information away. Trials at 140 and 160 minutes are treated as maximally different while trials at 60 and 140 are treated as identical. It also invites the analyst to try several cut-points and report the one that "works," which is a form of multiple testing that rarely appears in the methods section. Treating the covariate as continuous uses all the information and yields a slope: the expected change in effect per unit change in the covariate. In metafor: dat$min60 <- dat$minutes / 60 # rescale: one unit = 60 minutes per week res_dose <- rma(yi, vi, mods = ~ min60, data = dat, method = "REML") Rescaling is not cosmetic. A slope expressed per minute will be a tiny number that is hard to interpret; a slope per 60 minutes, or per 10 years of mean age, is readable. Centering the covariate, by subtracting a meaningful value such as the median or a clinically relevant reference dose, makes the intercept interpretable as the expected effect at that value rather than at zero minutes, which may lie far outside the observed range. Centering matters even more when interaction terms are included, because it changes what the main-effect coefficients mean. Suppose the hypothetical dose model gives a slope of minus 0.12 per 60 minutes, with a standard error of 0.05 under the default Wald-type test, and an intercept (at the median of 150 minutes after centering) of minus 0.46. The model predicts an effect of about minus 0.34 at 90 minutes and about minus 0.58 at 210 minutes. That is a meaningful gradient, if it is real. The chapters that follow are, in large part, about how to tell. Linearity is an assumption, not a given. Dose-response relationships frequently plateau, and a straight line fitted to a curved relationship can misstate both the size and the shape of the effect. With few studies, the data rarely support much flexibility, but the eighth chapter discusses restricted cubic splines and simple checks for non-linearity that are feasible in larger syntheses. Categorical moderators with several levels Many moderators have more than two categories. Suppose the exercise trials were delivered in three settings: clinics, community facilities and participants' homes. The meta-regression represents setting with two indicator variables, one for community and one for home, with clinic as the reference category. The intercept is then the expected effect in clinic-based trials, and each of the other two coefficients is the difference between that setting and the clinic. The choice of reference category does not change the fit of the model or the omnibus test, but it determines which comparisons are reported directly, and it is easy to misread the output as a set of independent tests. The first question, whether setting matters at all, is answered by the omnibus test of both indicator coefficients together, which has two degrees of freedom. Individual comparisons, such as community versus home, which does not appear directly when clinic is the reference, can be obtained as linear contrasts. res_set <- rma(yi, vi, mods = ~ factor(setting), data = dat, method = "REML", test = "knha") anova(res_set, btt = 2:3) # omnibus test: does setting matter? anova(res_set, X = c(0, -1, 1)) # contrast: home versus community With three or more categories and a modest number of studies, some categories will usually contain very few trials, and their estimates will rest heavily on those few. Categories should be defined in the protocol, and sparse categories combined according to a pre-specified rule rather than after seeing which combination produces a difference. Pairwise comparisons among several categories multiply the number of tests, and the omnibus test should normally come first: if it gives no evidence that setting matters, pursuing individual pairwise contrasts is a search for noise. Allowing heterogeneity to differ between subgroups The common-tau-squared assumption of the standard meta-regression can be relaxed when there is reason to think that some kinds of study are more variable than others. In the exercise example, unsupervised home programs might plausibly produce more variable effects than supervised ones, because adherence varies more when nobody is watching. A model that estimates separate residual variances for the two subgroups tests that idea directly, and metafor's location-scale formulation handles it within a single model. res_ls <- rma(yi, vi, mods = ~ supervised, scale = ~ supervised, data = dat) The scale part of the output reports whether the log of the residual variance differs between subgroups. With fourteen and twelve trials, that test will have little power, and the separate variance estimates will be imprecise, which is exactly Borenstein and colleagues' reason for preferring a pooled estimate in small analyses. The model is best used when the difference in variability is itself a question of interest, or as a sensitivity analysis to check that the comparison of means is not distorted by assuming equal variances. Coefficients on ratio scales For binary outcomes, meta-regression is usually fitted on the log odds ratio or log risk ratio scale, and the coefficients are therefore differences in log ratios. They are easier to communicate after exponentiation. A slope of minus 0.10 on the log odds ratio scale per 10 years of mean age becomes, when exponentiated, a ratio of odds ratios of about 0.90: for each additional 10 years of mean age, the odds ratio is multiplied by 0.90. If the odds ratio at a mean age of 50 is 0.80, the model predicts about 0.72 at 60 and about 0.65 at 70. Ratios of odds ratios are unfamiliar to many readers, and predicted odds ratios at specific covariate values are almost always clearer. The predict() function in metafor accepts a transformation argument, so that predictions and their intervals can be reported on the ratio scale directly. Confidence intervals must be computed on the log scale and then transformed, never by exponentiating a point estimate and adding a symmetric margin. The same logic applies to prediction intervals, which on the ratio scale are asymmetric and often strikingly wide at the upper end. predict(res_age, newmods = c(50, 60, 70), transf = exp) Reading the output The printed output of a mixed-effects meta-regression contains more than a table of coefficients, and each part answers a different question. The test of moderators, labelled QM in metafor, is an omnibus test of whether the covariates, taken together, explain any variation in effects. With a single covariate it is simply the square of the Wald z statistic for that coefficient. With several, it asks whether any of them matter. Under the Knapp-Hartung adjustment discussed in the fourth chapter, it becomes an F test. The test for residual heterogeneity, labelled QE, asks whether there is variation left over after accounting for the covariates. It is the analogue of Cochran's Q for the regression residuals, and it follows a chi-squared distribution with k minus p degrees of freedom, where p is the number of coefficients including the intercept. A significant QE means the covariates have not explained all the between-study variation, which in practice is the usual finding. The residual tau-squared is the estimated between-study variance that remains after the covariates are included. It is, conceptually, the variance of the true effects around the regression line rather than around a single mean. The residual I-squared in metafor's output expresses that residual variance relative to the typical within-study variance, with the same caveats as ordinary I-squared. The pseudo-R-squared, which metafor labels as the amount of heterogeneity accounted for, is the proportional reduction in tau-squared from the model without covariates to the model with them. If the intercept-only model estimates tau-squared at 0.09 and the dose model estimates the residual tau-squared at 0.065, the pseudo-R-squared is (0.09 minus 0.065) divided by 0.09, or about 28 percent. What the pseudo-R-squared can and cannot tell you The pseudo-R-squared is intuitive and widely reported, and it is also one of the least reliable statistics in a meta-regression. Because it is a ratio of two noisy estimates of variance, it inherits and compounds their imprecision. A simulation study by López-López and colleagues, published in the British Journal of Mathematical and Statistical Psychology in 2014, examined how well this statistic estimates the true proportion of explained heterogeneity and found that its accuracy depends heavily on the number of studies: with the numbers typical of applied syntheses, and particularly with fewer than about twenty to forty studies, the estimates can be badly off and highly variable. It is truncated at zero, so a covariate with no explanatory power frequently shows zero, and one with modest power can show anything from zero to a large value by chance. It can also be misleading when the within-study variances are correlated with the covariate, because adding a covariate changes the weights. The practical advice is to report the pseudo-R-squared if it is conventional in the field, but never to treat it as a precise measure, and never to base a claim about the importance of a moderator primarily on it. A better way to convey what a moderator explains is to show the predicted effects at clinically meaningful values of the covariate, with their confidence intervals and prediction intervals, and to report how the residual tau compares to the original tau. In the hypothetical example, residual tau is about 0.25 against an original 0.30. The dose gradient is interesting, but most of the spread in true effects remains unexplained, and a new trial at any given dose should still be expected to vary substantially. Weights in meta-regression One feature that surprises newcomers is how the weights work. In a random-effects meta-analysis, each study is weighted by the inverse of its within-study variance plus tau-squared. In a mixed-effects meta-regression, the weights are the inverse of the within-study variance plus the residual tau-squared. Because the residual tau-squared is usually smaller than the unconditional tau-squared, adding a covariate that explains some heterogeneity shifts the weights back toward the larger studies. When residual tau-squared is large relative to the within-study variances, the weights become nearly equal, and the meta-regression behaves much like an unweighted regression of effect sizes on covariates. When residual tau-squared is near zero, the weights approach those of a fixed-effect analysis and a few large studies can dominate. This has consequences for diagnostics. A single large study with an extreme covariate value can pull the regression line toward itself when residual heterogeneity is small, and the bubble plot discussed in the sixth chapter is the primary way to see this happening. Missing covariate data A last practical matter concerns missing covariates. If only 19 of the 26 trials report mean participant age, a meta-regression on age uses only those 19 trials, and the results are not directly comparable to the intercept-only model on all 26. The pseudo-R-squared, in particular, should be computed against an intercept-only model fitted to the same 19 trials, which metafor does automatically when a covariate is missing but which is easy to get wrong in hand calculations or in comparisons across separately fitted models. More importantly, the trials that do not report a covariate may differ systematically from those that do. Older trials often report less detail, for example, and if older trials also had larger effects, the subset analysis is biased. Multiple imputation of missing study-level covariates is possible, but with small numbers of studies it rests heavily on assumptions, and the more common and defensible approach is to report the number of studies contributing to each analysis and to consider whether missingness itself is associated with the effect. With the model and its output in hand, the next question is the one that most sharply divides good meta-regression from misleading meta-regression: whether to include the residual between-study variance at all. Chapter 3: Fixed-Effect Versus Random-Effects Meta-Regression The single most consequential modeling decision in a meta-regression is whether to include a term for residual between-study variance. Leave it out, and you have a fixed-effect meta-regression, which assumes that the covariates in the model account for every difference among the true study effects. Include it, and you have a random-effects, or mixed-effects, meta-regression, which assumes that the covariates account for some of the variation and that the rest is random. The choice sounds technical. In practice it often determines whether a moderator is declared significant, and the wrong choice is a leading cause of false-positive findings in published syntheses. What the fixed-effect model assumes The fixed-effect meta-regression model writes each observed effect as a linear function of covariates plus sampling error alone: y_i = b0 + b1 * x_i + e_i, with the variance of e_i equal to v_i. The coefficients are estimated by weighted least squares with weights equal to 1 / v_i, the inverse of each study's within-study variance. The standard errors of the coefficients are computed on the assumption that the only source of scatter around the regression line is the known sampling error. If that assumption is right, the method is efficient and its tests have the stated error rates. The assumption is strong. It says that once you know a trial's weekly exercise minutes, you know its true effect exactly; any remaining difference between the observed effect and the regression line is pure sampling noise. For that to hold, every other source of variation among trials, including population, outcome measure, therapist, country, co-interventions, attrition and risk of bias, must either be absent or be perfectly captured by the covariate. In a synthesis of real studies conducted by different teams in different places, this is almost never plausible. The residual heterogeneity test, QE, usually confirms it: after adding a covariate, significant variation typically remains. Why ignoring residual heterogeneity produces false positives When residual heterogeneity exists but the model ignores it, the observed effects scatter around the regression line more than the within-study variances predict. The fixed-effect standard errors are computed as if the scatter were smaller than it is, so they are too small. The test statistic for the slope is the slope divided by its standard error, so it is too large. The p-value is too small. The result is a test that rejects the null hypothesis of no moderator effect far more often than its nominal 5 percent when there is in fact no moderator effect. The problem is worst precisely when meta-analysts are most tempted by it. If the included studies are large, their within-study variances are small, and the fixed-effect standard errors shrink correspondingly. Yet the between-study variance does not shrink with sample size. A synthesis of big trials can therefore have fixed-effect standard errors for a moderator that are a fraction of the true uncertainty. Any covariate that happens, by chance, to line up with the unexplained spread among those trials will appear highly significant. Thompson and Sharp, in a 1999 comparison of methods for explaining heterogeneity published in Statistics in Medicine, worked through this problem with examples and concluded that fixed-effect meta-regression is inappropriate when residual heterogeneity is present, because it can yield spuriously significant associations. Thompson and Higgins reiterated the point in 2002, recommending that meta-regression should allow for residual heterogeneity by using a random-effects formulation. Higgins and Thompson's 2004 simulation study, discussed in detail in the fifth chapter, showed how badly the false-positive rate of fixed-effect meta-regression can be inflated when true effects vary. The Cochrane Handbook's guidance follows this literature. A worked contrast Return to the hypothetical exercise synthesis. Suppose the fixed-effect meta-regression of effect size on weekly minutes (per 60 minutes) gives a slope of minus 0.13 with a standard error of 0.028. The z statistic is about minus 4.6, and the p-value is well below 0.001. The analyst might write that dose "strongly predicts" benefit. The random-effects meta-regression of the same data gives a slope of minus 0.12 with a standard error of 0.05, a z statistic of about minus 2.4 and a p-value of about 0.016. The point estimate barely moves; the standard error nearly doubles. With the Knapp-Hartung adjustment described in the next chapter, the standard error might grow further to about 0.062, the t statistic on 24 degrees of freedom falls to about minus 1.94, and the p-value rises to about 0.06. The same data, then, have been reported as overwhelming evidence, as moderate evidence, and as borderline evidence, depending on assumptions. Only the latter two acknowledge that trials with the same exercise dose still differ in their true effects, and the difference in apparent certainty between the first and the third comes almost entirely from ignoring that fact. The fixed-effect standard error of 0.028 answers the question "how uncertain would the slope be if dose explained everything?" That is not the question anyone should be asking. res_fe <- rma(yi, vi, mods = ~ min60, data = dat, method = "EE") # equal/fixed effect res_re <- rma(yi, vi, mods = ~ min60, data = dat, method = "REML") res_kh <- rma(yi, vi, mods = ~ min60, data = dat, method = "REML", test = "knha") Note that metafor now distinguishes method = "EE" (equal-effects) from method = "FE" (fixed-effects); for meta-regression the computations of the coefficients and standard errors coincide, and the distinction concerns interpretation, discussed below. A trap in general-purpose software A related error arises when analysts fit meta-regressions with general regression software, such as R's lm() function or a weighted regression procedure in a statistics package, using weights equal to 1 / v_i. These routines treat the weights as relative rather than absolute. They estimate a residual variance from the data and scale all the standard errors by it. The resulting model assumes that each study's total variance is proportional to its within-study variance, multiplied by a common factor estimated from the residuals. Thompson and Sharp described this as a multiplicative model of heterogeneity, in contrast to the additive model of random-effects meta-regression, in which the between-study variance is added to each study's within-study variance. The multiplicative model is not absurd, and it is in fact close in spirit to the Knapp-Hartung adjustment, which also rescales standard errors using the residual scatter. But when used naively it has two defects. First, the weights remain proportional to 1 / v_i, so large studies dominate as much as in the fixed-effect model, whereas an additive between-study variance would pull the weights toward equality. Second, when the residual scatter is less than expected from sampling error alone, the scaling factor is below one and the standard errors are made smaller than even the fixed-effect ones. The result can be either too liberal or too conservative, depending on the data. Meta-regressions should be fitted with software designed for them, in which the within-study variances are treated as known. Table 2 sets out the practical differences between the approaches. Table 2. Fixed-effect, multiplicative and random-effects meta-regression compared. Feature Fixed-effect Multiplicative (naive WLS) Random-effects (mixed) Residual between-study variance Assumed zero Scales within-study variance Added to within-study variance Weights 1 / v_i Proportional to 1 / v_i 1 / (v_i + residual tau-squared) Standard errors under heterogeneity Too small Can be too small or too large Appropriate, given tau-squared Influence of large studies Dominant Dominant Moderated Typical use today Rarely justified Not recommended as fitted Default choice Estimating the residual between-study variance Choosing a random-effects meta-regression immediately raises a second question: how to estimate the residual tau-squared. Several estimators are available, and they can give noticeably different answers in small syntheses. The DerSimonian-Laird estimator, from the 1986 paper that introduced the random-effects method to clinical meta-analysis, is a method-of-moments estimator. It sets the residual Q statistic equal to its expected value and solves for tau-squared. It is simple, non-iterative, and was for decades the default in most software. Its extension to meta-regression is straightforward. Its known weakness is a tendency to underestimate tau-squared when heterogeneity is substantial, particularly when studies are few or of very unequal size. Restricted maximum likelihood, REML, estimates tau-squared by maximizing a likelihood that accounts for the loss of degrees of freedom from estimating the regression coefficients. It is approximately unbiased in many settings and is the default in metafor. Maximum likelihood, without the restriction, tends to underestimate tau-squared in small samples because it does not make that adjustment, a problem that grows as more covariates are added. The Paule-Mandel estimator is another moment-based method, which chooses tau-squared so that the generalized Q statistic, computed with random-effects weights, equals its degrees of freedom. It has performed well in several simulation comparisons. Two large reviews of the evidence on these estimators, by Veroniki and colleagues in 2016 and by Langan and colleagues in 2019, both published in Research Synthesis Methods, concluded that DerSimonian-Laird should not be the automatic default and that REML and Paule-Mandel generally perform better, though no estimator is best in every scenario. Langan and colleagues noted that with small numbers of studies, all estimators are imprecise, which argues for sensitivity analysis rather than for faith in any single choice. Viechtbauer's 2005 study in the Journal of Educational and Behavioral Statistics, on the bias and efficiency of meta-analytic variance estimators, is also a foundational reference for REML's properties. For meta-regression specifically, REML is the most widely recommended choice, and it is a sensible default. When k is small, it is worth re-fitting the key models with the Paule-Mandel or DerSimonian-Laird estimator and reporting whether conclusions change. res_pm <- rma(yi, vi, mods = ~ min60, data = dat, method = "PM", test = "knha") res_dl <- rma(yi, vi, mods = ~ min60, data = dat, method = "DL", test = "knha") When the residual variance is estimated as zero A frequent practical event in small meta-regressions is that the estimated residual tau-squared is exactly zero. This happens because the estimators are truncated at zero: if the residual scatter is no larger than sampling error would predict, the estimate is set to zero. At that point, the random-effects meta-regression becomes numerically identical to the fixed-effect one, with the same weights and the same standard errors. It would be a mistake to read a zero estimate as proof that the covariate has explained all heterogeneity. With 15 studies, a residual tau-squared of zero is entirely compatible with a true residual variance that is moderate, because the estimate is so imprecise. The confidence interval for the residual tau-squared, which metafor can compute, will usually make this plain. The zero estimate also means that the random-effects model provides no protection against the over-confidence of the fixed-effect model in exactly the situation where the data are least informative. This is one of the reasons the adjustments in the next chapter matter: the Knapp-Hartung method, and permutation tests, provide some protection even when the residual variance estimate collapses. A Bayesian analysis with a weakly informative prior on tau, which avoids estimates of exactly zero, is another route, and it is increasingly practical with modern software. A Bayesian alternative Bayesian meta-regression addresses the zero-estimate problem, and the broader problem of uncertainty in tau-squared, in a different way. Instead of plugging a single estimate of the residual variance into the calculation, it places a prior distribution on the residual standard deviation and averages over its uncertainty. The resulting posterior distribution for each coefficient reflects the full range of plausible values of tau, which tends to widen intervals appropriately when the data are sparse. The choice of prior matters most when studies are few, which is precisely when the Bayesian approach is most attractive. A weakly informative prior, such as a half-normal distribution on tau with a scale chosen to make very large heterogeneity implausible for the effect measure in question, is a common and defensible choice. For standardized mean differences, a half-normal prior with a scale of about 0.5 allows substantial heterogeneity while ruling out values that would be implausible in most literatures; for log odds ratios, similar reasoning applies on the log scale. The prior should be justified in the protocol and its influence checked by refitting with alternatives. library(brms) fit <- brm(yi | se(sqrt(vi)) ~ min60 + (1 | study), data = dat, prior = c(prior(normal(0, 1), class = "b"), prior(normal(0, 0.5), class = "sd")), seed = 2026) summary(fit) In this formulation, the se() term supplies the known within-study standard errors, the (1 | study) term supplies the residual study effect, and the prior on class sd, which brms bounds at zero, is a half-normal prior on tau. The Bayesian approach does not solve the problems of multiplicity, ecological bias or confounding, and a posterior probability that a slope is negative is no more a proof of effect modification than a small p-value. But it offers a coherent way to express the uncertainty of small meta-regressions, and it avoids the artefact of residual variance estimates of exactly zero. Interpretation: conditional and unconditional inference There is a genuine conceptual question hiding behind the terminology. Hedges and Vevea, in a 1998 paper in Psychological Methods, distinguished between conditional inference, which is about the particular set of studies included, and unconditional inference, which is about the broader population of studies from which they are seen as a sample. A fixed-effects analysis, in the plural sense used in parts of the social sciences, does not necessarily assume that all studies share one true effect. It may simply confine its inference to the studies at hand, estimating the weighted average of their true effects without claiming anything about studies not included. The same distinction applies to meta-regression. A fixed-effects meta-regression can be read as describing how the effects in these particular studies relate to these particular covariates. That is a coherent question, and it is why metafor distinguishes the "fixed-effects" from the "equal-effects" label. But it is rarely the question asked. A review that concludes "higher exercise doses produce larger benefits" is making a claim about exercise programs in general, not about the 26 trials that happened to be published. That generalizing claim requires a model that acknowledges that other trials, at the same dose, would have different true effects. It requires the random-effects formulation. The rule of thumb that follows is simple enough to state as a default. Use random-effects meta-regression, estimate the residual variance with REML or a comparably well-behaved estimator, check sensitivity to the estimator, and do not rely on the standard z-test when the number of studies is modest. Fixed-effect meta-regression should appear, if at all, as a sensitivity analysis whose over-confidence is acknowledged. What the random-effects model does not fix It is important not to oversell the random-effects formulation. It corrects the most serious source of false positives, which is ignoring residual heterogeneity, but it leaves others untouched. It does not correct for the uncertainty in the estimate of tau-squared, which is the subject of the next chapter. It does not correct for testing many covariates, which is the subject of the fifth. It does not guard against influential studies, ecological bias, confounding among covariates or the peculiar problems of baseline-risk regression, which occupy the sixth and seventh. The random-effects model is the necessary foundation, not a finished building. One more subtlety deserves mention. The random-effects meta-regression assumes that the residual effects u_i are independent of the covariates and normally distributed with constant variance. If the amount of residual heterogeneity differs systematically by a covariate, for example if supervised trials are much more homogeneous than unsupervised ones, the constant-variance assumption fails. Location-scale meta-regression models, described by Viechtbauer and López-López and available in metafor through the scale argument of rma(), allow the residual tau-squared itself to depend on covariates. With the small numbers of studies in most syntheses, such models are hard to estimate well, but they are useful when the question of whether one kind of study is more variable than another is itself of interest. Hashtags: #MetaRegression #Heterogeneity #SubgroupAnalysis #ModeratorAnalysis #RandomEffectsMetaRegression #MixedEffectsMetaRegression #FixedEffectMetaRegression #CochranQ #ISquared #TauSquared #PredictionIntervals #KnappHartung #PermutationTests #RobustVarianceEstimation #EcologicalBias #StudyLevelModerators #MetaAnalysis #EvidenceSynthesis #BubblePlots #InfluenceDiagnostics #PseudoRSquared #REML #ModeratorMultiplicity #Metafor #FutureOfMetaRegression
- Managing University-Industry Research Collaborations (IP, Secrecy, and Publications)
Download the Book (PDF): Introduction A professor of chemical engineering receives a call from a mid-sized manufacturer. The company has a stubborn problem with catalyst degradation, it has read her papers, and it would like to fund two doctoral students for three years to work on it. The money is real, the problem is interesting, and the students need support. Within a week a draft agreement arrives from the company's legal department. It says that all results belong to the company, that nothing may be published without the company's written consent, that the students must sign individual confidentiality undertakings, and that the university warrants the work will not infringe anyone's patents. The professor forwards it to her university's research contracts office and asks, reasonably, whether this is normal. It is normal in the sense that companies send drafts like this every day. It is not acceptable in the sense that almost every clause, taken literally, would damage something the university exists to protect. Signing it would mean two doctoral theses that might never be examined in public, a professor whose own field of expertise could become partly off-limits to her, and a university promising something no research organisation can honestly promise. Refusing it outright would mean losing the funding, the problem, and the relationship. The work that sits between those two outcomes is the subject of this book. Collaboration between universities and firms is now a routine part of research life in almost every applied discipline and a good many basic ones. Companies fund laboratories, share proprietary materials and data, license university patents, hire graduates, sit on advisory boards, and co-author papers. Governments on every continent encourage it, measure it, and tie funding to it. Yet the practical knowledge needed to manage these relationships well tends to be held by a small number of specialists: contracts officers, technology transfer managers, university counsel, and the occasional veteran principal investigator who has learned the hard way. The researchers who actually do the work, and the students whose careers depend on it, often meet the key questions only when a deadline is close and a document is already on the table. Why the Tension Is Structural The difficulty is not that companies are grasping or that academics are naive. It is that the two institutions run on different economies of knowledge, and each economy is rational on its own terms. A university scientist earns standing by making knowledge public. Priority of discovery, established through publication, is the currency of academic careers: it decides hiring, promotion, grants, and reputation. The sociologist Robert Merton described the norms that sustain this system, among them what he called communalism, the expectation that findings belong to the scientific community once published, and organised scepticism, the expectation that claims will be exposed to scrutiny by others. Disclosure is not an afterthought in academic science. It is how the work becomes science at all. A firm earns its return by controlling knowledge long enough to profit from it. It may do this through patents, which require disclosure but grant a temporary right to exclude; through trade secrets, which require that information never be disclosed; or simply through speed and lead time. A company that pays for research and then watches its competitors read the results in a journal before it has had a chance to act on them has, from its point of view, subsidised its rivals. Neither economy is wrong. Most of the world's useful technology has depended on both: public science that establishes what is possible and private investment that turns possibility into products. The problem is that when the two meet inside a single project, they pull the same pieces of information in opposite directions. The researcher wants to publish; the sponsor wants to delay or withhold. The researcher wants to use what she has learned in her next project; the sponsor wants to own it. The student wants a thesis that anyone can read; the sponsor wants to keep the process conditions out of view. Every major clause in a research agreement is, in effect, a negotiated settlement of one of these pulls. What This Book Covers The book follows the life of a collaboration from its first conversation to its long afterlife in patents, licences, and careers. The first chapter sets out why universities and firms collaborate at all, the different forms that collaboration takes, and why the form matters so much for what follows. The second surveys the legal and policy frameworks that decide who owns publicly funded inventions in the first place, starting with the Bayh-Dole Act in the United States and moving through the quite different arrangements in the United Kingdom, continental Europe, Scandinavia, and East Asia. The third dissects the sponsored research agreement clause by clause. The middle of the book deals with the three assets that collaborations fight over. The fourth chapter examines ownership of results: inventions, software, data, and know-how, and the frequently misunderstood difference between who invented something and who owns it. The fifth turns to secrecy, covering confidentiality agreements, trade secrets, and the export control and research security rules that increasingly shape what can be shared and with whom. The sixth concerns the right to publish, including the interaction between publication and patent novelty, which differs sharply between the United States and Europe, and the cases in which sponsors have tried to suppress unwelcome findings. The last part of the book looks at what happens downstream and to the people involved. The seventh chapter covers licensing: exclusive and non-exclusive licences, options, royalties, equity, and the reserved rights that keep university research possible. The eighth addresses students and early-career researchers, and the financial conflicts of interest that can distort judgement when academics hold a stake in the outcome. The conclusion draws out what follows for institutions and individuals who want to collaborate without giving away what makes academic research worth funding. Who This Book Is For The intended reader is a researcher or research leader who is not a lawyer and does not want to become one, but who needs to understand the terms under which their work is being done. That includes principal investigators negotiating their first industrial contract, department heads and deans setting local policy, postdoctoral researchers considering an industry-funded position, graduate students wondering what they have signed, and research managers who want a coherent picture of a field they have learned piecemeal. Industry scientists and managers who work with universities will also find in it an explanation of why their academic partners resist certain terms, and which of those objections are negotiable. Because the subject touches law in several jurisdictions, a word of caution is needed. This book offers general information about how these arrangements typically work and why. It is not legal advice, and it cannot substitute for the judgement of your institution's contracts office, technology transfer staff, or legal counsel, who know your specific agreement, your funders' terms, and the law that governs them. Law and policy in this area also change. The United States government, for instance, has been re-examining how it uses its rights over federally funded inventions, and European countries have reformed their university invention rules as recently as 2023. Where the book describes a statute or rule, it describes its general shape; for any real decision, check the current text with someone qualified to read it. The Argument in Brief The central claim of this book is simple to state and harder to practise. A university-industry collaboration succeeds when the parties identify, early and explicitly, the small number of things the university cannot give up without ceasing to function as a university, and then are generous and precise about everything else. Those non-negotiable commitments are few. The university must be able to publish the results of its research after a reasonable delay. It must be able to use what it learns in further research and teaching. It must not let its students' education become hostage to a sponsor's commercial timetable. And it must not make promises, about outcomes, warranties, or secrecy, that a research organisation cannot keep. Almost everything else, including who owns patents, how licences are priced, how long confidential information stays confidential, and how much lead time a sponsor gets, can be traded in good faith. Most of the damage done in these relationships comes not from bad intentions but from ambiguity. An agreement that is vague about what counts as confidential, which results the sponsor has rights to, or how long a publication can be held back, becomes a source of conflict the moment the research produces something valuable or something unwelcome. A collaboration that has settled those questions in advance, in plain language the scientists themselves understand, can survive both kinds of surprise. The rest of the book is about how to reach that kind of clarity, and what is at stake when it is missing. Chapter 1: Two Economies of Knowledge Every negotiation between a university and a company is shaped by what each side thinks knowledge is for. The contracts officer and the corporate lawyer may argue about indemnities and governing law, but beneath those arguments lies a deeper disagreement about who should be able to know what, and when. Understanding that disagreement is the first step toward managing it, because it explains why certain clauses provoke resistance out of proportion to their apparent importance, and why others that look alarming turn out to be easy to settle. The Academic and the Industrial Bargains The modern research university rests on an unusual bargain. Society pays for research whose practical value is uncertain, and in return researchers make their results freely available. The researcher is paid partly in salary and partly in recognition, and recognition is awarded to whoever discloses a finding first. The sociologist Robert Merton, writing in 1942, set out the norms that he believed sustained this arrangement: universalism, meaning claims are judged on their merits rather than by who makes them; communalism, meaning findings become common property once published; disinterestedness, meaning scientists act for the advancement of knowledge rather than personal gain; and organised scepticism, meaning claims are exposed to systematic scrutiny. Merton was describing an ideal rather than a reality, and later sociologists showed how often it is breached. But the ideal remains powerful because the reward system still depends on it. Economists later gave the same idea a more mechanical form. In an influential 1994 paper, Partha Dasgupta and Paul David argued that "open science" and "proprietary technology" are two distinct institutional systems for producing knowledge. Open science solves the problem of rewarding people for producing something that is expensive to create and cheap to copy by granting priority, and therefore reputation, to whoever discloses first. The incentive to disclose quickly is built into the structure: a researcher who sits on a result risks being scooped and receiving nothing. Proprietary research solves the same problem the opposite way, by granting control, through patents or secrecy, to whoever can exclude others. What follows from this is that disclosure is not merely something academics like to do. It is the mechanism through which they are paid. A clause that delays publication by six months is, from the academic side, a clause that puts six months of the researcher's career at risk. A clause that prevents publication altogether is a clause that takes away the researcher's wages. This is why the right to publish carries so much weight in university negotiations, often more than the question of who owns patents, which can look to outsiders like the more valuable asset. Disclosure also does work beyond individual careers. Publication subjects claims to verification by others, which is how errors are caught. It allows students to learn from the whole field rather than only from their own laboratory. It lets other researchers build on results without repeating them. And it is the basis of the public's trust in academic research: an unpublished study cannot be checked, and a university known to suppress results on behalf of sponsors cannot be believed when it speaks as an independent authority. These functions are public goods, and they are the reason universities receive public money and tax privileges. A university that bargains them away is spending capital that does not belong to it alone. A company doing research faces a different problem. It must recover its investment through products or services, and it competes with other firms that would happily benefit from its work without paying for it. The core question for industrial research is appropriability: how the firm will capture enough of the value it creates to justify the spending. Firms have several tools for this, and it matters which ones they rely on in a given sector. Patents grant a right to exclude others from making, using, or selling an invention for a limited period, typically twenty years from filing, in exchange for public disclosure of how the invention works. Trade secrets protect information that derives value from not being generally known, for as long as the owner takes reasonable steps to keep it secret, which can be indefinitely. Lead time, the advantage of simply being first to market, can matter as much as either. So can complementary assets such as manufacturing capacity, distribution, and brand. Surveys of industrial managers carried out in the United States from the 1980s onward, notably the Yale survey led by Richard Levin and colleagues and the later Carnegie Mellon survey led by Wesley Cohen, Richard Nelson, and John Walsh, found that the importance of patents varies enormously by industry. In pharmaceuticals and some chemicals, patents are central, because a molecule once disclosed is easy to copy and expensive to develop. In many other industries, including much of electronics and machinery, firms reported relying more on secrecy, lead time, and complementary capabilities than on patents. This variation explains a good deal about why companies from different sectors bring such different demands to the negotiating table. A pharmaceutical sponsor will care intensely about patent rights and will usually accept publication once patent applications have been filed. A materials or process company may care less about patents and more about keeping specific operating conditions, formulations, or data entirely out of print. None of this makes firms hostile to science. Companies that fund university research generally do so because they want access to capabilities they cannot easily build in-house: specialised equipment, deep expertise, exploratory work that is too risky for a product budget, and, very often, the students who will become their future employees. Many industrial scientists publish, attend conferences, and share the academic values described above. The conflict arises at the specific points where the company's need to appropriate collides with the university's need to disclose. The Many Forms of Collaboration "University-industry collaboration" covers a wide range of arrangements, and much confusion comes from applying the expectations of one form to another. The form of the relationship largely determines what terms are appropriate, and a surprising number of disputes come down to the parties disagreeing, without realising it, about which form they are in. Sponsored research is the most common structured form. A company pays the university to carry out a defined programme of research, usually under a principal investigator's direction, with the results reported to the sponsor. The research is the university's in the sense that its staff design and conduct it, its students work on it, and its facilities host it. Because the work is university research, the university normally insists on its standard protections: the right to publish, retention of ownership of inventions its employees make or at least of rights to use them for research, and limits on confidentiality. The company typically receives reports, a right to review publications before submission, and some form of preferential access to intellectual property, often an option to negotiate a licence. Collaborative research differs in that both parties contribute intellectual effort, and sometimes both contribute money or in-kind resources. Company scientists may work alongside university researchers, or the project may be part of a larger publicly funded programme with industrial partners. Because inventions can arise from either side or from both jointly, ownership questions are more complex, and agreements need to deal carefully with jointly made results. Contract research or testing services are closer to buying a service. The company specifies the work, often a routine measurement, analysis, or test, and expects to own the results outright and keep them confidential. Many universities do such work through separate service units or subsidiaries, precisely because the terms are inconsistent with academic research. The key distinction is that contract research is not expected to generate new knowledge the university would want to publish. When a project is labelled as contract research but actually involves novel, publishable investigation, the label is doing the company's negotiating for it. Clinical trial agreements govern studies, frequently designed by the sponsor, in which university hospitals or academic investigators enrol patients. They raise particular issues about access to data, authorship, and publication of negative results, and they sit within a regulatory framework that includes trial registration and reporting obligations. Gifts and donations are unrestricted or loosely restricted funds given to support an area of research, with no deliverables and no rights to results. Because the donor gets nothing specific in return, a gift carries few legal complications. Problems arise when a so-called gift arrives with strings attached, such as a requirement to report results or an expectation of first access to inventions, which turns it in substance into sponsored research without the protections a sponsored research agreement would provide. Consulting is a personal arrangement between an individual academic and a company, typically permitted by university policy for a limited number of days, under which the academic provides advice outside their university duties. Consulting agreements are signed by the individual, not the institution, and they create some of the most serious conflicts in this field, because companies' standard consulting templates often claim ownership of any invention the consultant makes in the relevant field. That can collide directly with the academic's obligations to assign inventions to the university. Material transfer agreements govern the exchange of research materials such as cell lines, compounds, organisms, or data. They are small documents with large consequences, since they frequently contain terms about ownership of improvements, publication review, and rights to inventions made using the material. Licensing and spinouts come after research has produced something protectable. The university licenses its intellectual property to an existing company or to a new company formed around the technology, often with the inventors as founders or advisers. Open consortia are a deliberate alternative to the proprietary model. Several companies and public funders jointly fund pre-competitive research on the understanding that results will be published openly and not patented. The Structural Genomics Consortium, a public-private partnership founded in the early 2000s with pharmaceutical and public funding, has operated on this basis, placing protein structures, chemical probes, and related research outputs in the public domain. The model works when companies value shared foundational knowledge more than exclusive rights to it, which is most likely early in the research pipeline. Table 1 summarises how these forms typically differ on the three questions that matter most. Table 1. Typical default positions on ownership, publication, and confidentiality by form of collaboration. Form Who usually owns new IP Publication Confidentiality burden Sponsored research University, sponsor gets option or licence Yes, after short review Sponsor's information only Collaborative research Each party owns its own; joint results shared Yes, after review Both parties' information Contract research or service Sponsor Only with sponsor consent Broad, often results too Clinical trial Varies; sponsor often owns data Yes, with review and data access terms Protocol and sponsor data Gift University Unrestricted None Consulting Often company, limited by university policy Company controls Broad, personal obligation Open consortium No patents, or public domain Required, often rapid Minimal These are typical positions, not rules. Institutions vary, and every one of them is negotiated in practice. But the table illustrates the principle that should guide early conversations: the parties should agree first on what kind of relationship they are entering, and let the terms follow from that choice. Why Collaboration Has Grown It is worth understanding why the pressure to collaborate has intensified, because the policy motives affect how governments and universities judge success, and therefore how hard institutions push on terms. The most important change came in the United States in 1980, when the Bayh-Dole Act gave universities and other non-profit organisations the right to retain title to inventions made with federal funding, and encouraged them to patent and license those inventions. Before then, federal agencies often kept title themselves or required case-by-case approval before universities could own inventions, and relatively few federally funded inventions were licensed. After Bayh-Dole, American universities built technology transfer offices, and university patenting rose sharply through the 1980s and 1990s. The Act is examined in detail in the next chapter. Its broader significance is that it changed the university's own interests. A university that owns patents and receives licence income has something to gain from commercialisation, which aligns it partly with industrial partners and partly sets it against them, since both now want the same intellectual property. Other countries followed with their own reforms, often explicitly modelled on the American example, and national funding agencies increasingly asked universities to demonstrate economic impact. At the same time, companies in many research-intensive industries reduced the in-house basic research that large corporate laboratories had once performed, and looked to universities, start-ups, and partnerships instead. The phrase "open innovation", popularised by Henry Chesbrough in 2003, described this shift toward sourcing ideas from outside the firm. Industry money nonetheless remains a modest share of university research funding in most systems. In the United States, for example, business funding has long accounted for a single-digit percentage of total academic research and development spending, according to the National Science Foundation's Higher Education Research and Development survey, with federal and institutional funds far larger. But industry money is concentrated in particular fields, particular institutions, and particular laboratories, and for those groups it can be decisive. It also carries a symbolic weight beyond its size, because commercial relationships are what politicians and university leaders point to when they describe research as useful. Where the Collisions Happen The rest of this book takes up the specific points where the two economies collide. It helps to name them now, because they recur in nearly every agreement. The first is ownership. When university researchers make an invention using a company's money, who owns it? The company feels it paid for the result. The university points out that the company paid for effort, not results, that its fees rarely cover the full cost of the research, and that the invention draws on years of publicly funded expertise. Both positions have merit, and the resolution usually lies in separating ownership from access. The second is secrecy. Companies need to share confidential information for many projects to make sense, and they need assurance that it will be protected. Universities can protect specific confidential information supplied to them, but they cannot operate laboratories in which students and staff are sworn to broad, indefinite secrecy about their own work, and in some cases doing so would jeopardise legal exemptions on which they depend. The third is publication. The company wants the opportunity to review results for its own confidential information, to file patents before public disclosure, and sometimes to delay publication for competitive reasons. The university needs publication to proceed within a reasonable time and without the sponsor controlling content. The fourth is people. Students and junior researchers are not simply labour on a project. They are there to be educated, and their theses, publications, and job prospects are at stake in decisions about delay and confidentiality. Senior researchers, meanwhile, may have personal financial interests in the companies that fund them, which can compromise both their judgement and the institution's credibility. The fifth is risk. Companies are used to contracts in which a supplier warrants the quality of what it delivers and indemnifies the customer against claims. Universities doing research cannot warrant results that do not yet exist, and public institutions often cannot legally accept open-ended liabilities. Each collision can be resolved, and there is a well-developed body of practice for doing so. What that practice requires first is that both sides understand that the other is not being unreasonable. The company asking to own everything is protecting its investment within the logic of proprietary research. The university refusing is protecting its capacity to function within the logic of open science. A good agreement gives each side enough of what its logic requires, and no more. Chapter 2: Who Owns Publicly Funded Inventions Before any company signs a research agreement with a university, someone else has usually already staked a claim to the results. Governments fund the salaries, facilities, and earlier research on which industrial projects build, and most research-intensive countries have rules deciding who owns inventions made with public money and by university employees. Those rules set the baseline from which every negotiation starts. A sponsor who does not understand them may demand rights the university is legally unable to give, and a university that does not explain them may appear to be merely stubborn. The rules differ considerably between countries. They also change, sometimes sharply. What follows describes their general shape as of the mid-2020s, with the caution that anyone relying on them for a specific agreement needs to check the current law. The Bayh-Dole Act in the United States The Patent and Trademark Law Amendments Act of 1980, known after its sponsors, Senators Birch Bayh and Robert Dole, as the Bayh-Dole Act, is the most influential law in this field anywhere. Its core principle is that small businesses and non-profit organisations, including universities, may elect to retain title to inventions made in the performance of federally funded research. Before 1980, federal agencies followed a patchwork of policies, and many retained title to inventions they funded or required universities to seek approval to own them. The government, holding thousands of patents, licensed relatively few. The premise of Bayh-Dole was that the organisations closest to the research had the best incentive and knowledge to find companies willing to develop it. Retaining title comes with obligations, set out in the statute and its implementing regulations at 37 CFR Part 401. The university must require its employees to disclose inventions, and must disclose each "subject invention" to the funding agency within two months after the inventor discloses it in writing to the university's responsible personnel. It must elect whether to retain title within two years of that disclosure, which can be shortened where a publication or other disclosure has started a statutory clock. If it elects title, it must file a patent application within a year of election, subject to similar adjustments. The federal government receives a non-exclusive, non-transferable, irrevocable, paid-up licence to practise the invention for government purposes. Universities must share royalties with inventors and use the remainder, after expenses, for education and research. Exclusive licences to use or sell the invention in the United States generally require that products be manufactured substantially in the United States, unless the agency grants a waiver. And universities are expected to give preference to small businesses when licensing, where feasible. The government also keeps "march-in" rights. In defined circumstances, principally where the university or its licensee has not taken effective steps to achieve practical application of the invention, or where action is needed to meet health or safety needs not reasonably satisfied by the licensee, the funding agency may require the university to grant a licence to another applicant, or grant one itself. In more than four decades, no agency has exercised this power, despite a series of petitions, most seeking march-in on grounds of high drug prices. In December 2023 the National Institute of Standards and Technology published a draft interagency framework suggesting that price could be one factor agencies consider, which drew very large numbers of comments and strong reactions on both sides. In August 2025 the Department of Commerce went further in a different direction, announcing a review of Harvard University's compliance with Bayh-Dole obligations, including timely disclosure and the domestic manufacturing preference, and signalling that it might use march-in as a remedy. Senior officials also publicly floated proposals for the government to take a share of university patent revenue. How these initiatives develop is uncertain, and they illustrate that the settlement Bayh-Dole created is less fixed than it appeared for most of its history. Three points about Bayh-Dole matter especially for industrial collaborations. First, when a project is funded partly by the government and partly by a company, any invention "conceived or first actually reduced to practice" in the performance of the federal award is subject to the Act, and the government's licence and march-in rights attach to it. A sponsor cannot contract its way out of this. Sponsors sometimes discover this late, when they learn that the laboratory they funded also holds federal grants and that the invention they expected to own outright was made partly on federal effort. Careful universities separate projects in their accounting and records, and careful sponsors ask about overlapping funding at the start. Second, the Act does not itself transfer ownership from inventors to universities. In Board of Trustees of Leland Stanford Junior University v. Roche Molecular Systems, decided by the US Supreme Court in 2011, a Stanford researcher had signed a university agreement saying he "agree[d] to assign" his inventions to Stanford, and later, while working at a company then called Cetus, signed a document stating that he "do[es] hereby assign" his rights to the company. The Court held that Bayh-Dole does not automatically vest title in the federal contractor; inventors own their inventions first, and ownership passes only by assignment. The earlier-in-time promise to assign in the future lost to the later present assignment. After Stanford v. Roche, American universities widely revised their policies and agreements to use present-assignment language, and federal regulations were later amended to require contractors to obtain written agreements from employees assigning rights in subject inventions. The case also shows how the individual agreements academics sign, including with companies they visit or consult for, can override institutional expectations. Third, Bayh-Dole applies only to federally funded inventions. Work funded entirely by a company is governed by the university's own intellectual property policy and the contract. In practice most research universities in the United States have policies claiming ownership of inventions made by employees using university resources, regardless of funding source, and they apply the same framework to industry-funded work. But the ownership rule there is a matter of employment and institutional policy, not federal statute. The United Kingdom and Continental Europe In the United Kingdom there is no statute equivalent to Bayh-Dole, because none was needed. Under section 39 of the Patents Act 1977, an invention made by an employee belongs to the employer if it was made in the course of the employee's normal duties, or duties specifically assigned, in circumstances where an invention might reasonably be expected to result. University academics are generally treated as employees whose duties include research, so their inventions normally belong to the university. For a period after the Second World War, publicly funded inventions were channelled through a state body, the National Research Development Corporation, later part of the British Technology Group, but that monopoly ended in the 1980s and universities have since managed their own inventions. Sections 40 and 41 of the Act allow an employee inventor to claim compensation where a patent or invention has been of "outstanding benefit" to the employer. For decades this looked like a dead letter, but courts have made awards in a few cases, including Kelly and Chiu v GE Healthcare in 2009 and, most prominently, Shanks v Unilever, where the Supreme Court in 2019 upheld an award to an inventor employed by a large company. Universities generally operate royalty-sharing schemes that make these claims unlikely in practice. The United Kingdom's most distinctive contribution to collaboration practice is the Lambert Toolkit, a set of model research agreements developed after Richard Lambert's 2003 review of business-university collaboration for the Treasury. The toolkit offers standard agreements covering a range of ownership arrangements, from the university owning results and the company receiving a non-exclusive licence in a defined field, through arrangements in which the company may negotiate a licence or assignment, to arrangements in which the company owns results and the university keeps rights to use them for non-commercial research. One model, intended for contract research, requires the company's permission to publish. The toolkit's value lies less in any one template than in the way it makes the choice explicit: the parties decide which model fits, and the terms follow. That approach is taken up in Chapter 3. On the continent, the story has turned on a single doctrine. Several European countries traditionally applied a "professor's privilege": university teachers, alone among employees, owned their own inventions. The rationale was academic freedom. Professors chose their own research, were not directed by the institution, and so should not be treated like corporate employees whose inventions belong to their employer. The practical result, critics argued, was that many inventions were never patented or were patented by companies that had quietly arranged to own them through consulting or funding relationships, leaving universities with little role in commercialisation. Beginning around 2000, most of the countries that had this privilege abolished it, largely in imitation of Bayh-Dole. Denmark passed an Act on Inventions at Public Research Institutions in 1999, giving institutions the right to claim employees' inventions. Germany amended its Employees' Inventions Act in 2002, removing the special status of university teachers. Under the German rules, universities can claim inventions like other employers, but academics retain two notable protections: they keep the freedom not to disclose an invention at all if they choose not to publish or exploit it, and if they do intend to publish they must notify the university in advance so it can decide whether to file first. Inventors whose inventions the university exploits are entitled to thirty per cent of the gross income. Austria and Norway made similar changes in the early 2000s, and Finland did so through an act on university inventions that took effect in 2007. Italy took an unusual route. It introduced a professor's privilege by law in 2001, when most of Europe was moving the other way, and kept it for more than two decades. In 2023, a reform of the Industrial Property Code, adopted as Law No. 102 of 24 July 2023, reversed course and gave ownership of inventions made by researchers at universities and public research organisations to the institutions. The inventor must disclose the invention to the institution, which then has six months to file a patent application or declare that it is not interested; if it does neither, the inventor may proceed alone. A ministerial decree later in 2023 issued guidelines for research contracts with external funders. Sweden remains the principal European exception. Its "teacher's exemption", lärarundantaget, still gives university teachers and researchers ownership of their own inventions in the ordinary case, although agreements for externally funded research can change the position. A company collaborating with a Swedish university therefore needs to think about agreements with individual researchers, not only with the institution. France, the Netherlands, Spain, and most other European systems treat inventions of public researchers as belonging to their employing institution, usually with statutory or policy-based sharing of revenue with inventors. European Union research programmes such as Horizon Europe add their own rules for funded projects, generally providing that results belong to the participants that generate them, with obligations to exploit and disseminate results and to grant access rights to other participants on defined terms. East Asia and Elsewhere Japan explicitly borrowed from Bayh-Dole. Article 30 of the Act on Special Measures for Industrial Revitalization, enacted in 1999, allowed contractors to retain patent rights arising from government-commissioned research, and the provision became known as Japan's Bayh-Dole clause. It was later moved into other legislation. The incorporation of Japan's national universities in 2004 was equally important, because before then inventions by national university faculty had in many cases belonged to the individual researcher or to the state; after incorporation, universities could own inventions institutionally and build technology licensing capacity. China has used a series of laws and policies to encourage transfer of publicly funded research to industry, most notably the Law on Promoting the Transformation of Scientific and Technological Achievements, substantially amended in 2015. The amended law gave public research institutions greater autonomy to transfer and license their results and set a default for rewarding the researchers involved: in the absence of an institutional rule or agreement, not less than half of the net income from transferring or licensing a result should go to the personnel who made it. Some Chinese provinces have also piloted giving researchers partial ownership of their results. Readers dealing with Chinese partners should expect rapid policy change and substantial variation between institutions. Canada has no national rule on university invention ownership, and practice varies by institution. Some universities, such as the University of Waterloo, have long followed a creator-owned policy, while others claim institutional ownership or share it. Australia's major universities generally claim ownership of employee inventions under their intellectual property policies. India's government funding agencies have issued guidelines allowing institutions to retain rights in funded inventions, and several countries in Latin America, South Africa, and elsewhere have passed laws explicitly modelled on Bayh-Dole. Table 2 summarises the main patterns. Table 2. Default ownership of university inventions in selected jurisdictions. Jurisdiction Default owner Legal basis Notable feature United States Inventor first; university by assignment Bayh-Dole Act (federal funding); institutional policy Government licence, march-in rights United Kingdom University as employer Patents Act 1977, s.39 Employee compensation for outstanding benefit Germany University may claim Employees' Inventions Act, amended 2002 Inventor 30% of gross income; freedom not to disclose Italy Institution Industrial Property Code, reformed 2023 Inventor may file if institution does not act in six months Sweden Researcher Teacher's exemption Contracts can alter default Japan University 1999 Bayh-Dole clause; 2004 university incorporation Explicitly modelled on US China Institution 2015 transformation law Default of at least half of net income to researchers What the Baseline Means for Negotiation These frameworks do not settle every question a sponsor and a university will face, but they set limits that neither party can negotiate away. A few practical consequences follow. A university cannot assign or grant to a sponsor rights it does not hold. If a researcher owns her inventions by law, as in Sweden, or if a federal licence attaches because of mixed funding, as in the United States, the sponsor's rights are necessarily subject to those facts. Sponsors from abroad sometimes assume that the university's position reflects institutional preference when it reflects statute. The inventor's own agreements matter as much as the institution's. Stanford v. Roche shows how a document signed by an individual researcher, perhaps without much attention, can shift ownership in ways that surprise everyone later. Consulting agreements, visiting researcher agreements, and personal confidentiality undertakings all need to be checked against institutional obligations. Rules about inventor rewards affect incentives within the university. Where the law gives inventors a fixed share of revenue, as in Germany and by default in China, or where institutional policy gives them a generous one, researchers have their own financial stake in how agreements treat intellectual property. That can make them valuable allies in reaching workable terms, or it can create the conflicts of interest discussed in Chapter 8. Finally, the frameworks reflect a policy judgement that is still contested. Supporters of Bayh-Dole credit it with creating the modern technology transfer system and with helping to bring many products, especially in biotechnology, from university laboratories to market. Critics argue that it encouraged universities to patent too much, to treat licensing income as a revenue source it rarely becomes, and to erect barriers to the free flow of research tools and knowledge. Research by economists such as David Mowery, Richard Nelson, Bhaven Sampat, and Arvids Ziedonis suggested that some of the growth in university patenting reflected trends already under way before 1980, especially in biomedical science, and that patents were one of several channels, and often not the most important, through which university research reached industry. The same debate runs through the European reforms. Whatever one's view, the practical point for collaborations is that ownership rules are a policy instrument, not a law of nature, and a university's position on them is best explained to partners as a matter of public obligation rather than of institutional appetite. Chapter 3: Anatomy of a Sponsored Research Agreement The sponsored research agreement is the document in which the collision between the two economies of knowledge is finally settled, or, too often, merely postponed. It is usually drafted by lawyers and negotiated by contracts officers, but it governs what scientists may do every day of the project and for years afterward. Researchers who have never read one closely tend to discover its contents at the worst moment: when a student's thesis is ready, when a result looks commercially valuable, or when a paper reports something the sponsor would rather not see in print. This chapter walks through the main parts of a typical agreement, explaining what each clause does, why sponsors and universities tend to want different things from it, and where compromise usually lies. The details vary across institutions and countries, but the architecture is remarkably consistent. Framing the Project and Paying for It The most consequential decisions are often made before any lawyer is involved, in conversations between the researcher and the company's scientific staff. Those conversations decide what the project is for, what each side is contributing, and what each expects to get out of it. If they are clear, the contract becomes a matter of recording an agreement already reached. If they are vague, the contract becomes the place where two different projects discover each other. Four questions deserve explicit answers at this stage. First, is this research or a service? If the university is being asked to apply known methods to produce data the company will own and use privately, the right vehicle may be a service agreement, perhaps through a dedicated testing unit, not a research agreement. If the goal is new knowledge that the researcher and students will want to publish, it is research, and the university's standard protections should apply. Second, what is each side contributing? The company may contribute money, materials, data, proprietary know-how, or its own scientists' time. The university contributes expertise, facilities, students, and usually a substantial body of prior work. Third, what does the company actually need from the results? Often it is not ownership of everything but freedom to use particular outcomes in its own products, or early access to results in a field. Fourth, what does the researcher need? Almost always this includes the right to publish, the ability to use results in further work, and theses that can be completed on time. The statement of work records the answers. It should describe the research plan in enough detail to define the scope of the project, which matters later because many rights attach only to results arising "under the project" or "in the field". A statement of work that is too broad can give a sponsor rights over research the university would have done anyway, or that another sponsor funds. One that is too narrow can leave important results outside the agreement, causing disputes about whether they are covered. Good practice is to describe objectives and methods specifically, to acknowledge that research may change direction, and to provide a mechanism for amending scope by mutual agreement. Sponsored research is paid for, but the nature of what is being bought is often misunderstood. The company pays for effort, not results. A university undertakes to use reasonable efforts to carry out the research described, and to report what it finds. It does not promise to reach any particular outcome, because research whose outcome was known in advance would not be research. Agreements express this through terms such as "reasonable efforts" or "best efforts", and through a disclaimer that the university does not guarantee results. Sponsors used to procurement contracts sometimes find this unsettling, but it reflects the nature of the activity, and the same principle applies whether the funder is a company, a charity, or a government. The budget normally covers direct costs, meaning the salaries of staff and students working on the project, consumables, equipment, and travel, plus indirect costs, meaning a share of the facilities, administration, libraries, and services that make research possible. Universities differ greatly in how they calculate and apply indirect costs, and companies frequently push to reduce them. In the United Kingdom, universities cost projects on a "full economic cost" basis under a methodology known as TRAC, and they generally expect industrial sponsors of contract research to pay at least that full cost, while accepting some discount for genuinely collaborative projects where the university gains research value of its own. In the United States, public universities are often subject to policies requiring a standard indirect rate for industry-sponsored work. Because indirect costs are real expenses, discounting them means that the university's other funds subsidise the sponsor, a point that becomes relevant when the parties later argue about who "paid for" an invention. Payment terms matter more than they seem. Research universities hire students and staff, often on fixed-term commitments, before the work begins. Agreements that pay only on delivery of milestones, or allow the sponsor to terminate on short notice without covering commitments already made, shift risk onto the university and ultimately onto the people employed on the project. A standard protection is that on early termination the sponsor pays for costs incurred and for non-cancellable obligations, including, in many institutions, the remaining support of any doctoral student whose training depends on the project. Deliverables in research agreements are normally reports: periodic progress reports and a final report. Some agreements also provide for delivery of materials, samples, prototypes, or software developed in the project. The contract should say what form reports take and what rights the sponsor has to use them, since reports often contain results that are also going to be published. Intellectual Property, Confidentiality, and Publication Most agreements distinguish two kinds of intellectual property. Background intellectual property is what each party brings to the project: patents, know-how, software, materials, and data that existed before the project or were developed independently of it. Foreground intellectual property, sometimes called project IP or arising IP, is what the project creates. The distinction is simple in principle and easy to blur in practice, because researchers rarely stop at the boundary of a project to ask whether an idea is new or an extension of what they already knew. Background is usually straightforward. Each party keeps its own, and grants the other only such rights as are needed to carry out the project, typically a non-exclusive, royalty-free licence limited to the project's duration and purpose. Difficulties arise when a sponsor needs background to exploit the foreground. A company that obtains rights to a new method may find it cannot practise the method without also licensing an earlier university patent. Some agreements address this by giving the sponsor an option to negotiate a licence to relevant background on commercial terms; others expressly exclude background from any grant. Universities are usually wary of committing background broadly, because it may already be licensed to another company or may be needed for the university's future research. Foreground is where the major negotiation lies, and Chapter 4 examines the ownership options in detail. The typical positions are that the university owns results created by its staff and students, the sponsor owns results created by its own staff, jointly created results are jointly owned, and the sponsor receives an option to negotiate a licence, often exclusive, to university-owned foreground within a defined field and period. Variations include pre-agreed licence terms, a royalty-free non-exclusive licence for internal research, or, where the sponsor is paying a premium and the university accepts the consequences, assignment of foreground to the sponsor subject to the university's retained right to use it for research and teaching. The critical principle is that ownership and access are different questions. A sponsor that needs to commercialise a result can usually be given what it needs through a licence, without the university giving up ownership. A university that needs to keep researching in a field can usually be given what it needs through a reserved licence, without insisting on owning every result. The parties who understand this tend to settle intellectual property clauses quickly. The parties who treat ownership as a matter of principle tend not to. The confidentiality clause defines what information the parties must protect, how, and for how long. The publication clause defines how the university may disclose its results. The two must be read together, because the most common problem in research agreements is a confidentiality definition broad enough to swallow the research results, making the publication clause meaningless. A well-drafted agreement defines confidential information as specific information disclosed by one party to the other and identified as confidential, typically by marking or by written confirmation after oral disclosure. It excludes information that is already public, already known to the recipient, independently developed, or received from a third party without restriction. It limits the duration of the obligation, frequently to a period of three to five years after disclosure, while recognising that some trade secrets may merit longer protection. And it makes clear that results of the research are not the sponsor's confidential information simply because the sponsor paid for them. Chapter 5 considers these terms more closely. The publication clause normally gives the sponsor a right to review proposed publications for a fixed period before submission, commonly around thirty to sixty days. During that period, the sponsor may identify its own confidential information, which the university agrees to remove, and may identify patentable results, in which case the university agrees to delay submission for a further limited period, often around sixty to ninety days, to allow a patent application to be filed. The sponsor does not have the right to prevent publication or to change the scientific content, conclusions, or interpretation. Chapter 6 discusses why these limits matter and what happens when they are absent. Risk and the Supporting Clauses Commercial contracts allocate risk through warranties, which are promises that certain facts are true, indemnities, which are promises to cover another party's losses from specified claims, and limitations of liability, which cap how much a party can be required to pay. Sponsors' standard templates often reflect supplier relationships, in which the supplier warrants its deliverables and indemnifies the buyer. Research relationships are different, and universities typically resist these terms for principled reasons. Universities generally disclaim warranties that results will be accurate, fit for a purpose, or free from infringement of third-party intellectual property. They cannot promise accuracy about research outcomes still to be determined, and they cannot run freedom-to-operate searches on results that may never be commercialised. The sponsor, if it decides to commercialise, is in a better position to investigate infringement risk and to take the commercial decisions that create it. For the same reason, universities typically ask the sponsor to indemnify them against claims arising from the sponsor's use of results, particularly product liability claims arising from commercial products. Public universities often face legal limits on the liabilities they can accept. Many American state universities are constrained by state law, including sovereign immunity doctrines and limits on the obligations state agencies can incur, and cannot agree to indemnify private companies or submit to the law and courts of another state. Public universities in other countries face comparable rules. These are not negotiating postures, and experienced sponsors recognise them. The limitation of liability clause generally caps each party's liability, often at the value of the agreement, and excludes indirect or consequential losses. Universities usually ask that the cap and exclusions apply symmetrically, with carve-outs for matters like breach of confidentiality or personal injury where caps would be inappropriate. Several other provisions appear in almost every agreement and can cause unexpected problems. Use of names clauses prohibit either party from using the other's name or logo in advertising or promotion without consent. Universities care about this because their reputation for independence is valuable and easily misappropriated in marketing. Companies care because they do not want to be associated with research results they have not approved. Governing law and dispute resolution clauses decide which country's or state's law governs the contract and how disagreements are resolved. Public institutions are often required to use their own law, and multinational sponsors often prefer theirs. Common compromises include silence on governing law, the law of the defendant's jurisdiction, or neutral arbitration. Dispute resolution usually begins with escalation to senior management before any formal proceedings. Term and termination clauses decide how long the agreement lasts and how either party may end it. Rights that should survive termination, such as confidentiality, publication rights, licences already granted, and the sponsor's payment obligations for work performed, need to be expressly preserved. Export control clauses require the parties to comply with applicable laws and often require the sponsor to notify the university before providing controlled information or items, a point explained in Chapter 5. Assignment clauses decide whether a party may transfer the agreement, which matters when a sponsor is acquired or restructures. Making Negotiation Faster The time taken to negotiate research agreements is a constant source of frustration for scientists and sponsors alike. Several approaches have emerged to reduce it. The most widely used is the master research agreement, a framework contract agreed once between a university and a company that sets standard terms for intellectual property, confidentiality, publication, and liability. Individual projects are then launched through short statements of work that incorporate the master terms. Master agreements work well where a company funds many projects at the same institution, and they allow negotiation effort to be spent once, carefully. The second is the use of model agreements. The United Kingdom's Lambert Toolkit provides template agreements and decision guides that allow parties to select a model reflecting their chosen ownership arrangement. In the United States, the University-Industry Demonstration Partnership, a membership organisation of universities and companies formed under the auspices of the National Academies, has produced guidance and tools intended to help parties identify and resolve differences efficiently. Some universities publish standard "express" terms for smaller projects: a sponsor that accepts the university's standard agreement unchanged can start work quickly, while a sponsor that wants bespoke terms accepts a longer negotiation. The third, and in the end the most effective, is early alignment on the type of relationship. A negotiation in which the company believes it is buying a service and the university believes it is conducting research will take months, because every clause reopens the underlying disagreement. A negotiation that begins by agreeing whether the project is sponsored research, collaborative research, or a service, and by choosing a corresponding ownership model, usually goes smoothly, because most clauses then follow from that choice. The researcher, who is often the only person who understands both what the company needs and what the research involves, is in the best position to make that alignment happen before the lawyers begin. It helps to see how these principles work on a real draft. Return to the chemical engineer from the introduction, whose sponsor's template said that all results belonged to the company, that nothing could be published without consent, that her students must sign individual confidentiality undertakings, and that the university warranted non-infringement. A competent contracts office would not simply strike these clauses. It would first establish, through a conversation between the professor and the company's research manager, what the company actually needed. Suppose the answer is that the company wants to be able to use any improved catalyst formulation in its own plants without paying anyone, wants a head start over competitors on any patentable advance, and wants its own process data kept out of print. Each of those needs can be met without touching the university's core commitments. The company can receive a royalty-free, non-exclusive licence to use project results in its own internal operations, plus an option, for a defined period after each invention is disclosed, to negotiate an exclusive licence in its field on commercial terms. Its process data can be defined as confidential information, marked as such, and kept out of publications and theses, with the students' thesis committees able to see it under confidentiality where necessary. Publications can be submitted to the company for review a set period in advance, with a further short delay available for patent filing. The students need not sign individual undertakings, because the university takes responsibility for protecting confidential information through its own procedures, and the students are bound through their relationship with the university, not with the sponsor. The warranty of non-infringement is replaced by a mutual disclaimer and a statement that the company will make its own assessment before commercial use. None of this is exotic, and a sponsor that has worked with universities before will recognise it immediately. The revised agreement gives the company what its business requires and gives the university what its mission requires. What it does not do is give either side the abstract comfort of total control, which is what the original template was reaching for, and which is almost never what anyone actually needs. Hashtags: #ManagingUniversityIndustryResearchCollaborations #UniversityIndustryCollaboration #SponsoredResearch #CollaborativeResearch #ResearchContracts #IntellectualProperty #BackgroundIP #ForegroundIP #PatentOwnership #BayhDoleAct #TechnologyTransfer #ResearchCommercialization #ConfidentialityAgreements #TradeSecrets #PublicationRights #PublicationDelay #PatentNovelty #LicensingAgreements #ResearchLicensing #StudentProtections #ConflictOfInterest #ResearchSecurity #ExportControls #AcademicFreedom #FutureOfUniversityIndustryCollaboration
- Lipidomics Protocols (Extraction Chemistry, Mass Spectrometric Identification, and Signaling)
Download the Book (PDF): Introduction A lipid name in a results table is a claim. When a paper reports that "PC 16:0/18:1(9Z)" rises threefold in inflamed tissue, it is asserting, among other things, that the molecule was a phosphatidylcholine rather than an isomeric phosphatidylethanolamine, that palmitate sat on the first glycerol carbon and oleate on the second, that the double bond lay between carbons nine and ten in the cis geometry, that the molecule was actually present in the tissue rather than generated in the tube, and that its signal was not a borrowed isotope peak from a neighbour one double bond richer. Each of those assertions rests on a different step of the workflow. Most of them are rarely tested. The result is a literature full of lipid names written with more precision than the data behind them can bear. This book is about closing that gap. Its controlling idea is simple to state and demanding to practise: a lipidomics measurement is only as specific as the weakest step in the chain that runs from the freezer to the figure, and every lipid a laboratory reports should be named at exactly the level of structural and quantitative certainty that its extraction, separation, fragmentation and controls actually support. Everything that follows is an elaboration of that rule, applied first to the bench chemistry of getting lipids out of tissue, then to the mass spectrometry of recognising what came out, and finally to the most consequential and most error-prone application of the field, the mapping of lipid mediators that start, sustain and end inflammation. Why lipids are hard Proteins and nucleic acids are polymers of a small alphabet read in sequence. Lipids are not. They are a chemically heterogeneous family defined more by solubility than by structure: fatty acids and their oxidised derivatives, glycerolipids such as triacylglycerols, glycerophospholipids that make up most of every membrane, sphingolipids built on long-chain amino alcohols, sterols, prenols, and sugar-linked and polyketide lipids at the margins. The LIPID MAPS classification, set out by Eoin Fahy and colleagues in 2005 and updated since, divides this space into eight categories and hundreds of classes, and the number of individual molecular structures that a mammalian cell can build from the combinatorics of head groups, backbones, chain lengths, unsaturation, linkage types and oxidation runs into the thousands. Three properties of this diversity drive almost every protocol decision in the book. The first is dynamic range. A plasma sample contains cholesterol and its esters, triacylglycerols and phosphatidylcholines at concentrations in the high micromolar to millimolar range, alongside prostaglandins and resolvins that circulate, when they are detectable at all, in the picomolar range. No single extraction, chromatographic method or mass spectrometer tuning serves both ends well, and any protocol that claims to cover everything is making trade-offs it may not state. The second property is isomerism. Lipids built from the same atoms in different arrangements are the rule, not the exception. A single elemental formula can correspond to phosphatidylcholines and phosphatidylethanolamines of different total chain length, to dozens of chain combinations within one class, to different positions of each chain on the glycerol, to different double-bond positions and geometries, and to ether-linked or ester-linked chains. Mass alone cannot separate isomers, and even very high resolving power cannot separate them either; they require chromatography, ion mobility, informative fragmentation or chemical derivatisation. The third property is lability. Lipids are substrates for enzymes that remain active after a sample leaves the body, and polyunsaturated chains react with oxygen without any enzyme at all. Phospholipases keep cleaving acyl chains, lipoxygenases keep oxygenating arachidonate, platelets in clotting blood release thromboxane and 12-HETE in quantities that dwarf anything present in circulation, and air oxidises docosahexaenoic acid on the bench. A lipidomics laboratory therefore spends as much effort preventing the formation of lipids that were not there as it does detecting the ones that were. What the book covers The chapters follow the order of work. Chapter 1 treats the lipidome as an analytical object and deals with the pre-analytical phase: collection, quenching, anticoagulants, antioxidants, storage and the addition of internal standards, because the most expensive mass spectrometer cannot rescue a sample that was mishandled in the first hour. Chapters 2 and 3 are about extraction chemistry. The two-phase methods published by Jordi Folch and colleagues in 1957 and by E. G. Bligh and W. J. Dyer in 1959 remain the reference points against which everything else is measured, and they are explained here as phase diagrams rather than recipes: why the 2:1 chloroform to methanol ratio works, why the final 8:4:3 proportions matter, why Bligh and Dyer's economy of solvent becomes a liability in fatty tissue. Chapter 3 then turns to the methyl tert-butyl ether method published by Vitali Matyash and colleagues in 2008, to the butanol-based and single-phase methods that followed, and to the solid-phase extraction that mediator work demands, and it sets out how to choose between them. Chapters 4, 5 and 6 are the core of identification. Chapter 4 explains the shorthand notation developed by Gerhard Liebisch and colleagues and adopted by LIPID MAPS, and argues that its real value is not tidiness but honesty: the notation encodes how much is known. Chapter 5 covers ionisation, adduct formation and the class-specific fragmentation that allows a phosphocholine head group or a sphingoid base to be recognised. Chapter 6 confronts isobaric and isomeric overlap directly, with worked mass calculations showing when resolving power suffices and when it cannot, and with the modern techniques for locating double bonds and assigning chain positions. Chapter 7 compares shotgun lipidomics, in which the total extract is infused directly into the mass spectrometer, with liquid chromatography coupled to mass spectrometry, and explains why each is better at different questions. Chapter 8 addresses quantification, quality control and reporting in the framework developed by the Lipidomics Standards Initiative, including internal standard strategy, the use of reference materials such as NIST SRM 1950, and the Lipidomics Minimal Reporting Checklist. Chapters 9 and 10 apply all of this to inflammation. Chapter 9 maps the enzymatic pathways by which arachidonic acid and the omega-3 fatty acids become prostaglandins, leukotrienes, lipoxins, resolvins, protectins and maresins, and the receptors through which they act, and places sphingolipid and lysophospholipid signalling alongside them. Chapter 10 describes targeted mediator lipidomics as a protocol, from sample acidification through solid-phase extraction and multiple reaction monitoring to identification criteria, and it takes a clear position on the current dispute over whether specialised pro-resolving mediators can be reliably measured at the concentrations reported for them. How to read the protocols The procedures in this book give real ratios and realistic volumes, drawn from the original methods where those are well documented. They are written to be understood, not merely followed, because a protocol applied without understanding breaks the first time a sample differs from the one it was designed for. Where a laboratory must choose a parameter for itself, such as a centrifugation speed for a particular tube format or an internal standard concentration matched to its own samples, the book says so rather than inventing a universal number. Worked examples are either calculated from first principles, as with the exact masses in Chapter 6, or clearly framed as hypothetical scenarios. When a published study is cited, it is a real one, and the Notes at the end list the principal sources. Readers who want the primary literature should start with the papers listed in Further Reading, several of which are short enough to read in an afternoon and have shaped the field more than most textbooks. Throughout, the reader will find the same question asked in different forms: what does this step allow us to claim, and what does it not? Asked consistently, that question produces better protocols, smaller and more defensible results tables, and biological conclusions that survive the next laboratory's attempt to reproduce them. That is the standard this book sets out to teach. Chapter 1: The Lipidome as an Analytical Object Before any solvent touches a sample, a lipidomics study has already made its most consequential decisions: what tissue or fluid to take, how quickly to stop its chemistry, what to add to it, and how to store it. These pre-analytical choices do not merely add noise. They create and destroy specific lipids in predictable directions, and the resulting artefacts are often indistinguishable from biology in the final data. A laboratory that understands the lipidome as a set of chemically reactive molecules, rather than as a static inventory, designs its sampling accordingly. The LIPID MAPS consortium's classification, first published by Fahy and colleagues in the Journal of Lipid Research in 2005 and revised in 2009, organises lipids into eight categories defined by their biosynthetic origin and core chemistry. Fatty acyls (FA) include free fatty acids and the oxygenated eicosanoids and docosanoids that dominate the last two chapters of this book, along with fatty amides such as the endocannabinoid anandamide. Glycerolipids (GL) are glycerol esters without a phosphate head group, chiefly mono-, di- and triacylglycerols. Glycerophospholipids (GP) carry a phosphate at the sn-3 position of glycerol, usually esterified to a polar head group such as choline, ethanolamine, serine, inositol or glycerol, and they make up the bulk of cellular membranes. Sphingolipids (SP) are built on a sphingoid base such as sphingosine and include ceramides, sphingomyelins, glycosphingolipids and the signalling molecule sphingosine-1-phosphate. Sterol lipids (ST) include cholesterol, its esters, steroid hormones and bile acids. Prenol lipids (PR), saccharolipids (SL) and polyketides (PK) complete the scheme and are less central to mammalian inflammation work, though prenol lipids such as the ubiquinones and dolichols appear in any untargeted dataset. Within a category, classes are defined by head group and backbone, and it is at the class level that most analytical chemistry is organised. Phosphatidylcholine (PC), phosphatidylethanolamine (PE), phosphatidylserine (PS), phosphatidylinositol (PI), phosphatidylglycerol (PG), phosphatidic acid (PA), their lyso forms with a single chain, sphingomyelin (SM), ceramide (Cer), cholesteryl ester (CE), diacylglycerol (DG) and triacylglycerol (TG) together account for most of the lipid mass in mammalian cells and plasma. Their head groups determine how they partition between solvents, how they ionise, what fragments they give and where they elute on a column. Their chains determine how many distinct molecules each class contains. The combinatorial arithmetic explains why lipidomics is hard. A mammalian cell draws on perhaps twenty common acyl chains, differing in length from fourteen to twenty-two carbons and in unsaturation from none to six double bonds. A diacyl glycerophospholipid can place any of these at sn-1 and sn-2, and a triacylglycerol at three positions. Chains may be ester-linked, ether-linked (the plasmanyl lipids) or linked through a vinyl ether (the plasmalogens), and double bonds may sit at different positions within chains of the same length. The number of chemically distinct species in a single class therefore runs into the hundreds, even though only a subset is abundant in any tissue. The chemistry that continues after sampling When tissue is excised or blood is drawn, its enzymes do not stop working. Three families matter most. Phospholipases, especially the calcium-dependent cytosolic phospholipase A2 and the secreted phospholipases, hydrolyse the sn-2 ester of glycerophospholipids, releasing a free fatty acid and a lysophospholipid. Lipases act on triacylglycerols and diacylglycerols, raising free fatty acid and partial glyceride concentrations. Oxygenases, including cyclooxygenases and lipoxygenases, convert the released polyunsaturated fatty acids into eicosanoids. In blood, the lecithin:cholesterol acyltransferase continues to transfer a chain from phosphatidylcholine to cholesterol, generating lysophosphatidylcholine and cholesteryl ester. The practical consequences are consistent. Samples left at room temperature before processing tend to show rising lysophospholipids and free fatty acids and, where blood cells or platelets are present, rising oxylipins. Tissue that sits warm after excision can undergo rapid hydrolysis of signalling lipids; brain is the classic example, where post-mortem delay alters levels of arachidonic acid, diacylglycerols and endocannabinoids. Clotting blood is an extreme case. When platelets activate during coagulation, they convert arachidonic acid through cyclooxygenase-1 and 12-lipoxygenase into thromboxane A2, which hydrolyses to thromboxane B2, and into 12-HETE. Serum therefore contains these products at concentrations that reflect the clotting process in the tube rather than anything in the circulation. For mediator work, plasma collected into anticoagulant is required, and serum values of platelet-derived eicosanoids should be interpreted as a measure of platelet capacity, not circulating tone. Non-enzymatic oxidation adds a second route. Polyunsaturated chains, especially arachidonate (20:4), eicosapentaenoate (20:5) and docosahexaenoate (22:6), contain bis-allylic methylene groups whose hydrogens are easily abstracted, initiating radical chain reactions that produce hydroperoxides, hydroxides, isoprostanes and truncated aldehydes. The isoprostanes, such as 8-iso-prostaglandin F2α, are useful biomarkers of oxidative stress precisely because they form without cyclooxygenase, but this also means they form in samples exposed to air, light, heat and transition metals. Racemic hydroxy fatty acids are a signature of such chemistry, because radical oxidation lacks the stereoselectivity of enzymes; Chapter 10 returns to this as a diagnostic. Collection, quenching, additives and storage A sampling protocol for lipidomics should state four things explicitly: the time from collection to quenching, the temperature throughout, the anticoagulant or buffer, and any additives. The following principles apply across sample types. For blood, EDTA plasma is the most common choice for lipidomics, because EDTA chelates the calcium needed by cytosolic phospholipase A2 and many other enzymes, and because it is the matrix of the NIST reference material SRM 1950 discussed in Chapter 8. Heparin is usable but can interfere with some downstream steps and activates lipoprotein lipase in vivo if given to the subject; citrate dilutes the sample by a fixed volume that must be corrected. Whatever the choice, it must be consistent across a study, because anticoagulant differences produce measurable shifts in lipid profiles. Blood should be kept cold and centrifuged promptly, within an hour where possible, and the plasma aliquoted and frozen. For eicosanoids, many laboratories add a cyclooxygenase inhibitor such as indomethacin to collection tubes to block ex vivo prostanoid formation, and some add an antioxidant. For tissue, the aim is to stop metabolism within seconds. Snap-freezing in liquid nitrogen is standard. Where rapid post-mortem changes are a concern, as in brain, head-focused microwave irradiation has been used in animal studies to denature enzymes in situ before dissection. Frozen tissue should be pulverised under liquid nitrogen, or homogenised directly in cold extraction solvent, rather than thawed and cut, because the thawing interval is when hydrolysis accelerates. The weighed frozen powder is the natural unit for normalisation. For cultured cells, the medium is removed quickly, cells are washed with cold isotonic buffer such as ammonium formate or phosphate-buffered saline to remove medium lipids, and metabolism is quenched with cold methanol, which also serves as the first solvent of most extractions. Scraping cells into methanol is preferable to trypsinisation, which takes minutes at 37 °C and activates signalling. Normalisation to cell number, total protein or DNA should be decided in advance and measured on a parallel sample or an aliquot of the same homogenate. Antioxidants are the most debated additive. Butylated hydroxytoluene (BHT), a radical scavenger, is commonly added to extraction solvents at low concentration, often around 0.01% w/v, to suppress autoxidation of polyunsaturated lipids during processing. It does not reverse oxidation already present and does not stop enzymes, and it appears as a contaminant in some mass spectra, but for any study concerned with oxidised lipids or with the relative abundance of polyunsaturated species it is a reasonable default. Some protocols also add a metal chelator and flush tubes with nitrogen or argon. Glassware is preferred over plastics wherever chloroform is used, because chloroform leaches plasticisers such as phthalates and slip agents such as erucamide that show up as prominent ions and can suppress lipid signals. Storage and freeze–thaw Lipids are relatively stable at −80 °C in intact plasma or frozen tissue, but not indefinitely and not uniformly. Polyunsaturated and oxidisable species, lysophospholipids and mediators are the most sensitive. The practical rules are to aliquot before first freezing so that each analysis uses a fresh aliquot, to record every freeze–thaw cycle, to store extracts under inert gas in solvent at −20 °C or below if they cannot be analysed promptly, and to randomise sample storage so that storage time is not confounded with experimental group. The last point matters more than it appears. In a longitudinal clinical study, samples collected early have been stored longest; if storage degrades a lipid, a false time trend emerges unless the design accounts for it. Dried extracts are especially vulnerable. Evaporation under nitrogen at modest temperature is standard, but leaving a dried lipid film exposed to air, or storing it dry for long periods, invites oxidation. Most laboratories reconstitute immediately in the injection solvent, or store extracts in a chloroform–methanol or isopropanol-containing solvent with antioxidant at low temperature. Internal standards: added first, chosen deliberately The single most important pre-analytical decision in quantitative lipidomics is when and what internal standards to add. The answer to the first question is almost always as early as possible, ideally to the sample or homogenate before any solvent partitioning. An internal standard added at this stage experiences the same extraction losses, phase partitioning, evaporation losses, ionisation conditions and matrix effects as the endogenous lipids of its class, and dividing endogenous signal by standard signal corrects for all of them together. A standard added after extraction corrects only for instrument variation and hides recovery problems entirely. The choice of standard depends on the quantitative ambition. For class-wide profiling, one non-endogenous standard per lipid class is the minimum that the Lipidomics Standards Initiative regards as acceptable, and the standard must share the head group of the class it represents, because head group governs both extraction and ionisation. Two strategies dominate. Odd-chain or unusual-chain standards, such as PC 17:0/17:0 or PE 17:0/17:0, are absent or rare in mammalian samples and inexpensive, but odd-chain lipids do occur at low levels from diet and microbial metabolism, so a blank matrix check is needed. Stable-isotope-labelled standards, typically with deuterium or carbon-13, are chemically near-identical to endogenous species; deuterated lipids may elute slightly earlier than their protium forms in reversed-phase chromatography, a small effect but one to account for when setting retention windows. Commercial mixtures now provide labelled standards covering the major classes at concentrations roughly matched to human plasma, which simplifies the work considerably. A single standard per class assumes that all species in the class respond identically. That assumption is approximately true for direct infusion of dilute extracts, where head group dominates ionisation, and it becomes weaker when chain length and unsaturation change the ionisation efficiency, particularly in gradient chromatography where different species elute into different solvent compositions. Chapter 8 discusses how to correct for this and how to report it. The point for pre-analytics is that the internal standard mixture, its concentration and the moment of its addition must be fixed in the protocol before the first sample is processed. Changing any of them mid-study breaks comparability. Variation, normalisation and study design Even perfectly handled samples carry variation that has nothing to do with the experimental question, and a lipidomics design must control it rather than hope it averages out. Nutritional state is the largest source in blood. Triacylglycerols, and to a lesser extent diacylglycerols and some phospholipid species, rise after a meal as chylomicrons and then very-low-density lipoproteins enter the circulation, and the fatty acid composition of those triacylglycerols reflects the meal itself. A study that samples fasting controls in the morning and non-fasting patients in the afternoon has built a difference into its data before any analysis. Fasting for a defined period, commonly overnight in human studies, and sampling at a consistent time of day are the standard remedies. In rodents, which feed at night, the time of sampling relative to the light cycle matters in the same way. Circadian rhythm, sex, age, body composition and medication all shift the lipidome. Statins lower cholesterol and alter the ratio of cholesteryl esters; fibrates and omega-3 supplements change triacylglycerol composition and the pool of eicosapentaenoate and docosahexaenoate available for mediator synthesis; non-steroidal anti-inflammatory drugs, including low-dose aspirin, suppress prostanoids and, in the case of aspirin, redirect cyclooxygenase-2 toward different products, as Chapter 9 explains. Recording these covariates is not optional in human work. In animal studies, diet is the equivalent variable: standard chow varies in fatty acid composition between suppliers and batches, and a switch in chow between cohorts can change tissue omega-3 content enough to alter mediator profiles. Haemolysis deserves a separate mention. Red cells contain lipids and enzymes, and lysed erythrocytes release material that alters plasma measurements, including some lysophospholipids and oxidised species. Visibly haemolysed samples should be flagged at collection, and a simple spectrophotometric haemoglobin index can be recorded for each plasma aliquot so that haemolysis can be tested as a covariate later. The quantity a laboratory divides by determines what its numbers mean, and it cannot be chosen after seeing the data without inviting bias. For biofluids, volume is the natural denominator: nanomoles per millilitre of plasma. For tissue, wet weight of the frozen powder is simplest, though water content varies between tissues and pathological states; dry weight or total protein can be more robust when oedema or fibrosis is expected. For cells, cell number is intuitive but hard to measure precisely at the moment of quenching, so total protein or DNA from the same homogenate is often better. Some laboratories normalise to the sum of all measured lipids or to total phosphatidylcholine, which converts absolute amounts into mole fractions of the measured lipidome. Mole-percent data are excellent for describing membrane composition but can create apparent changes in every species when one abundant class changes, a compositional artefact that must be recognised when interpreting results. A worked example makes the point. Suppose macrophages treated with an inflammatory stimulus accumulate triacylglycerol in lipid droplets, doubling total TG while phospholipid amounts per cell stay constant. Expressed per microgram of protein, the phospholipids are unchanged and TG doubles. Expressed as mole percent of total measured lipid, every phospholipid class appears to fall, because the denominator grew. Both presentations are arithmetically correct; only one answers the question "did membrane phospholipids change?" A protocol that fixes its normalisation in advance, and reports the raw denominators alongside the results, makes this kind of misreading much less likely. A laboratory planning a study of lipid changes in an inflammatory model can use these principles to write a sampling protocol that anticipates its own artefacts. Suppose a group wants to compare plasma lipids and mediators in mice at several times after an inflammatory challenge. The protocol should specify EDTA plasma, collected by a consistent method at a consistent time of day, placed on ice, centrifuged within a fixed interval, and split immediately into at least two aliquots: one for global lipidomics and one, with a cyclooxygenase inhibitor and antioxidant present, for mediators. Tissue, if taken, should be frozen in liquid nitrogen within seconds of excision. Internal standards for global lipidomics go into the extraction solvent at a known volume per microlitre of plasma; deuterated mediator standards go into the mediator aliquot before protein precipitation. Every sample carries a record of collection time, processing time and freeze–thaw history. Sample order for extraction and analysis is randomised across time points and groups. None of this is exotic, and none of it can be recovered afterwards. The rest of the book assumes it has been done, and returns repeatedly to the ways in which its absence shows up in the data: lysophospholipids that rise with processing delay, thromboxane in serum, racemic hydroxy acids from autoxidation, and batch effects that track storage time rather than biology. Chapter 2: Two-Phase Extraction with Chloroform: Folch and Bligh–Dyer Lipid extraction is a problem in phase equilibria. The aim is to disrupt the interactions that hold lipids to proteins and membranes, bring every lipid into a single solvent phase, and then split that solution into two immiscible layers so that lipids end up in one and sugars, amino acids, salts and denatured protein in the other or at the interface. The two methods that defined how this is done, published by Jordi Folch, M. Lees and G. H. Sloane Stanley in 1957 and by E. G. Bligh and W. J. Dyer in 1959, rest on the same ternary system of chloroform, methanol and water. They differ in the proportions used and in the path through the phase diagram, and those differences have consequences that a practitioner should be able to predict rather than discover. The chloroform–methanol–water system Chloroform dissolves neutral lipids such as triacylglycerols and cholesteryl esters readily but is a poor solvent for polar lipids bound to protein and a poor penetrant of hydrated tissue. Methanol is fully miscible with water, disrupts hydrogen bonding and electrostatic interactions between lipids and proteins, denatures enzymes, and dissolves the polar head groups of phospholipids. A mixture of the two has properties neither has alone: it wets and penetrates tissue, breaks lipid–protein associations, and dissolves lipids across the polarity range. The three solvents form a system with a region of complete miscibility and a region in which the mixture splits into two phases. At high methanol content relative to water, chloroform, methanol and water form a single phase; this is the monophasic extraction condition, in which lipids are solubilised. Adding water, or chloroform and water together, moves the composition across the phase boundary. The mixture then separates into a lower phase rich in chloroform, containing a little methanol and almost no water, and an upper phase of methanol and water containing very little chloroform. Most lipids partition strongly into the lower phase. Non-lipid solutes partition into the upper phase. Denatured protein, being soluble in neither, collects as a disc at the interface. The composition at which the two phases separate controls what goes where. If the final system contains too little water, the phases may not separate cleanly, and polar non-lipid contaminants remain in the lower phase. If it contains too much methanol in the lower phase, polar lipids are well retained but contaminants may follow. If the upper phase is too large or too polar, the most polar lipids, notably gangliosides, lysophospholipids, phosphoinositides and some acidic phospholipids, partition partly into it and are lost. Every variant of the chloroform–methanol extraction is a choice of where to sit on this diagram. The Folch method Folch, Lees and Sloane Stanley developed their procedure for brain, a tissue rich in lipids and particularly in complex polar sphingolipids and phospholipids. Their method homogenises tissue in chloroform–methanol 2:1 (v/v) at a ratio of twenty volumes of solvent per volume of tissue, so that one gram of tissue is extracted with twenty millilitres of solvent. The tissue water, taken as roughly equal in volume to its mass for soft tissues, is small relative to the solvent, and the homogenate forms a single phase with a residue of insoluble material. The extract is filtered or centrifuged to remove the residue. The crude extract is then washed by adding 0.2 volumes of water or of a dilute salt solution. The original paper examined water and several salt solutions, including sodium, potassium, calcium and magnesium chlorides, and showed that the presence of salts reduced the loss of acidic lipids into the upper phase; in current practice 0.9% sodium chloride or 0.88% potassium chloride is typical. With twenty millilitres of 2:1 solvent and four millilitres of aqueous wash, plus the water from the tissue, the overall proportions come close to chloroform, methanol and water at 8:4:3 by volume. At this composition the system separates into two phases. Folch and colleagues described the approximate compositions as chloroform–methanol–water 86:14:1 for the lower phase and 3:48:47 for the upper phase, and they introduced the practice of rinsing the interface and lower phase with "theoretical upper phase", a pre-mixed solvent of that upper-phase composition, so that washing does not disturb the equilibrium. Why does this work so well? The large solvent-to-tissue ratio ensures that tissue water never dominates, so extraction happens under conditions of excess organic solvent regardless of how much lipid the tissue contains. The 2:1 chloroform–methanol proportion is rich enough in chloroform to dissolve bulk neutral lipid and rich enough in methanol to strip polar lipids from protein. The salt in the wash suppresses the ionisation-dependent partitioning of acidic phospholipids into the aqueous phase by providing counter-ions. The resulting lower phase contains nearly all the glycerophospholipids, sphingomyelin, ceramides, neutral lipids and sterols. The arithmetic behind 8:4:3 is worth doing once, because it shows how tolerant the method is. Twenty millilitres of 2:1 solvent contain about 13.3 mL of chloroform and 6.7 mL of methanol. One gram of soft tissue contributes roughly 0.8 mL of water, and the wash adds 0.2 × 20 = 4 mL of aqueous solution, giving about 4.8 mL of water in total. The ratio 13.3 : 6.7 : 4.8 reduces to about 8 : 4 : 2.9, close to the nominal 8:4:3. If the tissue were drier or wetter by a few hundred microlitres, the ratio would barely move, because the solvent excess dwarfs the tissue water. That robustness is the Folch method's greatest virtue and the reason it remains the reference against which newer methods are validated. A modern bench version for a small tissue sample follows directly from these proportions: Weigh 20 mg of frozen tissue powder into a glass tube on dry ice. Add internal standards in a small volume of solvent. Add 400 µL of ice-cold chloroform–methanol 2:1 (v/v) containing 0.01% BHT, giving the twentyfold solvent excess. Homogenise with a bead mill or probe homogeniser while keeping the tube cold. Agitate for 15 to 30 minutes at low temperature, then centrifuge to pellet insoluble material, and transfer the supernatant to a clean glass tube. For quantitative work, re-extract the pellet with a further portion of the same solvent and combine. Add 0.2 volumes of 0.9% NaCl relative to the combined extract, vortex, and centrifuge at low speed to separate the phases. Remove the upper phase with a glass pipette, taking care not to disturb the interface. Optionally rinse the interface with a small volume of theoretical upper phase and remove it again. Collect the lower phase, passing the pipette through the interface, and dry it under nitrogen. The volumes scale linearly. What should not change are the 2:1 ratio, the twentyfold excess and the final 8:4:3 proportion, because those determine the phase behaviour. The method has drawbacks that matter today. Chloroform is toxic and a suspected carcinogen, and the lower phase must be collected by passing a pipette through the protein interface and the upper phase, which makes the method awkward to automate and invites contamination from the interface. The large solvent volume dilutes the extract and must be evaporated. Chloroform can also contain phosgene and hydrochloric acid when stored without a stabiliser, and those can damage plasmalogens and other acid-labile lipids; chloroform stabilised with ethanol or amylene, stored in the dark, is standard. The Bligh–Dyer method Bligh and Dyer worked at the Fisheries Research Board of Canada on fish muscle, a tissue with high water content and relatively low lipid content, and their explicit aim was economy: to extract lipid with a smaller volume of solvent in less time. Their key insight was to treat the water in the tissue as part of the solvent system from the start. In their original procedure, 100 grams of tissue, assumed to contain about 80 millilitres of water, were homogenised with 100 millilitres of chloroform and 200 millilitres of methanol. With the tissue water, this gives chloroform–methanol–water at 1:2:0.8, a composition inside the single-phase region, so that extraction occurs in a monophasic system. A further 100 millilitres of chloroform and then 100 millilitres of water were added with further homogenisation, taking the system to chloroform–methanol–water 2:2:1.8, which separates into two phases. The lower chloroform phase, containing the lipids, was collected. The total solvent volume was about four volumes of tissue rather than the twenty of Folch. For a 100 µL plasma sample, the modern scaled version runs as follows. Add internal standards. Add 375 µL of chloroform–methanol 1:2 (v/v); with the roughly 100 µL of water in the plasma, this approximates the monophasic 1:2:0.8 condition. Vortex and allow to stand for ten to fifteen minutes for extraction. Add 125 µL of chloroform and vortex, then 125 µL of water and vortex, reaching approximately 2:2:1.8. Centrifuge to separate the phases and collect the lower phase. Re-extraction of the upper phase and interface with a further portion of chloroform improves recovery. The economy has a price. Because the total solvent volume is small relative to the sample, the method works well only when the sample is wet and lean. Sara Iverson and colleagues compared the two methods in marine tissues in 2001 and found that the Bligh–Dyer method gave results equivalent to Folch in samples containing less than about 2% lipid by mass, but increasingly underestimated total lipid as lipid content rose. The cause is solvent capacity: in fatty tissues there is not enough chloroform to dissolve all the neutral lipid, and some remains in the residue or partitions poorly. For adipose tissue, liver from animals fed high-fat diets, or any sample in which triacylglycerols dominate, Folch proportions are safer, or the Bligh–Dyer volumes must be increased. The Bligh–Dyer method also depends on an assumption about sample water content. If a tissue is drier than the assumed 80%, or if a sample is supplied as a small volume of a dilute suspension, the actual composition at the monophasic step differs from 1:2:0.8, and the final composition may drift from 2:2:1.8. Practitioners should calculate the water contribution of each sample type and adjust the added water accordingly, rather than following volumes blindly. The monophasic step is where Bligh–Dyer does its real work, and it rewards patience. In a single phase, every lipid molecule is in contact with a solvent capable of dissolving it, and the methanol has time to denature lipid-binding proteins and penetrate lipoprotein particles and membranes. Shortening this step to a quick vortex before adding the second portion of chloroform reduces recovery, particularly of lipids tightly associated with protein. Conversely, a long monophasic incubation at room temperature gives residual enzymes and oxygen more time to act before methanol has fully denatured them, which is one reason to keep solvents cold and to include an antioxidant. Many laboratories settle on ten to thirty minutes on ice with intermittent mixing, and they hold that duration constant across a study, because extraction time is one of the quiet variables that differs between batches processed by different people. A frequent error when scaling Bligh–Dyer to plasma is to forget that plasma is not pure water. Its lipid and protein content is small enough that treating 100 µL of plasma as about 100 µL of water is adequate, but concentrated samples, such as lipoprotein fractions or homogenates prepared in small buffer volumes, may depart from that assumption. Writing the volume budget for water explicitly in the protocol, sample by sample type, prevents the drift. Recoveries, modifications and troubleshooting Neither classic method is equally efficient for every class, and the losses are systematic. The most polar and acidic lipids are the ones at risk. Phosphatidic acid, phosphatidylinositol and its phosphorylated forms, lysophosphatidic acid, sphingosine-1-phosphate and gangliosides all partition partly into the upper phase under neutral conditions, and polyphosphoinositides can bind to the protein interface. Adding acid, for example hydrochloric acid to a low final concentration in the Bligh–Dyer system, protonates phosphate groups and drives these lipids into the organic phase. Acidified extractions are standard for phosphoinositide and lysophosphatidic acid work. Their disadvantage is that acid hydrolyses the vinyl ether bond of plasmalogens, producing lysophospholipids and fatty aldehydes, and can promote acyl migration in lysophospholipids. A laboratory interested in both phosphoinositides and plasmalogens needs two extractions, or at least an awareness that an acidified protocol will understate plasmalogens and overstate some lyso species. Free fatty acids and oxylipins are extracted by both methods, but their partitioning depends on pH, because a carboxylic acid is largely ionised at neutral pH and partitions partly into the aqueous phase. Very low-abundance mediators are also swamped by bulk lipid in a total extract. For these reasons mediator work typically abandons liquid–liquid extraction altogether in favour of protein precipitation followed by solid-phase extraction, as Chapter 10 describes. Recovery should be treated as an empirical property of a protocol in a particular matrix, not a theoretical expectation. Ana Reis and colleagues compared five extraction solvent systems for human low-density lipoprotein in 2013 and found that the choice of method changed recovery of particular classes, with Folch and Bligh–Dyer performing well for most classes but with differences at the polar and neutral extremes. The general lesson, consistent across comparative studies, is that no method is universally best and that a laboratory should validate its chosen method by spiking representative standards into its own matrix before extraction and measuring recovery. Internal standards added at the start of extraction correct for recovery only if they partition like the endogenous lipids they represent; a PC standard does not correct PI losses. Two practical details improve every chloroform-based protocol. The first is glass: chloroform extracts plasticisers from polypropylene tubes and pipette tips, so glass tubes with polytetrafluoroethylene-lined caps and glass pipettes are used wherever the lower phase is handled. The second is careful phase collection. The lower phase is best withdrawn with a glass Pasteur pipette or a syringe inserted through the upper phase and interface while applying gentle positive pressure, so that no upper-phase material enters the pipette; alternatively, the upper phase and interface can be removed completely first. Either way, the act of penetrating the interface is the step where cross-contamination occurs, and it is the step that the methyl tert-butyl ether method, described next, was designed to eliminate. Troubleshooting the phase split Most failures of chloroform extractions show up at the phase separation, and they have recognisable causes. A cloudy lower phase usually means water is dispersed in it, either because centrifugation was too gentle or too short, or because the composition sits too close to the phase boundary; a longer spin, a slightly larger aqueous wash, or a few minutes at low temperature normally clears it. A persistent emulsion, common with plasma, milk, egg yolk or detergent-containing samples, reflects amphiphiles stabilising droplets at the interface; adding salt to the aqueous phase, increasing centrifugal force, or briefly chilling the tube helps. A thick, fluffy interface indicates a large mass of denatured protein that traps lipid; re-extracting the interface and pellet with fresh lower-phase solvent recovers much of it. An apparent third phase usually signals a gross error in proportions, most often a missed solvent addition, and the safest response is to recalculate what was actually added and correct the composition rather than to proceed. A useful habit is to record, for each sample type, the volumes of both phases after separation. If the upper phase is consistently larger or smaller than expected from the nominal proportions, the sample is contributing more or less water than assumed, and the protocol should be adjusted. This single measurement catches most of the silent drift that makes extraction a source of batch effects. Choosing between them The choice between Folch and Bligh–Dyer is less about tradition than about sample and question. For small, lean, aqueous samples such as plasma, cell pellets and most soft tissues, a properly scaled Bligh–Dyer extraction gives recoveries close to Folch with less solvent. For fatty tissues, brain, or any work in which total neutral lipid must be quantitative, the Folch excess protects against saturation. For acidic signalling lipids, either can be acidified, at the cost of plasmalogen integrity. For mediators, neither is ideal. Above all, the laboratory should write down the proportions it actually achieves, including sample water, rather than the nominal ones, because the phase diagram does not care what a protocol intended. Chapter 3: Beyond Chloroform: MTBE, Butanol and Single-Phase Extraction By the early 2000s, lipidomics had begun to ask more of extraction than the classic methods were designed to give. Studies involved hundreds or thousands of samples rather than dozens, robotic liquid handlers were replacing hand pipetting, and mass spectrometers were sensitive enough that contamination from the protein interface and from plastics became visible in every spectrum. Chloroform, dense and toxic, sat awkwardly in this new world: the lipid-containing phase was at the bottom of the tube, beneath the aqueous phase and the protein disc, and reaching it required penetrating both. The methods that followed kept the chemical logic of Folch and Bligh–Dyer but changed the solvent so that the lipids ended up on top. The Matyash MTBE method In 2008 Vitali Matyash, Gerhard Liebisch, Teymuras Kurzchalia, Andrej Shevchenko and Dominik Schwudke published in the Journal of Lipid Research a method that replaced chloroform with methyl tert-butyl ether (MTBE). MTBE is an ether with a density of about 0.74 g/mL, lower than that of water, and it is immiscible with water while dissolving lipids across a wide polarity range when combined with methanol. The consequence is that in the two-phase system, the lipid-rich organic phase forms the upper layer. Insoluble material, including denatured protein, sediments to the bottom of the tube as a pellet beneath the aqueous phase, rather than forming a disc between the two liquid phases. The upper phase can be collected by simply drawing it off the top, without passing through anything, which is faster, cleaner and far easier to automate. The published procedure for a 200 µL sample is compact. Methanol, 1.5 mL, is added to the sample in a glass tube and vortexed. MTBE, 5 mL, is added and the mixture is incubated for one hour at room temperature with shaking. Phase separation is induced by adding 1.25 mL of water; after ten minutes at room temperature the sample is centrifuged at 1,000 g for ten minutes. The upper organic phase is collected, and the lower phase is re-extracted with 2 mL of a solvent mixture of MTBE, methanol and water at 10:3:2.5 by volume, which approximates the composition of the upper phase. The combined organic phases are dried and reconstituted for analysis. The ratio of MTBE to methanol, 10:3, and the order of addition, methanol first to denature proteins and disrupt lipid–protein complexes, then MTBE to dissolve the lipids, then water to split the phases, are the structural features to preserve when scaling. For modern small-volume work, laboratories scale the method down proportionally. A version for 20 µL of plasma might use 150 µL of cold methanol containing internal standards, then 500 µL of MTBE, shaking for the extraction period, then 125 µL of water, a short incubation, and centrifugation, with the upper phase transferred to a fresh vial or well. Polypropylene tubes and plates are more tolerable with MTBE than with chloroform, though plastics still contribute background ions and the choice of consumables should be validated with blanks. Matyash and colleagues compared MTBE extraction with Folch and Bligh–Dyer across several sample types and reported recoveries of the major lipid classes that were similar or equivalent. Subsequent comparative studies have broadly agreed that MTBE extraction performs comparably for most glycerophospholipids, sphingolipids and neutral lipids. The weaknesses that have been reported sit at the polar edge. The MTBE upper phase carries more water and methanol than a chloroform lower phase, and the most polar lipids, including some lysophospholipids and highly polar glycosphingolipids, can partition partly into the aqueous phase, so their recovery should be checked in the laboratory's own hands. Because the upper phase contains more water, it also carries more polar non-lipid metabolites, which can be an advantage for combined lipid and metabolite workflows and a disadvantage for clean lipid extracts. MTBE brings its own hazards. It is highly flammable and volatile, it forms peroxides more slowly than diethyl ether but can still accumulate them on long storage, and its volatility means that the upper phase volume can change if tubes are left open, altering concentrations before internal standards have done their work. Keeping tubes capped and working in a fume hood with fresh solvent are the practical safeguards. Its toxicity profile is considerably more favourable than chloroform's, which is one reason it has become a default in many laboratories. Switching methods without breaking a dataset Laboratories rarely choose an extraction method on a blank slate. More often they have years of data generated by Folch or Bligh–Dyer and want to move to MTBE for throughput or safety. The change is reasonable, but it must be bridged. Suppose a laboratory holds a long-running cohort of plasma samples extracted by a scaled Bligh–Dyer protocol and plans to switch to MTBE for new samples. A sound bridging design takes a set of study-like samples, ideally several dozen spanning the range of the cohort, and extracts each by both methods in the same batch, with the same internal standards added at the same stage, then analyses all extracts in a single randomised run. For each lipid species, the paired results show whether the methods agree, whether there is a constant or proportional bias, and whether agreement depends on concentration. Classes that agree within the method's precision can be combined across the change; classes that show systematic bias, most likely at the polar edge, need either a correction factor derived from the bridging data or a statement that values from the two eras are not directly comparable. What the laboratory should not do is simply switch and rely on internal standards to absorb the difference. Internal standards correct for recovery only when they partition like the analytes, and a method change that alters partitioning of lysophospholipids, for example, alters it for the endogenous species and the standard in the same direction only if the standard is itself a lysophospholipid of similar structure. The bridging experiment is the only way to know. Automation and throughput The practical attraction of methods with an upper organic phase is that a liquid handler can aspirate the phase from a fixed height without ever touching the pellet. In 96-well format, a typical automated workflow dispenses methanol with internal standards into the plate, adds sample, adds the organic solvent, seals and shakes, adds water, centrifuges the plate, and transfers a fixed volume of the upper phase to a fresh plate for drying or direct injection. Transferring a fixed volume rather than the whole phase sacrifices a little recovery but gains reproducibility, because the internal standards correct for the fraction taken. The design choices that matter are the aspiration height, which must leave a safety margin above the interface; the plate seal, which must resist MTBE or butanol vapour; and evaporation control, since volatile solvents in open wells lose volume quickly. A liquid handler also makes extraction timing uniform across samples in a way that manual work struggles to match, which reduces one source of batch effects. Butanol-based and single-phase methods Two further developments pushed extraction toward high throughput. The first was the BUME method, published by Lars Löfgren and colleagues in 2012, which uses butanol and methanol for the initial monophasic extraction and heptane and ethyl acetate to create the second phase. In outline, sample is extracted with butanol–methanol 3:1, then heptane–ethyl acetate 3:1 is added, followed by an aqueous solution of acetic acid to induce phase separation. The lipids partition into the upper heptane-rich phase. The method was explicitly designed for automation in 96-well format and avoids chlorinated solvents entirely. The mildly acidic aqueous phase improves recovery of acidic lipids, with the same caveat as any acidified extraction regarding plasmalogen stability, though the conditions are much milder than those of strong-acid protocols. The second development was the abandonment of phase separation altogether. In single-phase extractions, a sample is mixed with a water-miscible organic solvent that precipitates protein and dissolves lipids, the precipitate is removed by centrifugation, and the supernatant is analysed directly, often without drying. Isopropanol alone, butanol–methanol mixtures and methanol-rich mixtures have all been used. Zahir Alshehry and colleagues in 2015 described a single-phase extraction of plasma with 1-butanol–methanol 1:1 containing ammonium formate, using a very small sample volume, and showed recoveries across a broad range of lipid classes suitable for large-scale clinical lipidomics. Single-phase methods are fast, need minimal handling, and recover polar lipids that two-phase methods can lose to the aqueous layer. Their limitation is the flip side of that inclusiveness. Nothing is partitioned away, so salts, sugars, amino acids and other polar metabolites remain in the extract, and so does any hydrophilic contaminant. In liquid chromatography, where these polar components elute early and away from most lipids, this is often tolerable. In shotgun lipidomics, where everything enters the ion source together, salts and polar metabolites add to ion suppression and adduct complexity, and a two-phase extraction is usually preferred. Single-phase methods also recover very non-polar lipids less completely in some solvent systems, because a methanol-rich solvent is a poor medium for bulk triacylglycerol and cholesteryl ester; isopropanol and butanol are better in this respect, which is why they feature in the most successful single-phase protocols. Extraction for mediators: protein precipitation and solid-phase extraction Eicosanoids and related mediators are a special case that no total-lipid extraction serves well. They are present at concentrations many orders of magnitude below those of phospholipids, they are carboxylic acids whose partitioning depends on pH, and they sit in a matrix of structurally similar free fatty acids and oxidised lipids. The standard approach, developed and refined over decades in eicosanoid laboratories, is to precipitate protein with cold methanol, which also halts enzymatic activity, and then to capture the mediators on a reversed-phase solid-phase extraction (SPE) cartridge. In outline, the sample receives deuterated internal standards and several volumes of cold methanol and is held at low temperature to allow protein precipitation. After centrifugation, the supernatant is diluted with acidified water so that methanol is a minor component and the pH is around 3.5, at which the carboxylic acids of eicosanoids are largely protonated and bind to C18 sorbent. The diluted sample is loaded onto a conditioned C18 cartridge, which retains the mediators while salts and polar material pass through. A water wash removes remaining polar contaminants, and a hexane wash removes some neutral lipids. The mediators are then eluted with a solvent such as methyl formate or methanol, dried under nitrogen and reconstituted in a small volume of the starting mobile phase. Protocols from the Serhan laboratory, such as the one described by Romain Colas and colleagues in 2014, use methyl formate elution after the hexane wash; other laboratories use polymeric mixed-mode or hydrophilic–lipophilic balanced sorbents with their own wash and elution schemes. The important features are the early addition of labelled standards, the acidification step, the removal of bulk lipid, and the use of glass or low-binding consumables to minimise adsorption. Acidification is itself a hazard for some mediators. Cysteinyl leukotrienes and certain epoxides are acid-sensitive, and prolonged exposure at low pH can degrade them. The acid step should therefore be brief and cold, and samples should move promptly through loading. Automated SPE in 96-well format makes this easier to standardise. Recovery must be measured for each class of mediator using spiked standards, because it varies with polarity and with the sorbent. Choosing and validating a method Table 1 summarises the practical differences between the main extraction strategies. It is a guide to trade-offs rather than a ranking, and the recoveries it describes are qualitative patterns drawn from the original methods and comparative studies, not universal numbers. Table 1. Principal lipid extraction strategies compared. Method Solvent system Lipid phase Main strengths Main limitations Folch (1957) CHCl3/MeOH 2:1, 20 vol; final 8:4:3 with water Lower Robust to sample fat and water; reference method Chloroform; interface penetration; large volumes Bligh–Dyer (1959) CHCl3/MeOH/H2O 1:2:0.8, then 2:2:1.8 Lower Less solvent; fast for lean wet samples Underestimates lipid in fatty samples Matyash MTBE (2008) MeOH, then MTBE; MTBE/MeOH 10:3 Upper No chloroform; easy collection; automatable Some polar lipid loss; volatile, flammable BUME (2012) BuOH/MeOH 3:1, heptane/EtOAc 3:1, dilute acid Upper Chloroform-free; 96-well automation Acid step; more complex solvent set Single phase e.g. BuOH/MeOH 1:1 or isopropanol None Fast; good polar recovery; tiny volumes Salts and metabolites retained; matrix effects Protein precipitation + SPE MeOH, acidify, C18 or polymeric sorbent Eluate Concentrates trace mediators; removes bulk lipid Class-specific recovery; acid-labile analytes For global lipidomics of plasma, a scaled MTBE extraction or a single-phase butanol–methanol extraction are reasonable defaults, with MTBE preferred when the extract will be infused directly and single-phase preferred when throughput and polar coverage dominate and chromatography will separate the matrix. For tissues rich in triacylglycerol, such as adipose tissue or steatotic liver, Folch proportions protect against incomplete neutral lipid extraction, and the sample should be diluted so that the solvent excess is maintained. For brain and myelin-rich tissue, Folch remains the reference, particularly if gangliosides and other complex glycosphingolipids are of interest, and the upper phase should be kept for ganglioside analysis rather than discarded. For phosphoinositides and lysophosphatidic acid, an acidified extraction is required. For eicosanoids and other oxylipins, protein precipitation followed by SPE is the standard. When a study needs both global lipids and mediators from the same small sample, one option is to split the sample before extraction. Another is to perform a two-phase extraction for global lipids and route the upper phase, or a separate aliquot of the methanolic precipitate, to SPE. The second approach conserves sample but complicates recovery calculations, and it should be validated with spiked standards for both workflows. Validating a chosen method Whatever method is chosen, three validation experiments should be done before the first study sample is extracted. The first is a recovery experiment, in which non-endogenous or labelled standards for each class of interest are spiked into the matrix before extraction and, separately, into a blank extract after extraction; the ratio of the two responses measures extraction recovery independently of matrix effects on ionisation. The second is a matrix effect experiment, in which the same standards spiked into an extracted matrix are compared with standards in neat solvent, which measures suppression or enhancement of ionisation. The third is a reproducibility experiment, in which a pooled sample is extracted several times on different days by the people who will run the study, to measure the combined variance of the whole process. These experiments are routine in regulated bioanalysis and are increasingly expected in lipidomics. They also reveal problems that no reading of a protocol can: a batch of tubes that leaches a contaminant, a centrifuge that does not reach the specified force, an evaporator that overheats, a liquid handler that aspirates part of the interface. The phase diagram sets what is possible; validation shows what a laboratory actually achieves. A final practical point concerns reconstitution. The dried extract must be redissolved in a solvent compatible with the downstream analysis and capable of dissolving all the lipids that were extracted. A solvent too polar for triacylglycerols or cholesteryl esters leaves them on the vial wall; a solvent too strong for the chromatographic starting conditions distorts early-eluting peaks. Mixtures such as isopropanol with methanol or acetonitrile, or chloroform–methanol for direct infusion, are common. The reconstitution volume sets final concentration and must be delivered accurately, because it multiplies into every result. Recovery validation should include this step, since incomplete reconstitution looks exactly like incomplete extraction in the final numbers. Hashtags: #LipidomicsProtocols #Lipidomics #LipidExtraction #ExtractionChemistry #MassSpectrometry #LipidIdentification #LipidSignaling #LIPIDMAPS #Phospholipids #Sphingolipids #Glycerolipids #Oxylipins #Eicosanoids #FolchExtraction #BlighDyerExtraction #MTBEExtraction #SolidPhaseExtraction #InternalStandards #LipidIsomers #LipidFragmentation #ShotgunLipidomics #LCMSLipidomics #LipidQuantification #InflammatoryMediators #FutureOfLipidomics
- Laboratory Automation Engineering (Integrating Python, PLCs, and Scientific Hardware)
Download the Book (PDF): Introduction Most automated laboratory systems do not fail because someone wrote the wrong chemistry into them. They fail at three in the morning because a USB hub reset, a serial reply arrived half a line late, a stepper motor lost forty steps against a sticky syringe plunger, or a thermocouple wire came loose and the heater controller, reading an open circuit, decided the block was cold and drove it harder. The science in the protocol was fine. The engineering underneath it was not. This book is about that engineering. It is written for researchers who have outgrown manual pipetting and vendor software that does almost what they need, and who have decided to build their own automated workflows: a syringe pump that doses reagent in response to a pH reading, a plate handler that feeds a spectrometer overnight, a reactor rig whose temperature, stirring, and sampling are coordinated from a Python script, a small cell culture station supervised by a programmable logic controller. It is also written for the engineers and technicians who end up supporting those systems once the graduate student who built them has moved on. The controlling argument is simple to state and harder to practise. An automated laboratory system is trustworthy only when every layer of it, from the voltage on a wire to the state machine in the orchestration script, has been designed around how it fails rather than how it works. Getting an instrument to respond to a command is the easy part and usually takes an afternoon. Knowing what the system will do when the instrument does not respond, responds with garbage, responds to a command sent to a different instrument, or responds correctly but too late, is the actual work. A workflow that runs perfectly on the bench while you watch it is a demonstration. A workflow that runs unattended for a week, and stops safely and informatively when something goes wrong, is a system. Why researchers end up doing this Commercial laboratory automation is excellent at what it was designed for. Integrated liquid handling platforms, plate readers with robotic stackers, and turnkey bioreactor controllers encode years of engineering, and where one of them fits a lab's needs it is almost always the better choice. The trouble is that research rarely stays within the envelope a vendor anticipated. The assay changes. A new detector arrives with its own control software that cannot talk to the old one. A reaction needs a dosing strategy that no commercial controller offers. The budget covers a pump, a balance, and a temperature controller, but not a platform that integrates them. At that point researchers reach for Python, and for good reasons. The language has mature libraries for exactly the interfaces laboratory hardware exposes: pyserial for serial ports, PyVISA for test and measurement instruments, pymodbus and asyncua for industrial controllers, and the whole numerical and data ecosystem for what happens to the measurements afterwards. Scripts are quick to write and easy to read. Open-source projects such as PyLabRobot for liquid handlers and Bluesky for experimental orchestration at synchrotron facilities have shown that Python can sit credibly at the centre of serious automated experiments. What Python does not provide is determinism, electrical knowledge, or safety. A Python process running on a desktop operating system can be paused for hundreds of milliseconds by garbage collection, antivirus scanning, or a Windows update. It knows nothing about ground loops, cable capacitance, or the difference between a normally open and a normally closed contact. And it cannot be the thing that stops a heater from boiling a solvent dry, because nothing running on a general-purpose computer should be trusted with that job alone. The engineering discipline of laboratory automation is largely the discipline of knowing which responsibilities belong to Python, which belong to firmware, programmable controllers, and dedicated hardware, and how those layers should watch one another. What the book covers The chapters move roughly from the wire upward and then back down to safety, because safety constrains everything above it. Chapter 1 sets out an architecture for laboratory automation: the layers of a typical system, the difference between supervisory and real-time control, and the question of where a PLC earns its place alongside a Python host. Chapters 2 and 3 cover the physical and logical plumbing that most laboratory instruments still use. Chapter 2 treats serial communication, including RS-232, RS-422, and RS-485, character framing, flow control, and the design of robust request-and-response code in pyserial. Chapter 3 covers USB, which usually means a serial port in disguise, but also brings device enumeration, stable naming, latency timers, and the peculiar failure modes of hubs and power management. Chapter 4 turns to VISA and SCPI, the conventions that let a single Python library control oscilloscopes, source meters, and spectrum analysers from many manufacturers through one consistent interface, and explains how to use the status system of IEEE 488.2 instead of guessing with sleep calls. Chapter 5 connects the Python world to industrial controllers through Modbus and OPC UA, and describes the handshake patterns that keep a PLC and a host computer from misunderstanding each other. Chapters 6 and 7 are about the physical world. Chapter 6 covers stepper motors in liquid handling: step resolution, microstepping, acceleration profiles, homing, lost steps, and the arithmetic that turns microsteps into microlitres. Chapter 7 covers sensor interfaces: analogue current loops and voltages, thermocouples and resistance thermometers, digital buses such as I2C and SPI, sampling, filtering, calibration, and time. Chapter 8 addresses the orchestration layer: how to structure a Python program that coordinates many devices so that it can be tested without hardware, recovered after a crash, and audited after the fact. Chapter 9 covers safety watchdogs and interlocks, from the heartbeat that a PLC expects from its host to the hardwired emergency stop that does not care whether any software is running at all. How to read it Each chapter stands on its own, and a reader with an urgent problem, such as a serial device that works in a terminal program but not from a script, can go directly to the relevant chapter. But the chapters were written to accumulate. The patterns introduced for serial communication reappear in the handling of VISA instruments and Modbus registers. The state-machine thinking in the orchestration chapter depends on the failure analysis done for each device. The safety chapter assumes the reader now knows enough about each layer to see why none of them can be trusted in isolation. The code examples are short and deliberately plain. They use current versions of the common libraries and are meant to teach a pattern rather than serve as a finished driver. Where library interfaces have changed between versions, as they have for pymodbus in particular, the text says so. Instrument command strings are given as representative examples; every real instrument has a manual, and the manual always wins. A few scope decisions deserve stating plainly. The book assumes a working knowledge of Python and basic electrical literacy, meaning comfort with voltage, current, resistance, and a multimeter. It does not teach PLC programming in depth, although it explains enough about how controllers think to integrate them well. It does not cover regulated environments in detail. Laboratories operating under good manufacturing practice or clinical regulation face requirements for computer system validation and electronic records that go beyond anything here, although the habits of logging, versioning, and traceability recommended throughout are a good foundation for them. And it treats safety standards as a map of the territory, not a substitute for a qualified person's assessment of a specific machine. A word about attitude There is a particular mindset that separates automation that lasts from automation that has to be babysat. It is a mild, practical pessimism. Every cable will eventually be unplugged. Every instrument will eventually return an error the documentation does not mention. Every timeout that seems generous will eventually be too short, and every retry loop without a limit will eventually retry forever. Power will fail in the middle of a dispense. Someone will run the script twice at once. None of this is exotic; all of it happens in ordinary laboratories every month. The good news is that designing for these events is not especially difficult once it becomes a habit. It means writing down what the safe state of each device is. It means putting timeouts on every wait, bounds on every retry, and checks on every reply. It means logging enough to reconstruct what happened. It means letting hardware do what hardware is good at, which is reacting quickly and reliably without an operating system, and letting software do what it is good at, which is coordinating, recording, and deciding. And it means accepting that a system which stops and explains itself is doing its job, even if the experiment is lost, because the alternative is a system that carries on confidently in the wrong state. The rest of this book is an attempt to make that habit concrete, one layer at a time. Chapter 1: The Architecture of an Automated Laboratory Before any code is written or any cable is crimped, an automated laboratory system has an architecture, whether its builder chose one or not. The unchosen architecture is familiar to anyone who has inherited a rig: a single Python script of several thousand lines, talking directly to six devices over five different interfaces, with timing held together by calls to time.sleep, and safety provided by a sticky note on the fume hood. It works until the person who wrote it leaves. Choosing an architecture deliberately is mostly a matter of deciding who is responsible for what, and at what speed each responsibility has to be met. Layers and time scales A useful way to think about any automated system is as a stack of control loops, each running at its own characteristic time scale. At the bottom are the loops that must react within microseconds to milliseconds: the current regulation inside a stepper motor driver, the PID loop inside a temperature controller, the pulse generation that moves a motor at a smooth velocity. In the middle are loops that must react within milliseconds to a second or so: a PLC scanning its inputs and deciding whether a door has opened, a pump controller checking whether a pressure limit has been exceeded, an interlock deciding to drop power to a heater. At the top are loops that operate over seconds to days: the orchestration logic that decides which plate to move next, the adaptive algorithm that chooses the next reaction conditions, the scheduler that runs a protocol overnight. The central rule of laboratory automation architecture is that each loop should run on something that can meet its time scale reliably, not merely on average. A desktop Python process can usually respond to an event within a few milliseconds. It cannot guarantee that it will, because the operating system scheduler, memory management, disk activity, and other programs can all delay it by amounts that are rare but not negligible. That is perfectly acceptable for deciding which well to aspirate from next. It is not acceptable for generating step pulses, closing a temperature loop on a small thermal mass, or reacting to an overpressure. This leads to a layered picture that recurs in almost every well-built system, shown in Table 1. The names differ between laboratories, but the division of labour does not. Table 1. Layers of a typical automated laboratory system. Layer Typical hardware Time scale Main responsibility Orchestration PC or server running Python Seconds to days Scheduling, decisions, data, logging Supervisory control PLC or microcontroller Milliseconds Sequencing, interlocks, device coordination Device control Instrument firmware, motor drivers Microseconds to milliseconds Closed loops, motion profiles Safety Safety relays, hardwired circuits Fixed, deterministic Removing energy when limits are breached Physical Sensors, actuators, wiring Continuous Converting between signals and the world The table places safety as its own layer rather than as a feature of any other layer, and that is deliberate. Safety functions, which Chapter 9 treats in detail, must work even when every layer above them has failed. They are not the orchestration script's job, and in most research systems they should not even be the PLC's job unless that PLC is a safety-rated controller configured by someone who knows what that entails. Supervisory control and the case for a PLC Many research groups build their first automated system with nothing between the Python host and the instruments. Each device has its own controller already, whether that is a syringe pump with embedded firmware, a temperature controller with its own PID loop, or a spectrometer with its own acquisition electronics. The host sends high-level commands and reads results. For a large class of experiments this is entirely adequate. If the host crashes, each instrument stops wherever it was, generally in a benign state, and nothing dangerous happens. The architecture changes when the system acquires any of three properties. The first is coordination that must be faster or more reliable than the host can provide: a valve that must close within fifty milliseconds of a level sensor tripping, or a gripper that must not open until a position sensor confirms it is over the target. The second is a meaningful hazard whose control depends on logic rather than on a single hardwired limit, such as a sequence that must purge a vessel with nitrogen before a heater is allowed to energise. The third is longevity: a system expected to run for years, maintained by people who were not involved in building it, in a setting where the host computer will be replaced, patched, and occasionally reinstalled. In any of these cases a programmable logic controller, or a comparably robust embedded controller, earns its place. A PLC is a computer designed around a single job: read all inputs, execute the control program, write all outputs, and repeat, typically every few milliseconds, forever. The program is written in one of the languages defined by IEC 61131-3, most often ladder diagram or structured text, and it runs without an operating system in the ordinary sense. The hardware is built for electrically noisy environments, with isolated inputs and outputs rated for 24 volt industrial signalling. And crucially, the controller has a defined behaviour when things go wrong: a watchdog that faults the processor if the scan overruns, configurable output states on fault, and an explicit notion of whether the program is running or stopped. The price of this robustness is expressiveness. PLC programs are good at sequences, interlocks, timers, and counters. They are poor at string handling, complex data structures, numerical analysis, and anything involving files or networks beyond their fieldbus. That is exactly why the PLC and Python complement each other. The PLC owns the fast, safety-relevant, and repetitive logic; Python owns scheduling, decisions, and data. The interface between them, usually Modbus TCP or OPC UA, is covered in Chapter 5. A smaller system can apply the same principle with a microcontroller in place of a PLC. An Arduino-class board or a Raspberry Pi Pico running a tight firmware loop can generate step pulses, watch limit switches, and enforce a heartbeat timeout from the host. The trade-off is that the builder now owns the firmware, its electrical robustness, and its failure behaviour, all of which a commercial PLC provides by design. For a benchtop prototype that is often the right trade. For a system that will run unattended with hazardous materials it usually is not. Command, status, and the problem of shared belief Every layer boundary in an automated system is a place where two pieces of software must share a belief about the state of the world. The orchestration script believes the syringe pump is at position zero with its valve set to the reservoir. The pump's firmware has its own belief. The physical plunger is wherever it actually is. Most serious failures in laboratory automation are, at root, cases where these beliefs diverged and nothing noticed. The architectural defence is to be explicit about the difference between a command and a status. A command is a request that something change. A status is an observation of what is true now. A well-designed interface never assumes that a command succeeded because it was sent. It confirms by reading status: the pump reports that it is idle and at the requested position, the PLC reports that the valve's position switch has made, the temperature controller reports that its process value has entered the tolerance band. Where the hardware provides independent confirmation, such as a position encoder on a motor that is otherwise driven open-loop, the software should use it. A second defence is to make ownership of each piece of state unambiguous. If the PLC runs the sequence that fills a vessel, then the PLC owns the knowledge of whether the vessel is full, and the host asks rather than tracks it separately. If the host decides which reagent to dispense, then the host owns that decision and the pump merely executes. Duplicated state, where two layers each keep their own copy of the same fact, is a standing invitation to inconsistency. When duplication is unavoidable, as when the host caches instrument settings to avoid querying them constantly, the cache must be invalidated whenever anything happens that might change the truth, including a reconnection, a power cycle, or a fault. A third defence is to design every device interface around the question of what happens when communication stops. Suppose the host is halfway through a sequence and its process is killed. What does each device do? A stepper driver with no further pulses simply holds position, which is usually fine. A heater controller holding its last setpoint will continue heating indefinitely, which may not be fine. A PLC waiting for the next instruction from the host may wait forever, or, if it has been designed well, notice that the host's heartbeat has stopped and move to a defined safe state. Writing down the answer to this question for every device is one of the most valuable hours an automation engineer can spend. Choosing interfaces The physical and logical interfaces a system uses are rarely chosen freely. They are mostly dictated by the instruments the laboratory already owns. Still, where there is a choice, some general preferences hold. Networked interfaces, meaning Ethernet with TCP/IP, are generally preferable to serial or USB for anything that will run for a long time. They tolerate longer cable runs, are galvanically isolated by design through the transformers in every Ethernet port, survive host reboots without renumbering, and can be monitored with standard tools. An instrument with an Ethernet port speaking SCPI over raw sockets or HiSLIP, or a PLC speaking Modbus TCP, is usually easier to integrate robustly than the same device over USB. Serial interfaces remain ubiquitous in laboratory equipment and are not going away. Balances, pumps, stirrers, temperature controllers, and many older analytical instruments speak ASCII over RS-232, sometimes RS-485. These interfaces are simple enough to understand completely, which is a real advantage, but they offer no error detection beyond what the device protocol adds and no discovery of any kind. Chapter 2 covers how to use them well. USB is the most convenient interface to plug in and the most troublesome to keep running. It brings enumeration, power management, hubs, and drivers into the picture, each of which introduces failure modes that do not exist with a simple serial line or a network socket. Chapter 3 covers how to tame it. Where a device offers both USB and Ethernet, Ethernet is almost always the better long-term choice. Industrial fieldbuses such as EtherCAT, PROFINET, and EtherNet/IP appear when PLCs and servo drives enter the picture. Python can talk to some of them directly, but it usually should not. The PLC should own the fieldbus, and Python should talk to the PLC. A reference architecture Pulling these ideas together gives a reference architecture that suits a large range of research systems and can be scaled down by omitting layers. At the top sits an orchestration process written in Python. It runs protocols, schedules steps, logs every command and response, stores data with metadata, and presents some interface to the operator, whether that is a command line, a notebook, or a small web dashboard. It talks to devices through a layer of driver classes, one per device type, each of which hides the device's communication details behind methods such as aspirate(volume_ul) or read_temperature(). The drivers can be replaced by simulated versions for testing without hardware, a point developed in Chapter 8. Beneath the orchestration layer, devices fall into two groups. Smart instruments with their own embedded controllers, such as spectrometers, balances, and plate readers, connect directly to the host by serial, USB, or network. Actuators and sensors that need fast coordination, such as valves, motors, level switches, and door interlocks, connect to a PLC or microcontroller, which in turn connects to the host over a single well-defined interface. Beneath everything, a hardwired safety circuit monitors the conditions that could cause harm, including an emergency stop button, enclosure door switches, and independent over-temperature cutouts, and removes power from the relevant actuators when any of them trips. It does not depend on the host, the PLC, or any software. It can report its state to the PLC so that the software knows why the world has stopped, but it does not ask permission. This architecture has a useful property: every layer can fail without making the layers beneath it unsafe. If the host crashes, the PLC detects the missing heartbeat and brings the process to a controlled stop. If the PLC faults, its outputs go to their configured fault state and the safety circuit still holds. If a sensor fails, the PLC logic sees an out-of-range signal and stops rather than acting on a false reading. Each layer trusts the one below it to be simpler and more reliable than itself, and each layer watches the one above it. The rest of this book fills in the details of each layer. But the architecture comes first, because no amount of careful driver code will rescue a system in which the Python script is the only thing standing between a heater and a fire. Applying the architecture to a dosing rig Consider a hypothetical but typical project. A group wants to run pH-controlled reactions overnight in a jacketed glass vessel. The equipment is a pH meter with an RS-232 output, a syringe pump with its own serial command set, an overhead stirrer controllable over USB, a recirculating chiller-heater that speaks a simple ASCII protocol, and a balance under a reagent bottle to confirm how much has actually been dispensed. The first instinct is to plug all five into a laptop and write a loop that reads the pH, decides how much base to add, commands the pump, and sleeps. Walking through the questions above changes the design. What happens if the laptop crashes with the pump mid-dispense? The pump completes its current move and stops, which is benign. What happens if the laptop crashes while the chiller is heating? The chiller holds its setpoint, which is benign if the setpoint was sensible and the circulator has its own over-temperature cutout, which it should. What happens if the pH probe fails and reads a constant value that looks acidic? The loop keeps adding base indefinitely, which is not benign, and here the architecture needs something more than a script. The fix does not require a PLC. It requires limits that live somewhere other than in the dosing logic. The orchestration software can track cumulative dispensed volume and refuse to exceed a protocol maximum. The balance provides an independent check that the dispensed mass matches the commanded volume. A plausibility check can reject a pH reading that has not changed at all over a period in which base has been added, since a live electrode always drifts a little. And the reagent reservoir itself can be sized so that emptying it completely cannot produce a dangerous condition, which is a physical limit no software failure can override. Suppose the group later adds an exothermic step in which the reaction could run away if cooling fails. Now the question of what happens when things stop has a harder answer. A loss of chiller flow must stop the dosing within seconds, regardless of whether the laptop is alive. That requirement belongs below the orchestration layer: a flow switch on the coolant line wired to an interlock that cuts power to the pump, or a small controller that monitors flow and temperature and holds the pump's enable input low when either is out of bounds. The Python script still decides how much to dose and when. It simply no longer holds the only key to stopping. This pattern, in which each new hazard pushes a responsibility downward into a simpler and more reliable layer, is the essence of the architecture. It rarely requires expensive hardware. It requires asking, for every failure, which layer is the right one to notice it and act. Chapter 2: Serial Communication Done Properly Serial communication is the oldest interface in most laboratories and still the most common. Balances, syringe pumps, peristaltic pumps, stirrers, hotplates, temperature controllers, pH meters, mass flow controllers, and a long tail of analytical instruments all expose a serial port, and many of them expose nothing else. The protocols involved are simple enough to understand completely, which makes serial devices a good place to learn the habits that the rest of this book depends on: never trusting a reply you have not validated, never waiting without a timeout, and never assuming the other side is in the state you last left it in. The physical layer "Serial" describes a family of standards that share a way of sending characters one bit at a time but differ in how those bits are represented electrically. Confusing them is the source of a surprising number of laboratory mysteries, including devices that were damaged by being connected to the wrong kind of port. At the heart of every serial port is a UART, a universal asynchronous receiver-transmitter, which converts bytes to a sequence of bits and back. Inside a microcontroller or a USB adapter chip, the UART's signals are logic levels, typically 3.3 or 5 volts, with a high level meaning a logical one and the line resting high when idle. This is often called TTL serial. It is intended for connections of a few centimetres on a circuit board or a short cable between two boards, and it is what appears on the pin headers of development boards. RS-232, standardised as TIA-232, is the classic external serial port found on older computers and on most laboratory instruments with a nine-pin D connector. It inverts and amplifies the logic levels: a logical one, called a mark, is a negative voltage between about minus 3 and minus 15 volts, and a logical zero, called a space, is positive in the same range. Connecting an RS-232 port directly to a 3.3 volt microcontroller pin can destroy the pin, and connecting a TTL output to an RS-232 input usually just fails silently because the voltage never goes negative. A level-shifting transceiver chip, of which the MAX232 family is the classic example, sits between the two worlds. RS-232 is single-ended: each signal is a voltage measured against a shared ground. That makes it vulnerable to noise and to ground potential differences between instruments, and it limits practical cable lengths. The original standard was framed in terms of cable capacitance, which in practice allows around fifteen metres at modest baud rates, less at higher ones. In a laboratory full of pumps, heaters, and switching power supplies, shorter is better. RS-422 and RS-485 solve the noise problem by sending each signal as the difference between two wires, a twisted pair. Noise that couples equally into both wires cancels at the receiver. These standards support cable runs of up to around 1,200 metres at low data rates and, in the case of RS-485, multiple devices sharing one pair of wires. RS-485 is the physical layer beneath Modbus RTU, beneath many multi-channel pump systems, and beneath a good deal of industrial instrumentation. It requires attention to details that RS-232 does not: a 120 ohm termination resistor at each end of the bus to prevent reflections, bias resistors to hold the line in a defined state when no device is transmitting, and a common reference conductor despite the differential signalling. Table 2 summarises the differences. Table 2. Common serial physical layers compared. Standard Signalling Typical reach Devices per link Common laboratory use TTL UART Single-ended logic levels Centimetres Two Microcontroller boards, modules RS-232 Single-ended, bipolar voltages About 15 m Two Balances, pumps, older instruments RS-422 Differential, one driver Up to about 1,200 m One driver, up to 10 receivers Long point-to-point links RS-485 Differential, shared bus Up to about 1,200 m 32 standard unit loads Modbus RTU, multi-drop pump chains On RS-232 connectors, the most important practical fact is the distinction between data terminal equipment (DTE), such as a computer, and data communication equipment (DCE), historically a modem. On a nine-pin DTE connector, pin 2 receives, pin 3 transmits, and pin 5 is signal ground. A DCE device swaps pins 2 and 3 so that a straight-through cable connects transmit to receive. Laboratory instruments are inconsistent about which role they play. When an instrument is wired as DTE, connecting it to a computer requires a null-modem cable that crosses the data lines. The fastest diagnostic for a silent serial link is a multimeter: with the port idle, the transmit pin of an RS-232 device should sit at a negative voltage of several volts relative to ground. Whichever pin shows that voltage is that device's transmit line. Framing, baud rate, and flow control Once the electrical layer is right, both ends must agree on framing. Each character is sent as a start bit, five to eight data bits (almost always eight or seven in practice), an optional parity bit, and one or two stop bits. The common shorthand 9600 8N1 means 9,600 bits per second, eight data bits, no parity, and one stop bit, which makes ten bits per character and therefore about 960 characters per second. Some older instruments default to 7E1, seven data bits with even parity, and some balances offer a menu of settings that must be matched exactly. A mismatch in baud rate produces garbage; a mismatch in parity or data bits often produces text that is almost right, with some characters consistently wrong, which is a useful clue. Parity catches single-bit errors in a character but not much else, and it provides no way to detect a missing or extra character. Serious error detection must come from the device protocol, such as a checksum or CRC appended to each message. Where a protocol offers one, the host should always verify it. Flow control determines what happens when one side cannot keep up. Hardware flow control uses the RTS and CTS lines, which each side asserts to indicate that it is ready to receive. Software flow control uses the XON and XOFF characters, which is incompatible with binary data. Most laboratory devices use neither and rely on the host to send commands slowly and wait for replies, which works because messages are short. The danger is a device that does expect hardware flow control, with a cable that does not connect RTS and CTS: the host's transmissions may never be sent at all. When a device ignores commands, checking the flow control settings is a close second after checking the wiring. Several instruments also use the DTR or RTS lines for purposes other than flow control. Opening a port in pyserial asserts DTR by default, and on many microcontroller boards, DTR is wired to reset the processor, so opening the port reboots the device. This is a common reason why the first command sent after opening a port to an Arduino-class board is lost: the board is still starting up. Either wait for the firmware to announce that it is ready, or configure the port so that DTR is not asserted, depending on which behaviour is wanted. Messages, terminators, and the read problem Serial ports deliver a stream of bytes with no inherent message boundaries. A device that sends +0012.345 g\r\n might deliver it to the host in one read, or in three, or together with the beginning of the next reading. Every serial protocol therefore defines how messages are delimited. Most laboratory instruments use a terminator, typically a carriage return, a line feed, or both. Some binary protocols use a fixed length, a length field in a header, or special start and end bytes. Modbus RTU, covered in Chapter 5, uses silence on the line. The single most common bug in laboratory serial code is reading without regard for message boundaries. A script sends a command, sleeps for a fixed time, and reads whatever is in the buffer. This works until the device is slower than usual, when the read returns a partial reply, or until a stale reply from an earlier command is still sitting in the buffer, when the script parses the wrong message and continues confidently. The fix is to read until the terminator, with a timeout, and to treat anything else as an error. The following function, using pyserial, shows the core pattern for an ASCII request-and-response device. import serial class SerialDeviceError(Exception): pass def open_port(path, baud=9600): return serial.Serial(path, baudrate=baud, bytesize=8, parity=serial.PARITY_NONE, stopbits=1, timeout=1.0, write_timeout=1.0) def transact(port, command, terminator=b"\r\n"): port.reset_input_buffer() # discard stale bytes port.write(command.encode("ascii") + b"\r") reply = port.read_until(terminator) # stops at terminator or timeout if not reply.endswith(terminator): raise SerialDeviceError( f"timeout or partial reply to {command!r}: {reply!r}") return reply[:-len(terminator)].decode("ascii") Several details carry real weight. The timeout passed to the port applies to read_until, which returns whatever it has when the timeout expires; the function checks for the terminator and raises if it is missing, rather than returning a fragment. The write_timeout prevents the script from blocking forever if hardware flow control is holding the port. Clearing the input buffer before sending discards unsolicited or stale data, so that the reply read is the reply to this command. For devices that stream data continuously, clearing the buffer is the wrong move, and a different pattern is needed, described below. The terminator on the command and the terminator on the reply are often different, and the manual must be read carefully. Some devices echo each command before replying, which means the first line read is the command itself. Some reply with an acknowledgement character, such as ASCII ACK (0x06), or a negative acknowledgement, NAK (0x15), rather than text. Some send nothing at all in response to a set command, which means the host cannot distinguish success from a lost message without a follow-up query. When a protocol allows it, confirming every write with a read of the resulting state is worth the extra round trip. Validating replies A reply that arrived with the right terminator can still be wrong. It can be a reply to a different command, a garbled message that happens to end in a line feed, an error code the host has not anticipated, or a correct message reporting that the device is in an unexpected state. Robust code parses replies strictly and rejects anything that does not match the expected form. Suppose a balance, queried with a print command, returns a line such as S S 12.3456 g, in which the first field is a command echo, the second a stability flag, and the rest the value and unit. This is the general shape of the widely used MT-SICS protocol on many laboratory balances, although the exact format should always be checked against the specific balance's manual. A strict parser checks that the line has the expected fields, that the stability flag indicates a stable reading if a stable reading is required, that the value parses as a number, and that the unit is the one expected. It raises a specific exception otherwise. It does not strip everything that is not a digit and hope for the best, which is how a reading of S I (unstable) or an overload indicator ends up being logged as a mass. A good driver distinguishes three classes of failure. Transport failures, such as timeouts and partial messages, suggest a communication problem and may justify a bounded retry. Protocol failures, such as a malformed reply, suggest that the host and device are out of step, and the right response is usually to resynchronise, by clearing buffers and sending a harmless query, before retrying. Device errors, such as an explicit error code or a reply indicating that the device refused the command, must not be retried blindly: the device is telling the host something about the world, and the orchestration layer needs to hear it. Retries deserve a specific warning. Retrying a query is safe. Retrying a command that causes motion or dispensing is not, unless the host can determine whether the first attempt was executed. If a pump receives a dispense command, executes it, and the reply is lost, a naive retry dispenses twice. The safe pattern is to query the device state after a failed command and decide based on what it reports, or to use absolute commands, such as moving to a position, instead of relative ones, such as moving by a distance, whenever the protocol offers the choice. Absolute commands are idempotent: sending them twice has the same effect as sending them once. Streaming devices and background readers Some instruments send data continuously without being asked. Balances can be configured to stream readings several times per second, and many sensor boards print a line per measurement. For these devices, the request-and-response pattern does not apply. The host must read continuously, split the stream into messages, and do something with each one, without falling behind. The usual approach is a dedicated reader thread or asyncio task per port. It reads lines in a loop, timestamps each one as soon as it arrives using a monotonic clock, parses it, and places the result in a queue or updates a shared latest-value structure. The rest of the program consumes from the queue. This decouples the timing of the device from the timing of the orchestration logic, and it ensures that the operating system's receive buffer never overflows, which would silently drop bytes. When a streaming device must also accept commands, the reader thread must be the only code that reads from the port, and it must be able to route replies to whoever sent the command. This is a small but real design problem. One workable pattern is to have the reader recognise replies by their form and deliver them to a waiting future, while delivering streamed data elsewhere. Another is to stop the stream, send the command, read the reply, and restart the stream, which is simpler but interrupts data collection. Multi-drop buses and addressing On RS-485, several devices share a single pair of wires, and only one may transmit at a time. Protocols for multi-drop buses therefore include a device address in each message, and the host acts as the sole master, sending a request to one address and waiting for that device alone to reply. Syringe pump families commonly work this way, with each pump in a chain set to a distinct address by a switch on its body. So does Modbus RTU. The engineering details are unforgiving. Addresses must be unique. Termination must be present at the two physical ends of the bus and nowhere else. The host adapter must switch its transmitter off promptly after sending so that it does not collide with the device's reply, which on USB-to-RS-485 adapters is usually handled automatically by the adapter chip but occasionally requires configuration. A single device that babbles, by transmitting when it should not, can take down the entire bus, and finding it requires disconnecting devices one at a time. Debugging a silent link When a serial device does not respond, the temptation is to change settings at random until something works. A methodical sequence is faster. Start by confirming the host side in isolation with a loopback test: connect the transmit and receive pins of the host's port or adapter together, open the port, write a string, and check that it reads back. This takes a minute and proves that the port exists, the driver works, and the script is writing where it thinks it is. pyserial even supports a loop:// URL that simulates a loopback in software, which is useful for testing parsing code without hardware. Next, confirm the cable and electrical levels. With the device powered and idle, measure its transmit pin against ground and confirm a mark level, then identify whether a straight or crossed cable is required. If the device has a front-panel menu for its interface, write down every setting it shows rather than trusting the manual's defaults, since a previous user may have changed them. Then talk to the device by hand, using a terminal program such as the miniterm tool that ships with pyserial, before writing any code. Seeing the raw reply, including whether the device echoes, which terminators it uses, and how quickly it answers, settles most questions that the manual leaves open. Displaying the bytes in hexadecimal reveals invisible characters: a stray carriage return, a null byte, or a leading ACK. Finally, if the device responds by hand but not from the script, compare the two byte for byte. The usual culprits are a missing or wrong terminator on the command, a script that reads too soon or clears the buffer after the reply has arrived, a DTR-triggered reset on opening, or a second program, often a forgotten terminal session or a vendor utility running in the background, holding the port open. On Linux, lsof or fuser on the device node shows which process has it; on Windows, the port simply refuses to open with an access error. A final practical note applies to every serial system: label everything. Each cable, each adapter, each device's address and communication settings. A laminated card on the side of the rig that states the port, baud rate, framing, terminator, and address of every device will save more hours than any clever code. Chapter 3: USB and the Problem of Identity USB made connecting instruments to computers easy, and made keeping them connected surprisingly hard. The same cable that lets a researcher plug in a new spectrometer in thirty seconds also introduces a bus that can reset without warning, device names that change between reboots, power management that suspends a port in the middle of an experiment, and a layer of drivers that sit between the script and the hardware. None of these problems is difficult once it is understood. All of them are baffling the first time they appear at two in the morning. What a USB instrument actually is USB is a host-controlled bus. The computer's host controller polls every device; a device can never speak unless asked. When a device is plugged in, the host enumerates it, reading a set of descriptors that identify the device by a vendor identifier (VID), a product identifier (PID), and usually a serial number string, and that declare one or more interfaces, each belonging to a device class. The operating system then binds a driver to each interface according to its class or its VID and PID. For laboratory purposes, instruments fall into a handful of patterns according to which class they present. The most common is a USB serial device. Here the instrument contains an ordinary UART, and a bridge chip translates between USB and serial. The bridge is either a dedicated chip from a vendor such as FTDI, Silicon Labs, or WCH, or it is implemented in the instrument's own microcontroller using the standard Communications Device Class, Abstract Control Model (CDC ACM). Either way, the operating system presents a virtual serial port: a COM port on Windows, a device node such as /dev/ttyUSB0 or /dev/ttyACM0 on Linux, and /dev/cu.usbserial-* or /dev/cu.usbmodem* on macOS. Everything in Chapter 2 applies, and the baud rate setting may or may not matter; on a native CDC ACM device it is often ignored entirely, while on a bridge chip it must match the instrument's UART. The second pattern is USBTMC, the USB Test and Measurement Class. Oscilloscopes, digital multimeters, source-measure units, function generators, and power supplies from the major test equipment manufacturers commonly implement it. USBTMC carries SCPI messages in structured transfers with defined message boundaries, and it is accessed through VISA, covered in Chapter 4, not through a serial port. The third pattern is a vendor-specific interface. Many spectrometers, cameras, data acquisition modules, and motion controllers use their own protocol over raw USB transfers and come with a vendor library, typically a DLL on Windows or a shared object on Linux, with Python bindings of varying quality. Integration here is dictated by the vendor library, and robustness depends heavily on how that library handles disconnection. The fourth pattern is HID, the Human Interface Device class, which some simple instruments, foot pedals, and barcode scanners use because it needs no driver. Barcode scanners in keyboard-emulation mode deserve special mention: they type their data into whatever window has focus, which is fragile for automation. Most scanners can be switched to a serial or CDC mode, which should be done for any unattended use. Stable names for unstable devices The first practical problem with USB serial devices is that their names are assigned in order of enumeration. On Linux, the first USB serial adapter to appear becomes /dev/ttyUSB0, the second /dev/ttyUSB1, and so on. Reboot the machine, or unplug and replug a cable, and the order can change. A script that hard-codes /dev/ttyUSB0 for the pump and /dev/ttyUSB1 for the balance will one day send pump commands to the balance. With luck, the balance ignores them. The robust solution is to identify devices by something intrinsic to them rather than by their order of appearance. There are three levels of identity to work with. The VID and PID identify the bridge chip or device model, which is enough when only one of each model is present. The USB serial number identifies an individual device, provided the manufacturer programmed a unique one, which reputable adapters do and some cheap clones do not. The physical port path identifies where on the bus the device is plugged in, which is stable as long as nobody moves the cable. On Linux, udev exposes all three. The directory /dev/serial/by-id/ contains symbolic links named from the vendor, product, and serial number, and /dev/serial/by-path/ contains links named from the physical port. A script can open these paths directly. Better still, a udev rule can create a meaningful name: # /etc/udev/rules.d/99-lab.rules SUBSYSTEM=="tty", ATTRS{idVendor}=="0403", ATTRS{idProduct}=="6001", \ ATTRS{serial}=="A10K3XYZ", SYMLINK+="lab/pump_base", \ MODE="0660", GROUP="dialout" After reloading the rules and replugging the adapter, the pump appears as /dev/lab/pump_base regardless of enumeration order, and the configuration file for the experiment refers to that name. The rule also sets permissions so that users in the dialout group can open the port without root privileges. The serial number shown here is a placeholder; the real one can be read with udevadm info on the device node. On Windows, COM port numbers are assigned per device and usually persist, keyed on the adapter's serial number, but they are not guaranteed to, and devices without serial numbers can be assigned a new number when moved to a different port. The portable approach, on any operating system, is to search for the device at startup using pyserial's port enumeration: from serial.tools import list_ports def find_port(vid, pid, serial_number=None): matches = [p for p in list_ports.comports() if p.vid == vid and p.pid == pid and serial_number in (None, p.serial_number)] if len(matches) != 1: raise RuntimeError(f"expected one {vid:04x}:{pid:04x}, " f"found {len(matches)}") return matches[0].device The insistence on exactly one match is the important part. Zero matches means the device is missing, which should stop the experiment before it starts. Two matches means the identification is ambiguous, which is worse, because silently picking the first one reintroduces the original problem. Identity should also be confirmed at the protocol level where the instrument allows it. Most instruments can report a model and serial number in response to a query. A driver that opens a port and immediately checks that the device on the other end identifies as the expected model has closed the last gap: even if the wrong cable ends up in the wrong adapter, the driver will refuse to proceed. Latency, buffering, and timing USB transfers data in scheduled frames. At full speed, the rate used by most USB serial adapters, frames occur every millisecond; high-speed devices use microframes of 125 microseconds. On top of that, bridge chips buffer incoming serial data and send it to the host only when their buffer fills or a timer expires. FTDI chips are the best-known case. Their latency timer defaults to 16 milliseconds, which means a short reply from an instrument can sit in the chip for up to 16 milliseconds before the host sees it. For a command-and-response protocol with many short exchanges, this adds up. A syringe pump chain that needs twenty status queries per second will spend a significant fraction of each second waiting for latency timers. On Linux, the timer for an FTDI device can be read and set through sysfs, at a path such as /sys/bus/usb-serial/devices/ttyUSB0/latency_timer, and reducing it to 1 millisecond is a common and harmless optimisation for interactive protocols. On Windows it is set in the device's advanced port properties. For streaming devices that send large volumes of data, the default is usually better, because it reduces the number of USB transfers and the load on the host. The more general lesson is that USB serial timing is not wire timing. Protocols that depend on precise inter-character gaps, most notably Modbus RTU, which uses silent intervals of three and a half character times to mark the end of a frame, can behave erratically through USB adapters because the adapter and the host's USB stack can insert or remove gaps. Most modern adapters and Modbus libraries cope, but when an RTU link through a USB adapter produces intermittent CRC errors or timeouts, latency and gap timing are prime suspects. A dedicated RS-485 interface on a PLC or a serial-to-Ethernet gateway that handles RTU framing itself is often the most reliable fix. When the bus goes away The most important difference between USB and a traditional serial port is that USB devices can disappear. A cable can be knocked, a hub can brown out when a motor starts and pulls down the supply, electrostatic discharge from a person touching a metal enclosure can reset a device, and the operating system can decide to suspend a port that looks idle. When that happens, the operating system removes the device node. Any open file handle becomes invalid; reads and writes raise exceptions. When the device reappears, it may do so with a different name. A script that assumes its port will exist forever will crash at that point, or worse, catch the exception in a broad handler and loop, logging errors, while the experiment proceeds without its instrument. Robust drivers treat disconnection as an expected event with a defined response. That response depends on the device. For a passive sensor, the driver might mark its readings as unavailable, attempt to reconnect with backoff, verify identity on reconnection, and resume. For a pump in the middle of a dispense, the orchestration layer must be told, because after reconnection the pump's state is unknown: it may have completed the move, stopped partway, or reset entirely and lost its position reference. The only safe continuation is to query its state, and if the answer is anything other than a definite known position, to re-home it or halt. This is an instance of the principle from Chapter 1: whenever communication has been lost, cached beliefs about a device must be thrown away. A reconnection is a fresh start, not a resumption. Hubs, power, and the physical layer again Many USB reliability problems are really power and grounding problems. A USB port supplies nominally 5 volts, and bus-powered hubs share a single upstream supply among all their ports. An instrument or adapter that draws more current than expected, or that sits next to a stepper driver sharing the same ground, can see its supply dip or its ground reference jump, and reset. The symptoms are intermittent disconnections that correlate with motor moves, heater switching, or someone turning on a nearby piece of equipment. A few practices eliminate most of these problems. Use externally powered hubs of decent quality, not bus-powered ones, for any rig with more than a couple of devices. Keep USB cables short, and keep them away from motor and heater power cables. Where a USB device shares ground with high-current equipment, consider a USB isolator, a small inline device that galvanically isolates the data and power lines, which breaks ground loops between the computer and the instrument. For serial instruments, isolated USB-to-serial adapters achieve the same thing. For anything critical, consider replacing the USB link with Ethernet, whose transformer-coupled ports provide isolation by design. Operating system power management deserves attention too. Both Windows and Linux can selectively suspend USB devices they believe are idle, and some bridge drivers handle resume poorly. On a dedicated automation computer, disabling USB selective suspend is standard practice. The same applies to the computer's own sleep and hibernation settings, automatic update restarts, and screen savers that launch heavy processes. A laboratory automation host should be configured as an appliance: it does one job and nothing is allowed to interrupt it without a person's decision. Vendor libraries and their habits Instruments that ship with vendor libraries bring a different set of concerns. The library usually hides the transport entirely and exposes functions such as open_device, set_integration_time, and get_spectrum. This is convenient, but the library's error handling becomes the driver's error handling. Some vendor libraries block indefinitely when a device disconnects. Some crash the process. Some are not thread-safe and corrupt state if called from two threads. Some require being called from the thread that opened the device. The defensive response is to isolate such libraries. Wrapping a vendor library in a thin class whose every method runs with an external timeout is a start. For a library that can crash or hang the process, running it in a separate process, communicating with the main orchestration program over a pipe, a socket, or a lightweight remote procedure call mechanism, means that a crash kills only that process, which the orchestrator can detect and restart. This is more work up front and it pays for itself the first time a vendor library segfaults at the end of a twelve-hour run. Where an open-source alternative to a vendor library exists, such as the python-seabreeze package for Ocean Optics spectrometers, it is often worth evaluating, both because the source can be read when something goes wrong and because such projects often handle edge cases that users have reported over years. The trade-off is that the vendor will not support it, and new models may not be covered. A driver that expects to lose its device It helps to see what disconnection-aware code looks like. The sketch below wraps a USB serial instrument so that every transaction either succeeds on a verified connection or raises an exception that tells the caller the device's state is unknown. It reuses the transact function from Chapter 2 and the find_port function above, and assumes an instrument that answers an identification query with a model string. import time, serial class DeviceStateUnknown(Exception): """Raised when the link dropped; cached state must be discarded.""" class UsbInstrument: def init(self, vid, pid, sn, expected_model): self.ids = (vid, pid, sn) self.expected_model = expected_model self.port = None def connect(self, attempts=5): for n in range(attempts): try: self.port = open_port(find_port(*self.ids)) model = transact(self.port, "ID?") if not model.startswith(self.expected_model): raise RuntimeError(f"wrong device on port: {model!r}") return except (serial.SerialException, RuntimeError, OSError): time.sleep(min(2 ** n, 30)) # bounded exponential backoff raise RuntimeError("device did not come back") def query(self, cmd): try: return transact(self.port, cmd) except (serial.SerialException, OSError) as exc: self.port = None raise DeviceStateUnknown(cmd) from exc Three decisions in this sketch matter more than its details. The reconnection loop is bounded, so a device that has genuinely failed produces an error rather than an eternal wait. Identity is re-verified on every connection, not just the first. And a lost link raises a distinct exception type rather than a generic error, which lets the orchestration layer respond specifically: re-home a pump, mark a sensor's data gap, or pause the protocol for an operator. Note also that the query method does not reconnect by itself. Whether reconnection is safe depends on what the device was doing, and that is a decision for the layer that knows. The only way to know that this code works is to exercise it. Before a new rig runs unattended, pull each USB cable in turn during a run, cut power to each hub, and watch what the software does. A test that takes twenty minutes on a Friday afternoon is far cheaper than discovering on Monday that a weekend of samples was processed with a balance that stopped reporting on Saturday morning. When to leave USB behind For many devices the most reliable USB strategy is to stop using USB. Serial device servers, small network appliances with one or more RS-232 or RS-485 ports, let a host reach serial instruments over Ethernet. Many present each port as a raw TCP socket, which pyserial can open directly with a socket:// URL, so driver code barely changes. Others install a virtual COM port driver on the host. Either way, the instrument keeps a short serial cable to a device that sits beside it, the long run to the computer is Ethernet with its built-in isolation, and the port's address no longer depends on enumeration order or the state of a hub. For a rig with several serial devices, a single multi-port server in the instrument rack is often cheaper in debugging time than any number of USB adapters. The same reasoning favours instruments with native Ethernet interfaces when both options exist, and argues for placing a small, robust computer, such as an industrial single-board machine, close to a cluster of USB-only devices so that the long connection to the main orchestration host is a network link. Each of these choices moves the fragile part of the system into a smaller, more controlled space. A USB checklist The habits in this chapter reduce to a short checklist that is worth applying to every USB device in a system before it runs unattended. Identify the device by serial number or physical port, never by enumeration order. Confirm its identity at the protocol level after opening. Set the bridge latency timer to suit the protocol. Power hubs externally and route cables away from power wiring. Disable selective suspend and host sleep. Treat every disconnection as an event that invalidates the device's state, and decide in advance what the orchestration layer should do about it. And isolate any vendor library that has not proven that it can survive a pulled cable, because sooner or later someone will test that for you. Hashtags: #LaboratoryAutomationEngineering #LaboratoryAutomation #PythonAutomation #ProgrammableLogicControllers #PLCIntegration #ScientificHardware #SupervisoryControl #RealTimeControl #SerialCommunication #RS232 #RS485 #USBInstrumentation #PySerial #PyVISA #SCPI #Modbus #OPCUA #StepperMotors #SensorInterfaces #StateMachineDesign #DeviceDrivers #HardwareInterlocks #SafetyWatchdogs #FaultTolerantAutomation #FutureOfLaboratoryAutomation
- In Vivo Imaging Systems (IVIS): Bioluminescence, Fluorescence, and Tumor Tracking
Download the Book (PDF): Introduction A mouse lies on a heated stage inside a light-tight box, asleep under isoflurane. Twelve minutes earlier it received an injection of D-luciferin into the peritoneal cavity. Somewhere in its left flank, a few million tumor cells engineered to express firefly luciferase are oxidizing that substrate and releasing yellow-green photons. Most of those photons are absorbed by hemoglobin before they travel a centimetre. Some scatter sideways and are lost. A small fraction reach the skin, leave the animal, and pass through a lens onto a cooled charge-coupled device that has been counting for thirty seconds. The software paints the result in false colour over a grey photograph of the mouse, a red-to-yellow blob that everyone in the room will call "the tumor." It is not the tumor. It is a measurement of light that escaped. That distinction is the subject of this book. Optical imaging of small animals, and in particular the family of instruments sold under the IVIS name, has become one of the most widely used tools in preclinical biology. It is fast: a cage of five mice can be imaged in a few minutes. It is sensitive: a well-labelled population of a few thousand cells can be seen through skin. It is cheap relative to magnetic resonance or positron emission tomography, it uses no ionizing radiation, and it allows the same animal to be followed from the day cells are implanted until the day the experiment ends. For cancer research, where the central question is usually whether a treatment slows the growth of a tumor, those properties are transformative. An orthotopic pancreatic tumor, a disseminated leukaemia, or a brain metastasis that could once be assessed only at necropsy can now be followed week by week in a living animal. The ease of the method is also its hazard. Because the image appears within a minute and the software reports a number with four significant figures, it is tempting to treat that number as tumor size. Yet the photon flux reported for a region of interest is the product of a long chain of factors: how many cells express the reporter, how much enzyme each cell makes, how much substrate reached those cells at the moment of acquisition, how much oxygen and ATP they had, what colour of light the reaction produced at body temperature, how deep the cells lay and what tissue sat between them and the lens, how the animal was positioned, and how the camera was set. Change any link and the number changes, sometimes by an order of magnitude, with no change at all in the number of tumor cells. The argument of this book The controlling claim is simple. An optical imaging signal is a relative measurement produced by a chain of biological, physical, and instrumental steps, and it becomes a trustworthy longitudinal readout of tumor burden only when every link in that chain is either held constant or explicitly accounted for. Good optical imaging is therefore less about the camera than about protocol: a disciplined sequence of choices about reporter, substrate, timing, anaesthesia, positioning, acquisition, region drawing, and statistical modelling, each justified by what it does to the signal. The practical corollary is that most of the variance in published bioluminescence data is self-inflicted and avoidable. Imaging at a fixed five minutes after an intraperitoneal injection, when the peak time is itself drifting as the tumor vascularizes, builds a bias into every growth curve. Drawing regions that differ in size from week to week introduces noise that no statistical test can remove. Comparing a deep orthotopic tumor with a subcutaneous one on raw flux compares their depths as much as their sizes. None of these problems requires new technology to fix. They require understanding what the instrument measures and designing the protocol accordingly. Who this book is for This book is written for the people who actually run imaging sessions and analyse the data: graduate students and postdoctoral researchers setting up a first tumor study, core facility staff who train users, principal investigators who need to judge whether a figure in a manuscript can bear the weight of its conclusions, and reviewers of such manuscripts. It assumes basic familiarity with cell culture, mouse handling, and the idea of a reporter gene. It does not assume any background in optics or image analysis; the physics that matters is explained from the beginning. The book concentrates on the planar, two-dimensional imaging that makes up the great majority of IVIS use, with bioluminescence as the primary modality and fluorescence treated as a complementary one with its own distinct problems. Three-dimensional tomographic reconstruction, available on some instruments, is discussed where it changes the practical picture, but it is not the centre of gravity. Nor is this an operating manual for any particular software version. Menus change; the principles behind them do not. How the book is organized The first chapter establishes what the instrument actually measures: how a cooled camera turns photons into counts, why calibrated radiance units exist, and what the settings of exposure, binning, aperture, and field of view do. Chapter 2 turns to the biological source of light, comparing the luciferases and fluorescent proteins available for labelling tumor cells and the practical steps for building a reliable reporter line, including the immunological cost of expressing a foreign protein. Chapters 3 and 4 deal with the two factors that most often corrupt quantitative bioluminescence. Chapter 3 treats substrate delivery: the dose, route, and timing of D-luciferin and the reasons a kinetic curve must be measured rather than assumed. Chapter 4 treats tissue optics: absorption, scattering, the dependence of attenuation on wavelength and depth, and the limited set of corrections that are genuinely available to the experimenter. Chapter 5 considers fluorescence imaging, where the need for excitation light brings a different set of problems, above all autofluorescence and the geometry of illumination. Chapter 6 assembles the preceding material into a standard operating protocol for an imaging session, from anaesthesia and warming to acquisition settings and record-keeping. Chapter 7 addresses region-of-interest quantification: the difference between total flux and average radiance, how to set background, how to handle saturation and overlapping signals, and how to keep the analysis blind and reproducible. Chapter 8 takes the resulting numbers into longitudinal modelling, comparing bioluminescence with caliper and cross-sectional imaging, examining the common growth models, and setting out statistical approaches that respect the structure of repeated measures on individual animals. The conclusion draws these threads into a short argument about what optical imaging can and cannot claim, and about the reporting standards that would make published optical data comparable across laboratories. A glossary, notes, and a list of further reading follow. A word on welfare and reporting Every protocol described here is performed on living animals, and the ethical case for optical imaging rests heavily on the claim that it reduces the number of animals needed and refines the experiments done on them. That claim is true only when the data are good. A longitudinal study that uses ten imaged mice in place of forty mice killed at four time points has earned its reduction only if the images measure what they are said to measure. Poorly controlled imaging that produces noisy, uninterpretable growth curves wastes animals as surely as a badly designed terminal study, and often obscures the fact. The discipline argued for throughout this book is scientific and ethical at once. The same logic applies to reporting. The ARRIVE guidelines for reporting animal research ask authors to describe their procedures in enough detail that others can repeat them. For optical imaging that means stating the substrate dose, route, and time to acquisition; the anaesthetic; the instrument settings; how regions were drawn; and what was done with saturated or undetectable signals. Very few published studies report all of these. By the end of this book the reader should understand why each one matters, and should be able to design a study that could be reported in full without embarrassment. Chapter 1: What the Camera Actually Measures The modern bioluminescence imager descends from a surprisingly simple observation. In 1995 Christopher Contag and colleagues at Stanford infected mice with Salmonella engineered to carry the bacterial lux operon and showed that the light produced by the bacteria inside living animals could be detected from outside the body with an intensified camera. The paper, "Photonic detection of bacterial pathogens in living hosts," established that mammalian tissue, although opaque to the eye, is not opaque to a sensitive enough detector. Within a few years the Stanford group and the company that grew from it, Xenogen, had built a dedicated instrument around a cooled charge-coupled device, and the IVIS line was born. After a series of corporate acquisitions the instruments passed to Caliper Life Sciences, then to PerkinElmer, and are now sold by Revvity, but the underlying design has changed less than the badge on the front. Understanding that design is the first step toward trusting its numbers. The instrument does one thing: it counts photons arriving at each pixel of a sensor during an exposure. Everything else, including the colour map and the units in the results table, is computation layered on top of that count. The light-tight box and the cooled sensor An IVIS imager is, in essence, a camera mounted above a stage inside a box that excludes all ambient light. The box matters as much as the camera. Bioluminescent signals from inside an animal can be a few hundred photons per second per square centimetre at the skin, many orders of magnitude fainter than room light, so any leak through a door seal or any glowing indicator lamp inside the chamber becomes a source of background. The stage is heated, typically to about 37 degrees Celsius, because an anaesthetized mouse loses heat quickly and, as later chapters explain, body temperature affects both physiology and the colour of firefly luciferase light. A manifold delivers isoflurane to nose cones so that several animals can remain anaesthetized during acquisition. The camera uses a scientific-grade charge-coupled device, usually back-illuminated (sometimes called back-thinned), which means light enters through the thinned rear surface of the silicon rather than through the layer of electrodes on the front. This arrangement gives quantum efficiencies well above 80 percent across much of the visible and near-infrared range: of every hundred photons that strike the sensor at a favourable wavelength, most generate a photoelectron. The sensor is cooled to far below freezing, commonly around minus 90 degrees Celsius, by thermoelectric elements. Cooling suppresses dark current, the trickle of thermally generated electrons that would otherwise accumulate in each pixel during a long exposure and be indistinguishable from signal. At the end of an exposure the charge in each pixel is shifted off the chip and digitized. Two sources of noise attend this step. Read noise is a fixed penalty of a few electrons per pixel per readout, added regardless of exposure length. Shot noise is the statistical fluctuation inherent in counting discrete events: if a pixel receives on average N photoelectrons, repeated measurements will scatter with a standard deviation of about √N. Shot noise is not an instrument defect but a law of nature, and it sets the ultimate floor. A pixel that collects 100 photoelectrons has a relative uncertainty of about 10 percent; one that collects 10,000 has about 1 percent. The practical consequence is that a longer exposure, which collects more photons, improves precision in proportion to the square root of the gain. Doubling the exposure does not double the signal-to-noise ratio; it improves it by about 41 percent. This square-root law recurs throughout the design of acquisition protocols. Counts and calibrated radiance The raw output of the camera is a number of counts for each pixel. A count is a digitized unit of charge, related to photoelectrons by the gain of the analog-to-digital converter. Counts are a perfectly good measurement for a single image, but they are not comparable across images taken with different settings. Double the exposure time and the counts double. Open the aperture and the counts rise. Combine pixels into larger bins and the counts per bin multiply. Move the camera to a wider field of view and the light from a given patch of skin falls on fewer pixels. To make images comparable, the manufacturer calibrates each instrument against a light source of known output traceable to national standards. The software uses that calibration, together with the exposure time, binning, aperture setting, and field of view recorded for each image, to convert counts into radiance: the number of photons leaving each square centimetre of the subject's surface per second into each steradian of solid angle. The unit is written photons per second per square centimetre per steradian, often abbreviated p/s/cm²/sr. When the software sums radiance over the area of a region and integrates over the relevant solid angle, it reports total flux in photons per second. This conversion is the reason modern optical imaging can support longitudinal studies at all. An image taken today with a 60-second exposure and small binning, and another taken next month with a 1-second exposure and large binning because the tumor has grown bright, can be compared directly in radiance units, provided both were well exposed. The calibration corrects for the camera. It does not, and cannot, correct for anything inside the animal. It is worth being clear about what radiance is not. It is not the number of photons produced by the tumor; it is the number emerging from the skin surface. It is not an absolute measure of cell number, enzyme amount, or tumor mass. It is also not a truly absolute physical measurement even of surface emission in every circumstance, because the calibration assumes a flat, Lambertian (uniformly scattering) emitting surface at a defined distance, and a mouse flank is neither flat nor uniformly placed. For relative comparisons within a well-controlled study these limitations are minor. For comparisons between laboratories, instruments, or anatomical sites they become important, and later chapters return to them. The four acquisition settings Every acquisition is defined by four adjustable parameters. Each trades one desirable property against another, and each is recorded in the image metadata so that the software can apply the radiance calibration. Table 1 summarizes their effects; the prose that follows explains the reasoning behind each row. Table 1. Acquisition settings and their effects on a bioluminescence image. Setting Increasing it does Main cost Corrected by radiance units? Exposure time Collects more photons; improves signal-to-noise by √ of the gain Longer session; risk of saturation; substrate kinetics change during exposure Yes Binning Combines adjacent pixels; raises counts per bin and sensitivity Coarser spatial resolution Yes Aperture (lower f-number) Admits more light through the lens Shallower depth of field Yes Field of view (wider) Images more animals at once Fewer pixels per animal; lower resolution Yes Exposure time is the most intuitive setting. For bright subcutaneous tumors, exposures of a second or less may suffice; for small numbers of cells deep in the body, exposures of several minutes may be needed. The software's automatic exposure mode takes a short test image and chooses an exposure intended to bring the brightest pixel into a target range. Automatic exposure is convenient and generally sound, but it means that each image in a series may have different settings, which is acceptable only because of the radiance calibration. A subtler problem arises with long exposures during rapidly changing substrate kinetics: a five-minute exposure starting at minute 8 after an intravenous injection averages over a period when the signal may be falling steeply. The image then reports a time-averaged value that depends on when the exposure started and how long it ran. Binning combines the charge from a block of adjacent pixels, such as 4 by 4 or 8 by 8, into a single value before readout. Because read noise is incurred once per binned pixel rather than once per original pixel, binning improves sensitivity substantially for faint signals. The cost is resolution. For whole-body tumor imaging, where optical scattering in tissue already blurs sources over several millimetres, moderate binning costs little real information. For imaging small superficial structures, such as individual lymph nodes or discrete metastases, it can merge sources that would otherwise be separable. Aperture is set by the f-number, or f-stop. A lower f-number means a wider opening and more light. The instruments usually offer settings from f/1 to f/8 or similar. The wide-open f/1 setting is standard for faint bioluminescence. The cost is depth of field: at f/1 the range of distances in focus is shallow, so a fat mouse's dorsal surface and a lean mouse's may not both be sharp. Because bioluminescent sources are blurred by tissue anyway, this rarely matters for quantification, but it matters for photographs and for any work at high magnification. Field of view is set by moving the camera or the stage, and is usually designated by letters that correspond to distances. The widest settings accommodate five mice; the narrowest image a small area of one mouse at higher resolution. Changing the field of view changes how many pixels a given anatomical region occupies. The calibration corrects the radiance, but the spatial sampling changes, which can affect how regions of interest capture blurred edges. The simplest safeguard is to use the same field of view for every session of a longitudinal study. Saturation and the lower limit of detection Every pixel has a maximum capacity. When the charge exceeds it, or when the digitized value reaches the top of the converter's range (65,535 for a 16-bit converter), the pixel saturates and further light is simply not recorded. A saturated image understates the true signal, and the understatement cannot be corrected afterwards. The software flags saturation, and the rule is absolute: a saturated image is not quantifiable. It must be reacquired with a shorter exposure, a higher f-number, or less binning. At the other end, an image in which the brightest pixel contains only a few dozen counts is dominated by read noise and background. The manufacturer's long-standing guidance is that a quantifiable image should have peak counts of at least several hundred, and the figure commonly quoted in training material is 600. The principle is more important than the number: an image should use a meaningful fraction of the sensor's range, far above the noise floor and safely below saturation. Automatic exposure is designed to achieve this. A related issue is the instrument's dark background. Even with the sensor cooled and the chamber sealed, an image taken with no animal present will show a small, spatially varying signal from residual dark current, readout offsets, and stray luminescence. The software subtracts a stored dark frame, and many users also take periodic background images of the empty chamber or of a non-luminescent mouse. A useful habit, discussed further in Chapter 7, is to include in every study at least one animal or region that should produce no signal, so that the effective floor under real conditions is measured rather than assumed. Cosmic rays and other artefacts Long exposures reveal a class of artefact that surprises new users: isolated bright pixels or short streaks that appear in one image and not the next. These are the tracks of cosmic-ray muons and of natural radioactivity in the surroundings, depositing charge directly in the silicon. The software includes a cosmic-ray correction that removes isolated outliers. It should be left on for quantification. On rare occasions a genuine very small, very bright source, such as a luminescent bead used for calibration, can be mistaken for a cosmic ray, but for tissue signals, which are always blurred by scattering, the correction is safe. Other artefacts come from the subject rather than the instrument. Fur scatters and absorbs light; dark fur absorbs strongly. Many laboratories use albino or nude strains for this reason, and others shave or depilate the imaged region. Urine containing excreted luciferin can glow if any luciferase-expressing material is present nearby, and contaminated bedding or gloves can transfer light-producing material. Luciferin itself on the skin at an injection site is not luminescent without enzyme, but leaking tumor cells or luciferase from a ruptured tumor can create signal in unexpected places. Keeping the instrument honest Radiance calibration is performed by the manufacturer at installation and at service visits, and users generally cannot alter it. That does not mean the instrument's behaviour can be taken on trust for the life of a study. Sensors age, filters degrade, stage heaters drift, and a camera that has been serviced midway through a twelve-week experiment may not report identical radiance for an identical source before and after. For a study whose conclusions rest on changes of a factor of two or three this is rarely decisive, but for studies seeking to detect modest treatment effects it can be. The remedy is a stable reference source imaged at every session. Several options exist. Sealed radioluminescent or phosphorescent sources provide steady output over months and can be placed in the field of view beside the animals. Commercially available light-emitting phantoms, sometimes built with embedded sources at known depths in tissue-simulating material, serve the same purpose and additionally allow users to check how the instrument handles depth. Some facilities keep a small aliquot of a fixed cell line or of purified luciferase with excess substrate for a quick check, although enzymatic sources are themselves variable and are better suited to confirming that the system works than to detecting small drifts. Whatever the reference, the practice is the same: image it with fixed settings at the start of each session, record its radiance in a log, and look at the log. A reference whose apparent output has fallen by ten percent over three months is telling the user something about the instrument that would otherwise be absorbed silently into the biological data. In core facilities that serve many groups, a shared log of this kind is one of the cheapest quality measures available and one of the least often implemented. It is also worth recording which instrument was used. Many institutions have more than one imager, sometimes of different models and generations. Radiance calibration is designed to make them comparable, and for bright signals imaged under identical conditions they usually agree reasonably well, but differences in filters, optics, and stage geometry mean they should not be used interchangeably within a single longitudinal study unless a direct comparison has been made. Comparisons of commercial optical imagers have found real differences in sensitivity and in the handling of fluorescence in particular, which reinforces the rule that one study should use one instrument wherever possible. What the instrument cannot tell you The camera captures a two-dimensional projection of light that emerged from a three-dimensional, turbid object. It does not know how deep the source was, how much tissue lay between source and skin, or whether a bright spot is a large tumor at depth or a small one near the surface. It cannot distinguish a tumor that has doubled in cell number from one whose cells now receive twice as much substrate. It cannot tell whether a fall in signal after treatment reflects cell death, reduced perfusion delivering less luciferin, hypoxia starving the enzyme of oxygen, or suppression of the promoter driving the reporter. These are not criticisms of the instrument. They define the job the rest of the protocol must do. The chapters that follow take each link of the chain in turn, beginning with the reporter that produces the light in the first place. Instruments with multiple emission filters and structured-light surface topography can partially address the depth problem through tomographic reconstruction, and Chapter 4 explains what that approach assumes and where it helps. But the fundamental point stands for all planar imaging: the number in the results table is a relative measurement of emergent light, and its interpretation depends entirely on what was held constant in producing it. Chapter 2: Choosing the Light Source Every optical signal begins with a molecule inside a cell. In bioluminescence that molecule is an enzyme, a luciferase, that oxidizes a substrate and releases the energy of the reaction as a photon. In fluorescence it is a protein or dye that absorbs a photon from an external lamp and re-emits one at a longer wavelength. The choice of reporter determines the colour of the light, which governs how much of it survives passage through tissue; the biochemical requirements of the reaction, which determine what physiological states the signal is sensitive to; and the immunological visibility of the labelled cells, which can change the biology being measured. It is the first and least reversible decision in an imaging study. Once a cell line has been engineered and a study begun, the reporter cannot be swapped without starting over. The firefly luciferase reaction The great majority of tumor imaging uses firefly luciferase from the North American firefly Photinus pyralis, usually in a codon-optimized, modified form such as luc2 that is expressed more strongly in mammalian cells. The enzyme catalyses a two-step reaction. First, D-luciferin is adenylated using ATP in the presence of magnesium ions. Second, the luciferyl-adenylate is oxidized by molecular oxygen to an excited-state oxyluciferin, which relaxes by emitting a photon. Carbon dioxide and AMP are released. Three consequences follow from this chemistry, and each has a direct bearing on how the signal should be interpreted. First, the reaction requires ATP. Only metabolically active cells produce light; dead cells, and cells in severe energy crisis, do not. This is often presented as an advantage, because it means the signal reports viable tumor burden rather than the total mass including necrotic core. It is an advantage, but it also means that anything which lowers intracellular ATP without killing cells, including some drugs, will lower the signal and may be misread as cytotoxicity. Second, the reaction requires oxygen. In well-perfused tissue this is not limiting, but tumors are notoriously heterogeneous in oxygenation. Work from the Stanford group led by Edward Graves, published in Molecular Imaging in 2007, showed that luciferase reporter activity fell under hypoxic conditions even when reporter protein levels were barely changed, and that clamping tumors or treating them with a vascular-disrupting agent reduced the bioluminescent signal accordingly. A large tumor with a hypoxic core therefore produces less light per viable cell than a small, well-oxygenated one. Antivascular and antiangiogenic treatments are the obvious danger: they may lower the signal through oxygen and substrate deprivation well before they reduce the number of living tumor cells. Third, the colour of the emitted light depends on temperature and on the local environment. At room temperature firefly luciferase emits yellow-green light peaking around 560 nanometres. At 37 degrees Celsius the spectrum shifts toward the red, with a peak near 610 nanometres. A study by Hui Zhao, Bradley Rice, Christopher Contag, and colleagues in the Journal of Biomedical Optics in 2005 documented this red shift and its importance: because red light penetrates tissue far better than green, the warm enzyme inside a living mouse is a considerably better deep-tissue reporter than measurements in a room-temperature plate would suggest. The same work compared several luciferases and showed that emission spectrum, more than raw brightness in vitro, largely determined which reporter was detected best from depth. Alternatives to firefly luciferase Several other luciferases are used in preclinical imaging, each with a distinct substrate and set of properties. Table 2 compares the most common, with approximate peak emission wavelengths at body temperature where that differs from the in vitro value. Table 2. Common reporters for in vivo optical imaging. Reporter Substrate or excitation Cofactors needed Approx. peak emission Principal strength or limitation Firefly luciferase (luc2) D-luciferin ATP, Mg²⁺, O₂ ~610 nm at 37 °C Workhorse; good depth; reports viable cells Click beetle red luciferase D-luciferin ATP, Mg²⁺, O₂ ~615 nm Red-shifted; pairs with green click beetle for dual imaging Renilla luciferase (RLuc8) Coelenterazine O₂ only ~480 nm ATP-independent; blue light poorly transmitted NanoLuc Furimazine O₂ only ~460 nm Very bright in vitro; blue light limits depth Akaluc AkaLumine-HCl ATP, Mg²⁺, O₂ ~650 nm Near-infrared emission; high deep-tissue sensitivity iRFP713 (fluorescent) Excitation ~690 nm Biliverdin (endogenous) ~713 nm Near-infrared fluorescence; no substrate injection Sources: Zhao et al. 2005; Hall et al. 2012; Filonov et al. 2011; Iwano et al. 2018. Values are approximate and vary with construct and conditions. Click beetle luciferases, from Pyrophorus plagiophthalamus, use the same D-luciferin substrate but come in green- and red-emitting variants. The red variant is useful as a deep-tissue reporter, and the pairing of green and red click beetle luciferases permits two cell populations to be distinguished spectrally after a single substrate injection, though the separation depends on spectral unmixing and is degraded by tissue, which absorbs the green component more strongly. Renilla luciferase, from the sea pansy, uses coelenterazine and requires neither ATP nor magnesium. Its emission is blue, peaking around 480 nanometres, and blue light is strongly absorbed by hemoglobin, so Renilla signals from depth are weak. Coelenterazine is also a substrate for the multidrug resistance transporter P-glycoprotein, which pumps it out of cells that express it, and it oxidizes spontaneously in serum to produce background light. Renilla luciferase found its main in vivo role as a second reporter alongside firefly luciferase, since the two use different substrates and can be imaged sequentially in the same animal. Engineered variants such as RLuc8 are brighter and more stable than the native enzyme. Gaussia luciferase, from a marine copepod, also uses coelenterazine and is naturally secreted. Its secretion makes it a poor imaging reporter but an excellent blood reporter: a few microlitres of plasma assayed in a plate reader give a measure of total tumor burden that is independent of optical depth. Some groups combine Gaussia measurement in blood with firefly imaging to separate changes in cell number from changes in optics or substrate delivery. NanoLuc, engineered by Promega scientists from a deep-sea shrimp luciferase and described by Mary Hall and colleagues in ACS Chemical Biology in 2012, uses a synthetic substrate, furimazine, and is small, stable, and ATP-independent. In vitro it is far brighter than firefly or Renilla luciferase, by roughly two orders of magnitude in the original comparisons. In vivo its blue emission near 460 nanometres limits sensitivity from depth, and furimazine has limited solubility and bioavailability. Two lines of development address this: fusing NanoLuc to a red-shifted fluorescent protein so that energy is transferred and re-emitted at longer wavelengths, as in the Antares construct, and developing substrate analogues with better pharmacology. For superficial tumors and for the sensitive detection of protein-protein interactions through split-reporter designs, NanoLuc is valuable; for deep orthotopic tumors firefly or red-shifted systems usually remain preferable. Akaluc and AkaLumine represent the most successful attempt to push bioluminescence into the near-infrared. AkaLumine, a synthetic luciferin analogue described by Takahiro Kuchimaru and colleagues in Nature Communications in 2016, produces light near 677 nanometres when oxidized by firefly luciferase. Satoshi Iwano and colleagues then evolved an enzyme, Akaluc, optimized for this substrate; their 2018 paper in Science, titled "Single-cell bioluminescence imaging of deep tissue in freely moving animals," reported that the combined system could detect very small numbers of cells in deep tissue, including in the brains of mice and marmosets. For deep tumors and small metastatic burdens the gain can be considerable. The substrate is more expensive than D-luciferin, and subsequent users have reported background signal from the liver after AkaLumine administration, which complicates imaging of abdominal tumors. As with any newer system, a pilot comparison in the relevant model is worth doing before committing a study to it. Fluorescent reporters Fluorescent proteins need no substrate, which is a large practical advantage: there is no injection, no kinetics, and no dependence on ATP or oxygen for the light-producing step, although chromophore maturation of GFP-family proteins does require oxygen. The disadvantages, taken up in detail in Chapter 5, arise from the need to deliver excitation light into the animal. Tissue autofluorescence and the absorption of both the excitation and emitted light mean that green fluorescent protein is essentially useless for whole-body imaging of anything below the skin. Red fluorescent proteins such as tdTomato and mCherry do better, and far-red proteins such as Katushka, described by Dmitry Shcherbo and colleagues in Nature Methods in 2007, better still. The most important advance for whole-body fluorescence has been the development of bacterial phytochrome-derived near-infrared fluorescent proteins. iRFP713, reported by Grigory Filonov, Vladislav Verkhusha, and colleagues in Nature Biotechnology in 2011, incorporates biliverdin, a product of heme breakdown present in mammalian cells, as its chromophore. It is excited near 690 nanometres and emits near 713 nanometres, inside the window where tissue absorption is lowest. For tumor tracking without substrate, iRFP-family proteins are now the fluorescent reporters of choice. Their brightness in a given cell depends in part on biliverdin availability, which varies between cell types and tissues. A frequent practical choice is a fusion or bicistronic construct encoding both a luciferase and a fluorescent protein. The fluorescent partner allows cells to be sorted by flow cytometry and identified in tissue sections; the luciferase does the whole-animal imaging. This combination is well suited to tumor work, where sorting for uniform expression and confirming the identity of cells at necropsy are both valuable. Building a reliable reporter line A reporter cell line is a reagent, and like any reagent it must be validated. Five properties matter. Stable integration and expression. Transient transfection is unsuitable for longitudinal work because expression is lost as cells divide. Lentiviral or retroviral transduction, or site-specific integration, gives stable genomic insertion. The promoter matters as well: the cytomegalovirus immediate-early promoter, though strong, is prone to silencing in some cell types over time and in vivo, whereas promoters such as elongation factor 1 alpha, phosphoglycerate kinase, or ubiquitin C are often more durable. Whether expression is stable should be tested directly by culturing cells without selection for several weeks and measuring luminescence per cell at intervals. A linear relationship between cell number and signal. Before implantation, a dilution series of cells plated with excess luciferin and imaged in the instrument should give a straight line on a log-log plot with a slope close to one. This confirms that the signal is proportional to cell number in the absence of tissue effects and establishes the in vitro detection limit. It does not establish the in vivo relationship, which is always weaker, but a line that fails in vitro will certainly fail in vivo. Clonal or polyclonal population. A single clone gives uniform expression but may differ biologically from the parental line in growth, invasiveness, or drug sensitivity, because clonal selection captures one sample of a heterogeneous population. A polyclonal pool, sorted for expression in a defined range, preserves heterogeneity but may drift as high- or low-expressing subpopulations outgrow others. Neither choice is universally right. What matters is that the labelled line be compared with the parental line for growth in vitro and, ideally, in vivo, and that its identity be confirmed by short tandem repeat profiling and its freedom from mycoplasma tested. Absence of a growth penalty. Early concerns that luciferase expression or the light reaction itself might slow tumor growth were addressed by Jessamy Tiffen and colleagues, who reported in Molecular Cancer in 2010 that neither the level of luciferase expression nor the bioluminescent reaction impaired growth of breast cancer or melanoma cells in vitro or in mice. That result is reassuring for immunodeficient hosts. It does not settle the question for immunocompetent ones. Immunological consequences. Firefly luciferase and fluorescent proteins are foreign proteins. In an immunocompetent host, cells expressing them can present peptides from the reporter to T cells and be attacked. Vladimir Baklaushev and colleagues, writing in Scientific Reports in 2017, studied luciferase-labelled 4T1 mammary carcinoma cells in syngeneic BALB/c mice and found that the labelled clones produced far fewer metastases than the parental line, and that the mice mounted an interferon-gamma response against a dominant luciferase epitope. Their conclusion was that the reporter restricted tumor growth and metastasis through an immune response. This is not a marginal curiosity. Many immuno-oncology studies depend on syngeneic tumors in immunocompetent mice, and if the reporter itself provokes immunity, the baseline against which a checkpoint inhibitor is tested has been altered. Several responses are available. One is to confirm, in each model, that labelled and unlabelled tumors grow similarly in the immunocompetent host. Another is to use host strains made tolerant to the reporter. Chi-Ping Day and colleagues at the National Cancer Institute described "glowing head" mice in PLoS One in 2014, which express luciferase and GFP in the anterior pituitary and are consequently immunologically tolerant to both, allowing labelled tumors to grow in immunocompetent animals without reporter-directed rejection. A third is to use reporters derived from the host species, though few such optical reporters exist. Whatever the choice, a study in immunocompetent animals that does not address reporter immunogenicity leaves an obvious question unanswered. Constitutive and conditional reporters For tracking tumor burden the reporter should be driven by a constitutive promoter, so that every living tumor cell makes roughly the same amount of enzyme regardless of what the cell is doing. The aim is for light to depend on how many cells there are, not on their state. It is worth stating this plainly because the same instruments and substrates are widely used for a quite different purpose: reporting on biology rather than burden. In a conditional reporter the luciferase is placed under the control of a promoter or response element that responds to a pathway of interest, such as hypoxia-responsive elements, NF-kappa-B binding sites, or a cell-cycle-regulated promoter. Light then reports pathway activity multiplied by the number of cells. Split-luciferase designs go further, dividing the enzyme into two inactive fragments fused to two proteins whose interaction reconstitutes activity, so that light reports a protein-protein interaction. Degron-tagged luciferases, whose stability depends on a signalling event, report on proteolysis or kinase activity. These are powerful tools, but they pose an interpretive problem in tumor studies, because a change in signal could reflect a change in pathway activity, a change in cell number, or both. The standard solution is a second, constitutive reporter in the same cells, read with a different substrate or at a different wavelength, whose signal serves as a denominator. Firefly luciferase driven by a responsive promoter combined with a constitutive Renilla or NanoLuc reporter is a common pairing. The ratio of the two signals approximates pathway activity per cell, though only approximately, because the two reporters emit different colours and are therefore attenuated differently by tissue. A change in tumor depth or composition shifts the ratio even if the biology is unchanged. The same caution applies to any two-colour measurement in vivo, and Chapter 4 explains why. A related consideration is that even constitutive promoters are not perfectly constant. Treatments that alter global transcription or translation, such as inhibitors of the mTOR pathway, histone deacetylases, or protein synthesis, can reduce reporter expression per cell. A drug that reduces bioluminescence by half within a day of the first dose, before any plausible change in cell number, is more likely acting on reporter expression, substrate handling, or cellular ATP than on tumor burden. A short in vitro experiment, treating the labelled cells with the drug at relevant concentrations and measuring luminescence per viable cell after a few hours, identifies this problem before the animal study begins and costs almost nothing. Matching the reporter to the question No single reporter is best for every study. For subcutaneous tumors monitored for response to a cytotoxic drug in immunodeficient mice, firefly luciferase with D-luciferin is inexpensive, well characterized, and more than sensitive enough. For small orthotopic or metastatic lesions deep in the abdomen, lungs, or brain, a red-shifted system may be worth the added cost. For studies in which substrate kinetics would be confounded by the treatment, as with antivascular agents, a near-infrared fluorescent protein offers a readout that does not depend on substrate delivery, though it has its own sensitivity limits. For immunotherapy studies, the immunological question must be resolved first. The next chapter turns to the step that, more than any other, introduces avoidable variance into firefly luciferase imaging: getting the substrate to the cells. Chapter 3: Getting the Substrate There Firefly luciferase does nothing without D-luciferin, and the animal does not make any. Every bioluminescence image of a firefly-labelled tumor is therefore an image of a pharmacological event: a small molecule was injected somewhere, absorbed, distributed through the circulation, delivered across the capillary wall and into tumor cells, and was simultaneously being cleared by the kidneys and liver. The light captured at any moment reflects the concentration of substrate inside the tumor cells at that moment, and that concentration rises, peaks, and falls over tens of minutes. An image taken at the wrong time, or at a time whose relationship to the peak changes over the course of the study, is a biased measurement no matter how carefully everything else was done. This chapter covers preparation of the substrate, the choice of dose and route, the measurement of kinetic curves, and the strategies for choosing when to acquire. It is the most practical chapter in the book because this is where most avoidable variance enters. Preparing D-luciferin D-luciferin is supplied as the free acid or, more commonly for in vivo use, as the potassium or sodium salt, which dissolve readily in aqueous buffer. The conventional stock for mice is 15 milligrams per millilitre in Dulbecco's phosphate-buffered saline without calcium and magnesium, sterile-filtered through a 0.2 micrometre filter. Injected at 10 microlitres per gram of body weight, this delivers 150 milligrams per kilogram, the dose that appears in most published protocols and in the instrument manufacturer's guidance. A 20-gram mouse receives 200 microlitres. Luciferin in solution is light-sensitive and slowly degrades. Good practice is to prepare a single batch sufficient for the study or a substantial part of it, divide it into single-use aliquots, and store them frozen and protected from light. Thawing a fresh aliquot for each session and discarding the remainder avoids the gradual loss of potency that comes from repeated freeze-thaw cycles or storage at 4 degrees. Using a single lot of substrate for an entire longitudinal study removes one more source of drift. If a new lot must be introduced, testing it side by side with the old one on a plate of labelled cells takes a few minutes and documents any difference. The dose should be calculated from each animal's measured weight on the day of imaging, not from a nominal weight. Tumor-bearing mice may lose weight as disease progresses, and a fixed volume given to an animal that has lost 15 percent of its body weight is a 15 percent higher dose per kilogram. Whether that matters depends on where on the dose-response curve the protocol sits, which is itself worth knowing. Dose and the question of saturation It is commonly assumed that 150 milligrams per kilogram saturates the luciferase in tumor cells, so that small variations in dose do not affect the signal. The evidence suggests the assumption is not generally safe. Markus Aswendt and colleagues, optimizing brain bioluminescence in transgenic mice and reporting in PLoS One in 2013, compared doses of 15, 150, 300, and 750 milligrams per kilogram and found that photon emission rose with dose without reaching saturation across that range, while the time to peak lengthened at higher doses. Their optimized protocol for brain imaging used 300 milligrams per kilogram injected before isoflurane anaesthesia, which roughly tripled the signal relative to the conventional 150 milligrams per kilogram injected after induction. The brain is an extreme case, because the blood-brain barrier restricts luciferin entry and efflux transporters actively remove it. Yimao Zhang, Martin Pomper, and colleagues at Johns Hopkins showed in Cancer Research in 2007 that D-luciferin is a substrate for the ATP-binding cassette transporter ABCG2, also known as breast cancer resistance protein, and that xenografts expressing the transporter produced substantially less bioluminescence as a result. Joshua Bakhsheshian, Matthew Hall, and colleagues at the National Cancer Institute later exploited the same property, reporting in the Proceedings of the National Academy of Sciences in 2013 that the brain bioluminescence of luciferase-expressing transgenic mice increased when ABCG2 inhibitors, including the kinase inhibitors gefitinib and nilotinib, were given alongside the substrate. That finding carries a warning for drug studies: a treatment that inhibits ABCG2 can raise tumor bioluminescence through substrate retention, masking a genuine reduction in tumor burden, and a treatment that induces the transporter can do the opposite. Tumors elsewhere in the body may be closer to saturation at standard doses. But the general lesson holds: dose is a variable, it is probably not fully saturating in many models, and it should therefore be held precisely constant on a per-kilogram basis. A pilot comparing two doses in the model of interest will show whether a higher dose buys useful sensitivity. Route of administration Three routes are in common use. Their properties differ in ways that matter for quantification, as Table 3 summarizes. Table 3. Routes of D-luciferin administration in mice. Route Time to peak signal Peak intensity Repeatability Practical notes Intraperitoneal (IP) Typically ~10–20 min; broad plateau Lower than IV Poorer; occasional failed injections Easy; standard in most protocols Intravenous (IV, tail vein or retro-orbital) A few minutes; faster decline Highest (5.6-fold IP in one study) Better Technically demanding; narrow imaging window Subcutaneous (SC) Variable; depends on site and tumor vascularity Intermediate Intermediate Easy; avoids gut injection; kinetics drift as tumor grows Sources: Keyaerts et al. 2008; Inoue et al. 2010; Aswendt et al. 2013. Times vary with model, dose, anaesthetic, and tumor site. Intraperitoneal injection is the most widely used route because it is quick, requires little skill, and gives a broad peak that is forgiving of small timing errors. Its drawbacks are real, however. The needle occasionally enters the intestine, the bladder, abdominal fat, or the subcutaneous space instead of the peritoneal cavity, and the resulting absorption can be much slower or much lower. Such misinjections are usually not apparent at the time. Their signature in the data is an animal whose signal collapses at one time point and recovers at the next, an outlier that is easily mistaken for biology. Absorption from the peritoneum also depends on the state of the abdomen: ascites, peritoneal tumor deposits, or bowel distension change it. Intravenous injection, usually into a lateral tail vein or the retro-orbital sinus, delivers the full dose into the circulation at once. The peak arrives within a few minutes and the signal then declines as luciferin is cleared. Marleen Keyaerts and colleagues at the Vrije Universiteit Brussel compared intravenous and intraperitoneal administration in mice bearing subcutaneous luciferase-expressing rhabdomyosarcoma, reporting in the European Journal of Nuclear Medicine and Molecular Imaging in 2008. Peak photon emission was 5.6 times higher after intravenous injection, time to peak was shorter and less variable, and repeated measurements four hours apart were more repeatable, with a coefficient of repeatability of 80.2 percent for intravenous against 95.0 percent for intraperitoneal injection. Those coefficients are large in both cases, which is itself a sobering finding about the day-to-day reproducibility of bioluminescence, but the intravenous route was clearly better. The cost is technical: tail-vein injection in a small or dark-tailed mouse requires practice, and repeated intravenous injection over weeks can damage veins. Subcutaneous injection, typically into the loose skin of the neck or back, is almost as easy as intraperitoneal injection and avoids the risk of injecting into the gut. Absorption is slower and depends on local blood flow. It is a reasonable default when intraperitoneal imaging is complicated by abdominal tumors, but its kinetics are sensitive to the site of injection and should be characterized in the model. A fourth route, direct intratumoral injection, maximizes local substrate but damages the tumor, distributes unevenly, and is not suitable for repeated measurement. It has a place in terminal experiments but not in longitudinal ones. Whatever route is chosen, it must be the same for every animal at every time point in a study. Switching routes midway, even between intraperitoneal and subcutaneous, changes the kinetic curve and therefore the relationship between the acquired image and the peak. The kinetic curve The central practical fact about luciferin is that the signal after injection follows a curve rather than reaching a steady value. Its shape reflects absorption into the blood, delivery to the tumor, uptake into cells, consumption by the enzyme, efflux, and clearance. Biodistribution studies with carbon-14-labelled luciferin by Frank Berger, Sanjiv Gambhir, and colleagues, published in 2008, found that after intravenous injection the substrate concentrated early in the kidneys and liver and later in the bladder and small intestine, reflecting its routes of elimination, and that uptake kinetics differed profoundly between intravenous and intraperitoneal routes. They found no clear trapping of the substrate in luciferase-expressing tissue, which means tumor cells do not accumulate a reservoir; the signal tracks the concurrent supply. A kinetic curve is measured by injecting the animal and acquiring a series of short images, for example every one or two minutes, from shortly after injection until the signal has clearly passed its peak and begun to decline, often 30 to 40 minutes for intraperitoneal injection. The software can display the total flux of a region across the sequence, from which the time to peak, the peak value, the width of the plateau, and the rate of decline can be read. Every new model, meaning every combination of cell line, tumor site, host strain, route, dose, and anaesthetic, deserves a kinetic curve before the study begins. Sites differ markedly. Subcutaneous flank tumors, orthotopic mammary tumors, intracranial tumors, bone metastases, and lung lesions can all peak at different times after the same injection. Why a fixed time point can mislead The simplest protocol is to inject every animal and image at a fixed interval, say 10 minutes. This is acceptable only if the time to peak is stable across animals and across the duration of the study. There is good evidence that it is not always stable. Yusuke Inoue and colleagues at the University of Tokyo imaged mice bearing subcutaneous tumors repeatedly after subcutaneous luciferin injection, reporting their findings in the International Journal of Biomedical Imaging in 2010. The time to peak was longer shortly after cell inoculation and shortened progressively over the first days, reaching a plateau of about 10 minutes by day 10. The authors attributed the change to the gradual establishment of tumor vasculature. Although signals at fixed time points correlated strongly with the peak signal overall, the signal at 5 or 10 minutes represented a smaller fraction of the peak early in the study than later. The effect was to underestimate early tumor burden and therefore overestimate the rate of tumor growth. The mechanism generalizes beyond this one model. Anything that changes tumor perfusion changes the kinetic curve. Tumors grow new vessels as they establish, develop necrotic and poorly perfused regions as they enlarge, and respond to many treatments, above all antiangiogenic and vascular-disrupting agents, with changes in blood flow. A treatment that slows luciferin delivery shifts the peak later and lowers it, and an image taken at a fixed early time will register a larger fall in signal than has occurred in tumor burden. The error runs in the direction of exaggerating treatment effect. Zain Paroo, Ralph Mason, and colleagues at UT Southwestern, in an early validation study published in Molecular Imaging in 2004, found a highly dynamic kinetic profile after both intraperitoneal and intratumoral substrate administration, but also strong correlations, with r greater than 0.8, between caliper-measured tumor volume and peak signal, area under the signal curve, and signal at specific time points. Their conclusion was that single-time-point imaging could serve for quantitative assessment of tumor burden where appropriate precautions were taken. Both findings are true together. A fixed time point can correlate well with tumor volume across a wide range while still carrying a systematic bias large enough to distort a growth rate or a treatment effect. Strategies for choosing when to image Three strategies are defensible, in increasing order of rigour. A fixed time on a verified plateau. Measure kinetic curves in a subset of animals at the start of the study and again at an intermediate time and near the end, including treated animals. If the plateau reliably spans the chosen time in all conditions, a fixed acquisition time is justified. Choose a time in the middle of the plateau rather than at its leading edge, so that small timing errors and small shifts in kinetics have minimal effect. For intraperitoneal injection in many subcutaneous models this ends up somewhere between 10 and 20 minutes, but the point is to measure rather than to assume. Peak signal from a short kinetic series at every session. Acquire a sequence of images, for example every two minutes from 6 to 24 minutes, and take the maximum for each animal. This costs more instrument time and more anaesthesia but is robust to shifts in kinetics between animals and over time. With five animals imaged together, a sequence of this kind adds perhaps fifteen to twenty minutes per cage. Area under the kinetic curve. Integrating the signal over a defined window uses all the data from a kinetic series and is less sensitive to noise in any single frame. It measures total light output over the window rather than peak rate, and is somewhat less intuitive to interpret, but for studies where treatment effects on perfusion are expected it is a strong choice. In all cases the interval between injection and acquisition must be recorded precisely for each animal. When five animals are injected in sequence at 30-second intervals and imaged together, the first injected animal is imaged two minutes later in its kinetic curve than the last. On a broad plateau this is harmless; on a steep curve it is not. Staggering the injections to match the position of each animal in the field, or injecting all animals within a short and consistent interval, keeps the timing uniform. Anaesthesia and substrate timing Most imaging is done under isoflurane, and the relative timing of anaesthesia and injection influences the signal. Isoflurane depresses cardiac output and respiration and lowers body temperature, and all three affect substrate delivery and the reaction itself. Aswendt and colleagues found a large gain in signal when luciferin was injected before induction of isoflurane anaesthesia rather than after, presumably because absorption and distribution proceeded under normal circulation for the first minutes. Injectable anaesthetics such as ketamine with xylazine or pentobarbital did not improve peak emission in their comparison, and they have longer recovery times. The practical rule is again consistency: choose a sequence, whether injection before induction or after, and apply it identically at every session. Many protocols inject awake animals intraperitoneally, return them to the cage for a few minutes, then induce anaesthesia in a chamber and transfer them to the imaging stage, timing acquisition from the injection. Others induce first for the sake of accurate injection. Either works if it is held constant and the timing is recorded. Repeated imaging on the same day Occasionally an animal must be imaged twice in one day, for example to repeat a failed acquisition or to image with two substrates. Luciferin clears over hours, and residual substrate from the first injection will add to the second. Keyaerts and colleagues repeated their acquisitions after four hours for their repeatability analysis. A delay of several hours, and preferably a check that the signal has returned to background before reinjection, is prudent. When two luciferases with different substrates are used, such as firefly and Renilla or NanoLuc, the substrate whose signal decays faster is usually imaged first, and the second is given only after the first signal has fallen away. Cross-reactivity between substrates is limited but not always zero, and should be checked in cells. What to take from this chapter The luciferin injection turns every imaging session into a small pharmacokinetic experiment. That experiment has to be run identically each time, and it has to be characterized well enough that the chosen acquisition time sits where the signal is least sensitive to the things that vary. When the treatment under study is expected to affect tumor perfusion, the substrate step becomes a direct confounder of the outcome, and acquiring kinetic data at every session is the only way to separate a change in tumor burden from a change in delivery. Hashtags: #InVivoImagingSystems #IVIS #BioluminescenceImaging #FluorescenceImaging #TumorTracking #OpticalImaging #FireflyLuciferase #LuciferaseReporters #DLuciferin #PhotonFlux #RadianceQuantification #LongitudinalImaging #ReporterGeneImaging #TumorBurden #SubstrateKinetics #TissueAttenuation #OpticalScattering #NearInfraredImaging #FluorescentReporters #RegionOfInterestAnalysis #AcquisitionSettings #ReporterValidation #PreclinicalImaging #SmallAnimalImaging #FutureOfOpticalImaging
Latest Book Releases:










































